* When available, the pilot will now send the job definition id with the tracing report. Requested by Angelos Molfetas
* The looping job killer has been rewritten. The old version used the find command of a known output file to locate output directory which is not optimal. The new version instead uses the find command to look for all files updated within the looping limit
* Reverted to 3h looping limit for analysis jobs. We need to keep a look on how that works out in practice
* Using PRODDISK space token when job.cloud is not the same as schedconfig.cloud
* Added exception handling in pilot TCP server handler in case of unusual/bad TCP messages. Previous pilot release contained a fix for too long TCP messages. This update will improve further on that (pilot will not crash due to an exception as seen before when TCP message contains bad data)
* Major refactoring of several essential functions in main pilot module related to pilot-Panda server communications and job log creation (new modules created: PandaServerClient, JobLog). This was a requirement for the new job recovery code which is now being developed in a separate module (JobRecovery). Further refactoring is needed to complete the new job recovery algorithm but only related to actual job recovery. The current job recovery code is still integrated with pilot.py but will soon be removed and retired
* Commented-out dumping of time-ordered pilot log to the batch stdout. This will essentially reduce the batch log with a factor two
* Now sending countryGroup to dispatcher again (if pilot was launched with '-o
* Added traceback to pilot time-out handler timed_command() call in get_data() and put_data() (dCacheSiteMover). Requested by Charles Waldman and Asoka De Silva, needed for debugging a rare dCache problem at TRIUMF
* Dispatcher error codes updated (by Jose Caballero)
* runJob module now imports stat module. Requested by Jose Caballero (needed for glExec testing)
* Any found corrupt output files are now reported to Consistency server. Note: this feature needs at least DQ2 Client 0.1.35 which now has also been deployed in the US (Xin Zhao et al). Requested by Cedric Serfon
* If local optional environmental variable NON_LOCAL_ATLAS_SCRATCH is set to true, the pilot will only perform space/size checks (user work dir size etc) every 30 minutes instead of the default 10 minutes. The space/size checks are using 'find' and 'du' commands which are very heavy for especially Lustre file systems. Requested by Leslie Groer et al (CSCS). Note: during reviewing of the site mover code, it was discovered that the find command is used to get the mod time of a file, which is unnecessary. It will be corrected in a later pilot version
Version info
Revert to 3h looping limit for analysis jobs (pilot)
Added jobDefId to getInitialTracingReport() (Mover)
Added cloud to Job class (Job)
Getting cloud in setJobDef() (Job)
Added jobCloud to mover_put_data() calls in moveLostOutputFiles(), transferLogFile(), transferAdditional() (pilot)
Added jobCloud to mover_put_data() calls in stageOut() (runJob)
Added jobCloud to mover_put_data() (Mover)
Added jobCloud to getProperSpaceTokenList() (Mover)
Using PRODDISK space token when job.cloud is not the same as schedconfig.cloud in getProperSpaceTokenList() (Mover)
Defaulting to dat extension in case json file is not available in getQueuedataFileName() (pUtil)
Updated looping job killer to find all modified files within N minutes (loopingJobKiller() in pilot)
Added check boolean to getQueuedataFileName(), used in getQueuedata() (pUtil)
Corrected wrong filename in replaceQueuedataField() (pUtil)
Resetting None value to empty string in readpar() [JSON issue] (pUtil)
Reverting to pickle file extension in getFilename(). JSON not compatible with Node.Node structure (JobState)
Corrected missing value in getpar() (pUtil)
Added cloud to displayJob() (Job)
Added exception handling to updateHandler:handle() in case of pilot server failure (pilot)
Moved getMetadata(), makeJobReport() from pilot to pUtil
Renamed transferAdditional() to transferAdditionalFile() to avoid possible name conflict with boolean in postJobTask (pilot)
Now only adding replica to skipped.xml if it is the last attempted replica in mover_get_data() (Mover)
Cleaned up verifySpaceToken() (Mover)
Commented-out dumping of pilot log to the batch stdout in runMain() (pilot)
Added logFile argument to mover_put_data() (Mover)
Added logFile to mover_put_data() call in stageOut() (runJob)
Added logFile to mover_put_data() calls in moveLostOutputFiles(), transferLogFile(), transferAdditionalFile() (pilot)
Added logFile to mover_put_data() calls in transferLogFile(), transferAdditionalFile() (JobLog)
Added logFile argument to chirp_put_data() (Mover)
Added logFile in put_data() calls in chirp_put_data(), mover_put_data() (Mover)
Added logFile handling in put_data() (lcgcpSiteMover)
Using job.logFile in verifyOutputFileSizes() instead of .log. (pilot)
Added logFile argument to genSubpath() (SiteMover)
Added logFile to genSubpath() call in put_data() (BNLdCacheSiteMover, SRMSiteMover)
Cleaned up put_data() a bit (SRMSiteMover)
Renamed getFileInfo() to getOutputFileInfo() (pUtil)
Using logFile instead of .log. in getOutputFileInfo() (pUtil)
Updated getFileInfo() to getOutputFileInfo() in createFileMetadata() (runJob)
Now sending countryGroup to dispatcher again (if pilot was launched with '-o
Added traceback to timed_command() call in get_data() and put_data() (dCacheSiteMover)
Created safe_call() (pUtil)
Dispatcher error codes updated (by Jose) in getNewJob() (pilot)
Created getDispatcherErrorDiag() (pUtil)
getNewJob() (pilot) and toServer() (pUtil) are now using getDispatcherErrorDiag()
Importing stat into runJob (requested by Jose, needed by glExec)
Removed transferLogFile(), transferAdditionalFile(), createLogFile(), removeLockFile() (pilot)
Created reportFileCorruption() (SiteMover)
mover_put_data() is now using reportFileCorruption() to report corrupted files to consistency server (Mover)
put_data() is now using reportFileCorruption() to report corrupted files to consistency server (lcgcpSiteMover)
Added surl to put_data_retfail() (SiteMover)
Now returning surl with put_data_retfail() for corrupted file (xrootdSiteMover, xrdcpSiteMover, lcgcp2SiteMover, LocalSiteMover,
dCacheSiteMover, dCacheLFCSiteMover, rfcpLFCSiteMover)
Accepting lcgcp2 as lcg-cp2 in getSiteMover() (SiteMoverFarm)
Created removeTree() (JobLog)
Created setUpdateFrequency() (pilot)
If local env variable NON_LOCAL_ATLAS_SCRATCH is true, then update_freq_space will be set to 30*60 (pilot)
Added more debugging info to get_data() and put_data(), error handling (dCacheSiteMover)
No comments:
Post a Comment