- Corrected bad overwriting of error code. Problem reported by R. Walker and discussed in JIRA ticket https://its.cern.ch/jira/browse/ATLASPANDA-317
- Fixed direct access issue for some ANALY sites (requesting DBRelease files not from CVMFS and configured to use non xrdcp copytools)
- Sitemover retry policy updates: Implemented random sleep time for stagein/stageout in case of failures set sleep time to (20..50) secs for stage-in and (4.9..5) minutes for stage-out
- Changes in stage-in workflow (ignore NO_REPLICA error and continue stage-in tries for other available protocols)
- New MV sitemover implementation (simple mover for ND)
NERSC updates:
- Changed path for batchid file
- Removed logic that failed job after some number of failed calls to polling command. Failure should be decided by the HpcManager module or above, not by the sheduler interface module
- Replaced sys.exit calls with exception handling. We do not want to use sys.exit in Yoda because MPI will not exit properly if one rank calls sys.exit
- Moved the logic to kill the job to HPCManager for when the scheduler commands fail. Gave the command 120 chances to succeed with 60 second wait time. So 2 hrs.
Code contributions from Wen Guan, Alexey Anisenkov, Taylor Childers and Paul Nilsson.
Version info
General updates:
- Corrected bad overwriting of error code in getNodeStructure(). Problem reported by R. Walker and discussed in JIRA ticket https://its.cern.ch/jira/browse/ATLASPANDA-317 (PandaSe\
rverClient)
Updates from Alexey Anisenkov:
https://github.com/PanDAWMS/pilot/pull/77
Commit Summary
- Implemented random sleep time (within min and max values) for stagein/stageout in case of failures
- Set sleep retry time values: (20..50) secs for stage-in, (4.9..5) minutes for stage-out
File Changes
M Mover.py (2)
M movers/base.py (14)
M movers/mover.py (43)
https://github.com/PanDAWMS/pilot/pull/80
Commit Summary
xrdcpSiteMover: switch to use legacy "-adler" option for upload since at some site newer --cksum does not work ([3013] query chksum is not supported error)
xrdcpSiteMover: reverted back xrdcp to use --cksum for proper server side remote checksum validation (will not work until xroot server provides remote checksum calc feature)
xrdcpsitemover: make a call to get xrdcp client version
xrdcp clean up
xrdcpSiteMover: implemented proper CRC verification in case xrootd does not not support server-side checksum calculation of fly
cosmetic verbose print
Merge branch 'main-dev' of https://github.com/PanDAWMS/pilot into main-dev
Merge branch 'main-dev' of https://github.com/PanDAWMS/pilot into main-dev
File Changes
M PilotErrors.py (6)
M SiteInformation.py (6)
M movers/base.py (4)
M movers/xrdcp_sitemover.py (19)
Updates from David Cameron:
https://github.com/PanDAWMS/pilot/pull/78
Commit Summary
mv implementation of new site mover
File Changes
A movers/mv_mover.py (71)
https://github.com/PanDAWMS/pilot/pull/79
Commit Summary
Fixed pilot error when killed with signal
File Changes
M pUtil.py (1)
Updates from Taylor Childers:
https://github.com/PanDAWMS/pilot/pull/81
Commit Summary
Change path for batchid file
Removed logic that failed job after some number of failed calls to polling command. Failure should be decided by the HpcManager module or above, not by the sheduler interface modu\
le
Replaced sys.exit calls with exception handling. We do not want to use sys.exit in Yoda because MPI will not exit properly if one rank calls sys.exit.
Moved the logic to kill the job to HPCManager for when the scheduler commands fail. Gave the command 120 chances to succeed with 60 second wait time. So 2 hrs.
Merge remote-tracking branch 'upstream/main-dev'
Fixed typo
Changed batchid path
File Changes
M HPC/HPCManager.py (21)
M HPC/HPCManagerPlugins/slurm.py (8)
M HPC/pandayoda/yodacore/Yoda.py (17)
M RunJobHpcEvent.py (5)
No comments:
Post a Comment