Thursday, June 8, 2017

69.0

Summary

Memory Monitoring
- Fix for memory monitor setup on Nordugrid (already used in production on ND)

Zip maps
- Pilot can extract ZIP_MAP info from jobPars to create archives of output files
- Note: unzip should be implemented separately, eg. as a new option for rucio download – it could of course be implemented in pilot as well, but needs discussion

Support for simplified payload setup
- If job def noExecStrCnv is set, jobPars is expected to contain payload setup (i.e. including ‘source asetup.sh ..’)
- Note: Only tested trivially on pilot side since no support yet in other tools

Initial container support
- Pilot can execute payload and stage-in/out with Singularity if activated via schedconfig.catchall by specifying singularity options
- Pilot uses cmtconfig/platform info to determine image file name (located in cvmfs)
- Stage-in/out currently only with rucio and gfal_copy so far

Test versions of asetup and xrootd
- Bug fixes for testing new versions of asetup and xrootd (ALRB_asetupVersion=testing)

Event Service
- Support for Prefetcher (from release 21.0.21 / non-digit releases - requires 'useprefetcher' in schedconfig.catchall since Prefetcher is still in test mode)
- Now allowing for remote input files if local SE is down
  Pilot checks local SE first
  Note: only relevant for AES in this release, schedconfig.catchall must include “allow_remote_inputs”
- Using normal SE:s with AES (i.e. instead of OS:s)
  Fix for using normal SE:s with merge jobs

Contributors: P. Nilsson, W. Guan, D. Drizhuk, A. Anisenkov, D. Cameron

Version info

69.0

General changes:
- Created handleAdditionalOutFiles() to cleanup main function a bit (RunJob)
- Corrected error messages (grammatical) in stageout(), do_put_files() (mover)
- Removed useless lfchost from getFileTransferInfo() (Experiment)
- Removed useless lfchost from getAnalysisRunCommand() (ATLASExperiment)
- Removed sourcing of copysetup in getAnalysisRunCommand() (ATLASExperiment)
  Tadashi's comment: IIRC, it was required for some trfs to read input files directly from BNL dcache which was unavailable in login shell by default.
  (BNL queues have used the new movers for a while and schedconfig.copysetup has been set to OLDMOVERDEPRECATED with no reported problem)
- Removed deprecated function getSetupFromCopysetup() (Experiment)
- Removed deprecated function transferAdditionalCERNVMFiles() and its call (JobLog)
- Removed deprecated function transferAdditionalFile() and its call (JobLog)
- Removed deprecated variable transferAdditional from postJobTask() (JobLog)
- Removed deprecated usage of getChecksumCommand(), replaced with "adler32" string in createMetadataForOutput(), 2 times (JobLog)
- Removed deprecated usage of getChecksumCommand(), replaced with "adler32" string in getXML() (PandaServerClient)
- Removed deprecated usage of getChecksumCommand(), replaced with "adler32" string in createFileMetadata() (RunJob)
- Removed deprecated usage of getChecksumCommand(), replaced with "adler32" string in createFileMetadata(), createFileMetadata4EventRange() (RunJobEvent)
- Removed deprecated function getChecksumCommand() (pUtil)
- Removed getSiteMover() and its import in createMetadataForOutput() (JobLog)
- Using fixed time-out instead of time-out from old site movers in for looping stage-outs in __loopingJobKiller() and removed usage of getSiteMover() (Monitor)
- Removed usage of deprecated function getSiteMover() in getXML() (PandaServerClient)
- Removed deprecated use of file stager variables in getAnalysisRunCommand() (ATLASExperiment)
- Removed deprecated lfchost in updateRunCommandList() (RunJobUtilities)
- Added a new function for calculating adler32 (same as Harvester), calc_adler32() (SiteMover)
- Avoid using the NordugridATLASExperiment::getJobExecutionCommand() (NordugridATLASExperiment)

Work directory
- Work directory naming has been simplified to make it simpler for especially checkpointing
- Removed time stamp and panda id from job work directory in mkJobWorkdir() (Job)

Memory monitor
- Created new getUtilityCommand() for Nordugrid since default function has problems with complex ND release (NordugridATLASExperiment)

Event Streaming Service
- Now waiting for Prefetcher to finish in __main__() (RunJobEvent)
- Using self.__usePrefetcher variable in adaptEventRanges(), if true add pfn field to event range (RunJobEvent)
- Renamed adaptEventRanges() to addPFNsToEventRanges() (RunJobEvent)
- Created addPFNToEventRange() (RunJobEvent)
- Added 'prefetch' support in get_input_files() (Job)
- Updated setUsePrefetcher() (RunJobEvent)

- Added updateJobParsForBrackets() for testing (RunJobEvent)

jobPars based payload setup:
- Added noExecStrCnv (Job)
- Added asetup boolean to getModernASetup() (ATLASExperiment)

ZIP_MAP support:
- Added zipmap data member (Job)
- Processing jobPars for ZIP_MAP info in new functions removeZipMapString() and populateZipMap() (Job)
- Created addArchivesToOutput() to update output file info for archives (Job)
- Created createArchives() used with ZIP_MAP info (RunJob)

Singularity support:
- Created new module (Singularity)
- Added cmtconfig to idat dictionary in setJobDef() so that it will be propagated with the file specification down to the movers (Job)
- Added the singularity wrapper to the stage-in command in stageIn() (rucio_sitemover)
- Added the singularity wrapper to the stage-out command in stageOut() (rucio_sitemover)
- Added the singularity wrapper to the stage-in command in stageInFile() (gfalcopy_sitemover)
- Added the singularity wrapper to the stage-out command in stageOutFile() (gfalcopy_sitemover)
- Sending cmtconfig to FileSpec in do_put_files() (mover)
- Added cmtconfig to _infile_keys and _outfile_keys in FileSpec (Job)
Updates from D. Drizhuk:

  https://github.com/PanDAWMS/pilot/pull/129

Commit Summary

Fix for https://its.cern.ch/jira/browse/ATLASPANDA-322?focusedCommentId=1413036&page=com.atlassian.jira.plugin.system.issuetabpanels%3Acomment-tabpanel#comment-1413036
One more param

File Changes

M ATLASExperiment.py (7)
M ATLASSiteInformation.py (2)
M Experiment.py (4)
M SiteInformation.py (13)

  https://github.com/PanDAWMS/pilot/pull/130

Commit Summary

Fix for https://github.com/PanDAWMS/pilot/pull/129#issuecomment-297464358

File Changes

M ATLASSiteInformation.py (2)
M Experiment.py (2)

Updates from W. Guan:

  https://github.com/PanDAWMS/pilot/pull/131

Commit Summary

1) allowRemoteInputs,
   When a site's storage is in down time, if this option is set, the site can still run es jobs which will get inputs files from other sites(based on rucio to select the sites). B\
efore we did an es test when bnl dcache was in down time, the test failed because we were not allowed to get remote inputs files.

2) normalize objectstore as normal RSE.
In the new mover structure, objectstore is not a special storage anymore. With this patch, we can update agis to define es output/logs storages. I already tested it to run both es\
 jobs and es merge jobs.

1. add allowRemoteInputs. If it's set, pilot will check local replicas at first. If there is no local replica, pilot will read/download remote replicas.
2. enable to use activity for es outputs.
pilot was designed to use special functions to call objectstores in es jobs. new updates make pilot be able to use activity in agis to use storages more than objectstores and diff\
erent copytools.
3. fix mover part to be able to merge es files from normal storages.
4. fix objectstore sitemover to be able to generate surl for files which are in normal grid storages but not registered in rucio.

to avoid calling stageout function two times
add allowRemoteInputs option
update activity pes to es_events
fix mover to support allowRemoteInputs and normal RSE as es storage
update objectstore sitemover to support normal rse
normalize os as normal rse
add log when using associating storages
fix issues in normalizing objectstore as normal RSE
removing test logs
add log
fix stagout log part
fix objectstoreId when file is registerred

File Changes

M Job.py (11)
M Mover.py (2)
M RunJobEvent.py (75)
M RunJobHpcEvent.py (5)
M movers/mover.py (70)
M movers/objectstore_sitemover.py (53)

  https://github.com/PanDAWMS/pilot/pull/132

Commit Summary

patch to support new file related event position since Athena 21.0.21

File Changes

M Job.py (9)
M RunJobEvent.py (32)

Updates from David Cameron:

  https://github.com/PanDAWMS/pilot/pull/133

Commit Summary

mv mover: stage out to TURL rather than SURL

File Changes

M movers/mv_sitemover.py (2)

Updates from Alexey Anisenkov:

  https://github.com/PanDAWMS/pilot/pull/134

Commit Summary

lsmSiteMover: added gsiftp and root protocols to the list of supported schemes
Merge branch 'main-dev' of https://github.com/PanDAWMS/pilot into main-dev

File Changes

M movers/lsm_sitemover.py (2)
M movers/mover.py (4)

Updates from Taylor Childers:

  https://github.com/PanDAWMS/pilot/pull/135

Commit Summary

I've changed two parts of the cmd line creation for the ATLAS job.

If the environment has RUCIO_ACCOUNT set, it overrides the default of 'pilot'.
The local site can override the Database setup by ATLASLocalRootBase by setting the environment variable 'PILOT_DB_LOCAL_SETUP_CMD'. This avoids having to edit the pilot directly \
for local environment changes.

- adding ability to change local database environment before transform is executed. Also added ability to override RUCIO_ACCOUNT via the environment

File Changes

M ATLASExperiment.py (4)

Updates from Wen Guan:

  https://github.com/PanDAWMS/pilot/pull/136

    copy jobMetrics-yoda.json to top working directory for ARC.
    ignore sqlite200 in ES logs copying(when copying logs from rank_* directory to pilot directory).
    update to download number of events from corecount to corecount*2
    update to support new mover in RunJobHpcEvent.py
    update to use storageId(rse id) instead of objectstoreId, in this case we can use normal grid storages.
    when normal stagein failed, ES will be able to try to stage in files from remote storages with rucio sitemover.
    stop eventservice when there are too many continuous stageout failures.

Commit Summary

    copy jobmetrics file to top directory
    ignore to copy sqlite200 to logs
    download 2 times corecount events
    to support new mover in hpc es
    merge
    update objectstoreid to storageid, disable old createpoolcatalog in runjobevent
    try remote inputs for failed stagein in eventservice
    stop ES when two many stageout failures
    able to stagein remote inputs for es
    to allow stagein from all other sites for es
    small fixes
    small fix

File Changes

    M Job.py (41)
    M Mover.py (11)
    M RunJob.py (3)
    M RunJobEvent.py (134)
    M RunJobHpcEvent.py (78)
    M movers/mover.py (57)
    M movers/rucio_sitemover.py (15)

No comments: