Monday, June 6, 2016

65.2

Summary

Direct i/o updates
- Update to handle direct access in reconstruction jobs, using an input file list handed over to the trf
  * The input file list is created early in the job and later the pilot updates it by translating the local file paths inside the file to TURLs
  * Requested by R. Walker et al
- FAX related direct i/o bug fix, --directIn not set in user jobs for case transferType=fax, direct_access_wan=True
  * Requested by I. Vukotic (HC test still pending)
- Relaxed the input file verification (that the file is actually present) in the mv site mover, to allow/prepare for direct access on Nordugrid
  * Enabling direct i/o on ND requires additional changes on the aCT and possibly more changes in the pilot. To be tested
  * Requested by A. Filipcic and D. Cameron

Only allow alternative OS stage-outs if "allow_alt_os_stageout" is present in catchall field
- Requested by A. Di Girolamo

Debugging info about thrown exception added to failed list_replica() call
- Sometimes list_replica() fails and the pilot reports it as “Rucio returned an empty replica dictionary” – now more debugging info has been added
- Requested by M. Lassnig

Memory monitoring: now enforcing that the payload memory usage is within the allowed limit (on non-CGROUPS sites)
- i.e. that the measured maxPSS stays within 2 times the schedconfig maxRSS value – otherwise the job will be killed  
- Requested by A. Forti et al

AES updates from W. Guan
- Yoda accounting and monitoring
  * Report cpu hour info and processed events to pilot, pilot then reports this info to PanDA with jobMetrics
- AES zip: Yoda AES zip outputs
  * Grid job AES zip outputs (RunJobEvent), unzip ES when merging (RunJob)
  * Controlled by “es_to_zip” in catchall field
- Report MPI job id to PanDA in jobMetrics
- Added bulk update event ranges to RunJobEvent
- Fixed os_bucket_id in Yoda
- Fixed race condition in RunJobEvent
  * When asynchronousOutputStager, updateEventRange and main thread getEventRange are running at the same time, the curl.config will be overwritten and asynchronousOutputStager may crash

Additional changes (requested since the ADC Weekly last week)
-  Added &schemes=https,http to all occurrences of ?select=geoip when the schedconfig.httpredirector is set
  * Requested by R. Walker, V. Garonne
- Updated memory monitor from release area 20.1.5.2 to 20.7.5.8
  * Requested by J. Elmsheuser and Tony Limosani

Version info

General changes:
- Added debugging (sys.exc_info()) info when list_replicas() fails with 'empty dictionary' in getRucioReplicaDictionary() (Mover)
- Removed duplicate definition of objectstore variable in mover_put_data() (Mover)
- Added &schemes=https,http to all occurrences of ?select=geoip when the schedconfig.httpredirector is set, requested by R. Walker, V. Garonne (aria2cSite\
Mover)

Bug fixes:
- Removed call to getDDMStorage() which overwrote alternative bucket id (Mover)
- Adding --directIn for transferType fax and direct_access_wan = True, in getAnalysisRunCommand() (ATLASExperiment)
- Setting old/newPrefix="" in getFileAccessInfo() (SiteInformation)

OS log transfers:
- Ignoring failed OS log transfer and resetting os_bucket_id, in transferActualLogFile() (JobLog)
- Only allow alt OS stage-out if "allow_alt_os_stageout" is present in catchall field, in mover_put_data() (Mover)

Direct Access on ND:
- Relaxed the input file verification in get_data() to allow for direct access, requested by D. Cameron and A. Filipcic (mvSiteMover)

Direct access updates:
- Now accepting writeToFile directive for non-ES jobs in setJobDef() (Job)
- Sending eventservice variable to writeToInputFile() from setJobDef() (Job)
- Added eventservice argument to writeToInputFile() (pUtil)
- Selecting --inputHitsFile or --inputAODFile depending on value of eventservice variable in writeToInputFile() (pUtil)
- Updated updateDispatcherData4ES(), updateJobPars() to work not only with ES jobs (pUtil)
- Created getPooFilenameFromJobPars(), updateInputFileWithTURLs() (pUtil) [not merged with main-dev]
- Generating the LFN_to_TURL_dictionary in getTURLs() (Mover)
- Receiving LFN_to_TURL_dictionary from PFC4TURLs() from _mover_get_data_new(), mover_get_data() (Mover)
- Returning LFN_to_TURL_dictionary from getTURLs(), createPFC4TURLs(), PFC4TURLs() (Mover)
- Importing updateInputFileWithTURLs (Mover)
- Calling updateInputFileWithTURLs() from _mover_get_data_new() and mover_get_data() (Mover)

- Added dummy return value from PFC4TURLs() call in getPoolFileCatalog() (RunJobEvent)
Memory monitoring:
- Now enforcing maxPSS from memory monitor is less than 2*maxRSS from schedconfig, otherwise kill job for non-CGROUPS sites, in __check_memory_usage() (Mo\
nitor)
- Updated memory monitor from release area 20.1.5.2 to 20.7.5.8 (ATLASExperiment)

Updates from Wen Guan:
- Yoda accounting and monitoring: report cpu hour info and processed events to pilot. Pilot then reports this info to panda with jobMetrics
- ES zip: Yoda ES zip outputs; Grid job ES zip outputs(RunJobEvent), unzip ES when merging(RunJob)
- Report MPI job id to panda in jobMetrics
- Added bulk update event ranges to RunJobEvent
- Fixed os_bucket_id in Yoda
- Fixed RunJobEvent with race condition: When asynchronousOutputStager, updateEventRange and main thread getEventRange are running at the same time, the c\
url.config will be overwritten and asynchronousOutputStager may crash

No comments: