Memory Monitoring update
- Now using release 21.0.18 (I/O values are stored in absolute values and I/O rates are reported as integer)
- Requested by Johannes Elmsheuser
Benchmarking
- 1/100 jobs run cern-benchmark in offline mode
- Chained benchmarks: Whetstone and fastBmk
- Using freetext = “Whetstone+fastBmk” for ES use
- Executed during stage-in to minimize wasting time
- Waiting for benchmark to finish before executing payload
- Output dictionary is added to jobReport.json
- Note: After the ADC Weekly meeting we added IP and hostname to the output dictionary as well. The hostname is however available elsewhere in the jobReport (machine section). We might want to avoid duplication of info.
Fix for missing traces when using remote i/o
- Caused data to be deleted even though it had recently been accessed
Support for (mainly) new HC options
- --useTestASetup
Sets special env variable (ALRB_asetupVersion) that will activate test version of asetup
- --useTestXRootD
Use xrootdsetup-dev.sh instead of xrootdsetup.sh for the LRS
- Improved handling of HC option --overwriteQueuedata
Supports new and more powerful ways to overwrite schedconfig parameters, incl. sending dictionaries (--overwriteQueuedata might become --overwriteAGISdata)
- Requested by Asoka de Silva (JIRA ticket not yet updated)
Event service updates
- Periodical tarring of event outputs
- Disabled long monitor sleep on ND because of failed heartbeat
- Using normal SE for AES: not yet tested, needs AGIS updates (e.g. currently not possible to add normal SE to ddmendpoints)
- Optimized S3 objectstore site mover to have less HEAD operations
- Updated Rucio site mover to support OS upload
- Fixed traces for pre-merged files (file size and event type)
- Protection against corruption while downloading/updating event ranges
Some DQ2 API cleanup (especially around deprecated code)
- Old site movers still rely on DQ2/ToA functionality; but can be avoided once all sites have migrated to use the new site movers; all DQ2 ref. will/can be removed at that point
Ddmendpoint fix for storm site mover
Avoiding all process kills on BOINC
No call to dispatcher for ND pilots
Note: Earlier today Johannes Elmsheuser alerted me to a problem with using nightlies releases. The fix was trivial so it made it for the release as well. Previously only VAL releases would have worked (i.e. if release was e.g. '21.0.X-VAL' but now also '21.0.X' should work).
Contributors: A. Anisenkov, W. Guan, M. Lassnig, D. Cameron, D. Drizhuk, P. Nilsson
Version info
General changes:
- Updated dump to pretty print written json dictionary, in writeJSON() (FileHandling)
Removal of remaining dq2 API usage:
- Deprecated and removed large parts of getTURLs() related to lcg-getturls, which also called getRSEType() and getRSE(), which in turn used the dq2 API (Mover)
- Removed isDPMSite(), not used any longer (Mover)
- Removed getRSEType(), not used any longer (Mover)
- Removed getRucioPath(), not used any longer (Mover, SiteMover)
- Removed getRucioFileList(), not used any longer (Mover)
- Cleaned up getPoolFileCatalog() a bit (Mover)
NOTE: getRSE() need to be rewritten, or logic changed to set the RSE, since it is used by all old site mover
Nightlies:
- Added .X to getCacheInfo() to fix a case where nightlies were not setup correctly. Requested by Johannes Elmsheuser (ATLASExperiment)
Benchmarking:
- Only executing benchmark tool once out of a hundred starts in shouldExecuteBenchmark() (ATLASSiteInformation)
- Created getJobReportFileName(), addToJobReport() (FileHandling)
- Created getBenchmarkFileName() (SiteInformation, ATLASSiteInformation)
- Created getBenchmarkDictionary() (JobLog)
- Created getBenchmarkSubprocess() (RunJob)
- Added benchmark process, running during stage-in (RunJob, RunJobEvent)
- Deprecated and removed unnecessary executeBenchmark() (SiteInformationm ATLASSiteInformation, Node)
Memory monitoring:
- Updated the memory monitor version to 21.0.18 in getUtilityCommand() (ATLASExperiment)
Copytools:
- Corrected logical bug when updating file state (direct_access -> remote_io) in stagein_real() (mover)
- Corrected loging bug when checking file states after transfer (direct_access -> remote_io) in get_data_new() (Mover)
Event Service:
- PREFETCHER IS ENABLED for late releases
- Added __yamplChannelNamePrefetcher, renamed __yamplChannelName to __yamplChannelNamePayload (RunJobEvent)
- Setting __yamplChannelNamePrefetcher in __init__() (RunJobEvent)
- Added prefetcher field to Job class (Job)
- Updated schemes for prefetcher and adding turl to fileState file in stagein_real() (mover)
- Created isPrefetcherReady(), setPrefetcherIsReady() (RunJobEvent)
Updates from Wen Guan:
https://github.com/PanDAWMS/pilot/pull/112
tar event outputs periodically
2)disable long monitor sleep on ND cloud because of failed 'send' (failed heartbeat)
fix to use different yampl channel name in yoda
payload may includes the same input files for more than one time, fix to stagein it only one time.
5)change updateEventRanges to support new version which supports to tar events periodically
You can view, comment on, or merge this pull request online at:
Commit Summary
fix files starting with zip and duplicated files
disable long monitor sleeping on ND cloud
using different yampl channel name for different athenamp
not show duplicated files when preparing inputs
fix to get correct jobid when naming curl config
fix Yoda to use different yampl channel name
make updateEventRanges to support different versions
RunJobEvent to support periodically tar and upload
fix python problem to pop events from list
to configure time gap between tar/zip functions
File Changes
M EventRanges.py (4)
M HPC/EventServer/EventServerJobManager.py (10)
M Job.py (6)
M Monitor.py (4)
M RunJobEvent.py (315)
M RunJobHpcEvent.py (15)
M RunJobUtilities.py (2)
M pUtil.py (5)
https://github.com/PanDAWMS/pilot/pull/114
1) normalized objectstore as a normal rse
2) optimized s3objectstore sitemover to have less HEAD operations: Dan reported that we had 4~5 times of 'HEAD' operations than 'GET' and 'PUT' operations. It caused load problems\
on objectstore.
3) updated rucio sitemover to support os upload
4) set objectstore keypair in environment, rucio site mover will use it.
Commit Summary
fix to remove tar/zip es files in the log
optimize s3objectstore sitemover to have less HEAD operation
update rucio sitemover to support os upload
normalize Objectstore as a normal RSE
set objectstore keypair in environment which rucio mover will use
normalize os as a normal rse
File Changes
M ATLASExperiment.py (4)
M Mover.py (13)
M RunJobEvent.py (64)
M S3ObjectstoreSiteMover.py (82)
M movers/mover.py (87)
M movers/rucio_sitemover.py (14)
https://github.com/PanDAWMS/pilot/pull/116
fix tracer report:
when es merge job stagein premerge files, the filesize is not filled.
when es merge job stagein premerge files, the eventtype is 'get_sm', it should be 'get_es'
Commit Summary
fix filesize in S3ObectStoreSiteMover
fix trace report
File Changes
M Mover.py (2)
M S3ObjectstoreSiteMover.py (17)
M movers/mover.py (3)
M movers/rucio_sitemover.py (2)
https://github.com/PanDAWMS/pilot/pull/117
Commit Summary
to protect corruption in downloading/updating eventranges
File Changes
M EventRanges.py (135)
Updates from Daniel Drizhuk:
https://github.com/PanDAWMS/pilot/pull/113
Introduced the new way of presenting HammerCloud parameter --overwriteQueuedata, added two more parameters: --useTestASetup and --useTestXRootD.
The proposed way for the new --overwriteQueuedata is to use common shell syntax.
In the new syntax --overwriteQueuedata is a multiargument parameter, that receives after it a set of parameters represented in key=value form. The set is ended when next parameter\
starts with - or with parameter --, that will be stripped.
The value in the parameter may be either a string or a valid JSON.
The escape sequences are posix shell compatible, so JSON should be probably wrapped into single quotes.
The argument string is parsed by shlex.
Examples of --overwriteQueuedata (assuming TRF is echo, lines do not include TRF):
Stripped end-of-parameters
The line this is --overwriteQueuedata key1 key2=null key3='{"a":1,"b":2}' -- test
will result in queuedata modification key1=True, key2=None, key3={a:1,b:2}
and command echo this is test
End-of-parameters is a dash-prefix of the next parameter
The line this is --overwriteQueuedata key1 key2=null key3='{"a":1,"b":2}' -test
will result in queuedata modification key1=True, key2=None, key3={a:1,b:2}
and command echo this is -test
Second occurrence and EOL as an end-of-parameters
this is --overwriteQueuedata key1 key2=null key3='{"a":1,"b":2}' -test --overwriteQueuedata key4
will result in queuedata modification key1=True, key2=None, key3={a:1,b:2}, key4=True
and command echo this is -test
Commit Summary
Merge pull request #2 from PanDAWMS/main-dev
Merge pull request #3 from PanDAWMS/main-dev
Testing parameters for HammerCloud
Merge remote-tracking branch 'origin/main-dev' into main-dev
File Changes
M ATLASExperiment.py (3)
M ATLASSiteInformation.py (5)
M SiteInformation.py (113)
https://github.com/PanDAWMS/pilot/pull/119
Commit Summary
Fixed issue with queuedata parameters logging when fixing.
File Changes
M SiteInformation.py (2)
Updates from Mario Lassnig:
https://github.com/PanDAWMS/pilot/pull/115
Commit Summary
fix ddmendpoint handling for storm sitemover
File Changes
M movers/storm_sitemover.py (17)
Updates from David Cameron:
https://github.com/PanDAWMS/pilot/pull/118
Commit Summary
do not kill anything on BOINC
File Changes
M processes.py (6)
https://github.com/PanDAWMS/pilot/pull/120
Commit Summary
do not call dispatcher for ND pilots
File Changes
M pilot.py (14)
Updates from Alexey Anisenkov:
https://github.com/PanDAWMS/pilot/pull/121
[Two conflicts resolved by Paul - RunJob type was already fixed but not committed; mover had code for prefetcher that should be outcommented in this version]
Commit Summary
fixed typo in RunJob
added traces for direct access jobs (plus prefetcher mode)
bugfix: protected ip/host resolve functions in trace report initialization
File Changes
M RunJob.py (4)
M movers/mover.py (25)
M movers/trace_report.py (11)
No comments:
Post a Comment