Wednesday, July 1, 2015

62.0

Summary


* new version numbering scheme: 62.0 (instead of what would have been 62a
in the old scheme)

* DQ2 API to Rucio API migration continued
- Main change: now using rucio.list_replicas() instead of
dq2.bulkFindReplicas(). Aside from replicas, list_replicas() also returns
the RSE and file type, which are used by the pilot and also means that the
old corresponding DQ2/ToA functions can be avoided.
- Reworked/improved input file handling; Copy tool to be used is now a
property of the input file in question which leads to an improvement with
FAX, as it solves a difficulty with switching to FAX when the replica is
missing in the catalog (previously this could not easily be done in all
catalog lookup failures). Also makes it simple to verify that file is not
on TAPE when e.g. FAX is used (request from Ilija Vukotic)

* Changed confusing error message 'Encountered an empty SURL-GUID
dictionary' to 'Rucio returned an empty replica dictionary'

* Support for additional output files
- In case an output file reaches its max size, a new output file will be
created - not listed in jobs' output file list. Pilot will now discover
these files and adds them to the stage-out queue

* Tracing report updates
- Additional parameters sent with report (url, stateReason). Requested by
Sergey Below
- JSON report stored in log (last one at the moment, can be extended if
necessary)

* glExec patches for problem with permissions on ARC+HTCondor affecting
RAL sites
- Already deployed on relevant glexec queues
- Code from Edward Karavakis

* Protection against triple slashes in FAX redirector

* Object store paths now generated from new AGIS JSON field 'objectstores'
- Corresponding queuedata file copied primarily from CVMFS and if CMVFS is
not available, then it is downloaded from AGIS (HPC:s)

* Object store site mover updates
- Support for new AGIS parameters related to object stores (e.g. is_secure
field)

* Setup changes for HLT jobs (some special handling was out-commented).
Requested by Rod Walker

* Now skipping killing a job when the maximum batch system time limit has
been reached and the NO_PILOT_TIME_LIMIT_KILL keyword is present in
catchall field. Requested by Rod Walker for special cases (useful for
short queues)

* Corrected a few cases where LFN:s were extracted from PFN:s instead of
using the job def LFN list. Requested by Rod Walker after a problem seen
at CERN-PROD with unexpected chars in PFN

* Skipping killing bash process, fixing a problem at e.g. RAL where batch
logs went missing when PID namespacing was used. Pilot essentially
committed suicide at the end of a job, killing its bash process -> no
stdout. Requested by Iain Bradford Steers et al

* Now verifying size of work dir for all jobs, not only for user jobs as
before. Requested by Alessandra Forti et al

* Introducing the concept of error priorities
- The highest priority error will/should be reported, which should
correspond to the most relevant error. E.g. fixing a problem seen by Rod
Walker where a job ran out of disk space, but the error reported was a
secondary problem (job killed)

* Send startTime with jobUpdate to fix a complication on aCT. Requested by
Andrej Filipcic

* Support for memory monitor developed by Nathalie Rauschmayr
- Runs in parallel with the trf and produces a JSON summary file sent to
the PanDA server with job updates (as soon as it is available)
- Memory monitor located in release 20 area; pilot tries to set it up with
same release as job, and defaults to rel 20 in case of failure

* Pilot now uses environment variable EVENT_INDEX_URL
- If not set, pilot falls back to default .es Event Index URL =

Version info

Sending gpfn (replica) and sitemover to isFAXAllowed() from mover_get_data() (Mover)
Added surl and sitemover arguments to isFAXAllowed() (Mover)
Removed fax_mode argument from getCatalogFileList(), and its call in getPoolFileCatalog() since it is not needed (Mover)
Now verifying that a replica is not on TAPE before FAX can be used, in isFAXAllowed() (Mover)
Updated error message "Encountered an empty SURL-GUID dictionary" to "Rucio returned an empty replica dictionary" in verifySURLGUIDDictionary(\
) (Mover)
Protection against triple slashes in updateRedirector() (FAXSiteMover)
Moved getExtension() from pUtil to FileHandling
Importing getExtension() from FileHandling, in moveToExternal() (pUtil)
Importing getExtension() from FileHandling instead of from pUtil (SiteInformation, ATLASSiteInformation, JobState, Mover, RunJob, RunJobEvent,\
 SiteMover)
Removed import of unused getExtension (NordugridSiteInformation)
Moved readJSON and writeJSON from FAXSiteMover to FileHandling
Created discoverAdditionalOutputFiles() (RunJob)
Now using discoverAdditionalOutputFiles() in __main__() (RunJob)
Corrected several readpar -> self.readpar (SiteInformation)
Created getField(), getNewQueuedata(), getObjectstoresList(), getObjectstoresField(), getObjectstorePath(), getObjectstoreName() (SiteInformat\
ion)
Replaced all occurances of str(e) with e (pUtil)
Now using local root setup in setup() (xrootdObjectstoreSiteMover)
Out-commented HLT related code in extractAppdir() (ATLASSiteInformation)
Created getOSTransferDictionaryFilename() (FileHandling)
Created FileHandling::addToOSTransferDictionary() used by RunJobEvent::transferToObjectStore() (FileHandling, RunJobEvent)
Now skipping pilot job kill when maximum batch system time limit has been reached and NO_PILOT_TIME_LIMIT_KILL is present in catchall, in __mo\
nitor_processes() (Monitor)
Added getter and setter functions for taskID (RunJobEvent)
readJSON() is now using openFile() and is more similar to getJSONDictionary() (FileHandling)
Created addToOSTransferDictionary(), getObjectstoresList(), getObjectstoresField(), getObjectstorePath(), getOSNames() (FileHandling)
Created getNewQueuedataFilename(), getNewQueuedata() (FileHandling)    MOVE??????
Created getField(), getObjectstorePath()
Created getOSJobMetrics() (PandaServerClient)
Skipping killing bash process in killOrphans(), fixing a problem at e.g. RAL where batch logs went missing (processes)
Now using getLFN() instead of basename of PFN in getFileInfo(). Requested by Rodney Walker after a problem seen at CERN-PROD with unexpected c\
hars in PFN (Mover)
Now verifying size of work dir for all jobs, not only for user jobs, in __check_remaining_space() (Monitor)
Now using getLFN() instead of basename of PFN in createPoolFileCatalog() (pUtil)

Added lfns argument to createPoolFileCatalog() (pUtil)

Sending lfns to createPFC4TURLs() from PFC4TURLs() (Mover)
Sending lfns to createPoolFileCatalog() from createPFC4TURLs() (Mover)
Sending self.__job.inFiles to createPoolFileCatalog() from prepareHPCJob() (RunJobHpcEvent)
Rewrote getLFN() to better support PFNs whose basename can contain additional characters than the LFN (pUtil)
Added copytool field to fileInfoDic in getFileInfo() (Mover)
Now returning copytool_dictionary from getCatalogFileList(), received in getPoolFileCatalog() (Mover)
Created getCopytoolDictionary() used by getCatalogFileList() (Mover)
Now setting copytool_dictionary on a replica by replica basis in getCatalogFileList() (Mover)
Now returning copytool_dictionary from getPoolFileCatalog(), received in getFileInfo() (Mover)
Created extractCopytoolForPFN() used in getFileInfo() (Mover)
Renamed copytool() to getCopytool() (Mover)
Extracting and returning copytool in extractInputFileInfo(), received in mover_get_data() (Mover)
Updating sitemover object dynamically depending on which copytool to use in mover_get_data() (Mover)
Removed sitemover argument from isFAXAllowed() (Mover)
Removed readpar import from pUtil (SiteInformation)
Added version argument to getQueuedataFileName(), readpar() (SiteInformation)
Sending version to getQueuedataFileName() from readpar() (SiteInformation)
readpar() is now using getField() if new queuedata version is requested (SiteInformation)
Sending lfnList to getPoolFileCatalog() from getPoolFileCatalog() (RunJobEvent)
Using setPilotlogFilename() in cleanup() to prevent problem with writing the pilotlog after workdir has been removed (pUtil)
Added version argument to readpar() (pUtil)
Now using getNewQueuedata() from handleQueuedata() (pUtil)
Sending si object to getFilePathForObjectStore() from getFileInfo() and getDDMStorage() (Mover)
Sending si object to getDDMStorage() from mover_put_data() (Mover)
Added si argument to getDDMStorage() (Mover)
Sending computingElement (queuename) to mover_get_data() from get_data() (Mover)
Added queuename argument to mover_get_data(), getFileInfo() (Mover)
Sending queuename to getFileInfo() from mover_get_data() (Mover)
Added getter and setter for __queuename (SiteInformation)
Using si.setQueueName() in mover_put_data() (Mover)
Sending site.computingElement (queuename) to mover_put_data() from transferActualLogFile(), transferAdditionalFile() (JobLog)
Sending site.computingElement (queuename) to mover_put_data() from moveLostOutputFiles() (pilot)
Sending site.computingElement (queuename) to mover_put_data() from stageOut() (RunJob, RunJobEvent)
Added queuename argument to mover_put_data() (Mover)
Now using si.getObjectstorePath() in getFilePathForObjectStore() (Mover)

Changed log order message in verifySetup() and stageOutFile(). Requested by Rodney Walker (xrdcpSiteMover)
Removed unused argument 'experiment' from transferActualLogFile(), and its calls in transferLogFile() (JobLog)
Added argument 'experiment' to getLogPath() and its call in transferLogFile() (JobLog)
Now using si.getObjectstorePath() instead of mover.getFilePathForObjectStore() in getLogPath() (JobLog)
Created getObjectstoreBucketEndpoint() (SiteInformation)
Created getHash(), getHashedBucketEndpoint() (FileHandling)
Added log_transfer argument to getDDMStorage(), sent from mover_put_data() (Mover)
Created isLogTransfer(), used by mover_put_data() (Mover)
Created shouldCleanupOS() used by __main__() (RunJob)
Created test method cleanupOS() (RunJob)
Removed downloadEventRanges() since it is an imported function from EventRanges module (RunJobEvent)
Returning cerr instead of cout if set in timedCommand() (pUtil)
Will not set status=False when special SE log transfer fails, in transferLogFile() (JobLog)
Created getPilotErrorReportFilename(), updatePilotErrorReport(), getHighestPriorityError() (FileHandling)
Using updatePilotErrorReport() for nine failure scenarios, high priority errors (Monitor)
Now using getHighestPriorityError() in getNodeStructure() to look for reported high priority errors (PandaServerClient)
Removed vmPeakMax, vmPeakMean, RSSMean from jobMetrics since they are no longer needed (PandaServerClient)
Updated __init__() and setup() to support new objectstores fields os_access_key, os_secret_key, os_is_secure (S3ObjectstoreSiteMover)
Created getUtilityCommand() (Experiment, ATLASExperiment, AMSTaiwanExperiment, NordugridATLASExperiment)
Created getUtilityInfo() used by getNodeStructure() (PandaServerClient)
Created getUtilitySubprocess() used by executePayload() (RunJob)
Created getUtilityJSONFilename() used by getUtilityCommand() (NordugridATLASExperiment, ATLASExperiment, Experiment)
Created shouldExecuteUtility() (NordugridATLASExperiment)
Added support for memory monitor in __main__() (RunJobEvent)
Added format argument to timeStampUTC() (pUtil)
Created writeTimeStampToFile() (pUtil)
Using writeTimeStampToFile() in getNewJob() (pilot)
Copying the time stamp file from the init dir to the work dir in __createJobWorkdir() (Monitor)
Reading back start time from file and adding it to node structure in getNodeStructure() (PandaServerClient)
Created getEventIndexURL() (EventService)
Sending url to getTokenExtractorProcess() from __main__() (RunJobEvent)
Added handling of url in getTokenExtractorProcess() (RunJobEvent)


No comments: