Changes between Version 2 and Version 3 of Workshops/JobParallelism/AfterYourJobHasCompleted


Ignore:
Timestamp:
09/02/2026 02:36:44 AM (43 hours ago)
Author:
Carl Baribault
Comment:

Upgraded to the local seff command

Legend:

Unmodified
Added
Removed
Modified
  • Workshops/JobParallelism/AfterYourJobHasCompleted

    v2 v3  
    4242=== On Cypress ===
    4343
    44  In the following we'll need to use the '''sacct''' command for analyzing completed jobs on Cypress. (Cypress uses an older version of SLURM (v14.03.0) with insufficient support for the seff command.)
     44==== Local version of seff for SLURM 14 ====
    4545
    46  Here are the relevant outputs that we'll be using from '''sacct'''.
     46We've provided a local version of the '''seff''' command in order to compensate for the earlier version of SLURM (v14.03.0) on Cypress where we've encountered insufficient support for the seff command as implemented on LONI clusters.
    4747
    48 ||='''sacct''' output column=||='''Description'''=||='''Format'''=||='''Notes'''=||
    49 ||'''TotalCPU'''||Total core hours used||[DD-[hh:]]mm:ss)||Needs conversion to seconds||
    50 ||'''CPUTimeRAW'''||Total cores hours allocated||Seconds||No conversion needed||
    51 ||'''REQMEM'''||Requested memory||GB or MB ||Defaults to 3200MB per core||
    52 ||'''MaxRSS'''||Maximum memory used||GB per node||Sampled every 30 seconds on Cypress||
     48More specifically, this version of SLURM makes certain job performance values available only after the given job as completed.
    5349
    54 == Cumulative core efficency: (total core hours used) / (total core hours allocated) ==
     50==== The sacct command ====
    5551
    56 === Ideal case ===
     52The local version of seff available on Cypress uses the SLURM command '''sacct''' with more information available from the man page. (See '''man sacct'''.)
    5753
    58  Ideally we have '''TotalCPU''' = '''CPUTimeRAW''' such as the following.
     54==== Result ====
    5955
    60  * TotalCPU=20 hours, CPUTimeRAW=20 hours - using all 20 requested cores, full time for 1 hour
    61  * Core efficiency = (20 hours TotalCPU / 20 hours CPUTimeRAW) = 1
    62 
    63 === Actual case ===
    64 
    65 ==== Using sacct ====
    66 
    67  Here is the sacct command used to for a completed job where we've masked the job ID XXXXXXX
     56Here is the result for seff on the previously running job which has now completed.
    6857
    6958{{{
    70 [tulaneID@cypress1 ~]$sacct  -P -n --format JobID,AllocCPUS,TotalCPU,CPUTimeRaw,REQMEM,MaxRSS -j XXXXXX
    71 XXXXXXX|10|11-04:18:08|1213660|128Gn|
    72 XXXXXXX.batch|1|11-04:18:08|121366|128Gn|3860640K
     59[tulaneID@cypress1 ~]$seff 3336943
     60Job ID: 3336943
     61Cluster: cypress
     62User/Group: cbaribault/hpcstaff
     63State: COMPLETED (exit code 0:0)
     64Cores: 16
     65CPU Utilized: 0:00:20
     66CPU Efficiency: 1.8% of 00:01:11 core-walltime
     67Memory Utilized: 0.08 GB
     68Memory Efficiency: 0.2% of requested per-CPU memory x 16
     69------------------------------------------------------------
    7370}}}
    7471
    75  In the following we'll use the values TotalCPU=11-04:18:08 and CPUTimeRAW=1213660 from the 2nd line, the XXXXXXX.batch step, in the above.
     72==== Summary - in contrast with the result for the running job ====
    7673
    77 ==== Converting TotalCPU to seconds ====
     74Admittedly, the above result for the completed job is stark contrast with the favorable result we obtained using our '''seffRunningPrototype''' while the job was still running.
    7875
    79  We'll use the following shell function to convert '''TotalCPU''' in format [DD-[hh:]]mm:ss) to seconds.
     76Even though we've made an attempt at optimization in the R script '''bootstrapFutureApply.R''' (using R package '''future.apply''') compared to the R script '''bootstrap.R''' (using '''doParallel'''), the result above shows that we still have room for improvement of the overall efficiency of the job.
    8077
     78==== Fewer requested resources = faster job starts
    8179
    82 {{{
    83 [tulaneID@cypress1 ~]$convert_totalcpu_to_seconds() {
    84    seconds=$(echo "$1" | awk -F'[:-]' '{
    85       if (NF == 4) {
    86           # Format: D-HH:MM:SS
    87           total = ($1 * 86400) + ($2 * 3600) + ($3 * 60) + $4
    88       } else if (NF == 3) {
    89           # Format: HH:MM:SS or MM:SS (assumes HH:MM:SS)
    90           total = ($1 * 3600) + ($2 * 60) + $3
    91       } else if (NF == 2) {
    92           # Format: MM:SS
    93           total = ($1 * 60) + $2
    94       } else {
    95           total = $1 # Assume only seconds if no separators found
    96       }
    97       print total
    98    }')
     80In general, jobs that request fewer resources (memory, number of cores, number of nodes, or run time) will start more quickly others.
    9981
    100    echo "$seconds"
    101 }
    102 [tulaneID@cypress1 ~]$convert_totalcpu_to_seconds 11-04:18:08
    103 965888
    104 }}}
    105 
    106 === Compute cumulative core efficiency ===
    107 
    108  Now that we have the job's '''TotalCPU''' in seconds, we can calculate the job's cumulative core efficiency.
    109 
    110 {{{
    111 [tulaneID@cypress1 ~]$bc <<< "scale=2; 965888 / 1213660"
    112 .79
    113 }}}
    114 
    115 === Summary for this job ===
    116 
    117 ==== Fewer requested resources = faster job queueing
    118 
    119  In general, whenever a job can nonetheless run to completion in a comparable elapsed time but with less memory and/or fewer processors (cores and/or nodes) requested, then easier the resource manager SLURM will find an earlier time slot - if not immediately so - to queue (start and run) the job.
    120 
    121 ==== Suggestions for requested processor count and RAM
    122 
    123  * With the above result of 0.79, we conclude that not all 10 requested cores were in use throughout the duration of the job.
    124    * We may be able to request fewer cores depending on the requirements of the parallel segments of the computation.
    125    * We should consult the software provider's information.
    126  * Also, the job used ~3.9GB ('''MaxRSS''') out of the requested 128GB ('''REQMEM''') of RAM
    127   * We could easily expect to have the job run in the same amount of time requesting 10 cores and greatly reduced memory, say, '''!--mem=32000''' or 32GB.
     82In other words, the SLURM resource manager will more frequently find an earlier - if not immediate - time slot to start the job as compared to other jobs requesting more resources.