Monitoring jobs

CommandPurposeExample
squeue --user=$USERList your running and pending jobssqueue --user=sam123
scontrol show job <jobid>Detailed job informationscontrol show job 123456
sacct -j <jobid> -lDetailed stats for a completed or running jobsacct -j 123456 -l
seff <jobid>Efficiency (CPU, memory) of completed jobsseff 123456

Modifying jobs

Place a job on hold to prevent it from running, and release it later:

scontrol hold 123456
scontrol release 123456

Canceling jobs

scancel 123456                          # a specific job
scancel --user=$USER                    # all your jobs
scancel --name=my_job                   # all jobs with this name
scancel --user=$USER --state=PENDING    # all your queued jobs

Checking efficiency & resource usage

Understanding CPU load

CPU load is the average number of processes trying to use the CPU. On a fully utilized 40-core node, expect a load around 40. Much lower means underutilization; much higher means oversubscription, which degrades performance. The statistic shows 1-, 5-, and 15-minute averages.

Efficiency and accounting

seff 123456
sacct -j 123456 --format=JobID,JobName,State,Elapsed,MaxRSS,CPUTime

SeaWulf also provides a built-in script that summarizes CPU and memory usage per job:

/gpfs/software/hpc_tools/get_resource_usage.py

Real-time monitoring on the node

  1. Find your node.

    squeue --user=$USER

    Note the node name in the NODELIST column (e.g. dn045).

  2. SSH to the node.

    ssh dn045
  3. Monitor resources with one of:

    module load glances && glances   # recommended: richest interface
    module load htop && htop         # classic detailed view
    top                              # basic, no module needed

Optimizing resource usage

  • Monitor CPU load values โ€” they should match the core count for full utilization
  • Check memory usage โ€” avoid over-requesting resources
  • Consider shared queues โ€” for jobs that don't need a full node
  • Adjust job scripts โ€” based on actual usage from seff and monitoring tools

Screenshots

Example output of get_resource_usage.py showing CPU and memory utilization for a compute node (inefficient use at 26.5% CPU with 10.58 load)
Example output of get_resource_usage.py showing CPU and memory utilization for a compute node (inefficient use at 26.5% CPU with 10.58 load)
glances monitoring interface showing real-time system resource metrics on a compute node
glances monitoring interface showing real-time system resource metrics on a compute node
htop interface displaying CPU, memory, and process information for resource monitoring
htop interface displaying CPU, memory, and process information for resource monitoring
top command output showing real-time CPU, memory, and process data overview
top command output showing real-time CPU, memory, and process data overview
Applies to All clusters