Every SeaWulf job runs in a queue (also called a partition). The queue decides which kind of node you get, how long the job may run, and how many nodes it can use. Pick the smallest, shortest queue that fits your job and it will usually start sooner.

On this page: How long can a job run ยท Quick picks ยท How to choose ยท Limits per person ยท 40-core Skylake ยท 96-core Milan ยท HBM ยท A100 GPU ยท 28-core Haswell (legacy) ยท Legacy GPU

Where you log in matters. The 40-core, 96-core, HBM, and A100 queues are reached from milan.seawulf.stonybrook.edu (or xeonmax). The 28-core and legacy GPU queues are reached from login.seawulf.stonybrook.edu. See Which SeaWulf login node?

How long can a job run? (maximum walltime)

The longest a single job can run depends on the cluster and the queue.

ClusterLongest runQueues
SeaWulf (CPU)7 daysextended-40core, extended-96core, hbm-extended-96core and extended-28core, up to 2 nodes. Their -shared versions allow 3.5 days, and the long queues 48 hours.
SeaWulf (A100 GPUs)48 hoursa100-long, 1 node. a100 and a100-large allow 8 hours.
NVwulf (H200 and RTX PRO 6000 GPUs)48 hoursThe -long queues: h200x4-long, h200x8-long and b40x4-long. Standard queues allow 8 hours and debug- queues 1 hour. See NVwulf queues.
ClinWulf10 daysThe "extra" queues. The long queues allow 8 hours. See Open OnDemand apps on ClinWulf.

Ask for the time your job needs plus a margin; shorter requests usually start sooner. If your work needs longer than the limit, save checkpoints so a new job can pick up where the last one stopped, or ask HPC support about options. To check a queue's limit on the cluster, run module load slurm and then sinfo -p <queue> -o "%P %l".

Quick picks

Your jobTry this queue
A quick test (under 1 hour)debug-40core
A few cores, doesn't need a whole nodeshort-40core-shared or short-96core-shared
Whole nodes, up to 4 hoursshort-40core or short-96core
Up to 12 hoursmedium-40core or medium-96core
Up to 2 dayslong-40core or long-96core
Up to 7 days, 1 to 2 nodesextended-40core or extended-96core
Many nodes at once (up to 8 hours)large-40core or large-96core
Limited by memory speedan HBM queue
Needs a GPUa100 (or NVwulf for large AI work, see Getting started on NVwulf)

How to choose

  • Start small. Run a short test to learn how much time and memory your job really needs.
  • Ask for what you'll use. Smaller requests usually start sooner.
  • Use shared queues for jobs that don't need a whole node. On shared queues, always set --mem; on the others, you get all of each node's memory automatically.
  • Check the architecture. Make sure your software was built for the node type (Skylake, Milan, Sapphire Rapids, or GPU).
  • See how busy a queue is with sinfo -p <queue>, after module load slurm.

In the tables, Time is the range of run times each queue is meant for; the upper number is the maximum you can request.

Limits per person. On SeaWulf you can use up to 32 nodes at once across all your running jobs, unless you're using one of the large queues, and you can have up to 100 jobs in the queue at one time.

40-core Intel Skylake

AVX-512 support, 192 GB memory per node. Reached from the milan login nodes.

QueueTimeMax nodesSharedBest for
debug-40core1 hr8NoTesting and debugging
short-40core1 to 4 hr8NoStandard jobs
short-40core-shared1 to 4 hr4YesSmaller jobs
medium-40core4 to 12 hr16NoMedium jobs
long-40core8 to 48 hr6NoLong jobs
long-40core-shared8 to 24 hr3YesShared long jobs
extended-40core8 hr to 7 days2NoVery long jobs
extended-40core-shared8 hr to 3.5 days1YesShared extended jobs
large-40core4 to 8 hr50NoLarge parallel jobs

96-core AMD EPYC Milan

AVX2 support, 256 GB memory per node. Reached from the milan login nodes.

QueueTimeMax nodesSharedBest for
short-96core1 to 4 hr8NoParallel jobs
short-96core-shared1 to 4 hr4YesModerate shared jobs
medium-96core4 to 12 hr16NoParameter sweeps
long-96core8 to 48 hr6NoLong parallel jobs
long-96core-shared8 to 24 hr3YesShared long jobs
extended-96core8 hr to 7 days2NoVery long jobs
extended-96core-shared8 hr to 3.5 days1YesShared extended jobs
large-96core4 to 8 hr38NoLarge parallel jobs

High-bandwidth memory (HBM) nodes

Intel Sapphire Rapids with AMX and AVX-512, 384 GB per node (256 GB DDR5 plus 128 GB HBM). Best for work limited by how fast data moves in and out of memory. Reached from the milan or xeonmax login nodes.

QueueTimeMax nodesNotesBest for
hbm-short-96core1 to 4 hr8High-bandwidth memoryMemory-intensive jobs
hbm-medium-96core4 to 12 hr16Faster memoryLarge datasets
hbm-long-96core8 to 48 hr62 to 4 times faster memoryMemory-bound simulations
hbm-extended-96core8 hr to 7 days2Longest HBM runsLong memory-bound jobs
hbm-large-96core4 to 8 hr38Many HBM nodes at onceVery large memory-bound jobs
hbm-1tb-long-96core8 to 48 hr11 TB memory plus 128 GB HBM cacheExtremely large datasets

NVIDIA A100 GPU nodes

Four A100 80 GB GPUs per node, on 64-core Intel Ice Lake nodes with 256 GB memory. Request GPUs with --gres=gpu:N, and check that your software supports CUDA first. These queues are shared.

QueueTimeMax nodesBest for
a1001 to 8 hr2GPU workloads, AI and ML
a100-long8 to 48 hr1Long GPU jobs
a100-large1 to 8 hr4Multi-node GPU jobs

28-core Intel Haswell (legacy)

AVX2 support, 128 GB memory per node. Reached from the login1 and login2 nodes (login.seawulf.stonybrook.edu).

QueueTimeMax nodesSharedBest for
debug-28core1 hr8NoTesting and debugging
short-28core1 to 4 hr12NoSmall and medium jobs
medium-28core4 to 12 hr24NoMedium parallel jobs
long-28core8 to 48 hr8NoLong simulations
extended-28core8 hr to 7 days2NoVery long simulations
large-28core4 to 8 hr80NoLarge parallel jobs

Legacy GPU queues

Older GPUs, reached from login.seawulf.stonybrook.edu. For new GPU work, the A100 queues or NVwulf are usually a better fit.

QueueGPUGPU memoryTimeBest for
gpu4ร— K8024 GB1 to 8 hrBasic GPU work
gpu-long4ร— K8024 GB8 to 48 hrLong GPU jobs
p1002ร— P10016 GB1 to 24 hrScientific computing
v1002ร— V10032 GB1 to 24 hrML and AI

Related

Applies to SeaWulf