Script structure

A well-structured sbatch script follows a consistent layout that makes it readable, maintainable, and less error-prone. Every script has four parts:

  1. Shebang line โ€” the interpreter directive
  2. SBATCH directives โ€” resource requirements and job configuration
  3. Environment setup โ€” module loading and variable definitions
  4. Job execution โ€” the actual commands to run

Complete template

#!/bin/bash

# Job identification
#SBATCH --job-name=my_job
#SBATCH --output=results_%j.out
#SBATCH --error=results_%j.err

# Resource allocation
#SBATCH --nodes=1
#SBATCH --ntasks-per-node=40
#SBATCH --mem=64GB
#SBATCH --time=02:00:00

# Queue selection
#SBATCH -p short-40core

# --- Environment setup ---
module purge
module load intel/oneAPI/2022.2
module load compiler/latest
module load mpi/latest

export OMP_NUM_THREADS=1
export I_MPI_PIN_DOMAIN=omp

echo "Job ID: $SLURM_JOBID"
echo "Running on nodes: $SLURM_JOB_NODELIST"
echo "Number of tasks: $SLURM_NTASKS"
echo "Start time: $(date)"

# --- Job execution ---
cd $SLURM_SUBMIT_DIR

mpirun ./my_app input.dat > computation.log

echo "Job completed at: $(date)"
Shortcut: Prefer a form? The SLURM job script builder generates a script like this from your answers.

Directive best practices

Group related directives with comments

# Job identification
#SBATCH --job-name=protein_folding
#SBATCH --output=folding_%j.out
#SBATCH --error=folding_%j.err

# Resource requirements
#SBATCH --nodes=4
#SBATCH --ntasks-per-node=40
#SBATCH --mem-per-cpu=2GB
#SBATCH --time=12:00:00

# Job placement
#SBATCH -p long-40core

Use descriptive output names

#SBATCH --output=simulation_%x_%j.out
#SBATCH --error=simulation_%x_%j.err
# %x = job name, %j = job ID

Environment setup guidelines

  • Always purge modules first โ€” module purge ensures a clean environment and prevents conflicts.
  • Set threading variables โ€” e.g. export OMP_NUM_THREADS=$SLURM_CPUS_PER_TASK.
  • Print job information โ€” job ID, node list, working directory, and start time make debugging much easier.

Execution best practices

# Exit on any error
set -e

# Check inputs exist before starting
if [ ! -f "input.dat" ]; then
    echo "ERROR: input.dat not found"
    exit 1
fi

# Time the main computation
start_time=$(date +%s)
mpirun ./my_app input.dat
end_time=$(date +%s)
echo "Computation completed in $((end_time - start_time)) seconds"

Common directive reference

DirectivePurposeExample
-p, --partitionSpecify queue/partition#SBATCH -p short-40core
-N, --nodesNumber of nodes#SBATCH -N 2
-n, --ntasksTotal number of tasks#SBATCH -n 80
--ntasks-per-nodeTasks per node#SBATCH --ntasks-per-node=40
-c, --cpus-per-taskCPUs per task (threading)#SBATCH -c 40
--memMemory per node#SBATCH --mem=64GB
--mem-per-cpuMemory per CPU#SBATCH --mem-per-cpu=2GB
-t, --timeWall time limit#SBATCH -t 02:30:00
-J, --job-nameJob name#SBATCH -J my_job
-o, --outputStandard output file#SBATCH -o job_%j.out
-e, --errorStandard error file#SBATCH -e job_%j.err
--gresGeneric resources (GPUs)#SBATCH --gres=gpu:2
--arrayJob array indices#SBATCH --array=1-100
Pro tip: Keep a template script for each type of job you commonly run โ€” and always test with small resource requests first, using interactive sessions to debug before submitting large batch jobs.
Applies to All clusters