Script structure
A well-structured sbatch script follows a consistent layout that makes it readable, maintainable, and less error-prone. Every script has four parts:
- Shebang line โ the interpreter directive
- SBATCH directives โ resource requirements and job configuration
- Environment setup โ module loading and variable definitions
- Job execution โ the actual commands to run
Complete template
#!/bin/bash # Job identification #SBATCH --job-name=my_job #SBATCH --output=results_%j.out #SBATCH --error=results_%j.err # Resource allocation #SBATCH --nodes=1 #SBATCH --ntasks-per-node=40 #SBATCH --mem=64GB #SBATCH --time=02:00:00 # Queue selection #SBATCH -p short-40core # --- Environment setup --- module purge module load intel/oneAPI/2022.2 module load compiler/latest module load mpi/latest export OMP_NUM_THREADS=1 export I_MPI_PIN_DOMAIN=omp echo "Job ID: $SLURM_JOBID" echo "Running on nodes: $SLURM_JOB_NODELIST" echo "Number of tasks: $SLURM_NTASKS" echo "Start time: $(date)" # --- Job execution --- cd $SLURM_SUBMIT_DIR mpirun ./my_app input.dat > computation.log echo "Job completed at: $(date)"
Shortcut: Prefer a form? The SLURM job script builder generates a script like this from your answers.
Directive best practices
Group related directives with comments
# Job identification #SBATCH --job-name=protein_folding #SBATCH --output=folding_%j.out #SBATCH --error=folding_%j.err # Resource requirements #SBATCH --nodes=4 #SBATCH --ntasks-per-node=40 #SBATCH --mem-per-cpu=2GB #SBATCH --time=12:00:00 # Job placement #SBATCH -p long-40core
Use descriptive output names
#SBATCH --output=simulation_%x_%j.out #SBATCH --error=simulation_%x_%j.err # %x = job name, %j = job ID
Environment setup guidelines
- Always purge modules first โ
module purgeensures a clean environment and prevents conflicts. - Set threading variables โ e.g.
export OMP_NUM_THREADS=$SLURM_CPUS_PER_TASK. - Print job information โ job ID, node list, working directory, and start time make debugging much easier.
Execution best practices
# Exit on any error
set -e
# Check inputs exist before starting
if [ ! -f "input.dat" ]; then
echo "ERROR: input.dat not found"
exit 1
fi
# Time the main computation
start_time=$(date +%s)
mpirun ./my_app input.dat
end_time=$(date +%s)
echo "Computation completed in $((end_time - start_time)) seconds"Common directive reference
| Directive | Purpose | Example |
|---|---|---|
-p, --partition | Specify queue/partition | #SBATCH -p short-40core |
-N, --nodes | Number of nodes | #SBATCH -N 2 |
-n, --ntasks | Total number of tasks | #SBATCH -n 80 |
--ntasks-per-node | Tasks per node | #SBATCH --ntasks-per-node=40 |
-c, --cpus-per-task | CPUs per task (threading) | #SBATCH -c 40 |
--mem | Memory per node | #SBATCH --mem=64GB |
--mem-per-cpu | Memory per CPU | #SBATCH --mem-per-cpu=2GB |
-t, --time | Wall time limit | #SBATCH -t 02:30:00 |
-J, --job-name | Job name | #SBATCH -J my_job |
-o, --output | Standard output file | #SBATCH -o job_%j.out |
-e, --error | Standard error file | #SBATCH -e job_%j.err |
--gres | Generic resources (GPUs) | #SBATCH --gres=gpu:2 |
--array | Job array indices | #SBATCH --array=1-100 |
Pro tip: Keep a template script for each type of job you commonly run โ and always test with small resource requests first, using interactive sessions to debug before submitting large batch jobs.
Applies to