βEmbarrassingly parallelβ workflows are ones where each task runs independently β no communication between processes. This guide shows four ways to parallelize them on SeaWulf. For communication-dependent workloads, see the companion guide to OpenMP and MPI.
The example task
A small Python script (number_square.py) that squares an integer β requires module load anaconda/3:
#!/usr/bin/env python
import numpy as np
import argparse
parser = argparse.ArgumentParser(description="Take in a number and print out its square")
parser.add_argument('integer', metavar='N', type=int,
help='an integer for the calculation')
args = parser.parse_args()
result = int(np.square(args.integer))
print(f'The square of {args.integer} is {result}')Running python number_square.py 8 prints The square of 8 is 64. The goal: compute this for many inputs at once.
Approach 1 β GNU Parallel
A command-line tool built exactly for this. Job script running 100 tasks, 40 at a time on one 40-core node:
#!/usr/bin/env bash
#SBATCH --ntasks=1
#SBATCH --nodes=1
#SBATCH --cpus-per-task=40
#SBATCH --time=05:00
#SBATCH --partition=short-40core
#SBATCH --job-name=parallel_job
#SBATCH --output=squared_numbers.txt
module load anaconda/3
module load gnu-parallel/6.0
parallel --jobs 40 python number_square.py {} ::: {1..100}--jobs 40β run 40 processes simultaneously{}β placeholder for each input value:::β separates the command from its inputs;{1..100}is shell range syntax- Add
--keep-orderif results must come out in input order
Approach 2 β Python multiprocessing
Parallelize inside Python itself with a worker pool (number_square_mp.py):
#!/usr/bin/env python
import numpy as np
import multiprocessing as mp
def square_me(num):
result = int(np.square(num))
return(result)
p = mp.Pool(processes=40)
for num in range(1,101):
results = p.apply_async(square_me, [num])
print(f'The square of {num} is {results.get()}')
p.close()#!/usr/bin/env bash #SBATCH --ntasks=1 #SBATCH --nodes=1 #SBATCH --cpus-per-task=40 #SBATCH --time=05:00 #SBATCH --partition=short-40core #SBATCH --job-name=parallel_job #SBATCH --output=squared_numbers.txt module load anaconda/3 python number_square_mp.py
A pool of 40 workers processes 40 calculations at a time, and results keep their input order naturally.
Approach 3 β Slurm job arrays
Spread independent tasks across many jobs (and nodes). Each array task reads one line from an input file (100_numbers.txt, one integer per line):
#!/usr/bin/env bash
#SBATCH --job-name=array_test
#SBATCH --output=array_test.%A_%a.log
#SBATCH --ntasks=1
#SBATCH -N 1
#SBATCH -p short-40core
#SBATCH -t 04:00:00
#SBATCH --array=1-100
module load anaconda/3
echo "Starting task $SLURM_ARRAY_TASK_ID"
INPUT=$(sed -n "${SLURM_ARRAY_TASK_ID}p" 100_numbers.txt)
python number_square.py $INPUTEach task writes its own log (e.g. array_test.816678_74.log). Note there is no within-node parallelization here β this shape suits tasks that individually use a whole node.
Approach 4 β hybrid: job arrays + GNU Parallel
Combine both: 40 array tasks, each running 100 calculations with 40 parallel workers β 4,000 calculations in total:
#!/usr/bin/env bash
#SBATCH --job-name=array_test
#SBATCH --output=array_test.%A_%a.log
#SBATCH --ntasks=1
#SBATCH --cpus-per-task=40
#SBATCH -N 1
#SBATCH -p short-40core
#SBATCH -t 04:00:00
#SBATCH --array=1-40
mkdir temp
module load anaconda/3
module load gnu-parallel/6.0
START=$SLURM_ARRAY_TASK_ID
NUMLINES=100
STOP=$((SLURM_ARRAY_TASK_ID*NUMLINES))
START="$(($STOP - $(($NUMLINES - 1))))"
echo "START=$START"
echo "STOP=$STOP"
for (( N = $START; N <= $STOP; N++ ))
do
LINE=$(sed -n "$N"p 4k_numbers.txt)
echo $LINE >> temp/tasks_${START}_${STOP}
done
cat temp/tasks_${START}_${STOP} | parallel --jobs 40 --verbose python number_square.py {}
rm temp/tasks_${START}_${STOP}Which approach when?
- GNU Parallel β simple standalone tasks; the fastest to set up
- Python multiprocessing β Python-native workflows needing fine-grained control
- Job arrays β distribute independent tasks across multiple nodes
- Hybrid β maximum scaling across nodes and cores at once