Core HPC concepts

TermDefinition
ClusterA collection of interconnected computers that work together as a single system
High-Performance Computing (HPC)Computing that aggregates processing power to solve problems requiring significant computational resources
NodeAn individual computer within a cluster, containing processors, memory, and storage
Login nodesNodes providing user access โ€” for job submission and file management, not intensive computation
Compute nodesNodes where jobs execute; accessed through SLURM, not directly
CoreAn individual processing unit within a CPU; SeaWulf nodes range from 28 to 96 cores
Parallel processingDistributing tasks across multiple computers to work simultaneously on portions of a problem
petaFLOPSOne quadrillion floating-point operations per second; SeaWulf peaks at 1.86 petaFLOPS

Hardware

TermDefinition
CPUThe processor that executes instructions and performs calculations
GPUSpecialized hardware for parallel processing, used in machine learning and scientific simulations
High-Bandwidth Memory (HBM)Advanced memory with faster data transfer than traditional RAM; available on select SeaWulf nodes
InfiniBandHigh-performance networking technology interconnecting cluster nodes
InterconnectThe network infrastructure connecting nodes for communication and data transfer

Storage

TermDefinition
GPFSIBM's shared-disk file system used for SeaWulf storage; concurrent file access across all nodes
Parallel file systemDistributed storage allowing multiple nodes to access the same files simultaneously
Scratch spaceHigh-performance temporary storage (20 TB); files older than 45 days are removed automatically

Scheduling

TermDefinition
SLURMThe cluster management and job scheduling system
JobA computational task submitted to the cluster, queued until resources are available
Queue (partition)A logical grouping of nodes with similar characteristics, time limits, and priorities
Job scriptA file of resource requirements and commands, submitted with sbatch
SchedulerSoftware allocating resources based on priority and availability
Resource allocationGranted access to compute time and storage quotas

Performance

TermDefinition
ThroughputThe rate at which computational work is processed
ScalabilityMaintaining or improving performance as resources are added
Load balancingDistributing work across nodes to optimize utilization
BenchmarkingMeasuring system performance with standardized metrics
Fault toleranceContinuing operation despite component failures
Applies to All clusters