Core HPC concepts
| Term | Definition |
|---|---|
| Cluster | A collection of interconnected computers that work together as a single system |
| High-Performance Computing (HPC) | Computing that aggregates processing power to solve problems requiring significant computational resources |
| Node | An individual computer within a cluster, containing processors, memory, and storage |
| Login nodes | Nodes providing user access โ for job submission and file management, not intensive computation |
| Compute nodes | Nodes where jobs execute; accessed through SLURM, not directly |
| Core | An individual processing unit within a CPU; SeaWulf nodes range from 28 to 96 cores |
| Parallel processing | Distributing tasks across multiple computers to work simultaneously on portions of a problem |
| petaFLOPS | One quadrillion floating-point operations per second; SeaWulf peaks at 1.86 petaFLOPS |
Hardware
| Term | Definition |
|---|---|
| CPU | The processor that executes instructions and performs calculations |
| GPU | Specialized hardware for parallel processing, used in machine learning and scientific simulations |
| High-Bandwidth Memory (HBM) | Advanced memory with faster data transfer than traditional RAM; available on select SeaWulf nodes |
| InfiniBand | High-performance networking technology interconnecting cluster nodes |
| Interconnect | The network infrastructure connecting nodes for communication and data transfer |
Storage
| Term | Definition |
|---|---|
| GPFS | IBM's shared-disk file system used for SeaWulf storage; concurrent file access across all nodes |
| Parallel file system | Distributed storage allowing multiple nodes to access the same files simultaneously |
| Scratch space | High-performance temporary storage (20 TB); files older than 45 days are removed automatically |
Scheduling
| Term | Definition |
|---|---|
| SLURM | The cluster management and job scheduling system |
| Job | A computational task submitted to the cluster, queued until resources are available |
| Queue (partition) | A logical grouping of nodes with similar characteristics, time limits, and priorities |
| Job script | A file of resource requirements and commands, submitted with sbatch |
| Scheduler | Software allocating resources based on priority and availability |
| Resource allocation | Granted access to compute time and storage quotas |
Performance
| Term | Definition |
|---|---|
| Throughput | The rate at which computational work is processed |
| Scalability | Maintaining or improving performance as resources are added |
| Load balancing | Distributing work across nodes to optimize utilization |
| Benchmarking | Measuring system performance with standardized metrics |
| Fault tolerance | Continuing operation despite component failures |
Applies to