NVIDIA GB200 NVL72 Optimized with Slurm for AI Supercomputing
NVIDIA‘s GB200 NVL72, a cutting-edge rack-scale AI supercomputer, is now achieving optimized performance through topology-aware job scheduling with Slurm. This advancement is critical as AI models, particularly trillion-parameter large language models (LLMs), demand both unprecedented compute power and efficient resource allocation. The system, built on NVIDIA’s Blackwell architecture, delivers up to 130 terabytes per second (TB/s) of GPU communication bandwidth and supports training and inference for some of the most complex AI workloads. The GB200 NVL72 integrates 72 NVIDIA Blackwell GPUs and 36 Grace CPUs in a single rack, interconnected via NVIDIA NVLink. According to NVIDIA, this setup not only supports large-scale training but also accelerates real-time inference with over 1.5 million tokens per second for OpenAI GPT models. However, maximizing this performance in shared clusters requires strategic scheduling, as highlighted in NVIDIA‘s collaboration with SchedMD to enhance Slurm’s topology-aware capabilities. Why Scheduling Matters for Exascale Systems AI workloads often run on shared clusters, where multiple jobs must compete for resources. Without topology-aware scheduling, jobs may span across NVLink domains inefficiently, leading to resource fragmentation and reduced performance. The newly introduced Slurm topology/block plugin aligns jobs with the physical network layout of the GB200 NVL72, preserving locality and minimizing fragmentation. This ensures that