Scaling Large MoE Models with Wide Expert Parallelism on NVL72 Rack Scale Systems

59 · NVIDIA Corporation · Oct. 20, 2025, 4:05 p.m.
Summary
This post discusses advancements in AI workloads, specifically focusing on model parallelism techniques that allow for efficient computation across multiple GPUs in NVL72 Rack Scale Systems. It outlines the significance of wide expert parallelism in scaling large Mixture of Experts (MoE) models, highlighting practical implementations and potential benefits for developers working in AI and machine learning.