NVIDIA Mellanox MQM9790-NS2F in Action: Building a Low-Latency RDMA Fabric for AI and HPC Clusters

August 24, 2026

Dernières nouvelles de l'entreprise NVIDIA Mellanox MQM9790-NS2F in Action: Building a Low-Latency RDMA Fabric for AI and HPC Clusters

NVIDIA Mellanox MQM9790-NS2F in Action: Building a Low-Latency RDMA Fabric for AI and HPC Clusters

Background & The Challenge: When Scale-Out Becomes the Bottleneck

As large language models scale from thousands to tens of thousands of GPUs, and weather simulations push toward kilometer-scale resolution, the limitations of traditional Ethernet fabrics become glaringly apparent. In these environments, MPI collective communication patterns—such as all-reduce and all-to-all—place extreme demands on network latency and congestion control. Without a purpose-built interconnect, even the most powerful compute nodes spend a significant portion of their cycles waiting for data, effectively wasting expensive GPU resources.

This is precisely the challenge that the NVIDIA Mellanox MQM9790-NS2F addresses. As a 400Gb/s NDR InfiniBand switch with 64 OSFP ports, it is engineered to deliver deterministic, sub-microsecond latency across thousands of endpoints, enabling true linear scaling for distributed workloads.

Solution & Deployment: A Leaf-Spine Fabric Built for RDMA

In a typical deployment scenario for a large AI research lab, the MQM9790-NS2F InfiniBand switch forms the core of a two-tier leaf-spine topology. At the leaf layer, each compute rack is equipped with one or two MQM9790-NS2F units, providing 64 ports of 400Gb/s connectivity to GPU servers via OSFP-to-OSFP direct-attach copper cables or active optical cables. At the spine layer, additional MQM9790-NS2F switches interconnect the leaves, creating a non-blocking fat-tree fabric with full bisection bandwidth.

What makes this solution particularly compelling is the switch's native support for RDMA over Converged Ethernet (RoCE) when operating in InfiniBand mode—though here it leverages InfiniBand's native RDMA capabilities for even lower latency and CPU offload. The MQM9790-NS2F 400Gb/s NDR 64-port OSFP integrates NVIDIA's SHARPv3 (Scalable Hierarchical Aggregation and Reduction Protocol) technology, which offloads collective operations directly onto the switch fabric. In practice, this means that an all-reduce operation that would normally traverse the network multiple times is now completed in a single pass, slashing job completion times by up to 30% for communication-intensive workloads.

For network architects evaluating the MQM9790-NS2F InfiniBand switch solution, the deployment process is streamlined by NVIDIA's Unified Fabric Manager (UFM). UFM provides centralized visibility into fabric health, automated topology discovery, and proactive congestion monitoring—critical capabilities when managing fabrics with over 2,000 ports. Early adopters have reported that the MQM9790-NS2F compatible ecosystem, which includes a wide range of NVIDIA-certified cables and transceivers, simplifies procurement and reduces deployment risk.

Results & Measurable Gains: From Theory to Production Reality

In a recent production deployment for a 2,048-GPU cluster dedicated to large-scale transformer model training, the upgrade from HDR (200Gb/s) to NDR (400Gb/s) infrastructure—with the NVIDIA Mellanox MQM9790-NS2F at its heart—yielded measurable improvements across multiple dimensions:

  • Job completion time reduction: End-to-end training iterations for a 175B-parameter model decreased by 28%, largely attributed to SHARPv3 offloading and reduced tail latency.
  • Network utilization: Average fabric utilization increased from 62% to 89% without triggering congestion collapse, thanks to the switch's advanced adaptive routing and congestion control mechanisms.
  • Operational simplicity: With 64 ports per switch, the number of spine switches required was cut by half compared to a 32-port HDR design, significantly reducing cabling complexity and power consumption per rack.

According to the MQM9790-NS2F specifications, the switch delivers sub-100ns port-to-port latency and supports up to 51.2Tb/s of aggregate switching capacity—figures that align closely with the performance observed in this deployment. Network engineers also noted that the MQM9790-NS2F datasheet accurately reflected real-world performance, with no surprises during the validation phase.

From a TCO perspective, the ability to consolidate more endpoints per switch tier directly reduced capital expenditure. While initial MQM9790-NS2F price points are higher than equivalent 200Gb/s solutions, the per-GPU networking cost actually decreased due to the higher radix and reduced number of switch layers. This economic advantage, combined with the performance gains, made the business case clear for the lab's steering committee.

Summary & Outlook: The NDR Foundation for Exascale Computing

The NVIDIA Mellanox MQM9790-NS2F is more than an incremental upgrade—it represents a fundamental shift in how HPC and AI clusters are designed. By delivering 400Gb/s NDR bandwidth, 64-port density, and in-network compute acceleration in a single 1U form factor, it enables architects to build fabrics that scale to tens of thousands of endpoints without sacrificing performance or manageability.

Looking ahead, as workloads continue to demand higher bandwidth and lower latency, the MQM9790-NS2F InfiniBand switch will serve as the foundational building block for next-generation exascale systems. Its compatibility with existing NVIDIA Quantum-2 platforms and seamless integration with UFM ensure a smooth migration path for organizations upgrading from HDR fabrics. For IT managers and architects evaluating their next-generation networking investments, the MQM9790-NS2F offers a proven, production-ready solution that delivers on the promise of low-latency RDMA interconnect optimization.