Skip to main content
Foredge

Case Study

Regional Cloud Provider

4,000-GPU cluster fabric rollout.

Industry
Cloud & AI infrastructure
Region
Asia-Pacific
Deployment scale
4,000 gpus deployed
Products used
G6400, G4400-C, S6700-32CQ +3 more
Deployment overview — schematic

Challenge

The provider was scaling a GPU training cluster to 4,000 GPUs across multiple datacenter halls, but its fabric was assembled from three vendors: one for switching, one for optics, one for integration support.

Interoperability tickets between spine-leaf software versions and third-party cabling were consuming integration weeks ahead of every expansion wave, and there was no single accountability path when a fabric incident crossed vendor boundaries.

Solution

Standardized the fabric on the S6700 spine family running SONiC, with 400G MPO optics and 1.6T ACC in-rack harnesses from the same qualified optics line.

Matched GPU rows of G6400 air-cooled nodes with G4400-C liquid-cooled rows where hall power density demanded it, keeping one vendor across compute, fabric and cabling.

Deployed Apollo Cloud AI for live fabric telemetry, so cluster health and port-level diagnosis sit in one pane instead of three vendors' tools.

Results

  • 4,000 GPUs brought online on a single-vendor fabric bill of materials, with one warranty desk and one escalation path.
  • Expansion waves no longer block on cross-vendor interoperability sign-off; the qualified cable matrix removes per-wave optics validation.
  • Fabric telemetry via Apollo shortened mean-time-to-diagnosis for port-level incidents during training runs.

Next case study: University Campus Network40-building wireless + switching refresh.