Case Study
Regional Cloud Provider
4,000-GPU cluster fabric rollout.
- Industry
- Cloud & AI infrastructure
- Region
- Asia-Pacific
- Deployment scale
- 4,000 gpus deployed
- Products used
- G6400, G4400-C, S6700-32CQ +3 more
Challenge
The provider was scaling a GPU training cluster to 4,000 GPUs across multiple datacenter halls, but its fabric was assembled from three vendors: one for switching, one for optics, one for integration support.
Interoperability tickets between spine-leaf software versions and third-party cabling were consuming integration weeks ahead of every expansion wave, and there was no single accountability path when a fabric incident crossed vendor boundaries.
Solution
Standardized the fabric on the S6700 spine family running SONiC, with 400G MPO optics and 1.6T ACC in-rack harnesses from the same qualified optics line.
Matched GPU rows of G6400 air-cooled nodes with G4400-C liquid-cooled rows where hall power density demanded it, keeping one vendor across compute, fabric and cabling.
Deployed Apollo Cloud AI for live fabric telemetry, so cluster health and port-level diagnosis sit in one pane instead of three vendors' tools.
Results
- 4,000 GPUs brought online on a single-vendor fabric bill of materials, with one warranty desk and one escalation path.
- Expansion waves no longer block on cross-vendor interoperability sign-off; the qualified cable matrix removes per-wave optics validation.
- Fabric telemetry via Apollo shortened mean-time-to-diagnosis for port-level incidents during training runs.
Next case study: University Campus Network — 40-building wireless + switching refresh.
