Back to Insights
Enterprise8 min read

How to Source Colocation for AI Workloads in 2024

The rapid growth of AI workloads has fundamentally changed how enterprises approach data center strategy. Unlike traditional IT infrastructure, AI and machine learning workloads demand significantly higher power density, specialized cooling solutions, and robust interconnectivity.

Power Density Requirements

AI GPU clusters commonly require 30–50 kW per rack — far beyond the 5–10 kW per rack typical of conventional enterprise deployments. When sourcing colocation for AI, your first filter should be whether a facility can deliver the power density your deployment demands, both today and as you scale.

Key questions to ask providers:

  • Maximum power per cabinet: What is the contracted and burst capacity per rack?
  • Power redundancy: Is 2N or N+1 power distribution available at high densities?
  • Utility feed capacity: Does the facility have room to grow its utility allocation?

Cooling Considerations

High-density GPU deployments generate enormous heat loads that traditional air cooling cannot efficiently manage. Liquid cooling — whether direct-to-chip, rear-door heat exchangers, or immersion — is increasingly a baseline requirement.

Cooling options to evaluate:

  • Direct-to-chip liquid cooling: Most common for NVIDIA DGX and HGX platforms. Requires facility plumbing and CDU (Coolant Distribution Unit) support.
  • Rear-door heat exchangers: A retrofit-friendly option that supplements air cooling with liquid-assisted heat rejection.
  • Immersion cooling: Emerging for the densest deployments. Requires purpose-built tanks and specialized facility design.

Network and Interconnect

AI training workloads require ultra-low-latency, high-bandwidth interconnects between nodes. Evaluate InfiniBand or RoCE availability, and ensure the facility can support your fabric architecture without bottlenecks.

Choosing the Right Provider

Not all colocation providers are created equal when it comes to AI readiness. Look for operators who have invested in: – High-density power infrastructure – Liquid cooling capabilities – On-site fiber connectivity and cloud on-ramps – Flexible contract terms that accommodate rapid scaling

At Stacked AI, we help enterprises navigate these complexities by benchmarking providers against your specific AI workload requirements, ensuring you find the right facility at the best possible terms.