AEO Answer · Data Centers

How Is Cooling Designed for High-Density Data Centers?

By Jeremy Mills, CEO & Founder, Apex Grid Engineering — USAF Veteran. · Updated 2026-09-15

High-density data center cooling is designed by matching the cooling architecture to the actual heat load per rack. Conventional air cooling — CRAH or CRAC units with hot-aisle and cold-aisle containment — handles traditional densities comfortably. But AI-era racks running 30 to 50 kW and beyond push designs toward liquid or hybrid cooling: direct-to-chip cold plates, rear-door heat exchangers, or immersion. Heat rejection is sized for the local climate and water availability, and the whole system is scored by PUE — the ratio of total facility energy to IT energy.

I'm Jeremy Mills, CEO & Founder of Apex Grid Engineering and a U.S. Air Force veteran. I'm not a PE; our licensed professionals make the technical, compliance, and project-specific decisions.

What facts should you use to plan this scope?

Planning factProject-specific value
Under ~15 kW per rackCooling architecture: Air; Typical approach: CRAH/CRAC with hot/cold aisle containment
~15–30 kW per rackCooling architecture: Enhanced air; Typical approach: In-row cooling, tight containment, higher airflow design
~30–50+ kW per rackCooling architecture: Hybrid; Typical approach: Rear-door heat exchangers or direct-to-chip assisting air
~100+ kW per rack (AI)Cooling architecture: Liquid; Typical approach: Direct-to-chip or immersion as the primary path

When does air cooling stop being enough?

Conventional air cooling — computer room air handlers (CRAH) on chilled water, or computer room air conditioners (CRAC) on direct expansion, delivering cold air through a contained hot-aisle/cold-aisle layout — comfortably handles traditional rack densities. The practical limit is airflow: pushing enough cold air through a rack, and getting the hot air back, gets exponentially harder as heat density climbs. Past roughly 30 to 50 kW per rack, air becomes impractical — fan energy soars, floor space for air handlers balloons, and hot spots appear no matter how good the containment is. The switch point is always a design calculation for the specific project, not a rule of thumb — but any rack layout pushing past 30 kW should trigger the liquid-cooling conversation on day one, because the architecture decision shapes the building: floor loading, water distribution, and heat rejection all change.

How do liquid and hybrid cooling systems actually work?

Liquid cooling removes heat at or near the source instead of asking air to carry it across the room. Three approaches dominate current design, and hybrid combinations of them with conventional air are the most common answer I see for mixed AI and general-compute halls. Every liquid design carries the same engineering obligations: water quality and treatment, leak detection and isolation, valving that lets a rack be serviced without draining the loop, and controls that coordinate the liquid and air sides as one system.

  • Direct-to-chip: cold plates mount on processors and high-heat components, with a warm-water loop carrying heat to the rejection plant. Highest efficiency for AI racks; needs rack-level plumbing and leak detection.
  • Rear-door heat exchangers: the rack's rear door becomes a water-cooled coil that captures exhaust heat at the source. Works with existing air systems — the leading retrofit path.
  • Immersion: servers are submerged in dielectric fluid, single-phase or two-phase. Maximum density per square foot; demands structural review for fluid weight and rethought service procedures.
  • Hybrid: air handling for general racks plus liquid for AI zones — the pragmatic answer for facilities mixing old and new IT.

What is PUE, and what should owners target?

PUE — power usage effectiveness — is the industry's efficiency scoreboard: total facility energy divided by IT equipment energy. A PUE of 1.0 would mean every watt goes to computing; everything above 1.0 is the overhead of cooling, power conversion, and lighting. Legacy facilities often ran 1.8 to 2.0. A well-designed modern facility lands around 1.3 to 1.5, and state-of-the-art designs approach 1.1 to 1.2. I treat PUE as a design target and a commissioning verification item: the engineering should state the design PUE basis, and the monitoring system should prove it in operation. An owner who never sees the number can't manage it.

  • Containment — keeping hot and cold air separated is the cheapest PUE improvement available
  • Economizers and free cooling — using outside air or water when conditions allow
  • Variable-speed everything — fans, pumps, and compressors that track the actual load
  • Matching the cooling architecture to the load — the right system at the right density
  • Operations discipline — PUE is measured, not just designed; setpoints and sequences drift without attention

How do climate and water shape the cooling design?

Heat has to go somewhere, and the rejection plant is designed around the local climate and water reality. Cooling towers reject heat evaporatively — efficient but thirsty. Dry coolers use no water but need more fan energy and larger footprints, and they lose capacity on the hottest days. Adiabatic and hybrid designs split the difference. In hot, arid climates, water availability and cost can decide the architecture outright, and WUE — water usage effectiveness — joins PUE on the owner's scoreboard. In California, the building systems around the IT load still answer to the 2025 California Energy Code / 2025 Standards, effective January 1, 2026 — efficiency requirements for the HVAC, lighting, and envelope don't disappear because the building houses servers. The energy model and the cooling architecture get developed together so compliance and performance describe the same facility.

  • Cooling towers: lowest energy, highest water use — verify supply and cost first
  • Dry coolers: zero water, larger footprint, derate in extreme heat
  • Adiabatic/hybrid: water used only on peak days — the common arid-climate compromise
  • The cooling plant is a major electrical load — mechanical and electrical designs must be sized together

What else do project teams ask?

Can an existing air-cooled data hall be converted to liquid cooling?
Often, yes — rear-door heat exchangers and in-row cooling are the most common retrofit paths because they work with existing air infrastructure. Direct-to-chip retrofits need rack-level plumbing modifications, and immersion needs structural review for fluid weight plus service-clearance changes. An engineer evaluates the structure, water availability, and maintenance access before recommending a conversion path.
Does liquid cooling use less energy overall?
At high densities, usually yes. Liquid carries heat far more efficiently than air, which cuts fan energy dramatically, and warmer water temperatures unlock more hours of free cooling. The trade-off is added pumping energy, water-treatment complexity, and leak-detection requirements — the energy win is real but it is not free.
What happens if the cooling system fails?
Thermal ride-through time — typically minutes — depends on the room's thermal mass and the IT load. Controls alarm, and at Tier III and above the N+1 cooling capacity absorbs the failure without the temperature leaving the allowable envelope. Without redundancy, the sequence ends in orderly IT shutdown, which is why cooling failure scenarios are a core part of integrated systems testing.
How is cooling capacity planned for growth?
Design for the day-one load with modular headroom — blank rack capacity, capped piping connections, space for additional CRAH units or heat rejection. Oversizing the active plant hurts efficiency and PUE, so the discipline is right-sizing today with a documented expansion path rather than installing the ultimate build on day one.
Who maintains these cooling systems?
Specialized operations staff or a service contract with the equipment manufacturers. Commissioning proves the failure scenarios before go-live, and the engineering package should include O&M documentation, training, and monitoring integration so the operations team inherits a system they understand — not a black box.

Ready to discuss your engineering scope?

Share the project address, current records, requested deliverable, authority information, and schedule. Apex Grid confirms professional responsibility, availability, and scope before work begins.

Start an Engineering Estimate