Skip to main content
Plumbing, not hardware

How Meta Is Solving AI's Cooling Crisis

As artificial intelligence hardware grows more powerful, so does the heat it generates—pushing traditional air cooling to its limits. Meta's latest approach turns to engineered plumbing. During a tour of Meta's Texas AI facility, the company revealed how closed-loop liquid cooling is becoming essential infrastructure for data centers.
How Meta Is Solving AI's Cooling Crisis
How Meta Is Solving AI's Cooling Crisis

One of the consequences of the AI boom is that cooling servers has become a serious engineering challenge. The more powerful the AI hardware becomes, the more heat it generates during operation. At a certain point, simply blowing more air through a server rack stops being efficient.

That's why during a visit to Meta's AI Infrastructure in Texas, one of the most interesting technologies wasn't the AI hardware itself—it was the plumbing.

The Cooling Shift

Traditional data centers use air cooling to keep hardware at optimal temperatures. A few years ago, air cooling was sufficient even for AI hardware. Meta visited a data center in Altoona, Iowa, where racks of 16 Nvidia H100s were kept cool entirely through air cooling with minimal water usage—water was used only to cool the air during warmer months, never sent directly to the hardware.

Newer AI hardware designs have since created demand for more optimal cooling methods: closed-loop liquid cooling.

There's a common misconception that AI data centers are automatically big water users. The reality depends on the cooling design. Meta's data centers use a closed-loop system that recirculates water in a sealed loop, using very little on an ongoing basis. Most of Meta's newest AI-optimized data centers use closed-loop liquid cooling as the most efficient method for cooling GPU servers—both from a resources and infrastructure standpoint.

What Is Closed-Loop Cooling?

A liquid coolant (a mix of water and glycol) is passed through server hardware to move heat away from the racks. Instead of being expelled, the liquid is pumped through heat exchangers that dissipate and transfer the heat away. Once cooled, it circulates back to the server racks in a continuous loop, with Meta expecting to use the same coolant for up to a decade without replacement.

Heat transfer methods vary by location and environment. For facilities without built-in liquid cooling infrastructure, Meta uses Air-Assisted Liquid Cooling—racks containing pumps and heat exchangers that perform the same closed-loop cooling on a smaller, distributed scale.

Closed-loop liquid cooling is ideal because it's resource efficient. A typical AI-optimized data center using a closed-loop system with dry coolers uses less water annually than a couple of full-service restaurants. When compared to real use cases rather than raw numbers, the low usage is impressive.

03_WaterCooling_Carousel-01.mp4 1 / 2

Optimizing for Efficiency

Beyond water savings, closed-loop liquid cooling is more efficient in terms of rack and data center space. Cooling the same servers with air would likely require nearly double the server tray size to accommodate required air cooling equipment—resulting in a much bigger tray but the same compute capacity, with diminishing returns as solutions grow larger.

With direct-to-chip closed-loop liquid cooling, engineers can fit more GPUs in the same server rack, requiring fewer racks overall. A facility of the same size can now scale capacity without needing additional space.

Meta's Open-Source Liquid Cooling Infrastructure

Meta designs and develops systems across its entire infrastructure stack, from custom chips to cooling and power infrastructure. Consistent with the Open Compute Project legacy, these advances are shared with the industry.

The Open Compute Project (founded 2011) is an open-source hardware and software initiative aimed at making data center infrastructure more efficient, scalable, and sustainable. In 2025, Meta announced IcePack, a liquid-cooled network rack platform being shared openly via the Open Compute Project.

Using AI to Optimize Data Center Cooling

Photo of Nvidia H100s in in a server rack using air cooling
Photo of Nvidia H100s in in a server rack using air cooling — Meta

Meta's engineering teams use reinforcement learning to optimize their cooling infrastructure. Cooling isn't as simple as setting a temperature and running the system—conditions and environments change, server workloads fluctuate, and cooling demands shift. The infrastructure must be flexible and purpose-built for each location.

Rather than experimenting on live data centers, where mistakes could disrupt operations, Meta's engineers built a physics-based simulator modeling variables such as weather conditions, server load, and cooling equipment behavior. This gives the reinforcement learning model a safe environment to test different decisions and learn how to reduce cooling requirements while keeping servers within optimal operating conditions.

This is no longer experimental. In a pilot at one of Meta's data centers, the reinforcement learning approach reduced energy consumed by air cooling supply fans by an average of 20% while reducing water usage by 4% across different weather conditions. Applied across an entire data center fleet, these reductions represent significant efficiency gains.

Jennifer friend

The greatest technological advancement is our ability to be truly present where life happens – Jenny F.