Thermal lag in CDU control systems is a hidden reliability risk in liquid-cooled AI data centers
PID-controlled cooling distribution units are tuned for steady-state loads. AI accelerators produce millisecond-scale power transients. The gap between those two timescales is an unaccounted-for reliability risk.

A detailed engineering analysis published by Data Center Dynamics on 27 July 2026 argues that the standard control architecture used in liquid-cooled data centers carries a structural reliability flaw that the industry has not yet priced into its uptime models[1]. The mechanism is thermal lag in coolant distribution unit (CDU) control systems - and the analysis contends it is not a tuning problem but a property of the underlying physics.
The mismatch between accelerator and controller timescales
A CDU is a closed-loop feedback control system that monitors supply and return coolant temperature, differential pressure, and volumetric flow rate, then adjusts pump speed and valve position to maintain defined setpoints[1]. In the large majority of deployed systems, this is implemented via proportional-integral-derivative (PID) control, with parameters tuned during commissioning for stable operation under expected steady-state load conditions[1].
The problem is that AI accelerators do not present steady-state loads. Inference workloads produce stochastic power spikes reaching 60 to 80 percent of peak thermal design power (TDP) within 30 to 50 milliseconds as the accelerator transitions between prefill and decode processing phases[1]. Training runs produce a step function in power demand, not a ramp[1]. A controller tuned for gradual load changes cannot respond to events that resolve in tens of milliseconds.
The lag chain
The DCD analysis traces the full sequence of events from a GPU power surge to the point at which the CDU delivers corrected coolant to the cold plate[1]. Each stage of the control chain adds delay:
- Sensor thermal response: PT100 resistance thermometers in standard industrial pocket-mount configurations take three to eight seconds to register a temperature change[1].
- Fluid transit through the secondary loop: With secondary loop volumes of 10 to 25 liters per rack row and design flow velocities of 0.2 to 0.5 meters per second, propagation delay runs 15 to 45 seconds[1].
- Heat exchanger equilibration: Based on plate heat exchanger water content of 10 to 40 liters at representative flow rates, equilibration takes 15 to 60 seconds[1].
The physics of PID control in a hydronic thermal system introduces inherent response lag at every stage of the control chain - not as a consequence of poor design, but as a property of the physical system the controller is acting upon[1].
Reliability consequences across two timescales
The analysis identifies what it calls a "thermal debt window" - the interval between a GPU power surge and the CDU's corrective response - and argues it carries consequences at two distinct timescales[1].
In the immediate term, junction temperatures on accelerators operating near their thermal design limit will exceed the steady-state baseline during the debt window[1]. For a rack-scale AI compute system operating at 95 percent of peak TDP with a supply coolant temperature of 35°C, a CDU response lag of 60 seconds can permit the junction temperature to rise 10 to 20°C above its steady-state operating point[1]. For accelerators tuned to operate near the firmware-level throttle threshold - a configuration common in deployments optimizing maximum utilization per watt of cooling energy - this exceedance triggers frequency capping and a measurable drop in compute throughput[1].
Over longer timescales, the damage is cumulative. A device that ceases to function after 18 months of operation in a liquid-cooled AI deployment may be exhibiting the accumulated mechanical damage of tens of thousands of thermal cycles, rather than a latent manufacturing defect[1]. That failure mode is almost impossible to distinguish from normal hardware attrition without correlating BMC logs against CDU response timestamps.
The monitoring blind spot
The analysis identifies a second problem layered on top of the control lag: operational monitoring platforms for liquid-cooled AI deployments are configured almost universally around steady-state thermal parameters[1]. The transient signal exists in the data - modern AI accelerators expose junction temperature telemetry through the baseboard management controller at sub-second polling intervals[1] - but in most deployed systems, that data is not being correlated against CDU response timestamps[1].
An accelerator operating at 35°C in steady state and reaching 52°C during a 75-second transient will register that excursion in BMC logs — if that data is being collected and correlated against CDU response timestamps. In most deployed systems, that correlation is not being made.[1]
What to watch
The analysis points to feedforward-augmented CDU control as the practical fix. Rather than waiting for return temperature to trigger corrective action, a feedforward system ingests GPU power telemetry directly via the baseboard management controller or the accelerator's power management interface, and pre-emptively adjusts pump speed and valve position before the thermal load arrives at the heat exchanger[1]. End-to-end latency from GPU power sensor to CDU actuator can be reduced to under two seconds in a well-implemented feedforward architecture, collapsing the thermal debt window from 60 to 120 seconds to under 10 seconds[1]. The capability is increasingly available as a software option in current-generation CDU control systems[1].
The question for operators commissioning new liquid-cooled AI halls is whether their CDU vendor's control software supports feedforward integration with GPU telemetry - and whether their monitoring stack is actually collecting the BMC data needed to detect thermal excursions when they occur.
The images and texts on this page were created with the help of AI.
Related
Markets & PolicyOfgem publishes final DCC remuneration guidance under Condition 20, tying senior pay to consumer outcomes ahead of 2 November transfer
Ofgem has finalised guidance linking DCC2 senior manager pay to operational performance and consumer value, issued under Condition 20 of the Smart Meter Communication Licence ahead of the 2 November 2026 handover.
8 Sept 2026
RenewablesShaun Campbell, Windpower Monthly editor from 2016 to 2020, dies at 67
Windpower Monthly has announced the death of former editor Shaun Campbell, who passed away in August after a short illness, aged 67.
8 Sept 2026
RenewablesIndustry coalition urges Denmark to publish a long-term offshore wind roadmap to 2050
The Offshore Wind 2050 Partnership says Denmark needs a durable planning framework and three new development tracks to unlock up to 80 GW of offshore capacity by 2050.
8 Sept 2026