AI load swings are outrunning the redundancy playbook
A DCD opinion piece separates the failure modes an engineer can name from the interactions between healthy systems that millisecond AI load swings can set off.
Data center engineers are very good at redundancy, a competence built on a history in which most operational problems traced to equipment failures, inadequate capacity, design deficiencies, or maintenance, and in which fault-current levels, protection coordination, generator performance, cooling capacity, and UPS autonomy could be modeled with precision. A Data Center Dynamics opinion piece published September 27 argues that being very good at redundancy is no longer the same thing as being good at resilience, and AI infrastructure is beginning to press on that distinction.
Two shifts land at once: the power demand of large AI deployments can swing within milliseconds, while the halls running them lean increasingly on UPS inverters, power-electronic converters, lithium-ion batteries, software-driven controls, automated switching logic, and digital operating platforms, a power-electronics-heavy stack threaded through multiple closed-loop control architectures. Individually each technology is a solid machine; collectively, the column argues, they form an interconnected dynamic system whose interactions may matter as much as the performance of its parts.
The discipline that answers for that is not the one these facilities staff for: mission-critical design has rested on steady-state performance — capacity planning, redundancy architecture, fault-current studies, thermal loading, protection coordination — and, though every item on that list remains necessary, none of it addresses how dozens of healthy systems behave together under highly dynamic operating conditions.
As the piece notes, mission-critical professionals rarely discuss poles and zeros, the roots of a system's transfer-function denominator, which determine whether a dynamic system settles smoothly, oscillates, or goes unstable; the teaching example is a vehicle suspension, a mass-spring-damper whose poles decide whether the car glides or bounces.
The gap shows up in the paperwork: redundancy is something an owner can spec, price, diligence, and test at commissioning, then write into an uptime agreement as availability, while dynamic stability is a property of the assembled system, and the contract package around an AI campus has nowhere to record it. Grid access is the deal currency now and the interconnection queue position is the asset, as PWD has argued; the dynamic-stability argument complicates that position because an executed interconnection agreement delivers megawatts, and damping is not in it. For anyone underwriting AI capacity, redundancy protects against the failure modes an engineer can name, while instability that emerges between two healthy subsystems never appears on an equipment schedule.
One limit is that this is an opinion column rather than a record of measured failures; the mechanism it describes — sustained utilization near infrastructure limits, millisecond load swings, control loops layered on control loops — is mechanical, so the measured evidence an underwriter would want is not in it.
Commissioning is where this surfaces: a load-step test run on the integrated power train, stepping demand on a millisecond profile and measuring recovery time instead of verifying components against nameplates, would surface interactions that equipment-level testing cannot see. That test is not in the standard script, and the redundancy record these assets are sold on describes a system nobody has modeled yet.
Save this analysis and keep the funds you follow together in My Desk.
Sign in to save articles or follow funds.