Cooling Failure Redraws the Data Center Risk Map
A Frankfurt outage that took Proton offline put a hard number on the industry's least-priced risk: 30 degrees in 20 minutes.
The graph Andy Yen posted after the outage belongs on the wall of every data center war room: in about 20 minutes, the temperature inside the Frankfurt facility climbed from a 30°C baseline to a 60°C peak, a rise that left alone would have taken Proton's servers past the point of permanent damage. The slow recovery that followed exposes an industry that has engineered for power redundancy while treating cooling failure as a tail risk too improbable to fund.
Proton, the Swiss company behind the privacy-focused email service ProtonMail, went down in the early hours of August 27 after what founder and CEO Andy Yen called an 'unprecedented total cooling failure' at a data center in Frankfurt. In a post published after midnight, the company said it had identified a critical cooling failure and was shifting traffic to backup sites; services were restored around 2 a.m. CEST, and Data Center Dynamics first reported the outage.
Proton's own account makes clear this failure mode had been considered. 'This particular failure mode was actually known to us,' Yen later explained, 'but mitigations were not prioritized because the scenario was judged to be extremely unlikely.' The calculus behind that sentence should worry every operator in the business, because it is not unique to Proton.
Yen was unsparing in response: 'We fucked up, and will do better,' he wrote, acknowledging the company had been slower than its service-level agreement allowed. The delay came from the facility being 'only partially dead' — automatic failover did not trigger because the loss was not total, and staff spent their first minutes trying to restore cooling before servers were permanently damaged. Yen called the temperature graph, with its 30°C baseline and 60°C peak, 'truly frightening how fast temperatures can rise in a high power data center.'
The 'partially dead' failure mode
Partially dead is the dangerous part. Redundant electrical paths are standard equipment — a generator can be started, a UPS can be switched, a feed can be isolated — but when heat removal stops in a high-density room, every rack is at risk at once, and the difference between a graceful transfer and a hardware catastrophe is measured in minutes, not hours. The automatic systems that handle power failovers have no equivalent for heat.
Frankfurt, for the bandwidth
Frankfurt was a connectivity decision. Proton's roots are in a former Swiss military bunker at Attinghausen, buried under 1,000 meters of granite and designed for a war the company's founders did not expect to fight, and the company previously ran infrastructure in Lausanne. The Frankfurt facility sits next to DE-CIX, Europe's largest internet exchange, and Proton launched there around 2021 after its Reddit community picked Germany in a poll; a second community poll three years ago chose Norway for another site that is now live.
Proton told a customer in the aftermath that Frankfurt was a bandwidth decision. The initial data centers in Geneva and Zurich remain running, Frankfurt was an addition, and Switzerland is still the legal jurisdiction regardless of where servers sit, reinforced by zero-access encryption that means Proton cannot read user data even when those servers are in Germany.
That compliance arrangement does not change the physical reality: a service that promises privacy now depends on a commercial colocation provider's cooling plant in an industrial zone in Frankfurt, a plant that just produced a 20-minute temperature spike. For a company that built its brand on the Swiss bunker, the move into the commercial market is a quiet acknowledgment that the privacy wars are won above the floor — in encryption and jurisdiction — not in the concrete of a mountain.
The industry's capital markets have not caught up. The standard data center financial model prices power redundancy aggressively — generators, dual feeds, battery strings — because a power event is a known, quantified financial risk, while cooling is treated as an operating expense, routine air handling. A total power failure buys ride-through time measured in minutes from batteries and seconds for a diesel start; a total cooling failure buys a 30-degree climb in 20 minutes, and the only response is humans running.
Power access has become the binding constraint on what gets built, as this publication has argued, but Proton's outage shows that power is only the first half of the equation. The second half is removing the heat that power creates, and in high-density racks the time constant is frighteningly short. Operators who price thermal integrity as a design-basis condition, with the same seriousness as power rights, will be the ones who own the next decade's uptime.
Yen said his company is working on changes to ensure this type of incident cannot recur, which is the right response. The harder truth is that 'extremely unlikely' is a probability guess, not a design-basis standard. The next operator to face a total cooling failure will not have the benefit of a warning this public. It will have a 20-minute clock.