In this explainer
  1. Start with what is actually protected
  2. Spare equipment and a working service are different promises
  3. Sometimes the answer is another location
  4. Ask what happens when something fails
  5. Watch the power connection
← All explainers

RELIABILITY · EXPLAINER

Why backup power doesn’t make a data center invincible

Keeping an AI service available means protecting more than the electricity entering its servers.

A data center can have backup power and still lose service. That sounds contradictory until we ask what, exactly, the backup protects.

Electricity is one requirement. The machines also need a working environment, functioning connections, and software capable of continuing or recovering. Protecting one part of the chain does not automatically protect the whole service.

Start with what is actually protected

AWS describes its data-center electrical systems as redundant and maintainable, with backup supplies supporting critical and essential facility loads during electrical failures. It separately describes temperature control, monitoring, and preventive maintenance. These are statements about AWS’s design and operating practices, not a specification for every AI facility. AWS: Data-center controls

The distinction matters. When someone says a facility “has backup,” the phrase leaves many questions unanswered. Which loads are covered? What failures were considered? What happens during maintenance? How is the complete arrangement tested?

An equipment list alone cannot answer those questions. A reassuring photograph of backup machinery is still only a photograph of machinery.

Spare equipment and a working service are different promises

Consider a deliberately simplified example: a service needs an electrical supply, a functioning server, and a network path. Providing an alternative electrical supply addresses a failure in that part of the arrangement. It does not, by itself, provide an alternative server or network path.

This is a logical illustration, not a claim about any named operator’s wiring. It explains why resilience has to be assessed across a system’s dependencies.

The same reasoning applies to maintenance. “We have a spare” and “we can take this component out of service without interrupting the workload” are different claims. To evaluate either, we need to know the paths and dependencies around the component, not simply how many units are installed.

Sometimes the answer is another location

AWS also describes physically separated Availability Zones and the ability to design applications to fail over between them. The crucial word is design: the existence of another location does not mean every application automatically uses it correctly. AWS: Availability and redundancy

For an AI service, the useful question is therefore broader than whether the original building stays powered. Can the application continue serving users elsewhere, or recover acceptably after a disruption?

Those are different goals. Continuing without interruption is not the same as restarting after a pause. A system can satisfy one requirement while failing another.

Ask what happens when something fails

A more useful conversation about a new campus would begin with its failure plan rather than a promise that failures will never happen.

Ask which services remain available, which operations may stop, and what evidence supports the recovery claims. Distinguish the operator’s documented design from an assumption based on the size of the building.

Backup power matters enormously. But the broader lesson is that reliability belongs to the complete service. A protected power supply is one part of that story; it is not the ending.