When an AI training run fails, does it start over?
The saved progress that can prevent an interrupted training job from losing all its work—and the infrastructure needed to preserve it.
If a computer crashes while you are writing, the first question is usually: when did it last save?
Large AI training jobs face their own version of that question. An interruption does not necessarily erase all previous progress. But recovery depends on having a usable saved state, and making that save is itself a piece of engineering.
The saved state is called a checkpoint.
More than a finished model
A model’s learned parameters describe its current state. Continuing training can require more information than that alone. PyTorch’s saving-and-loading guide explains that a training checkpoint should also preserve the optimizer’s state, and can include information about where training stopped. The optimizer is the mechanism that uses training results to adjust the model. PyTorch: Saving and loading models
A useful analogy is saving both a document and enough information about the work in progress to continue the task. A file intended only for using a trained model need not serve exactly the same purpose as a checkpoint intended for resuming training.
That is why “we saved the model” is not always a complete answer to “can we resume the job?”
Saving has a cost
PyTorch’s distributed-checkpoint tutorial identifies checkpointing as a potential training bottleneck. Its asynchronous approach copies state into CPU memory buffers and lets saving proceed separately, with additional memory requirements and a need to manage outstanding saves. Background saving therefore reduces a particular kind of interruption; it does not make the work or resource use disappear. PyTorch: Asynchronous checkpointing
The practical tradeoff follows from that. Saving more often can reduce the amount of progress exposed between saves, but adds more saving work. Saving less often does the reverse. There is no universal interval that is best for every model and system.
For a purely illustrative example, suppose a job saves successfully at noon and stops at 12:20. If recovery returns to that noon state, the later progress must be repeated. That is twenty minutes of lost progress, not the loss of the whole training run. This example assumes the saved state remains accessible and usable; it is not a measured incident.
Storage becomes part of the AI machine
Meta’s 2024 infrastructure account describes storage supporting synchronized checkpoint saves and loads across thousands of GPUs. It also identifies finding a problematic GPU that stalls a training job as a difficult operational task. Those details show why a recovery plan involves storage and diagnosis as well as replacement compute. Meta: Building GenAI infrastructure
There are therefore two questions to ask after an interruption: how much progress was preserved, and how long will it take to resume useful work? They are related, but they are not the same measure.
An AI cluster is built to calculate. It also needs to preserve enough of those calculations that a disruption does not destroy their value. The unglamorous save operation can be one of the most important jobs in the building.