In this explainer
  1. The conversation inside the building
  2. A bigger team needs a better meeting system
  3. What a GPU count leaves out
← All explainers

NETWORKING · EXPLAINER

Why more AI chips don’t always mean faster AI

The overlooked connections that turn a room full of processors into a coordinated machine.

A company can buy more AI chips and still end up disappointed by the speed of its training system. The chips may work perfectly. The problem is getting them to work together.

Imagine dividing a big calculation among a team. Everyone finishes their assigned portion, but the next stage requires results from the others. A faster calculator helps only until waiting becomes the dominant part of the job.

That is one reason an AI cluster is more than a collection of servers. Its connections are part of the computing system.

The conversation inside the building

In one common training method, different processors work with copies of the same model and different portions of the training data. They calculate proposed adjustments, then synchronize those adjustments so the copies continue learning together. PyTorch’s Distributed Data Parallel documentation describes this coordination through the exchange of gradients—the numerical information used to update the model. Other approaches split the model itself across devices, creating different communication needs. PyTorch: Distributed Data Parallel

This traffic is not simply someone downloading a chatbot response from the internet. It is communication among the machines performing the training.

One building block is called an all-reduce. Each participant contributes values; an operation combines them, such as by adding them, and makes the resulting values available to every participant. NVIDIA documents this alongside other collective operations that distribute or gather data in different ways. The important idea is coordination, not the acronym. NVIDIA: Collective operations

A bigger team needs a better meeting system

The practical inference is straightforward: adding workers does not guarantee a proportional increase in completed work when those workers depend on shared results. Some of the benefit can disappear into communication and waiting.

That does not mean every AI workload is network-limited. Nor does it mean a chip sits idle whenever a message travels. The balance depends on the workload and how the system schedules calculation and communication.

Meta offered a useful real-world example in its March 2024 account of its AI infrastructure. It described initially inconsistent communication performance in large clusters and changes to job placement, routing, and communication software that improved it. It also described using both Ethernet-based and InfiniBand networks. Those were specific engineering results, not proof that one network is universally best. Meta: Building GenAI infrastructure

What a GPU count leaves out

A headline announcing thousands of processors tells us something about the equipment purchased. It does not fully describe the useful work the installation can deliver.

The better questions are: What will those processors do? How much information must they exchange? Has the complete system been tested with that workload? Does performance hold up as the job grows?

The surprising part is that some improvements to an AI machine can come from arranging and coordinating the equipment already there. Buying faster chips is one route to better performance. Giving those chips fewer reasons to wait is another.