Skip to content
OpinionBuying & operations · 3 min read

Before buying more GPUs, find out why the current ones are waiting

Our view: capacity planning should start with completed useful work, not with a utilisation chart or a list of new accelerators.

Opinion · Our view, supported by the sources below.

Editorial illustration of active and idle computing tiles joined by a flowing connection
Installed capacity and productive capacity are different measures. Conceptual illustration; no utilisation data is represented.Editorial illustration · Our artwork, created with AI

An expensive machine can be busy and unproductive

GPU utilisation is a useful signal, but it is not the objective of an AI service. A machine can be heavily occupied by retries or work that does not meet the application's quality requirements. It can also look lightly used because requests arrive unevenly or another part of the system is delaying them.

Our position is that the first capacity question should be how much acceptable work the service completes. Only then should the team ask what limits that result and whether another GPU will improve it.

Follow the request through the system

Measure the time spent receiving data, preparing inputs, waiting in queues, running inference and completing downstream steps. Separate ordinary requests from unusually demanding ones. Record failures and tail latency instead of relying only on averages.

If storage or a CPU service is the bottleneck, additional accelerators may increase idle capacity. If memory limits concurrency, a scheduling change or different memory configuration may be more relevant than additional nominal compute.

NVIDIA Dynamo's support for context-aware routing and distributed inference illustrates how software can change the way hardware is used. It does not guarantee an improvement for every application, but it provides a reason to examine the serving architecture before assuming the only remedy is more devices. NVIDIA Dynamo: distributed inference architecture ↗

Utilisation must leave room for the service

There is a limit to this argument. A customer-facing system may need spare capacity for bursts, failures or maintenance. Driving every accelerator towards constant saturation can damage response times and resilience.

The target should therefore follow the service requirement. An offline batch queue can often tolerate different scheduling from an interactive product. Neither should be judged against a universal utilisation percentage.

Likewise, optimisation has a cost. Spending months of engineering effort to avoid one affordable hardware addition may be a poor business decision. The comparison should include staff time and the value of delivering the product sooner.

Buy the constraint you have identified

After measurement, the answer may indeed be more GPUs. That is a stronger purchase case when the team can explain the load the new capacity will carry and the conditions under which it will be needed.

Keep a simple forecast with ordinary demand, peak demand and a growth scenario. Revisit it after software and model changes, because those changes can alter the resource profile.

We support buying substantial compute when the workload warrants it. We also support fixing the data pipeline, adjusting scheduling or choosing a better-sized system first. The commercial objective is a service that performs reliably at an acceptable cost, not a rack that merely looks fully occupied.

Sources & further reading

Primary sources for the reported developments and technical context. Analysis and conclusions are our own; linked specifications and documentation can change.

Sources checked 29 September 2026.

How we cover the industry

Our editorial team writes about AI infrastructure, equipment procurement and the industry behind it. News analysis distinguishes reported developments from our conclusions; opinion articles are labelled as such.

Technical and industry references are linked within each article. Publication dates describe when an article was written, rather than implying that every specification or market condition remains unchanged.