Foundation Model Training Bottlenecks: Memory Bandwidth, Not GPU Count

Ana Pace

September 29, 2026

A training run falls behind schedule, and the default response is almost always to add GPUs. Procurement approves a bigger node, the cluster grows, and throughput barely moves. That's not a rare failure, it's the predictable result of sizing infrastructure around the wrong number. Past a certain point, the constraint on a foundation model training run isn't raw compute, it's how much memory each GPU has and how fast data moves between cards when gradients need to synchronize. A cluster with more GPU compute than its memory and networking can actually feed isn't a faster cluster, it's a more expensive one, paying full price for silicon that sits idle between synchronization steps. The three numbers that actually determine throughput, per-GPU memory capacity, interconnect bandwidth, and whether that interconnect is shared with other tenants, rarely show up on the spec sheet a procurement team compares against.

16 to 20 bytes per parameter: the memory math the spec sheet skips

GPU count gets the attention because it's the number on the invoice. The number that actually determines whether a model trains efficiently is what has to fit in memory alongside it. Standard mixed-precision training with the Adam optimizer, the default for most foundation model training, needs roughly 16 to 20 bytes of memory per parameter once everything is accounted for: 2 bytes for FP16 weights, 2 bytes for FP16 gradients, and 12 to 16 bytes for the FP32 optimizer states, momentum and variance, that Adam tracks individually for every parameter in the model. For a 7-billion-parameter model, that's 112 to 140GB before a single activation tensor or batch of training data enters the picture, more than an H100's entire 80GB of HBM3 can hold on its own. For a 70-billion-parameter model, optimizer state alone can exceed a terabyte spread across the cluster.

That number decides whether a model fits on a single card, needs sharding across GPUs with a framework like DeepSpeed ZeRO or fully sharded data parallelism, or can't be trained on a given node at all without offloading states to CPU memory or NVMe, a workaround that solves the capacity problem and creates a throughput one. An H200's 141GB of HBM3e versus an H100's 80GB isn't a marginal spec bump in this context, it's the difference between a 7B model training comfortably in memory and one that needs to be split across two cards just to hold its own optimizer state.

900GB/s vs. 128GB/s: what happens once the model has to be split

Once a model is large enough to require sharding, and most foundation-model-scale training does, the interconnect between GPUs matters as much as the memory on any single card. NVLink on Hopper-generation systems delivers up to 900GB/s of GPU-to-GPU bandwidth, according to NVIDIA's own architecture documentation. PCIe Gen5, the fallback when NVLink isn't available or isn't fully utilized for a given topology, tops out around 128GB/s, roughly a 7x gap. For data-parallel training, that gap determines how fast gradient updates propagate across the cluster after every batch. A system built around PCIe rather than NVLink turns every synchronization step, and a multi-week training run has thousands of them, into a wait the GPUs sit through instead of computing through.

That effect compounds with time, not just cluster size. One slow synchronization step is a rounding error. Thousands of them, repeated across weeks, are the difference between a training run that finishes on the timeline the roadmap assumed and one that runs weeks over budget with no single point of failure to blame, just accumulated wait time nobody priced in at signing.

184 samples to 20: what a 3x increase in model size does to batch size

This isn't theoretical. An independent academic study on LLM pretraining scaling, published on arXiv, ran controlled experiments varying node count and model size independently on identical hardware. Holding model size fixed and scaling node count, training performance scaled roughly linearly up to 128 data-parallel nodes, meaning the workload stayed GPU-bound rather than network-bound, exactly what a well-provisioned cluster should look like. Holding node count fixed and scaling model size told a different story: a 120-million-parameter model ran with a batch size of 184 samples per node. A 350-million-parameter model, same hardware, same node count, managed a batch size of 20, a roughly 9x drop from less than a 3x increase in parameters.

The GPU count didn't change between those two runs. The memory available to hold activations, gradients, and optimizer states per parameter did, and it became the binding constraint well before compute did. Extend that curve toward the 7B-to-70B parameter range most production foundation models actually sit in, and the gap between GPUs provisioned and GPUs actually kept busy gets worse, not better, which is exactly why sizing a cluster on GPU count alone tends to undershoot real-world throughput.

The ratio an AMD research team put a number on

Memory and interconnect constraints don't disappear once a team owns the hardware outright, but they get materially worse on shared infrastructure, where the interconnect a training run depends on is also carrying traffic from other tenants. A separate academic paper on training foundation models at scale, from an AMD-led research team, frames the constraint precisely: aggregate throughput during pretraining is bottlenecked by the ratio of GPU compute per node to interconnect bandwidth per node, not by GPU compute alone. On shared infrastructure, that ratio isn't fully in a team's control, other tenants' traffic on the same fabric competes for the same interconnect bandwidth gradient synchronization depends on. That contention rarely shows up as a clean error. It shows up as throughput that varies run to run for reasons a monitoring dashboard can't fully explain, which is the more frustrating failure mode, not the slowdown itself but the inability to plan around it.

On dedicated bare metal, the interconnect capacity a training run gets is the interconnect capacity provisioned for it, full stop. That's not a marginal difference, it's the difference between a bottleneck a team can measure and size a cluster against at signing, and one that shows up unpredictably mid-run and gets blamed on everything except its actual cause.

16,000 GPUs, and why the topology got more engineering attention than the chip

Meta's own published research on Llama 3.1 describes training the largest variant on a cluster of more than 16,000 H100 GPUs, and the accompanying technical detail is explicit about why network topology and parallelism strategy got as much engineering investment as the model architecture itself: at that scale, the interconnect is the resource that determines whether the rest of the cluster investment pays off. Interconnect topology and memory management at 16,000-GPU scale aren't implementation details buried in an appendix, they're the difference between a run that converges on schedule and one that doesn't converge at all after burning through a compute budget that size.

Few teams train at 16,000-GPU scale, but the mechanics are identical at 8, 64, or 512 GPUs, they just take longer to surface when the cluster is smaller, showing up as throughput sitting 20 to 30% below what the spec sheet implies rather than an outright failure. Throughput is bound by whichever resource runs out first, and for most foundation model training, that resource is memory bandwidth and interconnect capacity, long before it's raw compute.

Sizing by saturation, not by invoice

None of this argues against adding GPUs, it argues against adding GPUs as the default response to every throughput problem. Before scaling node count, the more useful diagnostic question is which resource is actually saturated. If GPU utilization is already high and memory and interconnect bandwidth still have headroom, more compute helps. If utilization drops during synchronization steps or optimizer updates, the fix is a larger-memory GPU, a faster interconnect, or infrastructure where that interconnect isn't shared with workloads outside a team's control, not a bigger invoice for more of the same node.

That's the case for evaluating training infrastructure on memory capacity and interconnect provisioning specifically, not GPU model and count alone. On dedicated bare metal, those numbers are fixed at signing and don't degrade based on what another tenant happens to be running that week, which makes the cluster provisioned the cluster actually available for the full length of a training run.

Talk to an engineer about what dedicated bare metal actually changes for a training run at your scale.

‍

SuscrĂ­bete a nuestro boletĂ­n