8 x 96GB: What Fits on an RTX Pro 6000 Blackwell Node

Ana Pace

October 6, 2026

A 70-billion-parameter model needs 140GB for its weights in FP16. That doesn't fit on a 96GB card. In FP8 it needs 70GB, which does, and in FP4 it needs about 35GB, which leaves more than half the card free. Those three numbers are most of the sizing exercise for LLM inference on an RTX Pro 6000 Blackwell node: 96GB of GDDR7 per GPU is large enough to hold a 70B model, and small enough that what sits next to the weights decides how many requests the card can serve. Here's the math, built on published architecture figures rather than vendor throughput claims.

140GB, 70GB, 35GB: what a 70B model's weights cost at each precision

Weight memory is parameter count times bytes per parameter: 2 bytes in FP16, 1 in FP8, 0.5 in FP4. For a 70B model that comes to 140GB, 70GB and 35GB. NVIDIA lists the RTX PRO 6000 Blackwell family at 96GB of GDDR7 with ECC per GPU, built on fifth-generation Tensor Cores that support FP4, so FP8 and FP4 are both on the table for a 70B model. FP16 isn't, at least not on a single card.

Lower precision isn't free. Dropping from FP16 to FP8 or FP4 trades some accuracy for memory, and how much depends on the model and the quantization method, so it's worth testing on your own evaluation set before committing to a node size. The arithmetic tells you which options exist, not which one to pick.

10.7GB per 32K-token sequence: where the rest of the 96GB goes

Weights are a fixed cost. The KV cache is the variable one: every token in every active sequence stores a key and a value vector at each layer, and that memory stays allocated for as long as the sequence is live. Meta's Llama 3 paper gives the 70B architecture: 80 layers, 64 attention heads across an 8,192 model dimension (128 per head), and 8 key/value heads. At 16-bit precision, that's 2 x 80 x 8 x 128 x 2 bytes, about 0.33MB per token. A 32,768-token context holds roughly 10.7GB. An 8,192-token context holds roughly 2.7GB.

Put those next to the weight budgets:

‍

These are upper bounds. They ignore activation memory, the CUDA context and framework overhead, and they assume a 16-bit KV cache: quantizing the cache stretches them, and allocator inefficiency shrinks them. The pattern holds either way. At FP8, long contexts rather than the model itself limit concurrency on one card. At FP4, the same card holds the model plus a working set of long sequences.

65B on 48GB: what QLoRA means for a 96GB card

Fine-tuning is the other half of the workload. The QLoRA paper showed that a 65B-parameter model can be fine-tuned on a single 48GB GPU while preserving the task performance of 16-bit fine-tuning. It gets there by quantizing the frozen base weights to 4-bit NormalFloat, quantizing the quantization constants as well, and using paged optimizers to absorb memory spikes. A 96GB card has twice the 48GB that result used. The extra room goes to what limits fine-tuning in practice, longer sequences and larger batches, rather than to whether the model loads at all.

Full fine-tuning and split models: where 96GB stops

The limit is full-parameter training. Our post on foundation model training bottlenecks put mixed-precision Adam at 16 to 20 bytes per parameter, which means even a 7B model needs 112 to 140GB for a full fine-tune. That's more than one 96GB card holds, so full fine-tuning at that size means sharding, and adapter-based methods like LoRA and QLoRA are the practical route on this hardware.

The same boundary applies to models whose weights exceed a single card. Splitting a model across GPUs makes the interconnect between them the bottleneck, which is a different sizing exercise from the one in this post, and the one we covered in the training post.

Eight pools of 96GB, not one pool of 768GB

A 1Legion RTX Pro 6000 Blackwell Max-Q node has 8 GPUs and 768GB of GPU memory in total. The useful way to read that number is as eight independent 96GB pools. Per-card workloads spread across them without needing GPU-to-GPU traffic: eight replicas of a 70B FP4 model behind a load balancer, eight separate QLoRA runs on eight datasets, or a split like four cards serving the current model while four fine-tune the next one. None of that depends on how fast the GPUs can talk to each other.

We rent the node whole, on dedicated bare metal, with a rate fixed for the term and a 12-month minimum. Why the card's price moved twice this year is covered in our earlier post. This one is the workload side.

The sizing rule: weights + KV cache + headroom must fit in 96GB

To check a model against one card, pick the precision, multiply parameters by bytes per parameter, then estimate KV cache as layers x KV heads x head dimension x 2 x bytes, multiplied by your target context length and concurrency. Leave room for activations and framework overhead. If the total fits in 96GB, the model is a per-card workload, and an 8-GPU node gives you eight of them to schedule. If it doesn't, the next question is how the model gets split, and that's a conversation about interconnect rather than memory.

Talk to an engineer about sizing an 8-GPU RTX Pro 6000 Blackwell node for your models.

‍

Subscribe to our newsletter