Insights / Private AI
Private AIPrivate LLM hardware requirements: how much GPU do you need?
· 9 min read

On this page
Getting private LLM hardware requirements right is the first step in running AI on your own premises. Buy too little GPU memory and the model will not load, or users will queue. Buy too much and capital sits idle.
This guide explains what uses GPU memory, how precision and context change it, and how to match the result to hardware. A sizing tool at the end of the method section recommends one of our three bundles.
What decides the hardware for a private LLM?
Three things drive the size of a private LLM server: the model, the number of people using it at once, and how much text each conversation holds.
The model sets the fixed cost. Its weights must sit in GPU memory the whole time it is serving. The users and their conversations set the variable cost. Each active session keeps its own working memory, called the KV cache.
The PagedAttention paper (SOSP 2023) shows the split on a real server. For a 13B model on a 40 GB GPU, about 65% of memory held the weights. Close to 30% held the KV cache for active requests.
Memory is not the only limit. Response speed and throughput can call for more GPUs even when the model fits. This guide sizes memory first, because a model that does not fit cannot run at all.
How do you calculate GPU memory for an LLM?
Add two parts: weights and KV cache.
Weights. Multiply the parameter count by the bytes per parameter. At 16-bit, each parameter takes 2 bytes. At 8-bit it takes 1 byte, and at 4-bit half a byte.
We then add 20% for activations, runtime buffers and memory fragmentation. That allowance is our planning assumption, not a fixed rule.
KV cache. For every token in a session, the model stores a key and a value in every layer. The PagedAttention paper gives the formula: 2 × layers × width × bytes per value. With grouped-query attention, the width is the number of key-value heads × head size. Its 13B example needs about 800 KB per token, or up to 1.6 GB for a 2,048-token request.
Interactive estimate
Private LLM hardware sizing: which bundle fits?
Everyone who will have access.
Open-weight dense model, in billions of parameters.
Share of users with a request in flight at the busiest moment.
Longer documents in RAG need longer context.
About 28.8 GB of GPU memory. Recommended: Department bundle.
Model weights
19.2 GB
KV cache (5 active)
9.6 GB
Total GPU memory
28.8 GB
Teal: weights. Amber: KV cache.
Recommended
Department bundle
Shared AI for a whole department, with several use cases on one platform.
- Typical users: 25–200
- GPU class: 4–8 data-centre GPUs
- Typical model size: Up to ~70B
Set by: team size, model size.
Assumptions
- Weights = parameters × bits ÷ 8 bytes, plus 20% for activations, runtime buffers and memory fragmentation (our allowance).
- KV cache per active session = 2 (keys and values) × layers × KV heads × head size × 2 bytes per token, at 16-bit (PagedAttention formula, SOSP 2023, generalised to GQA KV heads). Every active session is assumed to fill its full context window, which is conservative.
- KV sizes use published grouped-query-attention architectures with 8 KV heads of size 128: 32 layers at 8B, 80 at 70B and 126 at 405B (arXiv:2407.21783). Other sizes are interpolated. Dense models are assumed.
- Active sessions = users × concurrency, rounded up, minimum 1.
- Tier rules: users and model size follow our bundle ranges (Team up to 25 users and 14B; Department up to 200 users and ~70B). Memory limits assume up to 96 GB for Team (two 48 GB professional GPUs) and up to 640 GB for Department (eight 80 GB data-centre GPUs). The highest tier any rule needs wins.
- Throughput and latency targets are not modelled. They can raise the GPU count even when memory fits.
Indicative estimate only. Contact us for an engineered proposal.
The tool’s default case is 50 users on a 32B model at 4-bit, with 10% active and 8K context. That needs 19.2 GB for weights and 9.6 GB for the KV cache, or 28.8 GB in total. The memory fits a Team bundle, but 50 users and a 32B model point to the Department bundle.
How does precision change private LLM hardware requirements?
Precision is the number of bits used to store each weight. Fewer bits mean less memory. Here is a 70B-parameter model, including our 20% allowance:
| Precision | Bytes per parameter | Weights for 70B |
|---|---|---|
| 16-bit | 2 | 168 GB |
| 8-bit | 1 | 84 GB |
| 4-bit | 0.5 | 42 GB |
Quantization is how you get to 8 or 4 bits after training. The GPTQ paper (ICLR 2023) quantized a 175B-parameter model to 3 or 4 bits per weight. It reported negligible loss of accuracy compared with the full-precision model.
Quantization can also speed up answers. The GPTQ authors measured end-to-end generation about 3.25 times faster than 16-bit on high-end data-centre GPUs, and about 4.5 times faster on more cost-effective GPUs.
That result is encouraging but not universal. Quality can drop for some tasks, especially at the lowest bit widths. Test your own documents and questions at the precision you plan to run.
Why does the KV cache matter so much?
The weights are a fixed cost. The KV cache grows with every active user and every token they send. For busy systems it can rival the weights.
Modern open-weight models shrink it with grouped-query attention (GQA). The GQA paper (EMNLP 2023) shares each key and value head across a group of query heads. That cuts the cache while keeping quality close to standard attention.
Our tool uses published architectures that do this. Each has 8 key-value heads of size 128, with 32 layers at 8B, 80 at 70B and 126 at 405B. At 16-bit, that works out to about 0.13 GB per 1K (1,024) tokens of context at 8B, and 0.34 GB at 70B.
Take 150 users on a 70B model at 8-bit, with 10% active and 8K context. The weights need 84 GB and the KV cache about 40 GB, for 124.3 GB in total. Raise the context to 32K and the precision to 16-bit, and the total climbs to 329.1 GB.
The PagedAttention authors make the same point with their 13B example. A single GPU has tens of gigabytes of memory. Even if all of it held KV cache, only a few tens of requests would fit at once.
Two settings drive the cache: concurrency and context. Concurrency is the share of users with a request in flight at the busiest moment. Context is how much text each session holds, including documents pulled in by retrieval. Long documents need long context, and that multiplies the cache.
Serving software matters as well. The PagedAttention paper found that older systems used only 20.4% to 38.2% of their KV cache memory for actual token data. The rest was lost to reserved slots and fragmentation. The paper’s paged approach cuts this waste to near zero.
How many users are active at once?
Named users and active users are not the same. A department of 150 people rarely has 150 requests in flight at one moment. People read answers, edit drafts and attend meetings between questions.
The tool asks for peak concurrency: the share of users with a request running at the busiest moment. It then multiplies your user count by that share and rounds up. At 10%, 150 users means 15 active sessions.
There is no universal figure, so treat it as a planning input. Pilots give the best data. Log requests for a few weeks and look at the busiest hour. Until then, test a range of values to see how sensitive your sizing is.
Batch jobs change the picture. Summarising a large archive overnight can keep many sessions busy for hours. If you plan jobs like that, size for them separately or schedule them outside working hours.
What else does a private LLM server need?
GPU memory is the main constraint, but it is not the whole system. Each of our bundles also includes:
- Local NVMe storage for model files, the vector index and cached documents.
- Retrieval (RAG) integration so answers can cite your own files, with connectors to file shares and document systems.
- Access control with single sign-on, role-based permissions and audit logging.
- Monitoring of hardware health, utilisation and model serving.
Larger bundles add more. The Department bundle includes a rack power and cooling review for sustained full load. The Enterprise bundle adds liquid cooling, UPS and metered power, and a 400G/800G fabric between nodes.
Which hardware bundle fits your team?
Our three bundles map to clear ranges. The tool picks the highest tier that any one of team size, model size or memory calls for.
| Bundle | Typical users | Typical model size | GPUs |
|---|---|---|---|
| Team bundle | 5–25 | 7B–14B, quantized | 1–2 professional GPUs |
| Department bundle | 25–200 | Up to ~70B | 4–8 data-centre GPUs |
| Enterprise bundle | 200+ | Largest open-weight models | Multi-node cluster |
For memory, the tool assumes the Team tier offers up to 96 GB (two 48 GB professional GPUs) and the Department tier up to 640 GB (eight 80 GB data-centre GPUs). These are planning figures. The final hardware is set during the site assessment.
Here are three typical cases from the tool:
- A legal team of 20 on an 8B model at 4-bit, 20% active, 8K context: 9.1 GB. That is a Team bundle.
- A hospital department of 150 on a 70B model at 8-bit, 10% active, 8K context: 124.3 GB. That is a Department bundle.
- A bank with 1,000 users on the same model: 352.4 GB. Memory would fit an eight-GPU Department server with 80 GB per GPU (640 GB), but not a four-GPU one (320 GB). Either way, 1,000 users call for the Enterprise cluster.
Enterprise clusters span several servers. That brings in a high-speed fabric, covered in our guide to 800G Ethernet vs InfiniBand.
Why keep the model on your own hardware?
For many organisations, the reason is data. The UAE’s Federal Decree-Law No. 45 of 2021 sets out a framework for personal data. It includes requirements for cross-border transfer and sharing of personal data.
A private LLM keeps prompts, documents and answers inside your own facility. Our guide to private AI for sensitive data in the UAE covers the rules in more detail. For the cost side, see on-prem vs cloud AI TCO.
The rest of the stack matters too. Plan storage for models and document indexes, and check power and cooling for sustained load. Our partners confirm all of this during the private AI site assessment.
Key takeaways
- GPU memory = weights + KV cache. Weights = parameters × bits ÷ 8, plus an allowance for overhead.
- Quantizing to 8-bit halves the weights and 4-bit quarters them. Test quality on your own tasks.
- The KV cache grows with active users and context length. Long retrieval contexts multiply it.
- Team suits up to 25 users and 14B models, Department up to 200 users and ~70B, and Enterprise anything larger.
- Keeping the model on-premises keeps sensitive prompts and documents in-house.
Frequently asked questions
How much GPU memory does a private LLM need?
What is the KV cache?
Does quantization reduce answer quality?
Which bundle suits a 100-person department?
Why run an LLM on-premises instead of in the cloud?
Sources
- ACM SOSP (Kwon et al.), arXiv –Efficient Memory Management for Large Language Model Serving with PagedAttention (SOSP 2023)
- EMNLP (Ainslie et al.), arXiv –GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints (EMNLP 2023)
- ICLR (Frantar et al.), arXiv –GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers (ICLR 2023)
- arXiv –arXiv:2407.21783, open-weight model family technical report (Table 3: model architectures at 8B, 70B and 405B)
- UAE Government portal (u.ae) –Data protection laws (Federal Decree-Law No. 45 of 2021)


