What machine do you need to run an AI locally?
Video memory decides almost everything. Here is how to size a machine to run a language model on your premises, and how to check for yourself before buying.
It is the amount of video memory that decides, well ahead of processor power. A model must fit entirely in memory to answer quickly; if it overflows, speed collapses. The useful rule: estimate the size of the quantised model, add a margin for the context, and check on your own hardware before buying.
Why does video memory matter more than the rest?
A language model is a large table of numbers that must be traversed in full for every word produced. If it fits in the graphics card's memory, that traversal is fast. If it overflows into system memory, every word becomes a wait.
That is why a machine with a powerful processor but a small graphics card will answer more slowly than a modest, well-sized machine.
How do you estimate a model's size?
A model is described by its number of parameters — 7 billion, 14 billion, 70 billion. At native precision, each parameter occupies roughly two bytes. Quantisation reduces that precision and brings the footprint down to around half a byte per parameter for the most compact formats.
Then add room for the context: the more documents you send in a question, the more you need. Allow generously; the margin costs less than a deployment to redo.
What sizing for which use?
| Use | Model size | What constrains it |
|---|---|---|
| Management assistant, short questions | Small quantised model | Fits on a consumer card |
| Document retrieval over an internal archive | Mid-sized model | The context, more than the model |
| Long-form writing and analysis | Large model | Video memory |
| Several simultaneous users | Depends on the use | The number of parallel requests |
We do not publish a table of hardware references: it would age within months and mislead you. The method, on the other hand, stays valid.
How do you check before buying?
Install a local execution engine on a machine you already have, load a quantised model, and put ten questions to it that are representative of your real use — not demonstration questions.
Measure two things: the time before the first word, and the production speed thereafter. If the first exceeds a few seconds on a simple question, the model is too large for the machine.
That measurement takes an afternoon and is worth any number of comparison tables.
Do you really need a graphics card?
For a small model and a few users, a recent processor with plenty of RAM is often enough — the answer is slower but usable. As soon as several people query the system at the same time, the graphics card becomes necessary.
We assess the machine before any commitment, and we do sometimes say that a use case does not justify the investment. See how we deploy an AI on your own servers.
Updated Aug. 17, 2026