Running a language model locally: what to measure before buying the machine
Before choosing between a 32-billion and a 72-billion parameter model, you need to know what to measure. Here is the protocol we apply and the orders of magnitude to check on your own hardware.
The choice of a local model is not decided by a public leaderboard, but by three measurements taken on your own site: the memory actually consumed, the delay before the first word appears, and the quality of the answers on your own documents. A larger model that overflows memory is slower than a smaller one that fits.
Why are public leaderboards not enough?
They measure general capabilities on standardised test sets. Your requirement is specific: answering management questions in your own language, on your own documents, with a response time acceptable to a user waiting in front of a screen.
A model that tops a leaderboard may well be unusable on your site because it does not fit in memory.
What is the constraint that decides everything?
Memory. A model must be loaded in full to run at normal speed. If it does not fit, the system compensates using the disk and performance collapses — not by a few percent, but by an order of magnitude.
Compressing the model, known as quantisation, reduces the space it occupies by lowering the precision of the weights. Four-bit compression divides the footprint by roughly four compared with the original format, at the cost of slight degradation on most business uses.
An order of magnitude to check for yourself: allow around half a gigabyte of memory per billion parameters at four-bit compression, plus the memory needed for the context. A 32-billion parameter model therefore fits on a well-equipped machine; a 72-billion parameter model calls for a distinctly more generous configuration.
What exactly should you measure?
| Measurement | How | Why |
|---|---|---|
| Memory occupied | Read during a long generation | Determines whether the model fits |
| Delay before the first word | Timed from launch to display | This is what the user feels |
| Generation speed | Words produced per second | Comfort of continuous reading |
| Business quality | A set of 30 real questions, marked blind | The only criterion that really counts |
| Behaviour with several users | Three simultaneous requests | Reveals collapse under real use |
How do you build the test set?
Take thirty questions your teams genuinely ask, with their expected answers. Put them to each candidate model, and have the answers marked by somebody who does not know which model produced which.
It is tedious, and it is the only method that will stop you buying a machine for a model your users will not accept.
Should you always take the largest model possible?
No. On structured management questions, the quality gap between a mid-sized model and a very large one is often smaller than the comfort gap caused by response time.
Our rule: the largest model that fits comfortably in memory with room to spare for the context, never the one that saturates it.
Updated Aug. 11, 2026