Running a language model locally: what to measure before buying the machine

By Yamouni Noureddine · July 28, 2026 · 0 min read

Before choosing between a 32-billion and a 72-billion parameter model, you need to know what to measure. Here is the protocol we apply and the orders of magnitude to check on your own hardware.

The choice of a local model is not decided by a public leaderboard, but by three measurements taken on your own site: the memory actually consumed, the delay before the first word appears, and the quality of the answers on your own documents. A larger model that overflows memory is slower than a smaller one that fits.

Why are public leaderboards not enough?

They measure general capabilities on standardised test sets. Your requirement is specific: answering management questions in your own language, on your own documents, with a response time acceptable to a user waiting in front of a screen.

A model that tops a leaderboard may well be unusable on your site because it does not fit in memory.

What is the constraint that decides everything?

Memory. A model must be loaded in full to run at normal speed. If it does not fit, the system compensates using the disk and performance collapses — not by a few percent, but by an order of magnitude.

Compressing the model, known as quantisation, reduces the space it occupies by lowering the precision of the weights. Four-bit compression divides the footprint by roughly four compared with the original format, at the cost of slight degradation on most business uses.

An order of magnitude to check for yourself: allow around half a gigabyte of memory per billion parameters at four-bit compression, plus the memory needed for the context. A 32-billion parameter model therefore fits on a well-equipped machine; a 72-billion parameter model calls for a distinctly more generous configuration.

What exactly should you measure?

Measurement protocol on your own hardware
MeasurementHowWhy
Memory occupiedRead during a long generationDetermines whether the model fits
Delay before the first wordTimed from launch to displayThis is what the user feels
Generation speedWords produced per secondComfort of continuous reading
Business qualityA set of 30 real questions, marked blindThe only criterion that really counts
Behaviour with several usersThree simultaneous requestsReveals collapse under real use

How do you build the test set?

Take thirty questions your teams genuinely ask, with their expected answers. Put them to each candidate model, and have the answers marked by somebody who does not know which model produced which.

It is tedious, and it is the only method that will stop you buying a machine for a model your users will not accept.

Should you always take the largest model possible?

No. On structured management questions, the quality gap between a mid-sized model and a very large one is often smaller than the comfort gap caused by response time.

Our rule: the largest model that fits comfortably in memory with room to spare for the context, never the one that saturates it.

Updated Aug. 11, 2026

Sovereign zone

A project to scope ?

Describe your requirement in a few lines. We come back to you within 48 hours with a costed proposal.

Request a quote Free assessment — 30 min

The assessment is a thirty-minute conversation, with no commitment : we look at your processes and tell you frankly whether software is justified — including when the answer is no.

WhatsApp