Teaching a machine to read handwritten documents
Readings filled in by hand on the shop floor end up retyped into a spreadsheet. A vision model can read them — provided it can also say when it did not understand.
A convolutional neural network now reads handwritten shop-floor records reliably, on two conditions: restrict the expected vocabulary to what the form allows, and send anything the model is unsure about to a human operator. The difficulty is not the reading. It is catching what was misread without anyone noticing.
Why conventional OCR fails on shop-floor handwriting
Character-recognition engines were built for printed text: regular, on a clean background. A control report filled in during production brings together everything that defeats them — handwriting that differs from one operator to the next, digits crossed out and rewritten, grease marks, a form photocopied through several generations, boxes written over.
The result is not a failure to read, which would be easier to handle. It is a confident misreading: a 1 read as a 7, a 0 as a 6. The error enters the system without warning and surfaces weeks later as a traceability discrepancy.
What does a neural network add?
A trained vision model does not look up letter shapes in a catalogue. It learns, on your own documents, what a digit written by your teams looks like, on your forms, in your lighting. That specialisation is what makes the difference: a generic model knows every handwriting in the world and none of yours.
Training runs on documents already keyed in. This point is often missed: the years of archives you retyped by hand are exactly the material the model needs. That past work is not wasted — it becomes the reference.
How do you know the machine got it wrong?
This is the only question that matters in production, and the one demonstrations skip. A usable model does not return a value: it returns a value and its confidence.
You then set a threshold. Above it, the value goes straight into the system. Below it, the document is queued for human checking, with the doubtful area boxed on screen. The operator no longer retypes everything: they arbitrate the few cases the machine flags.
That setting is a business decision, not a technical one. A low threshold sends more documents through untouched and lets more errors past. A high threshold protects the data and lengthens the checking queue. It is set by looking at what an error actually costs in your process.
What to compare before deciding
| Criterion | Manual keying | Generic OCR | Model trained on your documents |
|---|---|---|---|
| Degraded handwriting | Reads correctly | Often fails | Reads correctly |
| Errors flagged | Rarely | Never | Yes, by confidence threshold |
| Time per document | High and constant | Instant | Instant, doubtful cases aside |
| Set-up | None | Immediate | A few weeks of training |
| Data required | None | None | Your already-keyed archives |
| Works offline | Yes | Depends on the tool | Yes, on your hardware |
Where to start
- One form type. Not "all the plant's documents": the one that costs the most hours of retyping.
- A batch of already-keyed archives. They train the model and measure what it is worth, by comparing its reading with yours.
- A confidence threshold and a checking queue. Decided with the teams who fill in the forms, not on their behalf.
- A machine on site. These models run locally: production documents do not leave the plant.
We delivered this kind of system for a forging manufacturer based in France, on control reports filled in by hand. The assignment in detail, and how a model runs on your own servers.
Updated Sept. 9, 2026