How do you make an AI answer on your own documents?

By Yamouni Noureddine · June 1, 2026 · 0 min read

A language model does not know your procedures. The method that works is to give it the right extracts at the moment of the question, rather than trying to teach them to it.

You do not teach a model your documents: you give them to it at the moment of the question. The system splits your files, retrieves the relevant passages, and asks the model to write the answer from those extracts alone. It is faster, cheaper, and the answer cites its source.

Why not simply train the model on your documents?

Because it is slow, expensive, and has to be redone at every update. One amended procedure would call for fresh training.

Above all, a trained model does not know where what it asserts came from. It cannot cite its source, so you cannot verify it — which rules it out of any use that carries consequences.

How does assisted document retrieval work?

Splitting. Your documents are broken into passages of reasonable size. Coarse splitting drowns the information; splitting that is too fine strips it of its context.

Indexing. Each passage is turned into a numerical representation capturing its meaning, then stored in a specialised database.

Retrieval. Given the question, the system finds the closest passages — close in meaning, not in the words used.

Writing. The model receives the question and those passages, with a strict instruction: answer from them alone, and say so when it cannot find the answer.

What makes these projects fail?

Rarely the model. Almost always the documents.

Unrecognised scans, multiple versions of the same file with no dates, tables turned to mush: a retrieval system cannot find what is not readable. The quality of the document archive decides the result, and it is the first thing we look at.

What should you check before starting?

PointQuestion to askIf the answer is no
ReadabilityAre the documents text, or images?Character recognition first
VersionsDo we know which one governs?Sorting comes before indexing
ScopeWho is entitled to read what?Partitioning to be defined
SourcesDoes the answer cite its source passage?System unusable where it matters
LocationDo the documents leave the company?Local deployment required

Should it be installed on your own premises?

As soon as the documents are confidential, yes. Querying an online service about contracts, patient records or internal procedures amounts to transmitting them to a third party, and stepping outside the legal framework that protects you.

DocuNova applies this method on your own servers, with every answer citing the document it came from.

Updated Aug. 17, 2026

Sovereign zone

A project to scope ?

Describe your requirement in a few lines. We come back to you within 48 hours with a costed proposal.

Request a quote Free assessment — 30 min

The assessment is a thirty-minute conversation, with no commitment : we look at your processes and tell you frankly whether software is justified — including when the answer is no.

WhatsApp