How do you make an AI answer on your own documents?
A language model does not know your procedures. The method that works is to give it the right extracts at the moment of the question, rather than trying to teach them to it.
You do not teach a model your documents: you give them to it at the moment of the question. The system splits your files, retrieves the relevant passages, and asks the model to write the answer from those extracts alone. It is faster, cheaper, and the answer cites its source.
Why not simply train the model on your documents?
Because it is slow, expensive, and has to be redone at every update. One amended procedure would call for fresh training.
Above all, a trained model does not know where what it asserts came from. It cannot cite its source, so you cannot verify it — which rules it out of any use that carries consequences.
How does assisted document retrieval work?
Splitting. Your documents are broken into passages of reasonable size. Coarse splitting drowns the information; splitting that is too fine strips it of its context.
Indexing. Each passage is turned into a numerical representation capturing its meaning, then stored in a specialised database.
Retrieval. Given the question, the system finds the closest passages — close in meaning, not in the words used.
Writing. The model receives the question and those passages, with a strict instruction: answer from them alone, and say so when it cannot find the answer.
What makes these projects fail?
Rarely the model. Almost always the documents.
Unrecognised scans, multiple versions of the same file with no dates, tables turned to mush: a retrieval system cannot find what is not readable. The quality of the document archive decides the result, and it is the first thing we look at.
What should you check before starting?
| Point | Question to ask | If the answer is no |
|---|---|---|
| Readability | Are the documents text, or images? | Character recognition first |
| Versions | Do we know which one governs? | Sorting comes before indexing |
| Scope | Who is entitled to read what? | Partitioning to be defined |
| Sources | Does the answer cite its source passage? | System unusable where it matters |
| Location | Do the documents leave the company? | Local deployment required |
Should it be installed on your own premises?
As soon as the documents are confidential, yes. Querying an online service about contracts, patient records or internal procedures amounts to transmitting them to a third party, and stepping outside the legal framework that protects you.
DocuNova applies this method on your own servers, with every answer citing the document it came from.
Updated Aug. 17, 2026