Sedang Bahagia
Start a projectStart a project

Document Q&A That Cites Its Sources: What It Takes to Build

Asking questions of your company documents is easy to demo and hard to trust. Five things that separate the two.

By ·· 5 min read

Key takeaways

  • A trustworthy document assistant answers only from your documents, cites the page, and says plainly when the answer is not there.
  • Most of the work is in preparing documents: reading scans, keeping page numbers, splitting by clause or section, and knowing which version is current.
  • Search should combine keyword and meaning-based matching, because people ask for clause numbers, product codes and names as well as concepts.
  • Answers must follow the same access rules as the shared drive, so nobody sees a document through the assistant they could not open themselves.
  • Test against a written set of real questions with expected answers, and re-run it on every change.

A useful document assistant does five things: it answers only from your own documents, cites the clause and page for every answer, says plainly when the answer is not there, shows each person only what they are already allowed to open, and is tested against a written set of real questions before anyone relies on it. The language model is the easy part. Preparing the documents and testing the answers is most of the work.

Asking questions of a contract or an SOP is one of the easiest AI features to demo. Drop a PDF into a chat tool, ask a question, get a fluent answer. The trouble starts when there are 400 contracts, three versions of each, scanned addenda in two languages, and a site manager who needs to know whether the answer is right before acting on it.

How document Q&A works

The common pattern is called retrieval-augmented generation, or RAG. It has two halves.

  1. Search. When a question arrives, find the few passages across all your documents most likely to contain the answer.
  2. Answer. Give only those passages to a language model, with instructions to answer from them, quote them, and say so if they do not contain the answer.

The model is not trained on your documents. That keeps them out of anyone else’s model, and means a new document is answerable as soon as it is indexed.

1. Prepare documents properly

Answer quality is limited by how well the documents were read and split. This step is where most of the effort goes.

  • Read scans with OCR and keep the layout. Tables, headings and clause numbers carry meaning. A penalty rate in a table is useless if the table is flattened into a line of numbers.
  • Keep page numbers and positions. Citations need them, and users trust an answer they can check in one click.
  • Split by structure, not by length. Split contracts by clause and SOPs by section, and keep the heading path with each piece, such as “Clause 14 › Retention › 14.2”. Fixed-length chunks cut clauses in half.
  • Record metadata. Project, counterparty, document type, effective date, language and which document it amends.

2. Handle versions and amendments

Contracts change through addenda, and SOPs through revisions. An assistant that quotes the original clause when an addendum replaced it is worse than no assistant.

  • Link every addendum to the contract it amends
  • Mark superseded SOP versions, and search only current ones by default
  • When an answer depends on a clause that has been amended, show both and say which applies

3. Search with keywords and meaning together

Meaning-based (vector) search finds passages about “late completion” when the contract says “delay in handover”. Keyword search finds “Clause 14.2”, “SOP-QA-017” or a supplier’s name exactly. Real questions need both, so we combine them and then re-rank the results before the model sees them.

Question What it tests
“What is the retention percentage for Tower B?” Finding a number in a table, scoped to one project
“Apa denda keterlambatan serah terima?” A Bahasa question about an English contract
“Has clause 9 been changed?” Following addenda
“Who approves a change order above the limit?” Finding an answer spread over two SOP sections

4. Respect access rights

If finance contracts are restricted to the finance team on the shared drive, they must be restricted in the assistant too. Filter search results by the user’s permissions before anything reaches the model, not after. Sync permissions from the source, such as SharePoint or Google Drive, rather than maintaining a second list by hand.

5. Make “I can’t find that” a good answer

The most important behaviour to test is refusal. When the documents do not contain the answer, the assistant should say so, show the closest passages it found, and offer to log the question. Logged questions show which documents are missing or unclear, and they become new test cases.

An assistant that sometimes says “not found” is trusted. One that always answers is not.

Extraction alongside Q&A

Some information should not wait for someone to ask. Milestone dates, penalty rates, renewal deadlines and named parties are better pulled out of every contract into a register as soon as it is signed. This is document extraction, the same technique used for invoices and claims, applied to clauses.

A sensible split:

  • Register (extraction): fields you need to track, report on or be reminded about
  • Assistant (Q&A): everything else, asked when it comes up

Have the document owner confirm each new register entry before it is used for reminders.

How to test it

Before launch, write down 50 to 100 real questions from the people who will use the assistant, with the correct answer and the source page for each, checked by someone who knows the documents. Include questions that should be refused because the answer is not in the documents.

For each run, mark every answer:

  • Correct, with the right citation
  • Correct, but cited the wrong place
  • Wrong
  • Should have refused but answered

Re-run the set after every change to documents, search or prompts. After launch, add every logged “not found” and every answer a user flags as wrong.

What to prepare before a project

  • A list of document sources and an owner for each
  • Which teams may see which documents
  • 20 real questions to start the test set, ideally collected from staff over a week
  • One person who can confirm whether an answer is right, as in our data readiness checklist

For an example built on contracts and SOPs across two languages, see our contract and SOP assistant with clause extraction. The same approach powers question answering over equipment manuals in our EAM and policy questions in our HR system.

If your team spends time hunting through documents for answers, tell us which documents and which questions. We will tell you what a first version could answer.

Frequently asked questions

Can we not just upload our documents to a general AI chat tool?

For a handful of documents and one person, often yes. It breaks down with hundreds of documents, several versions of each, different access rights per team, and a need to show where every answer came from. Those are the parts a proper build handles.

What is retrieval-augmented generation (RAG)?

It is the usual pattern behind document Q&A. When a question comes in, the system first searches your documents for the most relevant passages, then asks a language model to answer using only those passages and to cite them. The model does not need to be trained on your documents.

How do you stop the assistant from making things up?

Three ways together: it may answer only from the passages it retrieved, every answer must cite a source the user can open, and it is instructed and tested to say it cannot find the answer when the passages do not contain one. A test set of real questions checks all three.

Can staff ask in Bahasa about documents written in English?

Yes. Current models search and answer across English, Bahasa Indonesia and Bahasa Melayu, and the answer can quote the original clause while explaining it in the language the question was asked in.

Does the assistant have to run in the cloud?

No. The document index and search can run on your own servers, and the language model can be a cloud API with a data processing agreement or a model hosted inside your environment. The choice depends on how sensitive the documents are.

Have a product idea? Talk it through with our technical lead. We reply within 24 hours.

Book a call