OCR vs AI Document Extraction: Invoices, Receipts and Claims
OCR gives you text. Your finance team needs an invoice number, a total and a supplier. What sits between the two, and where projects go wrong.
Key takeaways
- OCR turns a scan into text. Extraction turns that text into named fields, such as invoice number, date, line items and total, that a system can post.
- A reliable pipeline has six steps: intake, split and classify, read, extract, check, and route to a person when needed.
- Business-rule checks, such as line items adding up to the total or an invoice matching its PO, catch more errors than model confidence alone.
- Measure accuracy per field on a labelled set of your own past documents, and start with one document type.
- Where a structured e-invoice exists, read that instead of the PDF.
OCR (optical character recognition) turns a scanned page into text. Document extraction turns that text into the named fields your system needs, such as supplier, invoice number, date, line items and total, and checks them before anything is posted. In practice you need both, plus business-rule checks and a review screen for the documents that fail them.
Most teams who ask for “OCR” actually want extraction. A finance clerk does not need the text of an invoice. They need the five or six numbers that go into the accounting system, and they need to trust them.
OCR, templates and AI extraction compared
| Approach | What it gives you | Where it breaks |
|---|---|---|
| Plain OCR | All the text on the page, in reading order | Someone still has to find and type the fields |
| Template extraction | Fields read from fixed positions on a known layout | Every new supplier or layout change needs a new template |
| AI extraction | Named fields and line items from any layout, with a confidence score | Can be confidently wrong, so it needs checks around it |
AI extraction is the right default for most SMEs today because supplier and hospital layouts change constantly. But it is one step in a pipeline, not the whole pipeline.
The six steps of a document pipeline
- Intake. Collect documents from where they actually arrive: an email inbox, a WhatsApp number, an upload page or a scanner folder. Record who sent each one and when.
- Split and classify. One PDF often holds an invoice, a delivery order and a credit note. Split it into documents and label each type before reading anything.
- Read. Run OCR with layout, so tables, stamps and handwriting keep their position on the page. Keep the page number and position of every word.
- Extract. Map the content into a fixed schema per document type. Every field carries a confidence score and a pointer to where it was found.
- Check. Apply business rules. This step does more for accuracy than any model choice.
- Route. Documents that pass every check go straight through. The rest go to a review screen with the scan beside the extracted fields, and corrections are logged.
Checks beat confidence scores
Confidence scores tell you how sure the model is, not whether it is right. A crisp, clean invoice with a mistyped total will often score high. Rules based on how your business works catch what the model cannot.
Checks worth adding
- Line items add up to the subtotal, and subtotal plus tax equals the total
- Tax is calculated at the rate that applies to that supplier and date
- The supplier exists in your master data, and the bank account matches the one on file
- The invoice matches its purchase order and goods receipt on quantity and price (a three-way match)
- The invoice number has not been seen before from that supplier
- Dates are plausible: not in the future, not before the PO
- For claims: the member, policy and treatment dates are valid, and the claimed amount is within the benefit
A check that fails is not an error. It is the system doing its job and sending the document to a person.
What makes documents in this region harder
Most extraction demos use clean English invoices. Real documents in Singapore, Malaysia and Indonesia look different.
- Number formats. Indonesian documents write Rp 1.250.000,00, where the full stop separates thousands and the comma marks decimals. Get this wrong and an amount is off by a factor of a thousand. Set the format per supplier or per country, and check it against the total.
- Dates. 03/04/2026 is 3 April here, not 4 March. Fix the date format per source rather than letting the model guess.
- Mixed languages. One hospital bill can have English headings, Bahasa Melayu item descriptions and a Chinese clinic stamp. Current models read this well, but your schema should store item descriptions as written.
- Photos, not scans. Claims and receipts often arrive as phone photos at an angle, with shadows, folds and faded thermal paper. Straighten and clean the image before reading, and expect these to be the main source of review work.
- Handwriting and stamps. Handwritten amounts and stamps across printed text are common on delivery orders and medical receipts. Treat these fields as review-required until accuracy on your own documents says otherwise.
Use the structured version where it exists
E-invoicing is spreading across the region. Malaysia’s MyInvois system, Singapore’s InvoiceNow network and Indonesia’s e-Faktur all produce invoices as structured data. Where a supplier sends one, read the structured data directly. Extraction from PDFs is for everything else, which for most SMEs is still the majority.
How to measure accuracy
Do not accept a single accuracy number from a vendor, including us. Measure it on your own documents.
- Take a few hundred past documents of one type, including the bad photos.
- Have someone who processes them today record the correct value for each field.
- Run the pipeline and compare field by field.
- Report accuracy per field, not per document. Supplier name might be near perfect while line-item quantities lag.
- Track the straight-through rate: the share of documents that pass every check and need no human touch.
Re-run this set every time the pipeline changes. It is the only reliable way to know whether a change helped.
Personal data in documents
Claims, payslips and identity documents hold personal data, often including identity numbers and health information. Decide before the build which fields the pipeline needs, mask the rest, and agree where documents are stored and for how long. Our data readiness checklist covers the data protection laws in each market.
Where to start
Pick one document type with real volume and a clear owner. Supplier invoices and medical claims are common first choices because the fields are well defined and the cost of manual entry is easy to see.
- For finance teams: supplier invoice capture with a three-way match is part of our ERP.
- For insurers and claims teams: see how we approached medical claims intake with OCR and extraction.
- For contracts and policies, where the goal is answering questions rather than filling fields, read document Q&A that cites its sources.
If you have a pile of documents someone types up every day, send us a sample with the personal data masked. We will tell you which fields can be read reliably and which will still need a person.
Frequently asked questions
What is the difference between OCR and document extraction?
OCR (optical character recognition) converts an image of a page into text. Document extraction takes that text and its layout and returns specific fields, such as the supplier name, invoice date, line items and total, in a structure your accounting or claims system can accept.
Do we still need templates for each supplier?
Usually not. Older tools needed a template per layout, which broke whenever a supplier changed its invoice. Current layout-aware and language models read new layouts without one, although a few high-volume layouts can still benefit from tuning.
Can extraction run without anyone checking it?
For some fields on some document types, once accuracy has been measured on your own documents. Most teams start with every document reviewed, then let through only the ones that pass every check with high confidence, and keep sampling those.
How many past documents do we need to start?
A few hundred real documents of the first type, covering the common layouts and the awkward cases, such as phone photos and handwritten amounts. They are used to measure accuracy, not to train a model from scratch.
Can it read documents in Bahasa and Chinese?
Yes. Current OCR and language models handle English, Bahasa Indonesia, Bahasa Melayu and Chinese on the same page. Local number and date formats need explicit rules, which we cover below.