In beta with first customers. Out of beta and taking new customers from 26 October. Talk to us
All insights

Quality

Why evaluations matter in AI document extraction

If a value is read wrong, every rule after it is wrong too. How to measure extraction instead of assuming it works.

Oct 2026 · 7 min read

PDFAI readingrun 5 of 5EXTRACTED VS VERIFIEDFIELDREADRESULTInvoice numberINV-20417ref INV-20417CorrectInvoice date04-10-2026ref 2026-10-04Format onlyUnit price, line 312.95ref 12.95CorrectVAT number—ref NL8231…B01MissingPO numberPO-88213ref —Invented5 RUNSConsistent, not right onceCOST VS RESULTModel AModel BTemplatecorrectcost

AI document extraction demos well. You drop in an invoice, the fields appear, the line items look right. Then the system goes live, and a few weeks later someone finds a unit price with a digit missing that went straight into the ERP. Nobody can say how often that happens, because nobody measured it.

Extraction is the foundation of every document workflow. Every rule, every match between documents and every routing decision starts from a value read off a page. If that value is wrong, everything built on it is wrong too, however good the rules are. That is why extraction needs evaluations: a repeatable measurement against answers you know are correct.

Why AI extraction has to be measured

Template-based extraction is predictable: a known layout gives the same answer every time, and when it fails it usually fails loudly. AI reading is more flexible, and that flexibility comes with different failure modes.

  • Quality varies with layout, scan quality, language and document length, often per supplier.
  • A model can invent a plausible value for a field that is not on the page.
  • The same document can give a slightly different result on a second run.
  • A change of prompt, model or provider can change results without anyone noticing.

"It looked right on the five documents I tried" is not evidence. Neither is a general accuracy figure from a vendor, measured on someone else's documents.

Start with verified answers

An evaluation needs a reference: for a set of sample documents, the correct result, confirmed by people who know the process. The set should look like your real inbox, not like a demo. Include good PDFs and bad scans, each important supplier or layout, multi-page documents and the edge cases your team knows by heart.

This does not have to be a big project. A few dozen well-chosen documents already tell you far more than impressions do, and the set grows as you go. Every correction a reviewer makes in production is a candidate for it.

Compare field by field, with fixed rules

Comparing extraction with the reference should itself be deterministic. Using another model to judge whether the answer "looks right" just moves the uncertainty. Compare every field and every line item, and tell apart four different results:

  • Correct: the value matches.
  • Formatting only: 04-10-2026 against 2026-10-04, or 1.250,00 against 1250. Not an error once normalised.
  • Missing: the value is on the page but was not read. Usually caught by a rule and sent to review.
  • Invented: a value that is not on the page, or a wrong value that looks right. The most dangerous result, because it can pass every check.

Line items deserve their own attention. A table with one row skipped or two rows merged can have every individual value right and still be wrong as a whole.

Run it more than once

One correct run proves little if the next one differs. Repeating extraction several times per document shows whether results hold up. A field that is right in four runs out of five is not reliable, and it is better to know that before production than after.

Weigh correctness against cost

The most expensive model is not always needed. Clean, known layouts may do well with a smaller model, or better still with a template. Difficult scans may need the stronger one. An evaluation that reports tokens, speed and estimated cost beside correctness lets you compare prompts, models and settings side by side, choose the best balance per document type and show that a change is an improvement.

Keep measuring

Evaluations are not a one-off acceptance test. Rerun them whenever a prompt, model, provider or setting changes, and when a supplier changes its layout. Model providers update their models; an evaluation set turns that from a silent risk into a measured change. It works the same way as regression tests for the workflow itself.

Evaluations and rules work together

Measuring extraction does not replace checking it. Business rules still catch what reading gets wrong at runtime: totals that do not add up, an unknown supplier code, quantities that differ from the packing list. Evaluations tell you how often you rely on that safety net, where the weak spots are and whether a change made them better or worse.

How we do it in StencilFlow

StencilFlow includes an evaluations suite. Your team confirms the correct result for sample documents. The StencilFlow agent runs extraction against them, repeatedly, compares every field and line item by fixed rules and reports the differences, together with tokens, speed and estimated cost. When the agent proposes a better prompt or a different model, the evaluation shows whether it really is better before anyone activates it.

Questions to ask any extraction vendor

  • Can I see results per field, on my own documents?
  • How do you tell an invented value from a missing one?
  • How consistent are results across repeated runs?
  • What happens to my results when the underlying model changes?
  • What does a better result cost per document?
Trust in extraction is not a feeling. It is a number per field, measured on your own documents, and measured again after every change.

Ready to look at your inbox?

Bring one process and a handful of real documents. We show you the path from inbox to decision, on your own data.