What "give the AI our invoices" actually means, technically
A concrete architecture for document extraction that a small Nepali services business can run on a budget, including the parts that go wrong.
2 min read
Applied AI for small teams · part 1“Use AI on our invoices” is the most common request I get and the least specific. This is my default starting architecture, with the honest failure modes.
The shape that has worked
- 01
Ingest to object storage
R2 or S3. Invoices live as files with an ID. Never as rows in a table — the file is the source of truth.
- 02
Render to images
Text-layer PDFs are cheap to parse. Scanned documents need rasterising at 200 DPI first.
- 03
Extract with a schema, not a prompt
Ask for vendor, invoice number, date, currency, subtotal, tax, total, due date. Structured output, enforced.
- 04
Validate arithmetically
Subtotal + tax should equal total. Reject and retry or escalate when it does not.
- 05
Confidence-gate into review
Above threshold: write straight through. Below: queue for a human. This is the whole design.
Why the confidence gate is the entire design
Every “we tried AI on our documents and it was garbage” conversation I have had traces back to the same omission: there was no threshold and no review queue, so the system confidently wrote bad data and nobody noticed for a month.
Arithmetic validation catches more than the model
The single highest-value check is not an LLM judging an LLM. It is multiplication.
from decimal import Decimal
def arithmetic_ok(f: dict) -> bool:
try:
subtotal = Decimal(str(f["subtotal"]))
tax = Decimal(str(f.get("tax", 0)))
total = Decimal(str(f["total"]))
except (KeyError, ArithmeticError):
return False
# Allow a small tolerance for rounding in source documents.
return abs((subtotal + tax) - total) <= Decimal("0.05")
If it does not balance, that is a strong signal the extraction is wrong even when every individual field looks plausible — which is exactly the case that sneaks past a model asked “is this correct?”
Cost, honestly
For a business processing roughly 1,500 invoices a month, on Claude with prompt caching for the schema and instructions:
1,500
Invoices / month
~$40
Model spend
~92%
Auto-cleared
~1,400
No human time
The model cost is rarely the interesting number. The interesting number is review-queue throughput: if 8% of 1,500 is 120 documents, and a person clears them at a minute each, that is two hours a month of human attention for a system that saved a great deal more.
What I build first, always
Not the extraction. A weekly report of the review queue: what failed, what was corrected, and the shape of the failures. It takes an hour to build, and it is how you find out whether the confidence threshold is calibrated or just optimistic.
On this page
Related reading
A six-month look at which AI tooling actually survived daily use
Six months of client work, tracking which tools stuck and which quietly stopped being opened. Includes the boring ones that won.
2 min readai · tooling