Invoice document extraction pipeline

R2 → rasterise → structured extraction → arithmetic validation → confidence-gated review queue. Running in production for a services business since early 2026.

Outcome1,500 invoices/month · ~92% auto-cleared · ~$40/month model spend
Role
IT consultant — design and build
Year
2026

The shape

Documents in, validated records out, with a human checkpoint where confidence is low. Not a chatbot — a narrow job with a defined correct answer.

The validation that does the real work

from decimal import Decimal

def arithmetic_ok(f: dict) -> bool:
    """A model asked 'is this correct?' agrees with itself. Multiplication does not."""
    try:
        subtotal = Decimal(str(f["subtotal"]))
        tax = Decimal(str(f.get("tax", 0)))
        total = Decimal(str(f["total"]))
    except (KeyError, ArithmeticError):
        return False
    return abs((subtotal + tax) - total) <= Decimal("0.05")

What I would build next

A calibration report. Not “how many passed” — a distribution of confidence scores against actual correctness, so the threshold moves with evidence instead of with my optimism.

The written version of this is in What “give the AI our invoices” actually means.

More work

Training2026

Building internal AI tooling on a budget

An interactive workshop for engineers at Nepali companies. Hands-on, tool-agnostic, and deliberately sceptical of autonomous pipelines.

Delivered to 4 cohorts · 90 participants

  • Claude Code
  • MCP
  • Python
  • GitHub Actions
Client work2026

Moving a legacy commerce platform to Cloudflare Workers

Retail client, Kathmandu · IT consultant — architecture, migration, handover

Static front-end on Pages, API on Workers, D1 for reads and R2 for files. p95 latency from 190ms to 68ms and the server contract ended.

p95 190ms → 68ms · 3 servers decommissioned

  • Cloudflare Workers
  • D1
  • R2
  • TypeScript
  • +2
Experiment2025

Runbook templates that actually get followed

The four-part structure, the confirm-before-acting rule, and the quarterly cold-reader test. Used on every engagement since 2024.

Median time-to-resolve 38min → 17min

  • Bash
  • Markdown
  • systemd
  • Prometheus

Published by Sudeep Dhakal, IT Consultant.