Deploying a Cloudflare Worker for a Nepali SaaS, and what surprised me

Latency numbers, D1 versus Postgres from Kathmandu, and the three assumptions that turned out to be wrong when I moved a client onto Cloudflare Workers.

2 min read

A client was paying for a VPS in Singapore to serve a small internal tool. That is the correct default for most people, and it was the wrong one for them. This is the migration, the reasoning, and the parts that didn’t match my priors.

The assumptions that turned out to be wrong

Assumption 1: Workers are slower than a VPS for regional traffic

I expected edge execution to win on latency and lose on CPU-heavy work. Both halves of that turned out to be irrelevant at their actual scale.

Their API p95 sat around 190 ms from a Singapore VPS, of which roughly 120 ms was network round trip to the origin region. Moving static assets and read endpoints to edge removed most of that. The remaining work was a handful of SQLite queries, which was never the bottleneck.

  • 190ms

    Before (p95)

  • 68ms

    After (p95)

  • 0

    Servers to maintain

Assumption 2: D1 would be too limited

D1 is SQLite at the edge. It is genuinely constrained compared to Postgres — no pgvector, limited concurrent writes, a strict transaction model. None of that mattered because their workload was 94% reads of a table with fewer than 40,000 rows.

The thing that did matter was backups, and that took real thought. D1’s Time Travel gives you point-in-time recovery, but I still export nightly to R2 because Time Travel is not a backup strategy, it’s an undo buffer.

# Nightly D1 export to R2 — run from cron, not from the Worker.
d1 export DB --remote --output ./nightly-$(date -u +%F).sql
rclone copy ./nightly-$(date -u +%F).sql r2:client-backups/d1/

Assumption 3: The developer experience would be a step down

It was a step up, but for a reason I did not anticipate: the local loop. No SSH into staging, no separate environment to keep in sync. wrangler dev runs the same runtime as production, including the D1 binding and the local SQLite file, which means a data migration is testable before it touches anything real.

The migration shape

The split that made this tractable was small and boring:

  1. 01

    Worker owns the API surface

    Auth checks, validation, and reads. D1 accessed via the Workers binding.

  2. 02

    R2 owns files

    Uploaded documents and generated exports. Never route file bytes through the database.

  3. 03

    Static front-end stays static

    The admin UI is a prebuilt bundle on Pages. The Worker never renders HTML.

  4. 04

    Secrets live in Workers, not in the repo

    wrangler secret put for everything. No .env on a server I have to remember to rotate.

What I’d warn a client about

The other warning is about lock-in. Moving back is entirely possible — it is plain TypeScript — but the D1 schema, the R2 lifecycle rules, and the bindings are Cloudflare-specific. That is a real exit cost even if it is a low one. I would rather name it than pretend it isn’t there.

The part I’m still evaluating

Search. They now have about 12,000 documents that nobody can find, and Cloudflare Vectorize is the obvious answer. I have not shipped it yet because the retrieval quality on a small, messy corpus is not something I want to promise before I have measured it. That is probably a future post here.

On this page

Backups: the checks that matter are the ones nobody runs

A restoration drill is the only backup that has ever saved anyone. How to run one in an afternoon, and the three failures it will find.

2 min readbackups · infrastructure

Upgrading a decade-old PHP app without stopping the business

A phased migration that kept the site up every weekend, what the strangler-fig pattern actually looks like in practice, and the two things that nearly went wrong.

2 min readmigration · php