Upgrading a decade-old PHP app without stopping the business

A phased migration that kept the site up every weekend, what the strangler-fig pattern actually looks like in practice, and the two things that nearly went wrong.

2 min read

The brief was “modernise the platform”. The reality was eleven years of accumulated business logic with no test coverage, one developer who understood it, and a shop that cannot afford a day of downtime.

The shape

  1. 01

    Put a reverse proxy in front, routing nothing new

    Every existing request still goes to the old app. This step alone is reversible and changes no behaviour.

  2. 02

    Move read-only routes first

    Catalogue, blog, static content. Low risk, visible progress, and it proves the deployment pipeline.

  3. 03

    Move authentication, carefully

    The highest-risk step. Dual-read sessions, keep the old session store authoritative until cutover is proven.

  4. 04

    Move one write path

    Pick the least painful write. Learn what the data model actually requires before the important ones.

  5. 05

    Never delete the old app until the boundary is empty

    Keep it running, frozen, for one full business quarter.

The strangler pattern, concretely

Routing is where this lives. Every migration step is one more prefix pointing at the new service instead of the old one.

# Phase 3 — read-heavy routes already migrated
location ~ ^/(catalogue|blog|about) {
    proxy_pass http://new_app;
    include /etc/nginx/proxy_params;
}

# Everything else, including all writes, still on the monolith
location / {
    proxy_pass http://legacy_app;
    include /etc/nginx/proxy_params;
}

Two things that nearly went wrong

The silent dual-write

Moving a write path means, briefly, two systems that both believe they own the data. We got this wrong for one deploy: the legacy app still had a hidden file_put_contents to a JSON cache that nobody had mentioned.

The symptom was not an error. It was an admin page showing data from four hours ago, which we initially assumed was a cache issue for about an hour.

Fix: search the old codebase for every write path before you claim one, and add a monitor that compares record counts between the two systems for a week.

Time zones

The legacy app stored local time in DATETIME columns with no timezone. The new service used TIMESTAMPTZ. Both were correct; together they were off by 5:45 hours, which only surfaced on a single daylight-saving boundary in a country I had not thought about.

What “modernised” ended up meaning

After four months the new service handled 71% of traffic. The monolith still ran, still frozen, handling checkout and the parts nobody wanted to touch. That is not a clean ending, and it is the correct one — the business stayed up the whole time and every step was reversible.

The remaining 29% is now a scheduled piece of work rather than an emergency.

On this page

Backups: the checks that matter are the ones nobody runs

A restoration drill is the only backup that has ever saved anyone. How to run one in an afternoon, and the three failures it will find.

2 min readbackups · infrastructure

Deploying a Cloudflare Worker for a Nepali SaaS, and what surprised me

Latency numbers, D1 versus Postgres from Kathmandu, and the three assumptions that turned out to be wrong when I moved a client onto Cloudflare Workers.

2 min readcloudflare · workers