A six-month look at which AI tooling actually survived daily use

Six months of client work, tracking which tools stuck and which quietly stopped being opened. Includes the boring ones that won.

I have been recommending AI tooling to clients for about six months. This is the retrospective: what survived contact with a normal week, and what didn’t.

What survived

Terminal-first agents, with a config file in git

The tools that stuck are the ones a developer can drive from a terminal, with prompts, tool permissions, and instructions committed to a repository. Reviewable diffs, reproducible behaviour, and a plain text file that explains what the agent is meant to do.

  • 4 of 6

    Tools still in use

  • 100%

    Config in git

  • 3

    Clients doing daily work

Retrieval over internal documents — when the corpus is clean

When a client’s documents were reasonably tidy, search over their own material beat general-purpose AI research immediately, because the answers were about their business rather than the internet’s.

The failure mode is a messy corpus. Extraction noise compounds, and the system becomes confidently wrong about their own procedures.

Plain scripts and a cron entry

The unsexy answer. Most client automation value came from a 200-line Python script that runs at 5am. Not an agent, not a workflow engine. A script.

What did not survive

Chat UIs without a task

An assistant someone can ask anything is a tool people ask nothing of. Every deployment that relied on open-ended chat saw engagement drop within three weeks.

The ones that worked had a task shape: “review these ten invoices and flag anomalies”, not “ask me anything”.

Multi-step autonomous pipelines

The pitch was a five-stage pipeline that ran unattended. In practice, stage three would do something confidently wrong, stage four would pass it on, and a human would find it four days later.

What worked instead was the same pipeline with a human checkpoint between stages two and three. Less autonomous, considerably more used.

Anything requiring the user to change behaviour first

The most common failure was not a bad tool. It was a good tool that required someone to write prompts, maintain context files, and review output — for a person whose actual job was unrelated to software. They did it for three weeks, then stopped.

What I now recommend by default

  1. 01

    Find the repetitive task with a known answer

    Not the interesting task. The one where a wrong answer is cheap to catch.

  2. 02

    Make it a script with a narrow prompt

    Version controlled, runs on a schedule, output goes somewhere a human already looks.

  3. 03

    Add the review queue before the automation

    Including the weekly report of what failed. This is the part that sustains adoption.

  4. 04

    Only then add autonomy

    Once you know the failure modes, because you have been watching them.

The honest caveat

Six months, three clients, and me choosing what to recommend. This is not a controlled comparison and it is not generalisable. But the boring tools won, the autonomous ones lost, and task specificity predicted adoption better than anything else I measured.

On this page
AI guideApplied AI for small teams · 1

What "give the AI our invoices" actually means, technically

A concrete architecture for document extraction that a small Nepali services business can run on a budget, including the parts that go wrong.

2 min readai · llm