San Lee

I run engineering the way a technical program manager runs a program, with AI agents as the build capacity.

Product at JPMorganChase, on collaboration platforms serving 300,000 employees; the first seven years there were the wires.

I set the direction, freeze the contracts between the parts, and gate every change on evidence. Nothing here was one-shotted.

The real calls behind the system. The ones that went the wrong way are here too, set with the same care as the wins, because that is the part worth reading.

A doubled rule marks a decision that was reversed.

Start here

A small AI system, built by directing Claude and run like a program: every change planned, built, reviewed, and iterated, with an eval gating the merge. If you have one minute, take one of these doors.

The measurement that lied to me: the Judge. The numbers, and how they were made honest: the Classifier. Six of my own checks reporting success for work that never ran: False Green. The program view: Product & Program.

The system

Three services wired into one loop: an API, an LLM classifier, and a RAG agent.

Notes API BackgroundTask Classifier Tags Knowledge base KB Agent Grounded answer

Select a component to see what it does and the decision behind it.

Read the full writeup

The decision log

Reversed. BM25 retiredEVAL · BM25

I added retrieval grounding, measured it, and cut it

I added BM25 lexical retrieval to ground each call, expecting a clear win. The first measurement said it barely helped. Then I found the measurement itself was unfair.

The result: I had been scoring the grounded arm against a stale baseline built on an older prompt. Rebuilt so both arms ran the same prompt, grounding fixed a domain call zero times and broke four across 162 classifications. It was retired outright, not left on as an option.

An unfair baseline flatters whatever you were hoping for, and comparing two things means changing exactly one of them. The retrieval code stays in the repo, dormant and reproducible, because the built-it-measured-it-cut-it trail is the artifact. Full eval →

Reversed. Model gated on evidenceSYS-002 · ADR-013

Gate model choice on evidence, not vibes

Sonnet by default across the system. Escalate to Opus only where an eval shows the quality gain pays for itself. I built the upgrade, measured it, and declined it.

The result: routing the uncertain cases to the premium model moved zero rows on both axes at about 1.97× the cost per article. Sending everything to the premium model scored identically, so there was no headroom for any router to capture.

A targeted prompt fix had already taken the ground routing was aimed at, so the cheapest fix won and that is now measured rather than asserted. The harness stays in the repo as the record. Full verdict →

Eval rebuiltEVAL · v1 → v3

I rebuilt the eval until it could tell me I was wrong

v1 graded the model on 300 snippets it generated itself, which measures consistency, not correctness. v2 replaced that with 54 hand-labeled real snippets, cross-checked by an Opus judge, and it has been catching me ever since.

The result: domain accuracy fell from 97.3% to 88.9% on real text, and that lower number is the trustworthy one.

Category accuracy actually rose, because real vocabulary separates more cleanly; domain fell because the synthetic set was trivially easy to grade. Full eval →

Eight more, in one line each. The full telling of every one is on the page it links to.

Classifier evidence

Where the classifier stands now, on 54 hand-labeled real snippets.

Category and domain land on the same figure here. Two axes, scored separately, not a copied cell.

Those are the numbers after the eval got harder, not before. The full trajectory — synthetic self-graded, then real hand-labeled, a model migration, a prompt fix and a third axis, with the ceiling I called wrongly and the misses that clustered on one boundary — is the table on the classifier page. See the full eval →

The classical baseline that lost the case for using an LLM at all is vendored into this site and runs in your browser, on the same rows. Run it yourself →

Field notes

Notes I keep while learning, in plain language, on the techniques behind all of this. Read them →

A selection; the full, growing set is on the notes site.

About, honestly

For most of my career, what to build was decided above me. I was close to the work: running systems, shipping pieces of them, living with the seams. I had not owned a full-loop system, so I built one myself, in the open, to get that experience — decisions recorded, failed attempts included. It’s a small system, but I took the process seriously.

I did not get here alone. Most of what I know came from people who were patient when I was the one who did not know yet, and generous with hard-won expertise. I try to pass that forward where I can.

It is a solo build, so the program layer is simulated — I lay out what that means on the Product & Program page.

Outside the work: traveling, hiking, a bit of photography, and a Scottish Fold named Sango who supervises. The good frames live in the gallery.

Sango, a Scottish Fold, asleep
Sango, Chief Nap Officer.