San Lee

Technical program manager building reliable AI systems.

In a product seat at JPMorganChase, on collaboration platforms serving 300,000 employees, after seven years in infrastructure and operations.

I set direction, define the interfaces between teams and systems, and gate delivery on evidence. Claude does most of the typing under those contracts. The evals and the postmortems on this site are the proof.

Calls that shipped, and calls that were reversed.

A doubled rule marks a decision that was reversed.

Selected work

Four views of the work: how I structured the system, measured its decisions, tested a judge, and managed the program. Deeper investigations remain linked throughout the record.

Run the classifier demo

Plate 01 Direction, then contracts, then an evidence gate. Direction set the aim Contracts freeze seams Evidence gate the merge eval gates the merge

The instrument

telltale

Five coding-agent CLIs, one brief, five columns — and gauges that render nothing when nothing was measured.

I ran a month-long evaluation of the AI coding agents that build this record. I wrote the test down before the evidence existed. The measurement changed my operating policy. telltale is the instrument that reads what five agent CLIs leave on disk.

v0.3.0 · 171 golden files · 2191 test functions

The writeup

The system

I wired three services into one loop: an API, an LLM classifier, and a RAG agent.

Map 02

Notes API BackgroundTask Classifier Tags Knowledge base KB Agent Grounded answer

Select a component to see what it does and the decision behind it.

Read the full writeup

The decision log

Three decisions, one row each. The full log is on Work. The decision log →

Classifier evidence

Where the classifier stands now, on 54 hand-labeled real snippets.

I score category and region separately. The two figures are equal here by coincidence.

Classifier v3.2.1 on the 54-snippet human-labeled gold set. Full eval →

The classical baseline lost the case for an LLM. I vendored it here, and it runs in the browser on the same rows. Run it yourself →

Field notes on the techniques behind this work sit on the work index. The field notes →

About

For the short version of how I came to this work and how I approach it, read About →

I run this record in the open. A measurement changed what I shipped, and a reversed decision carries a doubled rule on the page that recorded it. A figure caption states what the figure does not show, and the build fails on an empty one. The operating layer keeps its incident log in the open, postmortem-first.

It is a solo build, so I simulate the program layer. The Product & Program page states what that means.

I set the direction and the bar. Claude does most of the typing. The evals and postmortems are the proof. The colophon states the method in full.

A second record sits off this system: a physical lab that makes a factory-blank router provision itself. Zero-touch provisioning →

Outside the work I hike, take pictures, and live with a Scottish Fold named Sango. The rest is on About and in the gallery.

Sango, a Scottish Fold, asleep
Sango, Chief Nap Officer.