Complete index
Work
The homepage is curated. Nothing is hidden here.
Flagship case studies
- The SystemArchitecture, interfaces, asynchronous delivery, and a broken contract seam
- Defense News ClassifierHuman-labeled evaluation, reversals, and evidence-gated model decisions
- Faithfulness JudgeWhether an LLM judge can reliably identify unsupported claims
- Product & ProgramRoadmap, dependencies, risks, metrics, and the limits of a solo simulation
Evaluation and reliability
- Classifier baseline and live demo — run the TF-IDF logistic-regression baseline in the browser.
- False Green — six checks that passed when the required work never ran.
- Retrieval, Measured — a 27-query gold set and three paired retrieval experiments.
- The Autonomy Ladder — four levels of agent autonomy, with two shipped and two rejected.
- Loop Replay — interactive playback of two autonomous improvement runs.
- Context Architecture — an audit of relevance, sufficiency, isolation, economy, and provenance.
Systems and infrastructure
- One Note, End to End — one payload traced across the real service contracts.
- Zero-Touch Provisioning — a factory-blank router configured after one power cycle.
Looking for terminology used across the projects? Open the glossary →