Harness engineering
What it takes to run AI agents in production — orchestration, verification, context, guardrails and observability, from the systems we build and operate.
Context Engineering vs RAG
Teams build a retrieval pipeline, watch quality drop as sessions grow, and conclude the model is the problem. Retrieval was only ever one input to a larger decision.
16 July 2026 · 4 min read
80% of AI Initiatives Show No Benefit
Models got dramatically better while most companies got nothing. That gap is not a model problem, and the teams closing it all build the same unglamorous things first.
4 June 2026 · 4 min read
Making an Agent Show Receipts
The most common agent failure is not a wrong answer. It is a confident report of success on work that does not run — and no record of how it got there.
28 May 2026 · 4 min read
Harness Engineering
When an agent misbehaves the reflex is to blame the model. Almost always the fault is in the system around it. Here are the eleven layers that system is made of, and how to tell which one failed.
21 May 2026 · 8 min read