Harness engineering
What it takes to run AI agents in production — orchestration, verification, context, guardrails and observability, from the systems we build and operate.
The Software Factory Stack
Everyone is building the same machine and calling it something different. Here are its four layers, drawn from a factory that runs in production — and the one decision that determines how much of it you can still change later.
15 September 2026 · 14 min read
WebMCP: Give the Agent a Menu, Not a Map
Agents drive websites by screenshotting and guessing. WebMCP replaces the guess with a contract. Here is the standard as it stands this week, the code for both APIs, and an honest read on who should ship it now.
8 September 2026 · 24 min read
Making a Next.js Site Agent-Readable
Agents fetch, they do not browse — and most of what gets recommended for serving them has never been measured. We shipped all of it on this site, measured it, and separated the two halves.
27 August 2026 · 16 min read
Context Engineering vs RAG
Teams build a retrieval pipeline, watch quality drop as sessions grow, and conclude the model is the problem. Retrieval was only ever one input to a larger decision.
16 July 2026 · 4 min read
80% of AI Initiatives Show No Benefit
Models got dramatically better while most companies got nothing. That gap is not a model problem, and the teams closing it all build the same unglamorous things first.
4 June 2026 · 4 min read
Making an Agent Show Receipts
The most common agent failure is not a wrong answer. It is a confident report of success on work that does not run — and no record of how it got there.
28 May 2026 · 4 min read
Harness Engineering
When an agent misbehaves the reflex is to blame the model. Almost always the fault is in the system around it. Here are the eleven layers that system is made of, and how to tell which one failed.
21 May 2026 · 8 min read