or
How to Keep Agents On Track
“They’ve done studies, you know?
60% of the time, it works every time”
- Brian Fantana (Anchorman)


The Anthropic Moment
Feb 2025: Claude Code
Nov 2025: Opus
Holidays: Claude Code + Opus
It’s all about control
(…) a set of straps and fittings used to control an animal (…)
A horse Harness
Text in, text out
What agents are made of
Reasoning + Acting
Why this works?
Short answer: no
Four things you can add
Before it acts, after it acts
Who does the checking?
Fix the harness first, then the code
How to leverage the harness to make the agent smarter
Give the agent senses
Give the agent senses
Give the agent senses
Even OpenAI couldn't ship an empty box
Why work on a harness at all?
Wikipedia
“In the long run, general approaches that scale with available computational power tend to outperform ones based on domain-specific understanding.”
Rich Sutton, 2019
“Building in how we think we think does not work in the long run.”
Wikipedia
“Combines neural networks and symbolic AI … to create more robust, more reliable, and more trustworthy AI.”
Gary Marcus, 2026
“Claude Code isn’t better because of scaling. It’s better because it is neurosymbolic.”
Thariq Shihipar, Anthropic
“The models get better and better. And so the harness needs to become more and more complicated to allow the model to do more things.”
Both are right, about different layers
Direct it, don't remove it
“A good harness should not necessarily aim to fully eliminate human input, but to direct it to where our input is most important.”
And the tragedy of the commons
The tragedy of the commons
An economic and ecological concept describing how individuals, acting strictly in their own self-interest, deplete or spoil a shared resource, ultimately ruining it for everyone
Enables Speed and Cohesion
Dedicated Agent Infra
Experiment & Iterate
The talk in one slide
Experiment & Iterate
Read these next
or
How to Keep Agents On Track
My blog and newsletter