Harness Engineering

or

How to Keep Agents On Track

@iurysza
iurysouza.dev

About me

  • •
    Google Dev Expert
  • •
    Co-host @ Fragmented-AI
  • •
    Platform Engineering @ SumUp
Iury Souza
QR code for iurysouza.dev
@iurysza
iurysouza.dev

Non deterministic Coding

“They’ve done studies, you know?
60% of the time, it works every time”

- Brian Fantana (Anchorman)

Non deterministic Coding

  • •
    Agents are writing more and more code
  • •
    They write solid code, until they don’t
  • •
    New bottlenecks
    • •Keeping code quality
    • •Steering
    • •Containing blast radius

Non deterministic Coding

  • •
    Platform Teams attempts to keep quality standards for developers
  • •
    Harness Engineering does that for agents

Non deterministic Coding

  • •
    Prompt Engineering
  • •
    Context Engineering
  • •
    Harness Engineering

Non deterministic Coding

Days since last accident sign Made up word, scribbled over the sign
THIS TALK IS ABOUT LOOPSWhat is the work we do?INNER LOOPYour own cycle.Write, test, debug, repeat.OUTER LOOPThe team's delivery cycle.CI, review, deploy, operate.MIDDLE LOOPSupervising agents.Direct, evaluate, fix.CIreviewdeployoperatewritetestdebugINNER LOOPdirectevaluatefix
THE AGENTIC CODING LADDERHow much of the loop do you hand over?L1AI-ASSISTEDL2AI-GENERATED, HUMAN-REVIEWEDL3AI-GENERATED, AUTO-REVIEWEDL3.5SELECTIVE AUTO-MERGEL4MOSTLY AUTONOMOUSL5DARK FACTORYONE LOGIN BUG, SIX WAYSLEVEL L1LEVEL L2LEVEL L3LEVEL L3.5LEVEL L4LEVEL L5WHO DOES WHICH JOBPICKS WORKYOUYOUYOUYOUYOUYOUWRITESYOUAGENTAGENTAGENTAGENTAGENTREVIEWSYOUYOUAGENTAGENTAGENTAGENTMERGESYOUYOUYOUBY RISKAGENTAGENTVERY LIMITED MILEAGE, FOR NOW

How we got Here

The Anthropic Moment

The Anthropic Moment

Feb 2025: Claude Code

  • •
    Claude Code preview: Feb 25
  • •
    Claude 3.7 Sonnet launch
  • •
    TUI / CLI based agent

The Anthropic Moment

Nov 2025: Opus

  • •
    Opus 4.5
  • •
    Step change in quality
  • •
    Model can prompt itself
  • •
    Improved planning mode

The Anthropic Moment

Holidays: Claude Code + Opus

  • •
    The IDE starts losing ground
  • •
    The harness was the main improvement
  • •
    Competition shows up
    • •Codex · Opencode · Amp · Pi
    • •Gemini CLI · Copilot CLI · Cursor CLI · Droid · Crush
    • •Goose · Kiro · Qwen Code · Kimi CLI · Cline CLI
  • •
    It’s not magic
  • •
    The harness matters more than the model

What is a harness anyway?

It’s all about control

What is a harness

(…) a set of straps and fittings used to control an animal (…)

  • •
    Control
  • •
    Direction
  • •
    Constraints
A horse Harness

A horse Harness

A model alone can’t do anything

Text in, text out

  • •
    It can’t read your files
  • •
    It can’t run your tests
  • •
    It can’t call an API
  • •
    To act in the world, it needs tools

Stacking Loops

What agents are made of

YesNoCONTINUE · NEXT ITERATIONKV CACHEKeeps past keys & valuesso each new token skipsrecomputing them1TOKENIZATIONInput text is broken into tokens2PREFILLTokens pass through transformerlayers to build context3ATTENTION + LOGITSAttention finds relevant context andcomputes next-token probabilities4DECODINGSelect the next token usinggreedy or sampling methods5NEXT-TOKEN PREDICTIONAppend the token and repeatuntil completionSTOP?OUTPUT COMPLETEINNER LOOP · NEXT TOKEN GENERATIONITERATION 1INPUT TEXTTOKENSAgents72917·are553·made1903·of328TRANSFORMER LAYERS → KV CACHEL1L2L3L4prefill · every prompt token in parallel, one layer at a timeonly the new token is computed · the rest is a KV cache hit

ReAct

Reasoning + Acting

  • •
    One pattern emerged: think, then act
  • •
    Reason → act → observe → repeat
  • •
    The harness runs that loop
REACT LOOP · REASONING AND ACTINGCONTINUESTOP1USER GOALA question that needs an answer2REASONDecide the next step3ACTRun a command4OBSERVERead its output5REASON AGAINSo what next?6CONTINUE OR STOPAnother step, or an answer?ENDREACT · A LOGIN TEST INVESTIGATIONITERATION 1 / 2USER GOALWhy is the login test failing?THINK → RUN → READ OUTPUT01REASONStart with the test output.01ACT$npm test -- login01OBSERVEFAIL login: expected 200, got 40101REASON401 means auth failed. Is the header right?01CONTINUENot sure yet. Go and look at the code.02REASONFind where the header is set.02ACT$grep -n Authori src/auth.ts02OBSERVE12: headers['Authorisation'] = token02REASONTypo. The server expects "Authorization".02STOPRoot cause found. Answer.ANSWERThe request misspells the Authorization header.

Stacking Loops

Why this works?

  • •
    Non deterministic -> deterministic
  • •
    Code is verifiable
  • •
    Feedback loop
YESNORESULTS FEED BACK AS NEXT TURN INPUTUSER PROMPTLLM API CALL= one turnEXECUTE TOOLSCOLLECT RESULTSFINAL TEXTRESPONSETOOL REQUESTSIN RESPONSE?stopReason"toolUse"edit("src/components/card.tsx", …)edit("src/ui/Card.tsx", …)That's the whole agent. 13 lines.OUTER LOOP · THE AGENTIC LOOPTURN 1agent.tsasync function simpleLoop(messages, model, tools) { while (true) { const response = await callModel(model, messages, tools); messages.push(response); if (response.stopReason !== "toolUse") { return messages; } for (const toolCall of response.toolCalls) { const result = await executeTool(toolCall); messages.push(result); } }}MESSAGES[]length 0

Should I build a Harness?

Short answer: no

  • •
    You can customize it
  • •
    You can extend it
  • •
    You can build infrastructure for it
TWO HARNESSESWhose harness is it?BUILDER HARNESSClaude Code, Codex, Pi.YOUR HARNESSEverything you add for your codebase.YOUR HARNESSBUILDER HARNESSmodelprompt · loop · toolsAGENTS.mdskillsknowledge basetestslintlogsMCP · CLIs · hookshow they plug in

Your harness

Four things you can add

  • •
    AGENTS.md
  • •
    A skill
  • •
    A hook
  • •
    An extension

Guides and sensors

Before it acts, after it acts

  • •
    Guides steer before the agent acts
  • •
    They make the first try right more often
  • •
    Sensors check after the agent acts
  • •
    They make a second try possible

Computational or inferential

Who does the checking?

  • •
    Computational: deterministic
  • •
    Fast, cheap, same answer every time
  • •
    Inferential: non-deterministic
  • •
    Can judge meaning, but slower and can be wrong
YOUR HARNESS · TWO QUESTIONSBefore or after it acts? Code or a model?GUIDESfeedforward · before it actsSENSORSfeedback · after it actsCOMPUTATIONALdeterministicCPU · msINFERENTIALnon-deterministicGPU · s–minMake the right thing easyTools that shape the first attemptCLIsscriptstypesLSPTell it what good looks likeWords the model reads before actingAGENTS.mdskillsrulesreference docsknowledge baseCatch it, every timeCheap enough to run on every changetestslinttype checkhooksCI pipelineJudge what code can'tSemantic review, a model as judgereview agentself-reviewdrift review
YOUR HARNESSKeep quality leftThe sooner a check runs, the cheaper the fixBefore it actsAGENTS.mdskillstypesInside its looptestslinttype checkBefore commithooksbuildself-reviewBefore mergeCI pipelinereview agenthuman reviewAlwaysdead-code scanSLO alertsdrift reviewchecked by codejudged by a model

Same mistake twice?

Fix the harness first, then the code

Closing the Loop

How to leverage the harness to make the agent smarter

THIS TALK IS ABOUT LOOPSClosing the loopINNER LOOPYour own cycle.Write, test, debug, repeat.OUTER LOOPThe team's delivery cycle.CI, review, deploy, operate.MIDDLE LOOPSupervising agents.Direct, evaluate, fix.CIreviewdeployoperatewritetestdebugINNER LOOPdirectevaluatefix

Closing the Loop

Give the agent senses

  • •
    Observability enables self correction
  • •
    Design agent friendly tools, CLIs/MCPs
  • •
    Favor them over raw text instructions

Closing the Loop

Give the agent senses

  • •
    Machine Readable Output
  • •
    Self describing schemas
  • •
    Safety rails against hallucinations
  • •
    Instrument the repository for agents

Closing the Loop

Give the agent senses

  • •
    Token Efficiency
  • •
    Reliability
  • •
    Reproducibility
  • •
    Auditable
logstracesAPPOBSERVABILITY STACKVector DB / observability storeLOGS0TRACES0indexed & storedSYSTEM STATERuntime info, config, health, resourcesUI INSPECTIONDOM / component tree, layout, a11yAGENTquery / inspect / reason1. App emits logs & traces2. Stored & indexed3. Extra inspection data4. Agent queries all sourcesAGENT · INVESTIGATINGcheckout returns 500TRACE req-8f2a3.0sPOST /payauth.checkcart.loadpayments.chargedb.acquireROOT CAUSEDB pool exhausted → charge times out → 500Agents can't fix what they can't see.

Every app needs its own harness

Even OpenAI couldn't ship an empty box

  • •
    Teams add their own skills and plugins
  • •
    Experts bring the taste
  • •
    Docs live next to the code
  • •
    Your app is a domain-specific harness

Two camps

Why work on a harness at all?

  • •
    The bitter lesson: wait for a better model
  • •
    Neuro-symbolic: the model needs code around it

The bitter lesson

Wikipedia

“In the long run, general approaches that scale with available computational power tend to outperform ones based on domain-specific understanding.”

The bitter lesson

Rich Sutton, 2019

“Building in how we think we think does not work in the long run.”

Neuro-symbolic AI

Wikipedia

“Combines neural networks and symbolic AI … to create more robust, more reliable, and more trustworthy AI.”

Neuro-symbolic

Gary Marcus, 2026

“Claude Code isn’t better because of scaling. It’s better because it is neurosymbolic.”

Bigger, not smaller

Thariq Shihipar, Anthropic

“The models get better and better. And so the harness needs to become more and more complicated to allow the model to do more things.”

The middle ground

Both are right, about different layers

  • •
    Scaffolding that teaches it to think: the next model eats it
  • •
    Checks that connect it to your world: they stay
  • •
    Don’t teach the model to think. Give it ways to check.

Human input

Direct it, don't remove it

“A good harness should not necessarily aim to fully eliminate human input, but to direct it to where our input is most important.”

Platform Engineering

And the tragedy of the commons

Platform Engineering

The tragedy of the commons

An economic and ecological concept describing how individuals, acting strictly in their own self-interest, deplete or spoil a shared resource, ultimately ruining it for everyone

Platform Engineering

Enables Speed and Cohesion

  • •
    Company Scaling Implications
    • •Shared standards
    • •Tooling
    • •Guardrails

Platform Engineering

Dedicated Agent Infra

  • •
    Fast moving environment
  • •
    Agent Evals and benchmarking
  • •
    Skills versioning and distribution
  • •
    Safely expose company systems for agents
  • •
    Research and Development

Conclusion

Experiment & Iterate

What we covered

The talk in one slide

  • •
    A model alone can't act. The harness gives it a loop.
  • •
    Guides before it acts, sensors after
  • •
    Same mistake twice? Fix the harness first.
  • •
    There is now a new middle loop
  • •
    Someone needs to own the commons

Conclusion

Experiment & Iterate

  • •
    Mastering Harness engineering
    • •Enables shipping faster
    • •Fewer regressions
    • •Lower Organizational drift
  • •
    No one-size-fits-all
  • •
    Start experimenting and find out what works best for your team/org

Sources

Read these next

  • •
    Böckeler, "Harness engineering", martinfowler.com
  • •
    Yao et al., "ReAct", 2022
  • •
    Sutton, "The Bitter Lesson", 2019
  • •
    Marcus, "The biggest advance in AI since the LLM", 2026

Harness Engineering

or

How to Keep Agents On Track

@iurysza
iurysouza.dev

Keep reading

My blog and newsletter

  • •
    More on harnesses, agents and loops
  • •
    iurysouza.dev
  • •
    Link and QR on the last slide

Thank you!

QR code for iurysouza.dev
@iurysza
iurysouza.dev

Questions?

QR code for iurysouza.dev
@iurysza
iurysouza.dev