Harness Engineering

or

How to Keep Agents On Track

@iurysza
iurysouza.dev

About me

  • •
    Google Dev Expert
  • •
    Co-host @ Fragmented-AI
  • •
    Platform Engineering @ SumUp
Iury Souza on stage
QR code for iurysouza.dev
@iurysza
iurysouza.dev

Non deterministic Coding

“They’ve done studies, you know?
60% of the time, it works every time”

- Brian Fantana (Anchorman)

Non deterministic Coding

  • •
    Agents are writing more and more code
  • •
    They write solid code, until they don’t
  • •
    New bottlenecks
    • •Keeping code quality
    • •Steering
    • •Containing blast radius

Non deterministic Coding

  • •
    Platform Teams attempts to keep quality standards for developers
  • •
    Harness Engineering does that for agents

Non deterministic Coding

  • •
    Prompt Engineering
  • •
    Context Engineering
  • •
    Harness Engineering

Non deterministic Coding

Days since last accident sign Made up word, scribbled over the sign
THE AGENTIC CODING LADDERHow much of the loop do you hand over?each rung hands the agent one more jobL1AI-ASSISTEDL2AI-GENERATED, HUMAN-REVIEWEDL3AI-GENERATED, AUTO-REVIEWEDL3.5SELECTIVE AUTO-MERGEL4MOSTLY AUTONOMOUSL5DARK FACTORY (LIGHTS-OUT)ONE LOGIN BUG, SIX WAYSLEVEL L1LEVEL L2LEVEL L3LEVEL L3.5LEVEL L4LEVEL L5WHO DOES WHICH JOBPICKS WORKYOUYOUYOUYOUAGENTAGENTWRITESYOUAGENTAGENTAGENTAGENTAGENTREVIEWSYOUYOUAGENTAGENTAGENTAGENTMERGESYOUYOUYOUBY RISKAGENTAGENTYOUR ROLEAuthorReviewerApproverGatekeeper for risky changesSets goals, handles escalationsWrites the specLIGHTS OUT · NOBODY IS WATCHING THE FLOOR

How we got Here

The Anthropic Moment

The Anthropic Moment

Feb 2025: Claude Code

  • •
    Claude Code preview: Feb 25
  • •
    Claude 3.7 Sonnet launch
  • •
    TUI / CLI based agent

The Anthropic Moment

Sep 2025: Opus

  • •
    Opus 4.5
  • •
    Step change in quality
  • •
    Model can prompt itself
  • •
    Improved planning mode

The Anthropic Moment

Holidays: Claude Code + Opus

  • •
    The IDE starts losing ground
  • •
    The harness was the main improvement
  • •
    Competition shows up
    • •Codex
    • •Opencode
    • •Amp
    • •Pi
  • •
    It’s not magic
  • •
    The harness matters more than the model

What is a harness anyway?

It’s all about steering

What is a harness

(…) a set of straps and fittings used to control an animal (…)

  • •
    Control
  • •
    Direction
  • •
    Constraints
A horse Harness

A horse Harness

RUNTIMECAPABILITIESSAFETY & SCALELLMstateless modelUser requestenters hereTool callsgo outThe loopkeeps runningWHAT IS AN AGENT HARNESSWhat the LLM uses to navigate the world and complete tasks.The model is the engine.The harness is everything else.

Should I build a Harness?

Short answer: no

  • •
    You can customize it
  • •
    You can extend it
  • •
    You can build infrastructure for it

A minimal Agent

Agent: an LLM running inside a harness

  • •
    Instruction Layering
  • •
    Tools
  • •
    Permissions
  • •
    Context Management
  • •
    Loop control

A typical agent definition

EXTENSION MODULEOne TypeScript file. Pi loads it at startup.1TOOLSRegister, replace or wrap built-ins2COMMANDSAdd your own /slash commands3KEYBOARD SHORTCUTSBind any key to an action4EVENT HOOKSagent_start, tool_call, message_end,compact … observe, change or block5UI COMPONENTSWidgets, overlays, status,custom editor and footer6CONTEXT CONTROLModify messages before the LLM,custom compactionMAKE IT YOUR OWN · PI EXTENSIONS0 / 6 WIREDGUARD.TSPI SESSIONpi ~/dev/app · guard.ts loadedrun_tests("login")42 passed› /shiptests ✓ · build ✓ · PR #214ctrl+alt+T→ run_tests("all")bash("rm -rf ~/dev")✕ BLOCKEDtests ✓ 42 passingctx81%trimmed stale tool outputONE FILE · YOUR HARNESS

Stacking Loops

What agents are made of

REACT LOOP · REASONING AND ACTINGCONTINUESTOP1USER GOALA question that needs an answer2REASONDecide the next step3ACTUse a tool4OBSERVERead the result5REASON AGAINUpdate the plan6CONTINUE OR STOPAnother step, or an answer?ENDREACT · A LOGIN TEST INVESTIGATIONITERATION 1 / 2USER GOALWhy is the login test failing?THOUGHT → ACTION → OBSERVATION01REASONLook at the test output first01ACTrun_tests("login")01OBSERVEExpected 200, got 40101REASONCheck the auth header01CONTINUENeed one more observation02REASONRead the request code02ACTread_file("auth.ts")02OBSERVEHeader: "Authorisation"02REASONServer expects "Authorization"02STOPRoot cause foundANSWERThe request misspells the Authorization header.
YesNoCONTINUE · NEXT ITERATIONKV CACHEKeeps past keys & valuesso each new token skipsrecomputing them1TOKENIZATIONInput text is broken into tokens2PREFILLTokens pass through transformerlayers to build context3ATTENTION + LOGITSAttention finds relevant context andcomputes next-token probabilities4DECODINGSelect the next token usinggreedy or sampling methods5NEXT-TOKEN PREDICTIONAppend the token and repeatuntil completionSTOP?OUTPUT COMPLETEINNER LOOP · NEXT TOKEN GENERATIONITERATION 1INPUT TEXTTOKENSAgents72917·are553·made1903·of328TRANSFORMER LAYERS → KV CACHEL1L2L3L4prefill · every prompt token in parallel, one layer at a timeonly the new token is computed · the rest is a KV cache hit
YESNORESULTS FEED BACK AS NEXT TURN INPUTUSER PROMPTLLM API CALL= one turnEXECUTE TOOLSCOLLECT RESULTSFINAL TEXTRESPONSETOOL REQUESTSIN RESPONSE?stopReason"toolUse"read_file("src/user.ts")edit("src/user.ts", …)That's the whole agent. 13 lines.OUTER LOOP · THE AGENTIC LOOPTURN 1agent.tsasync function simpleLoop(messages, model, tools) { while (true) { const response = await callModel(model, messages, tools); messages.push(response); if (response.stopReason !== "toolUse") { return messages; } for (const toolCall of response.toolCalls) { const result = await executeTool(toolCall); messages.push(result); } }}MESSAGES[]length 0
LLMINFERENCELOOP1 lap = 1 tokenrun toolappend resultnpm testSTACKING LOOPSHow agents get things doneUSERRename getUser to fetchUserVERIFYnpm test→42 passed · safe to hand backINNER LAPS0AGENT TURNS·TEST RUNS·Each loop checks the one inside.

Stacking Loops

Why this works?

  • •
    Non deterministic -> deterministic
  • •
    Code is verifiable
  • •
    Feedback loop
NOYESLLM FIXES THE TOOL CALL AND RETRIES1USER REQUESTRequest a code change2LLM REASONSInspects request, picks the edit tool3EDIT TOOL CALLedit(path, edits)4EXECUTE EDITTool tries to apply the edit5. TOOL CALLSUCCEEDED?✗ no · attempt 1TOOL ERRORinvalid path, apply failed6. FINAL RESPONSEsummary of the changeWHY THIS WORKS · FEEDBACKATTEMPT 1TOOL CALL→ use the edit tool→ path was wrong, follow the hintRESULTThe model guesses.The tool checks.The error teaches.

Stacking Loops

Why this works?

E.G.: Opencode’s AskUserQuestion Tool Error

E.G.: Opencode’s AskUserQuestion Tool Error

Closing the Loop

How to leverage the harness to make the agent smarter

Closing the Loop

Give the agent senses

  • •
    Observability enables self correction
  • •
    Design agent friendly tools, CLIs/MCPs
  • •
    Favor them over raw text instructions

Closing the Loop

Give the agent senses

  • •
    Machine Readable Output
  • •
    Self describing schemas
  • •
    Safety rails against hallucinations
  • •
    Instrument the repository for agents

Closing the Loop

Give the agent senses

  • •
    Token Efficiency
  • •
    Reliability
  • •
    Reproducibility
  • •
    Auditable
logstracesAPPOBSERVABILITY STACKVector DB / observability storeLOGS0TRACES0indexed & storedSYSTEM STATERuntime info, config, health, resourcesUI INSPECTIONDOM / component tree, layout, a11yAGENTquery / inspect / reason1. App emits logs & traces2. Stored & indexed3. Extra inspection data4. Agent queries all sourcesAGENT · INVESTIGATINGcheckout returns 500TRACE req-8f2a3.0sPOST /payauth.checkcart.loadpayments.chargedb.acquireROOT CAUSEDB pool exhausted → charge times out → 500Agents can't fix what they can't see.

Context Management

Give the agent a memory

Context Management

Give the agent a memory

  • •
    Context is a scarce resource
  • •
    A well crafted AGENTS.md + docs goes a long way
  • •
    Progressive disclosure
  • •
    Agents ❤️ file system + find + grep
HARNESS FRAGMENTS1SYSTEM PROMPT0kBase rules, tone and safetyBase rules, tone and safety2AGENTS.MD + DOCS0kEvery doc pasted in up frontAn index that links to the docs3TOOL DEFINITIONS0k5 MCP servers, every schemaNames and one line each4FILE READS0kWhole files, just in casefind + grep, then read the lines5TOOL RESULTS0kThe full 4,000-line test logOnly the failing linesEAGER · LOAD EVERYTHING UP FRONTJust in case the model needs it.PROGRESSIVE DISCLOSURELoad an index.Fetch the detail on demand with find and grep.CONTEXT WINDOW · 200K TOKENSUSED 0KONE SQUARE = 1K TOKENSEvery fragment costs tokensNearly full. Little room left to think.Index only. The detail stays on disk.Detail fetched on demandMost of the window is free again0% used0% used0% used$find src -name "*auth*"3 paths · +1k$grep -n "Authorization" src/auth.ts2 matches · +2k$read src/auth.ts:40-8041 lines · +4kHEADROOM200k tokens free for the model to think

Platform Engineering

And the tragedy of the commons

Platform Engineering

The tragedy of the commons

An economic and ecological concept describing how individuals, acting strictly in their own self-interest, deplete or spoil a shared resource, ultimately ruining it for everyone

Platform Engineering

Enables Speed and Cohesion

  • •
    Company Scaling Implications
    • •Shared standards
    • •Tooling
    • •Guardrails

Platform Engineering

Dedicated Agent Infra

  • •
    Fast moving environment
  • •
    Agent Evals and benchmarking
  • •
    Skills versioning and distribution
  • •
    Safely expose company systems for agents
  • •
    Research and Development
BOUNDED WORK AGENT · ONE TASK, ONE BOXISSUE #482Login returns 401 for a valid tokenROUTESCOPE1 bug reportTOOLSreplyToIssueSKILLStriage·verifySANDBOXcloudflare()AGENT TRACEBUDGET0 / 12 stepsOUTPUTREPLYTOISSUE(#482)1 file · +1 −1 · tests passRoot cause: "Authorisation" header typo. One-line fix.DEDICATED AGENT INFRAagent.ts// Expose (and protect) your agents to the world:export const route = (c, next) => next();// Give agents the autonomy to solve complex tasks:const instructions = `Triage a bug report end-to-end: reproduce the bug,diagnose the root cause, verify whether the behavior isintentional, and attempt a fix.`;// Compose the context your agent needs to do real work,// complete with virtual, local, or remote container sandbox.export default defineAgent(() => ({ model: 'moonshotai/kimi-2-7', tools: [replyToIssue], skills: [triage, verify], sandbox: cloudflare(), instructions,}));Narrow the blast radius.scope · tools · sandbox · verify

Conclusion

Experiment & Iterate

Conclusion

Experiment & Iterate

  • •
    Mastering Harness engineering
    • •Enables shipping faster
    • •Fewer regressions
    • •Lower Organizational drift
  • •
    No one-size-fits-all
  • •
    Start experimenting and find out what works best for your team/org

Harness Engineering

or

How to Keep Agents On Track

@iurysza
iurysouza.dev

Thank you!

QR code for iurysouza.dev
@iurysza
iurysouza.dev

Questions?

QR code for iurysouza.dev
@iurysza
iurysouza.dev