All posts

Harness engineering in AI

/AI/4 min read

A model on its own can only turn text into more text. The harness is everything you build around it — tools, memory, guardrails, retries — and it is usually where most of the work lives.

A language model does one thing: text in, text out. It cannot look anything up, cannot remember yesterday, cannot check whether it was right, and cannot do anything at all in the outside world.

Harness engineering is everything you build around the model to close that gap. The model is one component. The harness is the system that makes it useful.

Why the model is not the product

Picture hiring someone genuinely brilliant, then giving them no logins, no access to your systems, no notebook, and no colleague to check their work. On day two they remember nothing about day one. They would still be brilliant, and still unable to do the job.

That is a raw model. Everything that turns it into something a user can rely on lives outside it:

  • reaching real data and real systems
  • remembering what was said earlier
  • coping when something fails
  • following instructions the same way every time
  • knowing whether the output was any good
  • seeing what is happening once real users arrive

A better model raises the ceiling. The harness decides how much of that ceiling you actually reach.

The parts of a harness

Most harnesses end up with the same seven pieces.

partwhat it does
PromptsAssembles the system instructions, templates and context sent on each call
ToolsDeclares what the model may call, and actually runs it when it asks
MemoryKeeps conversation history, and compresses it before it overruns the context
ErrorsHandles a failed tool, a malformed response, a timeout, a rate limit
Input and outputValidates what goes in, and parses and checks what comes back
GuardrailsBlocks unsafe or out-of-scope requests and responses
ObservabilityLogs, traces and metrics, so a production problem is diagnosable

None of this is model work. All of it is engineering.

The loop that makes an agent

An agent is a harness running a loop.

The model never executes anything itself. It reads what it has been given and says what it would like to happen next. Everything else is the harness's job:

build prompt model tool call? run the tool append result yes next turn answer no
The model only ever answers. Every other step is code you write.

Take a support agent asked: has order 4172 shipped, and if not, tell the customer when it will?

The harness assembles the prompt — the instructions, the available tools, the question — and sends it. The model replies asking for a tool:

lookup_order(id=4172)

The harness runs that against the real system, gets back status: packing, ships: 9 September, appends it to the conversation, and calls the model again. Now the model asks for the second tool:

send_message(to="customer", text="Your order is packed and ships on 9 September.")

The harness runs that too, appends the result, and calls the model once more. This time the model asks for nothing, so the loop ends and its final text is returned.

Three model calls, two tool executions, one answer. The model decided what to do. The harness did all of it.

Knowing whether it works

You cannot improve what you cannot measure, and "it seemed fine when I tried it" does not survive contact with real users.

An evaluation harness is a second harness whose job is grading the first. It holds a fixed set of test inputs, runs them through the system, and scores the results.

How you score depends on the task. When there is a right answer — a lookup, a classification, an extraction — compare against it directly and count. When there isn't — a summary, an explanation, a reply to a customer — the usual approach is to have another model grade the output against written criteria.

The point of fixing the test set is that you can change a prompt, swap a model, or rewrite a tool, and see whether things got better or worse instead of guessing.

What keeps it manageable

  • Keep the parts separate. Prompt assembly, tool execution, and memory should be independently replaceable. They change at different rates.
  • Log everything. Every prompt, every tool call, every response. When something goes wrong in production, the log is all you have.
  • Add guardrails early. Retrofitting safety onto a system that already has users is far harder than building it in.
  • Assume tools fail. Timeouts, bad input, rate limits. Decide what happens on each, rather than letting an exception escape.
  • Test the harness itself. It is ordinary code with ordinary bugs, and most production incidents come from it rather than from the model.
  • Watch it after launch. Real traffic finds inputs your test set never had.

The short version

  • A model only maps text to text. Everything else is the harness.
  • The harness handles prompts, tools, memory, errors, I/O, guardrails and observability.
  • An agent is the harness looping: call the model, run whatever tool it asks for, feed the result back, repeat until it stops asking.
  • An evaluation harness with a fixed test set is how you tell whether a change helped.
  • The harness is ordinary software, and it is usually most of the code.