Intelligence
Has a
Habitat
Same weights. Different environment. Different operator.
The
Filing
Signal + NoiseField Report
Oct · 2026
Same weights,
different operator
Back in May, we found three independent engineering teams converging on the same observation: the model and the harness together determine agent quality.
One team was already seeing the same model become faster, more accurate, and more efficient inside a tuned operating environment. The pattern was strong enough that I wrote: expect “harness engineering” to emerge as a named discipline within twelve months.
Five months later, there’s enough evidence to sharpen the call.
This is no longer a collection of anecdotes about better prompting or agent scaffolds. The environment is becoming part of the capability.
Signal + NoiseField Report
Oct · 2026
Five months
later
On SWE-bench Pro, the same GLM-5.2 weights scored 23% in one harness and 52% in another. A separate controlled study took GLM-5.1, GPT-5.4, and Kimi K2.6 and held the tasks constant while changing the harness. Each of the three models won under a different harness configuration.
And when models leave environments they’ve learned, some don’t merely get worse. They break. In the Orchard research, OpenSWE-32B scored zero on Terminal-Bench 2.0. Another system stopped producing valid tool calls outside its home harness entirely.
Meanwhile Poolside, Kimi, Qwen, Xiaomi, and Liquid AI are deliberately varying harnesses during training.
Swapping the harness moved scores more than swapping the model did.
Signal + NoiseField Report
Oct · 2026
Your benchmark has an environment attached
A harness is the software between a model and the task: the loop that manages context, exposes tools, executes calls, handles results, retries failures, and decides when the work is done.
For a chatbot, some of that can feel like plumbing. For an agent, it is the workplace.
Change the tool names and schemas. Change how history gets compressed. Change what happens after an error. Change who owns retries or stopping. You have changed the conditions under which the model operates.
The KwaiKAT team describes three forms of harness overfitting: action format, context structure, and control flow. The practical version is easier to see.
Writewrite.file_pathfilePath.The reasoning can be right and the task can still fail.
This also changes how we should read agent benchmarks. A score isn’t simply a property of the weights. It is a result produced by the model operating inside a particular system. That doesn’t make benchmarks useless. It makes the harness part of the disclosure.
If you’re evaluating an agent for production, “Which model scored highest?” is only the start. You need to know which model performs best inside the environment you’re actually going to run.
The model can learn the workplace instead of learning the work.
Signal + NoiseField Report
Oct · 2026
Model-agnostic doesn’t mean model-indifferent
I’ve argued for model-agnostic architecture for months. Own the orchestration. Rent the inference. Don’t build a company that has to be rebuilt every time the frontier changes. I still believe that.
But model-agnostic architecture does not mean model-agnostic behavior. Different models may need different context strategies, tool descriptions, prompts, retry logic, or execution patterns. The architecture should make those differences manageable rather than pretending they don’t exist.
Don’t lock the product to one provider. Own the orchestration and rent the inference, so a frontier change is a swap, not a rebuild.
Give each model the context strategy, tool descriptions, prompts, and retry logic it needs. Models aren’t interchangeable workers.
This is already showing up in training. These teams aren’t treating the environment as an implementation detail after the model is trained. They’re putting variation in the environment into the training process itself.
- Who is training across harnesses
- Poolside included 1.3 billion tokens of trajectories from OpenHands, OpenCode, and Mini-SWE-Agent in Laguna’s supervised training.
- Qwen3-Coder-Next generates agentic coding data across six harnesses.
- Kimi K3 trains against varying harness configurations, including versions modeled on Claude Code and Codex.
- Xiaomi uses a pool of task-adapted mini-harnesses.
- Liquid AI varies the harness during training for LFM2.5-2.6B.
Signal + NoiseField Report
Oct · 2026
Now the loop
can close
There has been an awkward engineering problem hiding underneath this shift. The harnesses people actually use weren’t designed as reinforcement-learning environments. Claude Code owns its loop. Codex owns its loop. OpenCode owns its loop. They manage their own tools, context, retries, compaction, and stopping behavior.
A traditional RL trainer wants to own that interaction itself. Reimplementing a production harness as a training environment solves the technical problem by creating another one: you may end up training the model inside a scaffold nobody actually ships.
Recent work from Hugging Face, Liquid AI, and collaborators takes a different route. Don’t rewrite the harness. Intercept the model boundary. Their OpenEnv implementation places a capture proxy where the harness expects to find a model API. The harness keeps operating normally while the proxy records the model interactions needed for training. Claude Code can remain Claude Code. Codex can remain Codex. OpenCode can remain OpenCode.
The environment you deploy can become the environment you learn from.
The team’s experiment is small—LFM2.5-2.6B, 1,000 training steps, one seed—and the authors are appropriately careful about the limits. But the pattern is useful.
| RL run | Where the gains landed | Overall |
|---|---|---|
| OpenCode only | OpenCode score rose from roughly 34% to 58%. Most of the gain concentrated in the environment where it trained. | ~52% |
| OpenCode + Claude Code + Codex + Mini-SWE-Agent | Gains spread across all four harnesses. | ~54% |
The authors explicitly say the gap between the two overall scores is within the noise.
The interesting part isn’t a two-point “win.” It’s where the learning traveled.
Signal + NoiseField Report
Oct · 2026
The eval becomes
the teacher
This is the piece I’ve changed my mind about most since we started tracking harnesses. In May, I wrote that the team that cracks scalable, domain-specific agent evaluation may have built something more valuable than the agent itself.
At the time, I was thinking primarily about operating the system. If an agent is going to touch financial records, production code, customer communication, or any other consequential workflow, somebody needs to know whether the work actually succeeded. That was the basis for the self-monitoring loops, quality gates, and completion criteria we were seeing in the field.
Now there is another reason to own the eval. A production run leaves a trajectory: what the model tried, which tools it chose, where it failed, whether it recovered, how long it kept going, and what state it left the system in. If you can reliably evaluate the end state, that trajectory can become more than an audit trail. It can become a training signal.
The verifier determines whether the work succeeded. The reward distinguishes useful behavior from less useful behavior. Training changes what the model is likely to do next time. That turns an operating loop into a learning loop:
- Work
- Trajectory
- Evaluation
- Reward
- Training
- Changed behavior
You don’t need reinforcement learning to build a good agent product. Most startups probably shouldn’t be training models yet.
If you don’t know whether the work succeeded, you don’t have a useful eval. And without the eval, there is no meaningful learning loop to close later.
Signal + NoiseField Report
Oct · 2026
What you reward shows up in the work
The multi-harness experiment has a small detail I think is more consequential than it looks. Earlier training runs rewarded correctness. Solve the task and get the reward. The problem was what happened after the model knew enough to solve it: nothing told the agent that continuing to explore was worse than stopping.
For the LFM experiments, the researchers added a small efficiency bonus for correct solutions that used fewer tool calls. Wrong answers still received zero. Reductions showed up across every harness, with roughly half as many calls under Codex.
The researchers didn’t run an LFM control without the efficiency bonus, so we can’t cleanly attribute the entire reduction to that reward. They don’t claim otherwise. But the mechanism exposes an important design question. What exactly counts as success?
An agent will inherit whatever definition the system can actually measure. Which brings us back to economics.
Signal + NoiseField Report
Oct · 2026
The harness changes the price of done
In September, I argued that price per token was becoming the wrong economic unit for agents. The useful number is cost per verified success. This research adds another variable to that equation.
If the same weights solve 23% of a benchmark in one harness and 52% in another, those aren’t just different leaderboard numbers. The system may have radically different economics depending on where the model runs. Same inference price. Different amount of useful work.
Then add tool efficiency. A model that reaches the same successful outcome with fewer calls sends less history back through the model, consumes fewer tokens, creates fewer opportunities for failure, and usually finishes faster.
Affects the completion rate.
Can affect the work required to reach completion.
Both end up in the price of done. This is why model pricing by itself is such a weak way to choose an agent stack. A cheaper model that repeatedly stalls, retries, or leaves work unfinished can be more expensive than a model with a higher rate card.
The unit is the work. And the work has to be verified.
Signal + NoiseField Report
Oct · 2026
The moat is getting wider than the harness
In April, I called the harness the moat. I wouldn’t retract that. I would make it more precise.
The harness is where a company can encode the parts of the system that don’t come from a general-purpose model: domain context, tool access, permissions, workflow logic, completion criteria, verification, recovery.
But once the system begins capturing what happens during the work, another asset appears. The trajectories.
- What a trajectory records
- What worked. What failed.
- Which context mattered.
- Which sequence of actions succeeded.
- Where a human intervened, and which correction fixed the run.
- Whether the final state matched what the agent claimed.
A pile of traces isn’t a moat. Neither is a pile of proprietary data by default. The interesting part is whether you can turn those traces into better performance on work customers care about. That’s where the old harness argument starts becoming a learning-system argument.
Signal + NoiseField Report
Oct · 2026
Three decisions
follow from this now
Signal + NoiseField Report
Oct · 2026
The evidence
is still early
Most of the cleanest work is in coding, where tasks are relatively easy to sandbox and grade. The Hugging Face/Liquid AI experiment uses a 2.6B-parameter model, a single seed, and unequal training exposure between its single- and multi-harness runs. Its authors are already preparing larger experiments.
Signal + NoiseField Report
Oct · 2026
The Close
Five months ago, I wrote that I expected harness engineering to emerge as a named discipline within twelve months. The terminology doesn’t matter much. The engineering does.
The field is moving from measuring models inside harnesses, to deliberately varying harnesses during training, to building infrastructure that can learn from the environments people actually use. The harness is still where the company defines the work. What’s changing is that the work can now produce the trajectories, evaluations, and rewards that teach a model how to perform it better.
Intelligence has a habitat. We’re starting to learn how much it matters.
See exactly how this impacts your specific industry and function. Upgrade to PRO to get bespoke tactical breakdowns generated instantly for your operating model.

