0
Intelligence Has a Habitat
FIG. 476σ 90
FIELD REPORT · APPLIED AI

Intelligence Has a Habitat

Field Report by Isaiah "Zay" Steinfeld (Neue Alchemy) on harness engineering: the same model weights perform very differently across agent harnesses, frontier teams now train across harnesses, and the eval is becoming the teacher.

Isaiah Steinfeld
Listen to Signal
0:00/0:00
Intelligence Has a Habitat
Neue AlchemySignal + Noise · Intelligence Desk
Field ReportHarness Engineering · Oct 2026

Intelligence
Has a
Habitat

Prepared byIsaiah “Zay” Steinfeld Filed toThe Desk FollowsThe Price of Done

Same weights. Different environment. Different operator.

GLM-5.2 · Harness A 23%
GLM-5.2 · Harness B 52%

The
Filing

01Same weights, different operator
02Five months later
03Your benchmark has an environment attached
04Model-agnostic doesn’t mean model-indifferent
05Now the loop can close
06The eval becomes the teacher
07What you reward shows up in the work
08The harness changes the price of done
09The moat is getting wider than the harness
10The Operator’s Lens
11What we’re watching
12The Close
01Neue Alchemy
Signal + Noise
Field Report
Oct · 2026
Where we left off

Same weights,
different operator

Back in May, we found three independent engineering teams converging on the same observation: the model and the harness together determine agent quality.

One team was already seeing the same model become faster, more accurate, and more efficient inside a tuned operating environment. The pattern was strong enough that I wrote: expect “harness engineering” to emerge as a named discipline within twelve months.

Five months later, there’s enough evidence to sharpen the call.

The call, sharpened

This is no longer a collection of anecdotes about better prompting or agent scaffolds. The environment is becoming part of the capability.

02Neue Alchemy
Signal + Noise
Field Report
Oct · 2026
The evidence

Five months
later

On SWE-bench Pro, the same GLM-5.2 weights scored 23% in one harness and 52% in another. A separate controlled study took GLM-5.1, GPT-5.4, and Kimi K2.6 and held the tasks constant while changing the harness. Each of the three models won under a different harness configuration.

23→52%
GLM-5.2 on SWE-bench Pro. Same weights, two harnesses.
13 pts
Swing from moving GLM-5.1 between harness configurations.
2.5–5 pts
Swing from swapping models with the harness held fixed.
62.4→3.6%
OpenSWE-32B, from OpenHands to Kimi-CLI, a harness it hadn’t seen in training.

And when models leave environments they’ve learned, some don’t merely get worse. They break. In the Orchard research, OpenSWE-32B scored zero on Terminal-Bench 2.0. Another system stopped producing valid tool calls outside its home harness entirely.

Meanwhile Poolside, Kimi, Qwen, Xiaomi, and Liquid AI are deliberately varying harnesses during training.

The read

Swapping the harness moved scores more than swapping the model did.

03Neue Alchemy
Signal + Noise
Field Report
Oct · 2026
Benchmarks

Your benchmark has an environment attached

A harness is the software between a model and the task: the loop that manages context, exposes tools, executes calls, handles results, retries failures, and decides when the work is done.

For a chatbot, some of that can feel like plumbing. For an agent, it is the workplace.

Change the tool names and schemas. Change how history gets compressed. Change what happens after an error. Change who owns retries or stopping. You have changed the conditions under which the model operates.

The KwaiKAT team describes three forms of harness overfitting: action format, context structure, and control flow. The practical version is easier to see.

A model learns to call Write
The next harness exposes write.
One expects file_path
Another expects filePath.
One harness automatically retries a failed command
Another hands the error back and waits for the model to decide.

The reasoning can be right and the task can still fail.

This also changes how we should read agent benchmarks. A score isn’t simply a property of the weights. It is a result produced by the model operating inside a particular system. That doesn’t make benchmarks useless. It makes the harness part of the disclosure.

If you’re evaluating an agent for production, “Which model scored highest?” is only the start. You need to know which model performs best inside the environment you’re actually going to run.

The model can learn the workplace instead of learning the work.

Field Report · Intelligence Has a Habitat
04Neue Alchemy
Signal + Noise
Field Report
Oct · 2026
Architecture

Model-agnostic doesn’t mean model-indifferent

I’ve argued for model-agnostic architecture for months. Own the orchestration. Rent the inference. Don’t build a company that has to be rebuilt every time the frontier changes. I still believe that.

But model-agnostic architecture does not mean model-agnostic behavior. Different models may need different context strategies, tool descriptions, prompts, retry logic, or execution patterns. The architecture should make those differences manageable rather than pretending they don’t exist.

Build for model optionality

Don’t lock the product to one provider. Own the orchestration and rent the inference, so a frontier change is a swap, not a rebuild.

Engineer for model specificity

Give each model the context strategy, tool descriptions, prompts, and retry logic it needs. Models aren’t interchangeable workers.

This is already showing up in training. These teams aren’t treating the environment as an implementation detail after the model is trained. They’re putting variation in the environment into the training process itself.

  • Who is training across harnesses
  • Poolside included 1.3 billion tokens of trajectories from OpenHands, OpenCode, and Mini-SWE-Agent in Laguna’s supervised training.
  • Qwen3-Coder-Next generates agentic coding data across six harnesses.
  • Kimi K3 trains against varying harness configurations, including versions modeled on Claude Code and Codex.
  • Xiaomi uses a pool of task-adapted mini-harnesses.
  • Liquid AI varies the harness during training for LFM2.5-2.6B.
05Neue Alchemy
Signal + Noise
Field Report
Oct · 2026
Infrastructure

Now the loop
can close

There has been an awkward engineering problem hiding underneath this shift. The harnesses people actually use weren’t designed as reinforcement-learning environments. Claude Code owns its loop. Codex owns its loop. OpenCode owns its loop. They manage their own tools, context, retries, compaction, and stopping behavior.

A traditional RL trainer wants to own that interaction itself. Reimplementing a production harness as a training environment solves the technical problem by creating another one: you may end up training the model inside a scaffold nobody actually ships.

Recent work from Hugging Face, Liquid AI, and collaborators takes a different route. Don’t rewrite the harness. Intercept the model boundary. Their OpenEnv implementation places a capture proxy where the harness expects to find a model API. The harness keeps operating normally while the proxy records the model interactions needed for training. Claude Code can remain Claude Code. Codex can remain Codex. OpenCode can remain OpenCode.

What that makes possible

The environment you deploy can become the environment you learn from.

The team’s experiment is small—LFM2.5-2.6B, 1,000 training steps, one seed—and the authors are appropriately careful about the limits. But the pattern is useful.

RL runWhere the gains landedOverall
OpenCode onlyOpenCode score rose from roughly 34% to 58%. Most of the gain concentrated in the environment where it trained.~52%
OpenCode + Claude Code + Codex + Mini-SWE-AgentGains spread across all four harnesses.~54%

The authors explicitly say the gap between the two overall scores is within the noise.

The interesting part isn’t a two-point “win.” It’s where the learning traveled.

06Neue Alchemy
Signal + Noise
Field Report
Oct · 2026
What changed my mind

The eval becomes
the teacher

This is the piece I’ve changed my mind about most since we started tracking harnesses. In May, I wrote that the team that cracks scalable, domain-specific agent evaluation may have built something more valuable than the agent itself.

At the time, I was thinking primarily about operating the system. If an agent is going to touch financial records, production code, customer communication, or any other consequential workflow, somebody needs to know whether the work actually succeeded. That was the basis for the self-monitoring loops, quality gates, and completion criteria we were seeing in the field.

Now there is another reason to own the eval. A production run leaves a trajectory: what the model tried, which tools it chose, where it failed, whether it recovered, how long it kept going, and what state it left the system in. If you can reliably evaluate the end state, that trajectory can become more than an audit trail. It can become a training signal.

The verifier determines whether the work succeeded. The reward distinguishes useful behavior from less useful behavior. Training changes what the model is likely to do next time. That turns an operating loop into a learning loop:

  1. Work
  2. Trajectory
  3. Evaluation
  4. Reward
  5. Training
  6. Changed behavior

You don’t need reinforcement learning to build a good agent product. Most startups probably shouldn’t be training models yet.

The prerequisite

If you don’t know whether the work succeeded, you don’t have a useful eval. And without the eval, there is no meaningful learning loop to close later.

07Neue Alchemy
Signal + Noise
Field Report
Oct · 2026
Reward design

What you reward shows up in the work

The multi-harness experiment has a small detail I think is more consequential than it looks. Earlier training runs rewarded correctness. Solve the task and get the reward. The problem was what happened after the model knew enough to solve it: nothing told the agent that continuing to explore was worse than stopping.

13→41
Tool calls per rollout in an earlier Qwen run rewarded on correctness alone.
−31%
Tool calls for the multi-harness LFM model vs. base, on tasks both solved, by step 1,000.

For the LFM experiments, the researchers added a small efficiency bonus for correct solutions that used fewer tool calls. Wrong answers still received zero. Reductions showed up across every harness, with roughly half as many calls under Codex.

The researchers didn’t run an LFM control without the efficiency bonus, so we can’t cleanly attribute the entire reduction to that reward. They don’t claim otherwise. But the mechanism exposes an important design question. What exactly counts as success?

Correct eventually?
Correct within ten steps?
Correct without exceeding authority?
Correct with a verified end state?
Correct at an acceptable cost?

An agent will inherit whatever definition the system can actually measure. Which brings us back to economics.

08Neue Alchemy
Signal + Noise
Field Report
Oct · 2026
Economics

The harness changes the price of done

In September, I argued that price per token was becoming the wrong economic unit for agents. The useful number is cost per verified success. This research adds another variable to that equation.

If the same weights solve 23% of a benchmark in one harness and 52% in another, those aren’t just different leaderboard numbers. The system may have radically different economics depending on where the model runs. Same inference price. Different amount of useful work.

Then add tool efficiency. A model that reaches the same successful outcome with fewer calls sends less history back through the model, consumes fewer tokens, creates fewer opportunities for failure, and usually finishes faster.

The harness

Affects the completion rate.

The training objective

Can affect the work required to reach completion.

Both end up in the price of done. This is why model pricing by itself is such a weak way to choose an agent stack. A cheaper model that repeatedly stalls, retries, or leaves work unfinished can be more expensive than a model with a higher rate card.

The unit

The unit is the work. And the work has to be verified.

09Neue Alchemy
Signal + Noise
Field Report
Oct · 2026
The moat

The moat is getting wider than the harness

In April, I called the harness the moat. I wouldn’t retract that. I would make it more precise.

The harness is where a company can encode the parts of the system that don’t come from a general-purpose model: domain context, tool access, permissions, workflow logic, completion criteria, verification, recovery.

But once the system begins capturing what happens during the work, another asset appears. The trajectories.

  • What a trajectory records
  • What worked. What failed.
  • Which context mattered.
  • Which sequence of actions succeeded.
  • Where a human intervened, and which correction fixed the run.
  • Whether the final state matched what the agent claimed.

A pile of traces isn’t a moat. Neither is a pile of proprietary data by default. The interesting part is whether you can turn those traces into better performance on work customers care about. That’s where the old harness argument starts becoming a learning-system argument.

The modelsupplies general capability.
The operating environmentshapes how that capability gets expressed.
The evaldefines what good means.
The trajectoriesrecord what happened.
Increasinglythose pieces can feed back into the intelligence doing the work.
10Neue Alchemy
Signal + Noise
Field Report
Oct · 2026
The Operator’s Lens

Three decisions
follow from this now

Build for model optionality. Engineer for model specificity.
Don’t lock the product unnecessarily to one model provider, but don’t design around the fiction that models are interchangeable workers. Run your own workload through the combinations you might actually deploy.
Own the eval before you worry about RL.
Define the end state of successful work, instrument it, and measure it. If your team can’t reliably tell whether the agent finished the job, training it against that workflow is premature.
Capture the work.
Tool calls, retries, failures, corrections, human interventions, and verified outcomes are becoming strategically different from ordinary application logs. You may not train against them today. Preserve the option.
11Neue Alchemy
Signal + Noise
Field Report
Oct · 2026
What we’re watching

The evidence
is still early

Most of the cleanest work is in coding, where tasks are relatively easy to sandbox and grade. The Hugging Face/Liquid AI experiment uses a 2.6B-parameter model, a single seed, and unequal training exposure between its single- and multi-harness runs. Its authors are already preparing larger experiments.

01
Whether the pattern holds for larger models and messier enterprise work, where “correct” isn’t always a test script returning 1.
02
How much harness diversity is enough. Training across every possible environment is neither practical nor necessary. The interesting question is which kinds of variation actually produce transferable behavior.
03
What happens when companies close this loop on proprietary workflows rather than public coding benchmarks. That is where this stops being an agent-training technique and becomes an operating advantage.
12Neue Alchemy
Signal + Noise
Field Report
Oct · 2026
The verdict

The Close

Filed on the record

Five months ago, I wrote that I expected harness engineering to emerge as a named discipline within twelve months. The terminology doesn’t matter much. The engineering does.

The field is moving from measuring models inside harnesses, to deliberately varying harnesses during training, to building infrastructure that can learn from the environments people actually use. The harness is still where the company defines the work. What’s changing is that the work can now produce the trajectories, evaluations, and rewards that teach a model how to perform it better.

The durable layer moves one step further out
From the model,
to the harness,
to the learning loop between the model and the environment.
Build for model optionality. Engineer for model specificity. Own the loop.

Intelligence has a habitat. We’re starting to learn how much it matters.

Signal
Harness choice can materially change observed model performance, models can overfit to the environments where they learn, and multiple frontier teams are now deliberately training across harnesses.
Noise
“Just use the best model.” Agent capability is increasingly a property of a model operating inside a particular environment, not a leaderboard number you can carry unchanged into production.
Field Note
The next durable advantage may not be access to a particular model or even the harness by itself. Watch the companies that can define successful work, capture what happened, verify the outcome, and feed what they learn back into the system.
Filed by Isaiah “Zay” Steinfeld, Founder & CEO of Neue Alchemy, to Signal + Noise. Field Report, October 2026. Follows the May harness reports and “The Price of Done” (September 2026). Research findings are characterized from the authors’ own published results and stated limits. Grading reserved: this is a live call, logged on the record for a future receipt.
Unlock the Operator's Lens

See exactly how this impacts your specific industry and function. Upgrade to PRO to get bespoke tactical breakdowns generated instantly for your operating model.

More from Signal + Noise

Daily Signal · Oct 6

Daily Signal — October 6, 2026

Daily Signal · Oct 5

Daily Signal — October 5, 2026

Weekly Signal · Oct 5

Weekly Signal — Sep 26–Oct 2, 2026