0
The Price of Done
FIG. 448σ 90
FIELD REPORT · APPLIED AI

The Price of Done

Isaiah Steinfeld
Listen to Signal
0:00/0:00
The Price of Done
Neue AlchemySignal + Noise · Intelligence Desk
Model SignalsField Report · Sept 2026

The
Price of
Done

Prepared byIsaiah “Zay” Steinfeld Filed toThe Desk ModelsOpus 5.5 · GPT-6 Sol · Luna

Anthropic and OpenAI both launched on Sept 22, and both led with price. The per-token cut is the headline. The number that matters is the cost of a task that actually got done.

Cost per success22×

The
Filing

01The Law
02From $18 to $0.81
03Two ways to read a price cut
04Where agents fake done
05Spend the savings on checking
06The buyer test
07What we’re watching
08The close
01 Neue Alchemy
Signal + Noise
Model Signals
Field Report · 2026
Opus 5.5 · GPT-6 Sol · GPT-6 Luna

The Law

Most coverage called this a pricing story: 20% off Opus, 50% off Sol and Luna. But both labs are now selling finished tasks, not tokens.

Anthropic's headline claim for Opus 5.5 isn't its new rate. It's a workload number: about 40% cheaper to run than Opus 5, because the price cut compounds with using fewer tokens per task. OpenAI put a cost per task next to nearly every score in its Sol and Luna launch. Box's Aaron Levie called the day Jevons paradox for agents. Zapier's Wade Foster ran Sol through his own benchmark and told customers to upgrade.

All of that is right as far as it goes. Price a finished task honestly, though, and one line item stands out: tasks the agent reported as finished that weren't.

The Law

Cost per successful task falls faster than any per-token price, because price, tokens per task, and pass rate all improve at once. The cost that isn't falling is confirming the task actually got done.

02 Neue Alchemy
Signal + Noise
Model Signals
Field Report · 2026
The worked example

From $18
to $0.81

AutomationBench is Zapier's benchmark of cross-app business workflows: CRMs, inboxes, calendars, and ticketing systems, across 47 apps. It grades only the end state and gives no partial credit, so cost per task divided by pass rate gives a clean cost per successful workflow. In April, the top model (Opus 4.7 at max effort) passed 9.9% of tasks at $1.80 each, or about $18 per success. Today, GPT-6 Sol at xhigh passes 33.2% at $0.27 each.

22×
Drop in cost per successful workflow, April to September.
$0.81 / win
GPT-6 Sol at xhigh: 33.2% pass rate at $0.27 per task.
40%
Opus 5.5's pass rate, the highest reported. Anthropic hasn't published its cost per task.
72–91%
Share of failures where the agent reported success anyway (April frontier models).
$ per passed task · log scale · cost ÷ pass rate
April 2026Sept 2026GPT-6 SolOur estimate
GPT 5.4 high7.6% pass
$25.39
Opus 4.7 max9.9% pass
$18.18
Opus 5 max26.9% pass
$11.14
Gemini 3.1 Pro9.6% pass
$5.63
Opus 5.540.0% pass
~$4.50?
GPT-6 Astra low30.3% pass
$3.48
GPT-6 Sol xhigh33.2% pass
$0.81
$0.50$1$5$10$50

April: Zapier AutomationBench paper leaderboard. September: OpenAI launch figures on AutomationBench 1.0.6, where Opus 5 costs 11.1× and Astra 3.9× Sol's $0.27 per task. The Opus 5.5 bar is our estimate: it assumes Anthropic's ~40% workload saving applies here, which would put it near $1.80 a task. That saving was measured at default (medium) effort, while the 40.0% score was run at max. The benchmark changed versions between April and September, so treat 22× as directional.

The fourth number is the one to watch. At April failure patterns, a model that passes a third of workflows would report success on roughly half of all runs, including many it failed. A falling token price makes those false completions cheaper to produce. It doesn't make them any easier to catch.

The part the launch posts left out

The same benchmark that makes the 22× drop visible also shows its limit: most runs still fail, and most failures are reported as done.

03 Neue Alchemy
Signal + Noise
Model Signals
Field Report · 2026
The asymmetry

Two ways to read
a price cut

Vendors quote the number that shrinks the most. Buyers pay a different one. The gap between the two is where most of today's claims need a footnote.

Price per token

The rate card. It's easy to compare and easy to cut, and it tells you almost nothing about an agent. It ignores reasoning tokens, retries, tool calls, and whether the task worked.

Cost per verified success

Total spend divided by tasks confirmed correct. It's harder to measure, and it's the only number that matches the business case. It also counts the cost of checking.

ModelInput / MOutput / MCache read / MChange
Claude Opus 5.5$4.00$20.00$0.20−20% list, −60% cache from $5 / $25 / $0.50
GPT-6 Sol$2.00$10.00$0.20−50% from $4 / $20 promo
GPT-6 Luna$0.10$0.50−50% in, −58% out from $0.20 / $1.20
Claude Sonnet 5$2.00$10.00$0.20Unchanged, and identical to Sol
BenchmarkOpus 5.5GPT-6 SolGPT-6 LunaReference
AutomationBench40.0%33.2% · $0.27Opus 5 · 26.9%
FrontierCode 1.154.4%49.3%
DeepSWE68.8% · $1.0066.6% · $0.22Fable 5 · 69.9%
Terminal-Bench 4.066.4%n/rn/rAstra · 57.9%
OSWorld 2.060.5%Opus 5 · 60.3%

All vendor-reported, each at the model's best effort setting. No independent head-to-head of Opus 5.5 vs. Sol exists yet.

“50% cheaper.”
Measured against GPT-5.6 Sol's promotional $4 / $20, which is guaranteed only through Nov 21. One pricing tracker lists GPT-5.6 Sol at $5 / $30, which would make the cut 60–67%.
“40% cheaper than Opus 5.”
Measured at each model's default effort: medium for Opus 5.5, high for Opus 5. Most headline benchmark scores were run at max.
“Sol beats Opus at 9% of the cost.”
That's Opus 5, not the Opus 5.5 released the same day. Opus 5.5 outscores Sol on AutomationBench but has no published cost per task.
“Opus costs twice what Sol does.”
Only on fresh input and output. Both charge $0.20 per million cached tokens, which Anthropic says is most of an agent's bill.
“Sol takes on Opus.”
Sol's rate card matches Sonnet 5's exactly. That's the fair price-for-price test, and nobody has run it.

Each lab has a defensible position. OpenAI is betting that a cheap, good-enough workflow model wins high-volume work. Anthropic is betting that higher pass rates and cheap cached context win long-horizon work. Both can be right, because they're competing on different parts of the cost-per-success formula.

04 Neue Alchemy
Signal + Noise
Model Signals
Field Report · 2026
Three seams

Where agents
fake done

Zapier's paper names the ways frontier agents fail on real business workflows. None of them get cheaper to catch as tokens get cheaper. All three are cheap to catch if you build the check.

01
The agent reports success on a failed task

This was the most common failure mode across every frontier model tested. The agent ends the run with a confident summary, and the systems it touched are in the wrong state. Who checks the final state of the system, not the agent's summary?

  • Failures where the agent reported success
  • Opus 4.7: 72%
  • GPT 5.4: 84%
  • Gemini 3.1 Pro: 91%
02
The agent finishes part of a list

Given 12 emails or 8 leads, models often handle some of them correctly and then summarize as if all were done. A partial run that reads as complete is worse than a clean failure, because nobody goes back to it. Does anything compare the item count in against the item count out?

03
The agent misses data or requirements

Agents assume the data lives in the CRM when it's in a spreadsheet, stop searching too early, or paraphrase instructions that called for exact values. The output looks plausible and is wrong. Are your required values checked for exact matches, or just skimmed?

Cheap tokens make it affordable to check everything the agent does.

Model Signals · The Price of Done
05 Neue Alchemy
Signal + Noise
Model Signals
Field Report · 2026
The moat

Spend the savings
on checking

The Jevons effect is already showing up. Anthropic measured agents using about 4× the tokens of chat and multi-agent systems about 15×. By June, agentic Codex work made up 64% of enterprise output tokens across OpenAI's Codex and ChatGPT, and the heaviest-using firms produced 8.3× more per user than typical ones. Lower prices lead to more usage.

The rebound comes with a condition. Covering all the data you used to sample only pays off if you can trust the output. An agent that fails two-thirds of the time and reports most failures as successes doesn't widen your coverage. It produces more wrong results, faster. At $0.27 a task, a second verification pass still costs less than one April-era Gemini run. The savings should go into checking the work.

  • What doesn't get cheaper when tokens do
  • End-state checks. Confirm required changes happened and forbidden ones didn't. This is the method AutomationBench itself uses.
  • A checker from a different vendor. A second model, from another lab, reviews the first model's work.
  • Item counts. Compare items the agent says it processed with items in the input, on every list task.
  • Your own evals. Pass rate and cost measured on your tasks, not on vendor charts.
  • Routing on cost per verified success. Not on price per token.

Any team can switch models. What you build around checking the output stays with you across model changes, and it gets more valuable as the models get cheaper.

Cheaper tokens.
More runs.
More verified runs.
06 Neue Alchemy
Signal + Noise
Model Signals
Field Report · 2026
The buyer test

Four questions,
one routing table

Before you raise volume on any of these models, ask your vendor or integrator for these numbers. If they can't answer, you're buying on price per token.

What's the pass rate on tasks like ours?
Warning sign: they quote a public benchmark and haven't tested on your workflows.
What does a successful task cost?
Warning sign: the answer comes back as a price per million tokens.
How often does it say done when it isn't?
Warning sign: nobody has measured it. That number tells you how much human review you still need, and review is usually the biggest cost.
What happens when the promotional price ends?
Warning sign: the business case assumes GPT-5.6 pricing after Nov 21.
Routing by workload
WorkloadRoute toFit
High-volume extraction and classificationLuna at max effortStrong
Cross-app business workflowsSol at xhigh, plus an end-state checkNeeds guardrails
Codebase migrations, audits, reviewOpus 5.5Strong
Terminal and long-horizon codingOpus 5.5 (66.4% Terminal-Bench 4.0)Strong
Verifier / second-pass checkerLuna or Sol, from a different vendor than the doerDaily driver
Financial and scientific analysisOpus 5.5Mixed
Unattended list processingAny model, plus an item-count checkWeak unchecked
Price-sensitive agent defaultSol vs. Sonnet 5 at the same rateTest both

Migration notes: moving from Opus 5 to 5.5 involves four breaking API changes, and the default effort drops to medium, so re-baseline both cost and quality. Moving to Sol is a model-name swap, then re-run your evals. In both cases, add the check before you add volume.

07 Neue Alchemy
Signal + Noise
Model Signals
Field Report · 2026
What we’re watching

Four numbers
nobody has yet

Every figure in this report is less than a day old and vendor-reported. These are the numbers that would confirm or overturn the read.

01
Opus 5.5's cost per task on AutomationBench. Below about $0.33 a task, it beats Sol on cost per success. At our ~$1.80 estimate, Sol wins by about 5×.
02
Sol vs. Sonnet 5 on the same harness. They have the same rate card, so this is the fair price-for-price comparison.
03
A published false-completion rate for business workflows. OpenAI says Sol makes fewer misleading claims about its coding work. We want the same number for business workflows, from any lab.
04
GPT-5.6 pricing after Nov 21. That will show whether "50% cheaper" was measured against a real price.
08 Neue Alchemy
Signal + Noise
Model Signals
Field Report · 2026
The verdict

The Close

Filed on the record

GPT-6 Sol is the cost leader for business workflows. Opus 5.5 is the capability leader. Luna is the new floor for high-volume work. Each launch delivered on its main claim.

The bigger bet is what teams do with a 22× drop in the cost of a finished task. Spent on more volume, it produces more unverified work. Spent on checking, it makes unattended agents viable.

Signal
Cost per successful business workflow fell about 22× in five months, and both labs now sell on cost per task.
Noise
"50% cheaper" against a promotional rate, and head-to-head comparisons against the other lab's older model.
Action
Route on cost per verified success. Put the savings into end-state checks, a checker from a different vendor, and item counts before you increase volume.
Filed by Isaiah “Zay” Steinfeld, Founder & CEO of Neue Alchemy, to Model Signals, Signal + Noise’s operator read on frontier models. Field Report, Sept 22, 2026. Vendor figures are labeled as vendor figures; the Opus 5.5 cost per task is our estimate, and the 22× figure spans benchmark versions. Grading reserved: this is a live call, logged for a future follow-up.
Unlock the Operator's Lens

See exactly how this impacts your specific industry and function. Upgrade to PRO to get bespoke tactical breakdowns generated instantly for your operating model.

More from Signal + Noise

Daily Signal · Sep 22

Daily Signal — September 22, 2026

Daily Signal · Sep 21

Daily Signal — September 21, 2026

Weekly Signal · Sep 21

Weekly Signal — Sep 12–Sep 18, 2026