Crafting Excellence in Software
Let’s build something extraordinary together.
Rely on Lasting Dynamics for unparalleled software quality.
Mohamed Boukrim
Sep 23, 2026 • 12 min read

On 22 September 2026 Anthropic shipped Claude Opus 5.5 at $4 per million input tokens and $20 per million output, and the same day OpenAI shipped GPT-6 Sol at $2 and $10, and GPT-6 Luna at $0.10 and $0.50. The day before, xAI had put Grok 4.7 out at $2 and $6. Within a day the web was full of pages ranking them on AI benchmarks, and every one of them answered a question the launch material cannot answer, because the two vendor tables were never built to be read against each other.
Anthropic's table runs nine benchmarks across five columns, the OpenAI models in it are GPT-6 Astra and GPT-5.6 Sol, and a note says both sets of figures are as reported by OpenAI. OpenAI's page runs five benchmarks and closes with the mirror-image disclaimer that evaluations of competitor models were taken from publicly available reports. Neither company ran the other's model, so each set of AI benchmarks cites the other vendor's press release.
Three benchmarks appear in both tables, and on two of them the same model carries a different number. GPT-6 Astra scores 41.4% on AutomationBench in Anthropic's table and 30.3% on AutomationBench 1.0.6 in OpenAI's, where it was run at low effort, while Claude Opus 5 scores 74.0% on OSWorld 2.0 in Anthropic's table and 60.3% in OpenAI's, where it was run at medium effort on the offline set. Neither figure is wrong, because each measures a different configuration, and that distinction, not the AI benchmarks themselves, decides which model you should be paying for.
A score on any of the current AI benchmarks is the output of four choices, and changing any one of them moves the number. The task set decides what is being asked and whether the model has seen it in training, which is why LLM benchmarks saturate once every frontier model clears 90%. The harness decides which tools the model gets, how many attempts it has and how long it may run, so a different scaffold is a different test even when the task set is identical.
The third choice, and the one this article turns on, is effort. Anthropic's effort documentation lists five levels for Opus 5.5 (low, medium, high, xhigh and max), with medium as the default, whereas Opus 5 defaulted to high, and launch tables are almost always run at the top of that dial. Anthropic says so under its own table: unless otherwise noted, all Claude Opus 5.5 results use adaptive thinking at max effort. The number you read in AI benchmarks is rarely the number you will be billed for at default settings.
The fourth is scoring, and with it the error bar that almost nobody prints. Anthropic reports a standard error of plus or minus 2.6 points for Opus 5.5 on Terminal-Bench 4.0 and plus or minus 3.5 to 5 points per model on Terminal-Bench-Science 0.1, so a lead of one or two points inside that band is noise, and several of the leads circulating this week are exactly that size.
The AI benchmarks that matter this week fall into a few families. Terminal-Bench and CursorBench measure agentic terminal and editor work, the SWE-bench family and DeepSWE measure repository bug fixing, OSWorld measures computer use, Humanity's Last Exam measures expert knowledge, and GDPval scores economically valuable tasks as an Elo rating rather than a percentage. A score from any of these LLM benchmarks, quoted without its harness, effort level and version, is a fact about a configuration and not a fact about a model.
Within a single model, the effort dial spans a wider range than the gap between competing models. On the Artificial Analysis Intelligence Index, Opus 5.5 scores 42 at low effort for $0.55 per task, 51 at medium for $1.34, 54 at high for $1.82, 56 at xhigh for $3.46 and 58 at max for $5.98. That is sixteen index points and an elevenfold cost spread inside one model, while the spread between Opus 5.5 at medium and GPT-6 Sol at max is three points.
The vendor tables admit this if you read to the bottom. Anthropic's note says Terminal-Bench 4.0 results are reported for Claude Opus 5.5 at xhigh effort and GPT-6 Astra at high effort, as reported by OpenAI; these represent each model's highest score, which is best-of-each rather than a matched comparison of AI benchmarks.
| Benchmark | Model | Score | Effort level | Reported by |
|---|---|---|---|---|
| AutomationBench | GPT-6 Astra | 41.4% | not stated, Zapier leaderboard | Anthropic, 22 Sep 2026 |
| AutomationBench 1.0.6 | GPT-6 Astra | 30.3% | low | OpenAI, 22 Sep 2026 |
| AutomationBench | Claude Opus 5 | 26.9% | max | both, from Zapier's leaderboard |
| OSWorld 2.0 (partial) | Claude Opus 5 | 74.0% | max | Anthropic, 22 Sep 2026 |
| OSWorld 2.0 offline (partial) | Claude Opus 5 | 60.3% | medium | OpenAI, 22 Sep 2026 |
| Terminal-Bench 4.0 | Claude Opus 5.5 | 66.4% | xhigh | Anthropic, own harness |
| Terminal-Bench 4.0 | GPT-6 Astra | 57.9% | high | Anthropic, as reported by OpenAI |
The row where both vendors report Claude Opus 5 at 26.9% on AutomationBench is the only one that reconciles, because both took it from the same Zapier leaderboard at the same effort, and everywhere the effort or the reporter changes, the number changes with it. Anthropic's own launch page says Opus 5.5 at default effort beats Opus 5 at max effort on Terminal-Bench 4.0 for about a fifth of the cost per task, and its early tester Factory reports medium effort matching Opus 5 at high while using 20 to 25% fewer output tokens, so the effort names do not mean the same thing across two versions of one model, and any AI benchmarks comparison that ignores them is comparing labels.
Let’s build something extraordinary together.
Rely on Lasting Dynamics for unparalleled software quality.
This is also where the widely quoted 40% saving comes from. Anthropic's wording is that at default settings it will cost 40% less than Opus 5 on typical workloads, a claim about two defaults, combining a lower per-token price with fewer tokens per task. It is real and it is useful, but it is not a claim about equal work at equal effort, and a team that reads LLM benchmarks that way will budget wrong. If you need a method for choosing the effort level per task, it is already on this site.
The upgrade is priced as a discount and delivered as a migration, and both halves belong in the same calculation. On Anthropic's pricing table, input falls from $5 to $4 per million tokens, output from $25 to $20, five-minute cache writes from $6.25 to $5 and cache reads from $0.50 to $0.20. For agentic workloads, where Anthropic says cache reads make up the majority of the cost, the 60% cut on reads matters more than the 20% headline.
The gains in Anthropic's own table, all at max effort, land on agentic and long-horizon work rather than knowledge. Opus 5 to Opus 5.5 moves Terminal-Bench 4.0 from 52.3% to 66.4%, with the higher figure run at xhigh, and Terminal-Bench-Science 0.1 from 29.0% to 58.7%, while Humanity's Last Exam with tools moves only from 63.6% to 67.7%. The same page says, in Anthropic's words, that benchmark margins have become a less reliable guide to real-world differences, and a vendor hedging its own AI benchmarks is rare enough to quote.
The migration cost is documented in Anthropic's migration notes and it is not trivial. Thinking can no longer be disabled, and a request that tries returns a 400 error. Forced tool use through tool_choice set to any or a named tool returns a 400 error too, and so does the token-counting endpoint. Thinking blocks are bound to the conversation prefix, enforced by default for API accounts created on or after 31 August 2026, so an integration that rewrites the system prompt mid-conversation works on an older account and breaks on a new one.
The quiet one is the response shape. Text the model wrote between tool calls used to arrive as text blocks, and on Opus 5.5 it arrives as thinking blocks that are empty at the default display setting, so nothing errors while an application that streamed that narration to users goes silent between tool calls, and the tokens are still billed at full length. Budget the migration in engineer-days and set it against the monthly token saving, because for a low-volume integration the 40% on the price sheet does not pay for the migration this quarter.
OpenAI positions Sol below GPT-6 Astra, so the tier-for-tier rival of Opus 5.5 is Astra at $10 and $50, and most comparison pages skip that sentence. Sol and Opus 5.5 sit at different points on the cost-quality curve, which makes the useful question which of your tasks belong at which point rather than which model wins the AI benchmarks.
Artificial Analysis is the only source that ran both models on the same ten evaluations at every effort level during launch week, so its figures are the only matched ones available. Opus 5.5 at medium scores 51 for $1.34 per task, while GPT-6 Sol scores 48 at max for $1.06, 44 at xhigh for $0.53 and 40 at medium for $0.25. Within that pair Sol is the cheaper route to any index score up to about 48, and Opus 5.5 only becomes necessary above it.
Across the whole leaderboard the picture is less tidy, because MiMo-V2.6-Pro reaches 46 for $0.13 and Grok 4.7 reaches 46 for $2.73 at high effort, which is the kind of fact a two-vendor reading of the AI benchmarks structurally cannot show you.

The same source records improvement and regression in one release, which is normal and rarely reported. GPT-6 Sol at max cuts its AA-Omniscience hallucination rate from 92% to 60% against GPT-5.6 Sol, but that metric is the share of non-correct answers that were confident and wrong rather than abstentions, and Sol gets there by attempting 83% of questions instead of 99%, which costs it five points of accuracy. Meanwhile Sol drops about 100 Elo on GDPval-AA v2.1, and Artificial Analysis describes the overall index as level with GPT-5.6 rather than improved.
From idea to launch, we craft scalable software tailored to your business needs.
Partner with us to accelerate your growth.
Luna deserves a paragraph rather than a footnote. At $0.10 and $0.50 it scores 37 at max effort for $0.07 per task, and OpenAI's own page puts it at 66.6% on DeepSWE 1.1 at max, which OpenAI describes as comparable to Opus 5 and Fable 5 at medium at a fraction of the cost. For classification, extraction and triage the question is whether Luna passes your acceptance test, not where it sits on the AI benchmarks, and the expensive mistake this week is running every call at a flagship's max effort because that is where the launch table was run.
Independent evaluators fix the comparability problem that vendor tables create, and then introduce a subtler one. Artificial Analysis runs every model on the same ten evaluations at every effort level, Epoch AI maintains its own hub with different design choices, and LiveBench limits contamination by refreshing its questions on a schedule, so any of them is a better starting point for AI benchmarks than a launch page.
An index is a mean, and a mean averages away the one capability your workload depends on. Sol's Intelligence Index held level while its GDPval-AA score fell by about 100 Elo, so a team whose work looks like GDPval would have read the headline and picked wrong. The Coding Agent Index on the same site is a separate number again, and at the time of writing it has no entry for Opus 5.5, a reminder that independent AI benchmarks lag launches by days or weeks.
Stanford HAI's framework on benchmark quality is the external reference for what separates a trustworthy benchmark from a marketing one, and the criteria it sets out, from documented methodology to statistical reporting and contamination control, are exactly the things the AI benchmarks in a launch table omit.
Price per token is an input, not a cost. Models differ in how many tokens they spend on a task, how often they fail it and how much human review their output needs, and only cost per completed, accepted task folds all three into one number that means the same thing before and after a switch, which is more than any of the public AI benchmarks can claim.
The token side is measurable and it moves. Artificial Analysis reports GPT-6 Sol using about 31,000 output tokens per index task against 29,000 for GPT-5.6 Sol, and Luna 51,000 against 41,000, so part of the price cut buys longer answers. CodeRabbit, testing Opus 5.5 in its production review pipeline on launch day, saw token usage 40% to 60% above its baseline in every configuration despite the lower price, and concluded that teams should verify whether the cheaper token produces a cheaper review.
The failure side compounds, because a failed attempt costs its own tokens, the retry's tokens and the time of the engineer who noticed it, and the review side is the part that does not fall when models improve, which is the argument this site already made about review time that does not shrink as models improve. The formula is short enough to hold in your head, and it is worth more than every line in the published AI benchmarks: total spend across all attempts, plus review minutes at a loaded rate, divided by the number of tasks a reviewer accepted.
The evaluation that settles the question is small, cheap and already sitting in your git history. Pick twenty merged pull requests from the last quarter that each had a clear acceptance test, spread across difficulty levels, with at least five that a junior engineer found hard. For each one, check out the parent commit, hand the model the original ticket text, and run the same harness for every candidate with the same tools, timeout and maximum attempts.
Run each candidate at two or three effort levels rather than one, because the effort curve is the finding, and no vendor publishing AI benchmarks will run that curve on your code.
We design and build high-quality digital products that stand out.
Reliability, performance, and innovation at every step.
Record six things per task and nothing else: pass or fail on the original tests, reviewer accept or reject, tokens in and out, dollar cost, wall-clock time, and review minutes. Cost per accepted task falls out of those columns directly, and so does the effort level at which each model stops improving on your work, which is the one figure none of the public AI benchmarks can give you.

CodeRabbit's launch-day evaluation is what this looks like at scale. On 80 known bug patterns their lower-effort Opus 5.5 configuration caught 51, their higher-effort one 50 and their production baseline 49, differences inside the noise of a set that size. On 13 harder cases higher effort caught 10 and lower effort 8, but once findings outside the changed lines were counted both reached 10 and missed different ones, and their conclusion was that turning up effort did not consistently produce a better review.
State the limit of your own run as honestly as they did, because twenty tasks gives you direction rather than statistical certainty, and a difference of one or two tasks between candidates is noise. What it gives you that no leaderboard can is the effort level at which each model stops paying for itself on your code. Keep the suite, because at the current cadence there is a frontier launch every few weeks, and a team that reruns twenty tasks each time has turned the question of which model to use from a debate into a regression test.
The next launch table will be as unreadable as this one, and it is not a defect the vendors are likely to fix. Read any of them with five questions in hand: which effort level each model ran at, which harness and how many attempts, which benchmark version, which comparison model and the date of its figure, and what standard error is printed. A table that answers all five is rare, and one that answers none is a press release with a grid on it, whatever AI benchmarks it claims to report.
The teams that make good model decisions this week are not the ones that read the AI benchmarks most carefully. They are the ones that stopped reading launch tables as evidence and started reading them as a trigger to rerun their own twenty tasks.
Choosing a model for production and want the 20-task test run on your own codebase? Tell us what you are building.
This article was written by Mohamed Boukrim at Lasting Dynamics, an EU-based custom software development company. His other writing on this blog runs from reasoning effort and what it costs to the productivity gains that AI coding tools promise and do not always deliver. The view taken here comes from client delivery rather than from a research lab, because a model that has to hold up inside someone else's repository, review process and deadline gets judged on the work that ships rather than on a leaderboard row, and that is the standard these AI benchmarks are held to.
AI benchmarks are standardized task sets with a scoring method, used to compare models. What most definitions leave out is that the score depends on the harness, the effort level and the benchmark version, so it describes a configuration rather than a model. Terminal-Bench measures agentic terminal work, SWE-bench measures repository bug fixing, and OSWorld measures computer use.
On Anthropic's own table at max effort, yes, most of all on agentic work, where Terminal-Bench 4.0 moves from 52.3% to 66.4%, the higher figure at xhigh effort, and it is 20% cheaper per token with cache reads 60% cheaper. It is not a drop-in upgrade: thinking cannot be disabled, forced tool choice returns a 400 error, and thinking blocks are bound to the conversation prefix.
They are different tiers, and OpenAI positions Sol below GPT-6 Astra. On the independent Artificial Analysis index Opus 5.5 at medium scores 51 for $1.34 per task, while Sol scores 48 at max for $1.06 and 40 at medium for $0.25, so Sol is the cheaper route to any score up to about 48. The right choice depends on task difficulty and comes from testing both on your own tasks.
Because AI benchmarks are run at different effort settings, in different harnesses, on different versions, and against comparison figures taken from the other vendor's press release. GPT-6 Astra scores 41.4% on AutomationBench in Anthropic's table and 30.3% on AutomationBench 1.0.6 in OpenAI's, where it ran at low effort. Both numbers are correct for their configuration.
Take twenty merged pull requests with a clear acceptance test, check out each parent commit, give every candidate the original ticket in the same harness with the same tools and timeout, and run each at two or three effort levels. Record pass or fail, reviewer acceptance, tokens, cost and review minutes, then compare cost per accepted task, and rerun the same suite at every launch.
Transform bold ideas into powerful applications.
Let’s create software that makes an impact together.
Mohamed Boukrim
I am a Software Engineer and Backend Developer with a relentless focus on software quality and robust architecture. I don't believe in shortcuts; outstanding results are the direct product of daily commitment and hard work. Alongside my team, I leverage Agile methodologies and continuous process evaluation to push boundaries and optimize performance. I approach every challenge with a highly competitive mindset: I work tirelessly to reach the top, and once there, I keep pushing to maintain the lead.