Claude Sonnet 5.5 Explained: Pricing, Benchmarks, Terminal-Bench
Two numbers define the launch of Claude Sonnet 5.5: 70.6% and 10.3%. Anthropic reports the first as the model’s Terminal-Bench 4.0 score and the second as the score of its predecessor, Sonnet 5, on the same test. A seven-fold jump between two models that share a price tag and a name family looks like either a breakthrough or a benchmark artifact, and the honest answer turns out to be a bit of both.
The model shipped on September 28, 2026, at unchanged list prices of $2 per million input tokens and $10 per million output tokens. That makes this a release about efficiency and behavior, not about a new price point. If you run agents, coding assistants, or document pipelines on Sonnet 5, you need to know what actually changed, which headline figures survive scrutiny, and which five API settings will now return errors.
By the end of this article you will be able to read the launch claims critically, estimate real cost per task, choose an effort level with a method rather than a guess, and plan a migration that does not break in production.
What this covers: the lineage and positioning, what Anthropic has and has not disclosed about architecture and training, every sourced benchmark with its caveats, real pricing arithmetic, the breaking API changes, safeguards, failure modes, and a decision matrix against Opus 5.5 and GPT-6 Sol.
Context and Background
Anthropic’s current generation is the Claude 5 line. Sonnet 5 arrived on June 30, 2026, positioned by Anthropic as the most agentic Sonnet yet, with performance it described as close to Opus 4.8 at lower cost, and permanent pricing of $2 and $10 per million tokens confirmed as of August 10, 2026 (source: Anthropic’s Sonnet 5 announcement). Claude Opus 5.5 followed as the flagship of the newer 5.5 generation, and Sonnet 5.5 is the second model in that family. Anthropic has said Claude Haiku 5.5 will join in the coming weeks; at the time of writing it is announced but not shipped.
The pattern is familiar to anyone who tracks the vendor. Opus is the ceiling model, Sonnet is the workhorse most production traffic actually runs on, and Haiku is the cheap, fast tier. What is different in 2026 is how much of the Sonnet tier’s value is now about tokens per task rather than dollars per token. When list prices are frozen across releases, the only lever left for improving economics is doing the same job with fewer output tokens, fewer tool calls, and less wall-clock time. Anthropic’s own framing for Sonnet 5.5 is exactly that: output generated more than 30% faster than Sonnet 5, and up to 30% lower cost per task.
The competitive frame also matters. OpenAI’s GPT-6 family, covered in our GPT-6 Sol and Luna breakdown, sits at the same list price for the Sol tier in Anthropic’s comparison, but the two vendors’ models consume very different numbers of output tokens for the same work. That gap is where the real budget decisions are made. For the immediate predecessor, see our earlier explainer on Claude Sonnet 5, and for the model above it in the lineup, the Claude Opus 5.5 flagship analysis.
A note on evidence quality before going further. Almost every benchmark number in circulation for this model is vendor-reported, and independent replication is thin so far. Where a figure comes from Anthropic’s announcement I say so; where an outside party measured something different, I flag the gap. Anthropic’s announcement is at anthropic.com/claude-sonnet-5-5 and the system card is published alongside it.
What Claude Sonnet 5.5 Actually Is: The Reference Picture
Claude Sonnet 5.5 is a closed-weights, proprietary large language model from Anthropic, released September 28, 2026 under the API identifier claude-sonnet-5-5. It offers a 1 million token context window, up to 128K output tokens, adaptive thinking with five effort levels, and list pricing of $2 input and $10 output per million tokens.

Figure 1: Where Claude Sonnet 5.5 sits in the Claude 5.5 family and which specifications and safeguards wrap it. Opus 5.5 pricing is from Anthropic’s comparison table; Haiku 5.5 is announced but not released.
Figure 1 summarizes the disclosed surface. Everything inside the box labeled Sonnet 5.5 comes from Anthropic’s launch materials or from documentation reporting them. Everything you might want to know about the internals, parameter count, expert routing, layer counts, tokenizer size, is absent, and that absence is itself a fact worth stating plainly.
What is disclosed
The confirmed specification list is short and useful. The context window is 1 million tokens and the maximum output per response is 128K tokens; secondary reporting adds that the Batch API allows up to 300K output tokens with a beta header, which I could not confirm in Anthropic’s primary text, so treat that one as reported rather than verified. The knowledge cutoff is June 2026. The model is available on the Claude Platform and through Amazon Web Services, Google Cloud, and Microsoft Azure, with zero data retention available, the same as for Opus 5.5 and Sonnet 5. Third-party gateways such as OpenRouter and Vercel AI Gateway list it too, under their own naming: for instance anthropic/claude-sonnet-5.5 on Vercel and anthropic.claude-sonnet-5-5 on Amazon Bedrock, per a developer guide I checked.
Anthropic also states that Sonnet 5.5 will be supported with a minimum retirement commitment out to September 28, 2027, according to a secondary summary of the launch documentation. If you plan a migration, that is your outer planning boundary for the model itself.
What is not disclosed
Anthropic has not published the parameter count, whether the model is dense or mixture-of-experts, the attention variant, the tokenizer vocabulary size, the pre-training data volume, or training compute for Sonnet 5.5. I have seen third-party blog posts that assert specific architectural numbers. I found no primary source behind them, and I am not repeating them. If a page tells you Sonnet 5.5 has a specific parameter count, ask for the citation.
What we can say is functional. The model uses what Anthropic calls adaptive thinking: the model decides how much internal reasoning to spend, bounded by an effort setting. The API exposes five effort levels: low, medium, high, xhigh, and max. The default is high on the Claude Platform API and medium in Claude Code and the Claude apps. That split default is the single most important thing to understand about the benchmark story later in this article.
Positioning against Opus 5.5
Anthropic describes Sonnet 5.5 as a faster, lower-cost complement to Opus 5.5. Its stated strengths are well-scoped everyday tasks, fixing bugs, and producing polished documents, slides, and spreadsheets with design sensibility. Opus 5.5 is positioned for complex, open-ended work requiring sustained judgment. Opus 5.5 costs exactly twice as much per token: $4 input and $20 output per million, with cache reads at the same $0.20 and cache writes at $5.
That two-to-one price ratio is the frame for every routing decision. If Sonnet 5.5 needs more than twice the tokens of Opus 5.5 to finish a task, the cheaper model is no longer cheaper. We will put numbers on that shortly.
Lineage: what changed versus Sonnet 5
Sonnet 5.5 keeps the same list prices and the same broad capability class as Sonnet 5 but changes behavior in four ways. It is faster at generating output, by more than 30% per Anthropic. It uses tokens more efficiently, with customers reporting reductions ranging from about 12% to much larger figures on specific workloads. It ships with a new safeguards layer, including the first Sonnet-tier reasoning-extraction classifiers. And it tightens the API, rejecting several parameter combinations that Sonnet 5 accepted.
Benchmark deltas are large. On Anthropic’s figures Sonnet 5.5 moves CursorBench 4.0 from 34.1% to 55.5%, GDPval-AA v2.1 from 1449 to 1844, AA-Briefcase v1.1 from 1359 to 1811, Humanity’s Last Exam with tools from 54.9% to 64.5%, and OSWorld 2.1 from 57.0% to 80.1%. Chartography without tools moves from 15.6% to 61.6%. We will examine each, because a jump that large deserves a look at what the benchmark versions and configurations actually are.
Training, Safeguards, and Behavior: What Anthropic Says
Anthropic has not disclosed Sonnet 5.5’s training data volume, compute budget, or post-training recipe in the material I could verify. The system card covers evaluation methodology, and the announcement states that a behavioral audit found the model “improves on or matches Sonnet 5 on most measures of alignment, resistance to misuse, and honesty.” I cannot tell you whether the gains come from reinforcement learning on verifiable coding tasks, distillation from Opus 5.5, or both. Any article that tells you so is guessing.
What is documented is the safeguards architecture, and it changes how you should design applications.
Cyber safeguards with visible fallback
Sonnet 5.5 is the first Sonnet deployed with cybersecurity safeguards comparable to those on Opus 5.5. When a request is judged higher risk for cybersecurity misuse, the system falls back to Sonnet 5 and does so visibly. Ordinary bug-fixing in user code stays on Sonnet 5.5. Anthropic describes the model’s cyber capability as comparable to Opus 5, and security teams can apply for an expanded Cyber Verification Program that offers tiered access; the announcement describes that program as coming soon.
Practically, this means a single conversation may be served by two different models. If you log model identifiers for cost attribution or evaluation, do not assume every response in a Sonnet 5.5 session was produced by Sonnet 5.5. Check the model field on each response.
Reasoning-extraction classifiers and account-bound thinking
Sonnet 5.5 is the first Sonnet with classifiers that block requests attempting to reproduce its reasoning, a defense against distillation. Thinking blocks are also bound to the account that created them. Per the developer documentation summarized in secondary sources, thinking blocks from older Opus and Sonnet models are rejected, and if a conversation moves between accounts the earlier thinking blocks are dropped. Editing prior conversation history with thinking blocks attached returns an HTTP 400 for accounts created after August 31, 2026. This is a real architectural constraint for anyone who stores and replays transcripts across tenants.
Biology safeguards
Biology safeguards are unchanged from Sonnet 5. Anthropic says routine research and education are unaffected, but secondary reporting notes that some microbiology and virology requests may trigger false positives, and a Life Sciences Verification Program exists for organizations that need broader access. If you work in that domain, test with your real prompts before committing.
Benchmarks: Reading the Numbers Without Getting Fooled
The table below collects the vendor-reported figures from Anthropic’s launch. Everything here is self-reported unless marked otherwise. “Opus 5.5” and “Sonnet 5” columns come from the same announcement. Dashes mean no figure was published in the sources I could verify.
| Benchmark | Sonnet 5.5 | Sonnet 5 | Opus 5.5 | GPT-6 Sol |
|---|---|---|---|---|
| Terminal-Bench 4.0 | 70.6% | 10.3% | 66.4% (xhigh) | not published |
| CursorBench 4.0 | 55.5% | 34.1% | 57.8% | not published |
| FrontierCode 1.1 (xhigh) | 52.1% | 42.4% | 54.4% | 49.3% |
| FrontierCode 1.1 (max) | 46.2% | not published | not published | not published |
| GDPval-AA v2.1 (Elo) | 1844 | 1449 | 1846 | 1487 |
| AA-Briefcase v1.1 (Elo) | 1811 | 1359 | 1822 | not published |
| OSWorld 2.1 | 80.1% | 57.0% | 81.8% | not published |
| Humanity’s Last Exam, tools | 64.5% | 54.9% | not published | not published |
| Chartography, no tools | 61.6% | 15.6% | 64.4% | 53.6% |
Two patterns stand out. First, Sonnet 5.5 sits within a few points of Opus 5.5 on nearly everything, and beats it on exactly one row, Terminal-Bench 4.0. Second, the gap to Sonnet 5 is enormous on rows where Sonnet 5 was weak. A move from 15.6% to 61.6% on Chartography, or 10.3% to 70.6% on Terminal-Bench, is not the shape of incremental model improvement. It looks like a fixed failure mode.
The Terminal-Bench 4.0 headline needs a footnote
The 70.6% figure is the score at max effort, not the score at the API default. Anthropic’s own footnote clarifies that the Opus 5.5 comparison figure of 66.4% is at xhigh, described there as that model’s highest effort. One secondary analysis reports that at the API default effort (high), Sonnet 5.5 scores 43.0% on Terminal-Bench 4.0 against 64.2% for Opus 5.5. I could not confirm that pair against Anthropic’s primary page, so treat it as reported, not verified. If it is right, the picture inverts: at the setting most API users will actually run, Opus 5.5 leads by more than 20 points.
Independent measurement complicates it further. Artificial Analysis reportedly measured about 64% on Terminal-Bench for Sonnet 5.5, roughly six points under the vendor claim, though harness differences can explain gaps of that size. The benchmark itself, according to one review, comprises 66 tasks with a standard error around plus or minus 2.5 points, which means a six-point gap is more than two standard errors but not implausible given different scaffolding.

Figure 3: Three different Terminal-Bench readings for Claude Sonnet 5.5, each tied to a different effort level or harness. Only your own evaluation on your own workload resolves the disagreement.
Figure 3 lays out the evidence chain. The vendor headline, the API-default figure, and the independent measurement are three different experiments. None is dishonest; they answer different questions. The vendor asks what the model can do at its ceiling. The default-effort figure asks what a developer who never touches the effort setting gets. The independent figure asks what a neutral harness measures.
Why did Sonnet 5 score only 10.3%?
Nobody outside Anthropic has explained it. One review states that independent testing confirmed Sonnet 5’s collapse on this benchmark, and describes the Sonnet 5.5 result as fixing a failure rather than an incremental gain. That reading fits the shape of the numbers: a model that scores 10.3% against a benchmark on which its own successor scores several times higher, while sitting close to Opus-class models elsewhere, probably had a specific behavior problem, such as terminating early, mishandling shell state, or failing on a tool protocol detail. That is my inference from the pattern, not a documented finding, and I label it as opinion.
The practical takeaway: if you evaluated Sonnet 5 for terminal-heavy agent work and rejected it, that evaluation is stale. Re-run it.
Benchmark versions matter
Several benchmark names above carry version numbers, and versions change what is measured. Anthropic’s Sonnet 5 announcement cited up to 78.5% on OSWorld-Verified. The 57.0% figure for Sonnet 5 in the Sonnet 5.5 table is on OSWorld 2.1, a different benchmark version marked “partial” in Anthropic’s materials. Do not compare a Sonnet 5 number from its own launch to a Sonnet 5 number from this launch as if they measure the same thing. The FrontierCode row has a similar wrinkle: the 46.2% at max effort is lower than 52.1% at xhigh, and Anthropic attributes this to increased code-review segmentation under max effort. Secondary reporting attributes it to timeout issues. Either way, more effort does not monotonically improve outcomes.
A further footnote from Anthropic: GDPval-AA and AA-Briefcase pre-release testing included a structured-output bug that has since been fixed, and the GPT-6 Sol comparison scores may not reflect the latest bug fixes for image understanding. Cross-vendor rows therefore carry more uncertainty than same-vendor rows.
The GDPval-AA result is the one to trust most
The GDPval-AA v2.1 score of 1844 against Opus 5.5’s 1846 is an Elo-style rating on knowledge-work tasks, and it is measured by Artificial Analysis’s framework rather than being purely internal to Anthropic. A two-point difference on an Elo scale is well inside noise. If your workload is document, spreadsheet, and slide generation, this is the closest thing to an apples-to-apples signal that Sonnet 5.5 matches Opus 5.5 at half the token price. The caveat is that the benchmark scores outputs, and your quality bar may differ.
Effort Levels, Tokens, and the Real Cost per Task
Sticker price is $2 and $10 per million tokens. The cost of a task is tokens times price, and tokens depend on effort. This is where the launch story either pays off or falls apart.
List price arithmetic
Take a concrete request: 100,000 input tokens and 20,000 output tokens. At Sonnet 5.5 list prices, input costs $0.20 and output costs $0.20, for $0.40 total. The same request on Opus 5.5 costs $0.80. Through the Batch API, which offers a 50% discount ($1 input, $5 output), the Sonnet 5.5 cost is $0.20. If 90,000 of those input tokens are served as cache reads at $0.20 per million, the input side drops from $0.20 to roughly $0.038 (90,000 tokens at $0.20 per million is $0.018, plus 10,000 fresh tokens at $2 per million is $0.02). The cache minimum is now 512 tokens, down from 1,024, so shorter prompts also qualify. Cache writes cost $2.50 per million for the five-minute tier and $4 per million for the one-hour tier, per a secondary summary.
These are illustrative calculations from published rates, not measured workloads.
What “up to 30% cheaper” means
Anthropic’s claim is that Sonnet 5.5 costs up to 30% less per task than Sonnet 5. Since list prices are identical, this is entirely a claim about token consumption. It is a ceiling drawn from Anthropic’s own testing, not a guarantee. Customer figures give a sense of the range: Slack reported 14% fewer output tokens, Box reported 12% fewer tokens and 2.4 times faster runs, Zendesk reported 20% faster ticket processing, and Lovable reported about one-third fewer tool calls and roughly half as many shell executions. Balyasny Asset Management measured about 121,000 tokens per answer versus 497,000 on Sonnet 5, a factor of roughly 4.1, which is far beyond 30% and presumably reflects a specific workload where Sonnet 5 was verbose. If all of those tokens were output at $10 per million, that is about $1.21 versus $4.97 per answer, an illustrative upper bound since the split between input and output was not published.
Effort changes cost by an order of magnitude
An independent hands-on test by ComputingForGeeks ran infrastructure-as-code generation tasks (Kubernetes manifests, bash scripts, OpenTofu configuration), validated with linters and a live k3s cluster deployment. It is a small sample, three runs per task, but the token economics are instructive:
| Setting | Lint pass | k3s deploy | Avg cost per task |
|---|---|---|---|
| Sonnet 5.5 medium | 9 of 9 | 3 of 3 | $0.067 |
| Sonnet 5.5 high | 9 of 9 | 3 of 3 | $0.123 |
| Sonnet 5 default | 9 of 9 | 0 of 3 | $0.065 |
Medium effort matched Sonnet 5’s token output at roughly 6,200 tokens on average while passing deployments Sonnet 5 failed. High effort, the API default, nearly doubled tokens to about 11,900 with no measurable quality gain on single-file work. Max effort generated about 22 times more tokens than medium, roughly 135,000, for identical deployment results, at a reported cost of $1.36 versus $0.07. That is a roughly 20-fold cost increase for zero observed benefit on this task class.
Kingy.ai reports a similar pattern on agentic coding: at max effort, Sonnet 5.5 consumed about 193,000 output tokens per task versus about 31,000 for GPT-6 Sol, a six-fold difference at identical list prices, and at max effort Sonnet 5.5 can cost more per task than Opus 5.5 despite the lower per-token price. Its CursorBench readings by effort were 39.2% at medium, 47.8% at high, 53.1% at xhigh, and 55.5% at max. So on hard agentic work, effort does buy accuracy. On simple generation, it buys nothing. The right effort level is a property of the workload, not the model.
A method for choosing effort
Anthropic’s guidance is to re-run effort sweeps rather than carry Sonnet 5 settings forward, start at medium for well-specified work, use high for harder or longer tasks, and reserve xhigh and max for cases where you have measured a gain. Figure 2 turns that into a loop.

Figure 2: A measured escalation loop for Claude Sonnet 5.5 effort levels. Escalate only when your evaluation fails, and treat reaching xhigh without passing as a signal to evaluate Opus 5.5 instead.
The loop has two exits. One is success at some level, which becomes your pinned setting. The other is failure at xhigh, at which point paying twice per token for Opus 5.5 may cost less per completed task than burning max-effort Sonnet tokens. Measure cost per successful task, not cost per attempt. A cheap model that succeeds 60% of the time and requires a retry loop is often dearer than an expensive one that succeeds first time.
Access, Deployment, and the Migration Checklist
Sonnet 5.5 is a hosted, closed-weights model. There is no self-hosting path, no VRAM requirement to plan around, and no quantization decision to make; those questions from open-weights deep dives do not apply. Access is through the Claude Platform API with the identifier claude-sonnet-5-5 (no date suffix), Amazon Bedrock, Google Cloud, and Microsoft Azure, plus gateways such as OpenRouter and Vercel AI Gateway. GitHub Copilot lists it for Pro, Pro+, Max, Business, and Enterprise plans with a gradual rollout, and Claude Code exposes it through the alias sonnet. Availability and rollout timing vary by platform, so check your provider’s console before assuming a region has it.
The breaking changes
Anthropic’s launch notes and the developer documentation describe a set of parameters that Sonnet 5 accepted and Sonnet 5.5 rejects. Different summaries group them differently, but the union is consistent:
- Thinking cannot be turned off.
thinking: {"type": "disabled"}returns an error. The replacement is{"type": "between_tools"}, which works at low, medium, and high effort but fails at xhigh and max according to one guide. - Forced tool use is rejected.
tool_choiceset toanyor a specific tool returns HTTP 400. Useautowithstrict: truetool definitions instead, with a reported limit of 20 strict tools per request. - Sampling parameters are locked. Non-default
temperature,top_p, ortop_kvalues return 400, and manualbudget_tokensare no longer accepted. - Assistant prefill is rejected, per one secondary source. If you rely on prefilled responses to force JSON or a particular opening, plan a replacement.
- Computer use toolset changed. The
computer_20251124tool version is rejected on the Claude API and Google Cloud; migrate tocomputer_toolset_20260801and drop thefine-grained-tool-streaming-2025-05-14beta header.
There is also a subtle change in what streaming UIs see. Notes the model writes between tool calls now surface as thinking blocks. If your front end renders text blocks and hides thinking blocks, users will suddenly stop seeing the running commentary during long agent runs. And the advisor-model feature now returns 400 if the advisor is Opus 4.x, Sonnet 4.6, or Sonnet 5; valid advisors include Opus 5 or 5.5 and Sonnet 5.5.

Figure 4: A typical Claude Sonnet 5.5 migration. Two of the most common failures, disabled thinking and forced tool choice, return HTTP 400 and each has a direct replacement.
The good news from Anthropic’s migration notes is that existing Sonnet 5 prompts should work unchanged. The bad news is that the parameters above are not prompts; they live in your client code, your SDK wrappers, and possibly in third-party frameworks that set temperature by default. Grep for them before you flip the model string.
Behavior changes that tests will catch
Beyond hard errors, secondary sources describe soft behavior shifts worth adding to your regression suite. At low and medium effort the model tends to stop and ask for confirmation on long agentic tasks. It has a tendency to add unrequested tests and documentation. JSON structured outputs may skip thinking at low or medium effort, and moving to high or enabling adaptive thinking may help. In Claude Code, one reviewer observed the model rewriting files through bash heredocs rather than using the Edit tool, which bypasses diff visibility if your review process depends on it. And two failure patterns from the IaC test are worth remembering: low-effort output that mounted a ConfigMap it never defined, and a missing writable volume for nginx under a non-root user. Both passed schema validation and failed on a real cluster. Static checks are not a substitute for execution.
A minimal, safe rollout
Run the new model in shadow mode on a sample of production traffic before cutover, compare on your own quality metric and on tokens per completed task, and pin an effort level per route rather than globally. Log the response model field so cyber-safeguard fallbacks are visible. Keep Sonnet 5 as a rollback target, noting that Sonnet 5 reaches end of default status with this release.
Limitations, Failure Modes, and What Goes Wrong
Every claim below is either documented by Anthropic, reported by named third parties, or marked as my inference.
Effort inflation is the main cost trap. The model’s headline scores are achieved at max effort, and the same model at max effort can burn twenty times the tokens of medium for the same outcome on easy tasks. A team that copies a benchmark configuration into production can multiply its bill without noticing, because the price per token never moved. Put a token budget and an alarm on every route.
The default may not be what you want. The API defaults to high effort; Claude Code and the apps default to medium. That means the same prompt can behave and cost differently across surfaces. If you prototype in Claude Code and deploy through the API, your latency and spend will shift.
Vendor benchmarks are not your workload. The 70.6% Terminal-Bench figure, the 43.0% default-effort reading, and the roughly 64% independent measurement are three different numbers from three different setups. Terminal-Bench 4.0 with 66 tasks has coarse resolution. A four-point difference is within a couple of standard errors.
Opus is still stronger on hard, open-ended work. Anthropic itself says so. On FrontierCode at xhigh, Sonnet 5.5 trails Opus 5.5 by 2.3 points; on CursorBench 4.0 by 2.3 points; on OSWorld 2.1 by 1.7 points. Small numerically, but these are the tasks where a wrong answer is expensive, and open-ended judgment is not what these benchmarks best capture.
Safeguard fallbacks can surprise you. A security-adjacent request may be served by Sonnet 5 instead of Sonnet 5.5. For penetration-testing tools or vulnerability triage products, that is a functional risk, and the Cyber Verification Program is the intended path, though Anthropic described it as coming soon.
Thinking-block portability limits multi-tenant replay. Because thinking blocks are bound to the originating account, architectures that share or migrate conversation state between accounts, or that edit history, need review.
Context length is not context quality. A 1 million token window is a capacity, not a promise of uniform recall. I did not find a published long-context retrieval benchmark for Sonnet 5.5 in the launch material; test needle-style retrieval and multi-document reasoning on your data before filling the window.
Unresolved unknowns. Model size, architecture, training data, and training compute are unpublished. Independent replication of most benchmarks is still pending. Haiku 5.5 is announced but not released, so the cheap tier of this generation is not available yet.
How Claude Sonnet 5.5 Compares
The matrix below reflects vendor-reported data and my judgment. It is a starting hypothesis for routing, not a verdict.
| Use case | Claude Sonnet 5.5 | Claude Opus 5.5 | GPT-6 Sol |
|---|---|---|---|
| Terminal and shell agents | Strong; 70.6% at max, lower at default | 66.4% at xhigh; steadier at default per one report | Not published |
| Document, slide, spreadsheet work | Near parity, 1844 vs 1846 Elo | Marginal edge, at twice the token price | 1487 Elo |
| Hard open-ended reasoning | Good; trails Opus slightly | Anthropic’s recommended choice | Insufficient data |
| Cost predictability | Depends heavily on effort | Higher per-token, often fewer tokens | Reported far fewer output tokens at max |
| Computer use | 80.1% OSWorld 2.1 | 81.8% | Not published |
Read the last row of the cost column carefully. A per-token comparison says Sonnet 5.5 and GPT-6 Sol cost the same. A per-task comparison from one secondary source says Sonnet 5.5 can consume six times the output tokens at max effort. Whether that reflects verbosity, more thorough work, or wasted reasoning depends on the outcome quality, which the token count alone cannot tell you. For a wider view of how agent benchmarks are converging and diverging this month, see our roundup of agentic coding benchmarks in September 2026.
Practical Recommendations
If you run Sonnet 5 in production, the case for upgrading is strong on paper: same price, faster output, and fewer tokens per task in several customer reports. The case for doing it carelessly is weak. Budget a sprint for the API changes, another for effort tuning, and keep a rollback.
If you are choosing between Sonnet 5.5 and Opus 5.5, do not decide on a single benchmark. Route by task type: Sonnet 5.5 for well-specified, high-volume work such as bug fixes, document generation, and agent inner loops at medium effort; Opus 5.5 for ambiguous, high-stakes tasks where a failure costs more than the token premium. Then measure cost per successful task on both, because the two-to-one price ratio can be overturned by token counts.
If you are comparing against GPT-6 Sol, compare on your data with token accounting enabled. List price parity hides large differences in verbosity.
Checklist before cutover:
- Search your code for
thinkingdisabled, forcedtool_choice, non-defaulttemperature,top_p,top_k,budget_tokens, and assistant prefill. - Update computer use to
computer_toolset_20260801if you use it. - Sweep effort levels low through high on your own evaluation set; add xhigh only if high fails.
- Set per-route token budgets and alerts.
- Log the model identifier on every response to catch safeguard fallbacks.
- Re-test structured JSON outputs at your chosen effort.
- Add execution-based checks, not only schema validation, for infrastructure-as-code and code generation.
- Re-run any terminal-agent evaluation that rejected Sonnet 5.
- Review multi-account or history-editing flows for thinking-block constraints.
Frequently Asked Questions
What is Claude Sonnet 5.5 and when was it released?
Claude Sonnet 5.5 is Anthropic’s mid-tier large language model in the Claude 5.5 family, released on September 28, 2026 under the API identifier claude-sonnet-5-5. It is positioned as a faster, lower-cost complement to Claude Opus 5.5 and succeeds Sonnet 5, which shipped June 30, 2026. It is closed-weights, available on the Claude Platform, Amazon Web Services, Google Cloud, and Microsoft Azure, and has a June 2026 knowledge cutoff.
How much does Claude Sonnet 5.5 cost?
List pricing is $2 per million input tokens and $10 per million output tokens, unchanged from Sonnet 5. Cache reads cost $0.20 per million and five-minute cache writes cost $2.50 per million. The Batch API halves input and output prices to $1 and $5. Opus 5.5 costs $4 input and $20 output. Actual spend depends on effort level, because higher effort produces many more output tokens, so cost per task can differ sharply from the list price.
What is the Claude Sonnet 5.5 context window?
The context window is 1 million tokens, with a maximum output of 128K tokens per response. Secondary reporting says the Batch API can extend output to 300K tokens with a beta header, which I could not verify in Anthropic’s primary text. The minimum cacheable prompt was reduced from 1,024 to 512 tokens. A large window is a capacity limit, so test retrieval quality on your own long documents before relying on it.
How does Claude Sonnet 5.5 compare with Sonnet 5?
Prices are identical. Anthropic reports output more than 30% faster and up to 30% lower cost per task, plus large benchmark gains, including Terminal-Bench 4.0 from 10.3% to 70.6% and CursorBench 4.0 from 34.1% to 55.5%. It also adds new safeguards and rejects several API settings that Sonnet 5 accepted, including disabled thinking and forced tool choice. Existing prompts should work unchanged, but client code often needs edits.
Is the 70.6% Terminal-Bench score what I will get?
Probably not at default settings. Anthropic’s headline is a max-effort result. One secondary analysis reports 43.0% at the API’s default high effort, and Artificial Analysis reportedly measured about 64% under its own harness. Benchmark tasks, harnesses, and effort settings differ from production workloads, so treat 70.6% as a ceiling and run a small evaluation on your own tasks before setting expectations or budgets.
Should I use Sonnet 5.5 or Opus 5.5?
Use Sonnet 5.5 for well-scoped, high-volume work like bug fixes, document generation, and agent inner loops, especially at medium effort. Use Opus 5.5 for complex, ambiguous tasks that need sustained judgment, where Anthropic says it remains clearly stronger. On GDPval-AA the two are within two Elo points. If Sonnet 5.5 needs more than about twice the tokens to finish a task, Opus 5.5 can be cheaper per completed task.
Further Reading
Internal:
- Claude Sonnet 5 explained: architecture and benchmarks, the predecessor this article compares against.
- Claude Opus 5.5: Anthropic’s flagship model, the model above Sonnet 5.5 in the lineup.
- GPT-6 Sol and Luna explained: architecture, pricing, benchmarks, the main cross-vendor comparison.
- Agentic coding benchmarks, September 2026, for context on Terminal-Bench, CursorBench, and FrontierCode.
External primary sources:
- Introducing Claude Sonnet 5.5, Anthropic
- Claude Sonnet 5.5 system card, Anthropic
- Claude Platform documentation, Anthropic
By Riju — about

Pingback: Typed-Output Decision Models: Jev and Solar Decide