Claude Sonnet 5.5 Explained: Pricing, Benchmarks, Terminal-Bench

Claude Sonnet 5.5 Explained: Pricing, Benchmarks, Terminal-Bench

Claude Sonnet 5.5 Explained: Pricing, Benchmarks, Terminal-Bench

Two numbers define the launch of Claude Sonnet 5.5: 70.6% and 10.3%. Anthropic reports the first as the model’s Terminal-Bench 4.0 score and the second as the score of its predecessor, Sonnet 5, on the same test. A seven-fold jump between two models that share a price tag and a name family looks like either a breakthrough or a benchmark artifact, and the honest answer turns out to be a bit of both.

The model shipped on September 28, 2026, at unchanged list prices of $2 per million input tokens and $10 per million output tokens. That makes this a release about efficiency and behavior, not about a new price point. If you run agents, coding assistants, or document pipelines on Sonnet 5, you need to know what actually changed, which headline figures survive scrutiny, and which five API settings will now return errors.

By the end of this article you will be able to read the launch claims critically, estimate real cost per task, choose an effort level with a method rather than a guess, and plan a migration that does not break in production.

What this covers: the lineage and positioning, what Anthropic has and has not disclosed about architecture and training, every sourced benchmark with its caveats, real pricing arithmetic, the breaking API changes, safeguards, failure modes, and a decision matrix against Opus 5.5 and GPT-6 Sol.

Context and Background

Anthropic’s current generation is the Claude 5 line. Sonnet 5 arrived on June 30, 2026, positioned by Anthropic as the most agentic Sonnet yet, with performance it described as close to Opus 4.8 at lower cost, and permanent pricing of $2 and $10 per million tokens confirmed as of August 10, 2026 (source: Anthropic’s Sonnet 5 announcement). Claude Opus 5.5 followed as the flagship of the newer 5.5 generation, and Sonnet 5.5 is the second model in that family. Anthropic has said Claude Haiku 5.5 will join in the coming weeks; at the time of writing it is announced but not shipped.

The pattern is familiar to anyone who tracks the vendor. Opus is the ceiling model, Sonnet is the workhorse most production traffic actually runs on, and Haiku is the cheap, fast tier. What is different in 2026 is how much of the Sonnet tier’s value is now about tokens per task rather than dollars per token. When list prices are frozen across releases, the only lever left for improving economics is doing the same job with fewer output tokens, fewer tool calls, and less wall-clock time. Anthropic’s own framing for Sonnet 5.5 is exactly that: output generated more than 30% faster than Sonnet 5, and up to 30% lower cost per task.

The competitive frame also matters. OpenAI’s GPT-6 family, covered in our GPT-6 Sol and Luna breakdown, sits at the same list price for the Sol tier in Anthropic’s comparison, but the two vendors’ models consume very different numbers of output tokens for the same work. That gap is where the real budget decisions are made. For the immediate predecessor, see our earlier explainer on Claude Sonnet 5, and for the model above it in the lineup, the Claude Opus 5.5 flagship analysis.

A note on evidence quality before going further. Almost every benchmark number in circulation for this model is vendor-reported, and independent replication is thin so far. Where a figure comes from Anthropic’s announcement I say so; where an outside party measured something different, I flag the gap. Anthropic’s announcement is at anthropic.com/claude-sonnet-5-5 and the system card is published alongside it.

What Claude Sonnet 5.5 Actually Is: The Reference Picture

Claude Sonnet 5.5 is a closed-weights, proprietary large language model from Anthropic, released September 28, 2026 under the API identifier claude-sonnet-5-5. It offers a 1 million token context window, up to 128K output tokens, adaptive thinking with five effort levels, and list pricing of $2 input and $10 output per million tokens.

Claude Sonnet 5.5 family position, specifications, and safeguards layer

Figure 1: Where Claude Sonnet 5.5 sits in the Claude 5.5 family and which specifications and safeguards wrap it. Opus 5.5 pricing is from Anthropic’s comparison table; Haiku 5.5 is announced but not released.

Figure 1 summarizes the disclosed surface. Everything inside the box labeled Sonnet 5.5 comes from Anthropic’s launch materials or from documentation reporting them. Everything you might want to know about the internals, parameter count, expert routing, layer counts, tokenizer size, is absent, and that absence is itself a fact worth stating plainly.

What is disclosed

The confirmed specification list is short and useful. The context window is 1 million tokens and the maximum output per response is 128K tokens; secondary reporting adds that the Batch API allows up to 300K output tokens with a beta header, which I could not confirm in Anthropic’s primary text, so treat that one as reported rather than verified. The knowledge cutoff is June 2026. The model is available on the Claude Platform and through Amazon Web Services, Google Cloud, and Microsoft Azure, with zero data retention available, the same as for Opus 5.5 and Sonnet 5. Third-party gateways such as OpenRouter and Vercel AI Gateway list it too, under their own naming: for instance anthropic/claude-sonnet-5.5 on Vercel and anthropic.claude-sonnet-5-5 on Amazon Bedrock, per a developer guide I checked.

Anthropic also states that Sonnet 5.5 will be supported with a minimum retirement commitment out to September 28, 2027, according to a secondary summary of the launch documentation. If you plan a migration, that is your outer planning boundary for the model itself.

What is not disclosed

Anthropic has not published the parameter count, whether the model is dense or mixture-of-experts, the attention variant, the tokenizer vocabulary size, the pre-training data volume, or training compute for Sonnet 5.5. I have seen third-party blog posts that assert specific architectural numbers. I found no primary source behind them, and I am not repeating them. If a page tells you Sonnet 5.5 has a specific parameter count, ask for the citation.

What we can say is functional. The model uses what Anthropic calls adaptive thinking: the model decides how much internal reasoning to spend, bounded by an effort setting. The API exposes five effort levels: low, medium, high, xhigh, and max. The default is high on the Claude Platform API and medium in Claude Code and the Claude apps. That split default is the single most important thing to understand about the benchmark story later in this article.

Positioning against Opus 5.5

Anthropic describes Sonnet 5.5 as a faster, lower-cost complement to Opus 5.5. Its stated strengths are well-scoped everyday tasks, fixing bugs, and producing polished documents, slides, and spreadsheets with design sensibility. Opus 5.5 is positioned for complex, open-ended work requiring sustained judgment. Opus 5.5 costs exactly twice as much per token: $4 input and $20 output per million, with cache reads at the same $0.20 and cache writes at $5.

That two-to-one price ratio is the frame for every routing decision. If Sonnet 5.5 needs more than twice the tokens of Opus 5.5 to finish a task, the cheaper model is no longer cheaper. We will put numbers on that shortly.

Lineage: what changed versus Sonnet 5

Sonnet 5.5 keeps the same list prices and the same broad capability class as Sonnet 5 but changes behavior in four ways. It is faster at generating output, by more than 30% per Anthropic. It uses tokens more efficiently, with customers reporting reductions ranging from about 12% to much larger figures on specific workloads. It ships with a new safeguards layer, including the first Sonnet-tier reasoning-extraction classifiers. And it tightens the API, rejecting several parameter combinations that Sonnet 5 accepted.

Benchmark deltas are large. On Anthropic’s figures Sonnet 5.5 moves CursorBench 4.0 from 34.1% to 55.5%, GDPval-AA v2.1 from 1449 to 1844, AA-Briefcase v1.1 from 1359 to 1811, Humanity’s Last Exam with tools from 54.9% to 64.5%, and OSWorld 2.1 from 57.0% to 80.1%. Chartography without tools moves from 15.6% to 61.6%. We will examine each, because a jump that large deserves a look at what the benchmark versions and configurations actually are.

Training, Safeguards, and Behavior: What Anthropic Says

Anthropic has not disclosed Sonnet 5.5’s training data volume, compute budget, or post-training recipe in the material I could verify. The system card covers evaluation methodology, and the announcement states that a behavioral audit found the model “improves on or matches Sonnet 5 on most measures of alignment, resistance to misuse, and honesty.” I cannot tell you whether the gains come from reinforcement learning on verifiable coding tasks, distillation from Opus 5.5, or both. Any article that tells you so is guessing.

What is documented is the safeguards architecture, and it changes how you should design applications.

Cyber safeguards with visible fallback

Sonnet 5.5 is the first Sonnet deployed with cybersecurity safeguards comparable to those on Opus 5.5. When a request is judged higher risk for cybersecurity misuse, the system falls back to Sonnet 5 and does so visibly. Ordinary bug-fixing in user code stays on Sonnet 5.5. Anthropic describes the model’s cyber capability as comparable to Opus 5, and security teams can apply for an expanded Cyber Verification Program that offers tiered access; the announcement describes that program as coming soon.

Practically, this means a single conversation may be served by two different models. If you log model identifiers for cost attribution or evaluation, do not assume every response in a Sonnet 5.5 session was produced by Sonnet 5.5. Check the model field on each response.

Reasoning-extraction classifiers and account-bound thinking

Sonnet 5.5 is the first Sonnet with classifiers that block requests attempting to reproduce its reasoning, a defense against distillation. Thinking blocks are also bound to the account that created them. Per the developer documentation summarized in secondary sources, thinking blocks from older Opus and Sonnet models are rejected, and if a conversation moves between accounts the earlier thinking blocks are dropped. Editing prior conversation history with thinking blocks attached returns an HTTP 400 for accounts created after August 31, 2026. This is a real architectural constraint for anyone who stores and replays transcripts across tenants.

Biology safeguards

Biology safeguards are unchanged from Sonnet 5. Anthropic says routine research and education are unaffected, but secondary reporting notes that some microbiology and virology requests may trigger false positives, and a Life Sciences Verification Program exists for organizations that need broader access. If you work in that domain, test with your real prompts before committing.

Benchmarks: Reading the Numbers Without Getting Fooled

The table below collects the vendor-reported figures from Anthropic’s launch. Everything here is self-reported unless marked otherwise. “Opus 5.5” and “Sonnet 5” columns come from the same announcement. Dashes mean no figure was published in the sources I could verify.

Benchmark Sonnet 5.5 Sonnet 5 Opus 5.5 GPT-6 Sol
Terminal-Bench 4.0 70.6% 10.3% 66.4% (xhigh) not published
CursorBench 4.0 55.5% 34.1% 57.8% not published
FrontierCode 1.1 (xhigh) 52.1% 42.4% 54.4% 49.3%
FrontierCode 1.1 (max) 46.2% not published not published not published
GDPval-AA v2.1 (Elo) 1844 1449 1846 1487
AA-Briefcase v1.1 (Elo) 1811 1359 1822 not published
OSWorld 2.1 80.1% 57.0% 81.8% not published
Humanity’s Last Exam, tools 64.5% 54.9% not published not published
Chartography, no tools 61.6% 15.6% 64.4% 53.6%

Two patterns stand out. First, Sonnet 5.5 sits within a few points of Opus 5.5 on nearly everything, and beats it on exactly one row, Terminal-Bench 4.0. Second, the gap to Sonnet 5 is enormous on rows where Sonnet 5 was weak. A move from 15.6% to 61.6% on Chartography, or 10.3% to 70.6% on Terminal-Bench, is not the shape of incremental model improvement. It looks like a fixed failure mode.

The Terminal-Bench 4.0 headline needs a footnote

The 70.6% figure is the score at max effort, not the score at the API default. Anthropic’s own footnote clarifies that the Opus 5.5 comparison figure of 66.4% is at xhigh, described there as that model’s highest effort. One secondary analysis reports that at the API default effort (high), Sonnet 5.5 scores 43.0% on Terminal-Bench 4.0 against 64.2% for Opus 5.5. I could not confirm that pair against Anthropic’s primary page, so treat it as reported, not verified. If it is right, the picture inverts: at the setting most API users will actually run, Opus 5.5 leads by more than 20 points.

Independent measurement complicates it further. Artificial Analysis reportedly measured about 64% on Terminal-Bench for Sonnet 5.5, roughly six points under the vendor claim, though harness differences can explain gaps of that size. The benchmark itself, according to one review, comprises 66 tasks with a standard error around plus or minus 2.5 points, which means a six-point gap is more than two standard errors but not implausible given different scaffolding.

Why the Terminal-Bench headline differs from what an application sees

Figure 3: Three different Terminal-Bench readings for Claude Sonnet 5.5, each tied to a different effort level or harness. Only your own evaluation on your own workload resolves the disagreement.

Figure 3 lays out the evidence chain. The vendor headline, the API-default figure, and the independent measurement are three different experiments. None is dishonest; they answer different questions. The vendor asks what the model can do at its ceiling. The default-effort figure asks what a developer who never touches the effort setting gets. The independent figure asks what a neutral harness measures.

Why did Sonnet 5 score only 10.3%?

Nobody outside Anthropic has explained it. One review states that independent testing confirmed Sonnet 5’s collapse on this benchmark, and describes the Sonnet 5.5 result as fixing a failure rather than an incremental gain. That reading fits the shape of the numbers: a model that scores 10.3% against a benchmark on which its own successor scores several times higher, while sitting close to Opus-class models elsewhere, probably had a specific behavior problem, such as terminating early, mishandling shell state, or failing on a tool protocol detail. That is my inference from the pattern, not a documented finding, and I label it as opinion.

The practical takeaway: if you evaluated Sonnet 5 for terminal-heavy agent work and rejected it, that evaluation is stale. Re-run it.

Benchmark versions matter

Several benchmark names above carry version numbers, and versions change what is measured. Anthropic’s Sonnet 5 announcement cited up to 78.5% on OSWorld-Verified. The 57.0% figure for Sonnet 5 in the Sonnet 5.5 table is on OSWorld 2.1, a different benchmark version marked “partial” in Anthropic’s materials. Do not compare a Sonnet 5 number from its own launch to a Sonnet 5 number from this launch as if they measure the same thing. The FrontierCode row has a similar wrinkle: the 46.2% at max effort is lower than 52.1% at xhigh, and Anthropic attributes this to increased code-review segmentation under max effort. Secondary reporting attributes it to timeout issues. Either way, more effort does not monotonically improve outcomes.

A further footnote from Anthropic: GDPval-AA and AA-Briefcase pre-release testing included a structured-output bug that has since been fixed, and the GPT-6 Sol comparison scores may not reflect the latest bug fixes for image understanding. Cross-vendor rows therefore carry more uncertainty than same-vendor rows.

The GDPval-AA result is the one to trust most

The GDPval-AA v2.1 score of 1844 against Opus 5.5’s 1846 is an Elo-style rating on knowledge-work tasks, and it is measured by Artificial Analysis’s framework rather than being purely internal to Anthropic. A two-point difference on an Elo scale is well inside noise. If your workload is document, spreadsheet, and slide generation, this is the closest thing to an apples-to-apples signal that Sonnet 5.5 matches Opus 5.5 at half the token price. The caveat is that the benchmark scores outputs, and your quality bar may differ.

Effort Levels, Tokens, and the Real Cost per Task

Sticker price is $2 and $10 per million tokens. The cost of a task is tokens times price, and tokens depend on effort. This is where the launch story either pays off or falls apart.

List price arithmetic

Take a concrete request: 100,000 input tokens and 20,000 output tokens. At Sonnet 5.5 list prices, input costs $0.20 and output costs $0.20, for $0.40 total. The same request on Opus 5.5 costs $0.80. Through the Batch API, which offers a 50% discount ($1 input, $5 output), the Sonnet 5.5 cost is $0.20. If 90,000 of those input tokens are served as cache reads at $0.20 per million, the input side drops from $0.20 to roughly $0.038 (90,000 tokens at $0.20 per million is $0.018, plus 10,000 fresh tokens at $2 per million is $0.02). The cache minimum is now 512 tokens, down from 1,024, so shorter prompts also qualify. Cache writes cost $2.50 per million for the five-minute tier and $4 per million for the one-hour tier, per a secondary summary.

These are illustrative calculations from published rates, not measured workloads.

What “up to 30% cheaper” means

Anthropic’s claim is that Sonnet 5.5 costs up to 30% less per task than Sonnet 5. Since list prices are identical, this is entirely a claim about token consumption. It is a ceiling drawn from Anthropic’s own testing, not a guarantee. Customer figures give a sense of the range: Slack reported 14% fewer output tokens, Box reported 12% fewer tokens and 2.4 times faster runs, Zendesk reported 20% faster ticket processing, and Lovable reported about one-third fewer tool calls and roughly half as many shell executions. Balyasny Asset Management measured about 121,000 tokens per answer versus 497,000 on Sonnet 5, a factor of roughly 4.1, which is far beyond 30% and presumably reflects a specific workload where Sonnet 5 was verbose. If all of those tokens were output at $10 per million, that is about $1.21 versus $4.97 per answer, an illustrative upper bound since the split between input and output was not published.

Effort changes cost by an order of magnitude

An independent hands-on test by ComputingForGeeks ran infrastructure-as-code generation tasks (Kubernetes manifests, bash scripts, OpenTofu configuration), validated with linters and a live k3s cluster deployment. It is a small sample, three runs per task, but the token economics are instructive:

Setting Lint pass k3s deploy Avg cost per task
Sonnet 5.5 medium 9 of 9 3 of 3 $0.067
Sonnet 5.5 high 9 of 9 3 of 3 $0.123
Sonnet 5 default 9 of 9 0 of 3 $0.065

Medium effort matched Sonnet 5’s token output at roughly 6,200 tokens on average while passing deployments Sonnet 5 failed. High effort, the API default, nearly doubled tokens to about 11,900 with no measurable quality gain on single-file work. Max effort generated about 22 times more tokens than medium, roughly 135,000, for identical deployment results, at a reported cost of $1.36 versus $0.07. That is a roughly 20-fold cost increase for zero observed benefit on this task class.

Kingy.ai reports a similar pattern on agentic coding: at max effort, Sonnet 5.5 consumed about 193,000 output tokens per task versus about 31,000 for GPT-6 Sol, a six-fold difference at identical list prices, and at max effort Sonnet 5.5 can cost more per task than Opus 5.5 despite the lower per-token price. Its CursorBench readings by effort were 39.2% at medium, 47.8% at high, 53.1% at xhigh, and 55.5% at max. So on hard agentic work, effort does buy accuracy. On simple generation, it buys nothing. The right effort level is a property of the workload, not the model.

A method for choosing effort

Anthropic’s guidance is to re-run effort sweeps rather than carry Sonnet 5 settings forward, start at medium for well-specified work, use high for harder or longer tasks, and reserve xhigh and max for cases where you have measured a gain. Figure 2 turns that into a loop.

Effort-level selection loop for Claude Sonnet 5.5

Figure 2: A measured escalation loop for Claude Sonnet 5.5 effort levels. Escalate only when your evaluation fails, and treat reaching xhigh without passing as a signal to evaluate Opus 5.5 instead.

The loop has two exits. One is success at some level, which becomes your pinned setting. The other is failure at xhigh, at which point paying twice per token for Opus 5.5 may cost less per completed task than burning max-effort Sonnet tokens. Measure cost per successful task, not cost per attempt. A cheap model that succeeds 60% of the time and requires a retry loop is often dearer than an expensive one that succeeds first time.

Access, Deployment, and the Migration Checklist

Sonnet 5.5 is a hosted, closed-weights model. There is no self-hosting path, no VRAM requirement to plan around, and no quantization decision to make; those questions from open-weights deep dives do not apply. Access is through the Claude Platform API with the identifier claude-sonnet-5-5 (no date suffix), Amazon Bedrock, Google Cloud, and Microsoft Azure, plus gateways such as OpenRouter and Vercel AI Gateway. GitHub Copilot lists it for Pro, Pro+, Max, Business, and Enterprise plans with a gradual rollout, and Claude Code exposes it through the alias sonnet. Availability and rollout timing vary by platform, so check your provider’s console before assuming a region has it.

The breaking changes

Anthropic’s launch notes and the developer documentation describe a set of parameters that Sonnet 5 accepted and Sonnet 5.5 rejects. Different summaries group them differently, but the union is consistent:

  1. Thinking cannot be turned off. thinking: {"type": "disabled"} returns an error. The replacement is {"type": "between_tools"}, which works at low, medium, and high effort but fails at xhigh and max according to one guide.
  2. Forced tool use is rejected. tool_choice set to any or a specific tool returns HTTP 400. Use auto with strict: true tool definitions instead, with a reported limit of 20 strict tools per request.
  3. Sampling parameters are locked. Non-default temperature, top_p, or top_k values return 400, and manual budget_tokens are no longer accepted.
  4. Assistant prefill is rejected, per one secondary source. If you rely on prefilled responses to force JSON or a particular opening, plan a replacement.
  5. Computer use toolset changed. The computer_20251124 tool version is rejected on the Claude API and Google Cloud; migrate to computer_toolset_20260801 and drop the fine-grained-tool-streaming-2025-05-14 beta header.

There is also a subtle change in what streaming UIs see. Notes the model writes between tool calls now surface as thinking blocks. If your front end renders text blocks and hides thinking blocks, users will suddenly stop seeing the running commentary during long agent runs. And the advisor-model feature now returns 400 if the advisor is Opus 4.x, Sonnet 4.6, or Sonnet 5; valid advisors include Opus 5 or 5.5 and Sonnet 5.5.

Migration sequence showing rejected settings and their replacements for Claude Sonnet 5.5

Figure 4: A typical Claude Sonnet 5.5 migration. Two of the most common failures, disabled thinking and forced tool choice, return HTTP 400 and each has a direct replacement.

The good news from Anthropic’s migration notes is that existing Sonnet 5 prompts should work unchanged. The bad news is that the parameters above are not prompts; they live in your client code, your SDK wrappers, and possibly in third-party frameworks that set temperature by default. Grep for them before you flip the model string.

Behavior changes that tests will catch

Beyond hard errors, secondary sources describe soft behavior shifts worth adding to your regression suite. At low and medium effort the model tends to stop and ask for confirmation on long agentic tasks. It has a tendency to add unrequested tests and documentation. JSON structured outputs may skip thinking at low or medium effort, and moving to high or enabling adaptive thinking may help. In Claude Code, one reviewer observed the model rewriting files through bash heredocs rather than using the Edit tool, which bypasses diff visibility if your review process depends on it. And two failure patterns from the IaC test are worth remembering: low-effort output that mounted a ConfigMap it never defined, and a missing writable volume for nginx under a non-root user. Both passed schema validation and failed on a real cluster. Static checks are not a substitute for execution.

A minimal, safe rollout

Run the new model in shadow mode on a sample of production traffic before cutover, compare on your own quality metric and on tokens per completed task, and pin an effort level per route rather than globally. Log the response model field so cyber-safeguard fallbacks are visible. Keep Sonnet 5 as a rollback target, noting that Sonnet 5 reaches end of default status with this release.

Limitations, Failure Modes, and What Goes Wrong

Every claim below is either documented by Anthropic, reported by named third parties, or marked as my inference.

Effort inflation is the main cost trap. The model’s headline scores are achieved at max effort, and the same model at max effort can burn twenty times the tokens of medium for the same outcome on easy tasks. A team that copies a benchmark configuration into production can multiply its bill without noticing, because the price per token never moved. Put a token budget and an alarm on every route.

The default may not be what you want. The API defaults to high effort; Claude Code and the apps default to medium. That means the same prompt can behave and cost differently across surfaces. If you prototype in Claude Code and deploy through the API, your latency and spend will shift.

Vendor benchmarks are not your workload. The 70.6% Terminal-Bench figure, the 43.0% default-effort reading, and the roughly 64% independent measurement are three different numbers from three different setups. Terminal-Bench 4.0 with 66 tasks has coarse resolution. A four-point difference is within a couple of standard errors.

Opus is still stronger on hard, open-ended work. Anthropic itself says so. On FrontierCode at xhigh, Sonnet 5.5 trails Opus 5.5 by 2.3 points; on CursorBench 4.0 by 2.3 points; on OSWorld 2.1 by 1.7 points. Small numerically, but these are the tasks where a wrong answer is expensive, and open-ended judgment is not what these benchmarks best capture.

Safeguard fallbacks can surprise you. A security-adjacent request may be served by Sonnet 5 instead of Sonnet 5.5. For penetration-testing tools or vulnerability triage products, that is a functional risk, and the Cyber Verification Program is the intended path, though Anthropic described it as coming soon.

Thinking-block portability limits multi-tenant replay. Because thinking blocks are bound to the originating account, architectures that share or migrate conversation state between accounts, or that edit history, need review.

Context length is not context quality. A 1 million token window is a capacity, not a promise of uniform recall. I did not find a published long-context retrieval benchmark for Sonnet 5.5 in the launch material; test needle-style retrieval and multi-document reasoning on your data before filling the window.

Unresolved unknowns. Model size, architecture, training data, and training compute are unpublished. Independent replication of most benchmarks is still pending. Haiku 5.5 is announced but not released, so the cheap tier of this generation is not available yet.

How Claude Sonnet 5.5 Compares

The matrix below reflects vendor-reported data and my judgment. It is a starting hypothesis for routing, not a verdict.

Use case Claude Sonnet 5.5 Claude Opus 5.5 GPT-6 Sol
Terminal and shell agents Strong; 70.6% at max, lower at default 66.4% at xhigh; steadier at default per one report Not published
Document, slide, spreadsheet work Near parity, 1844 vs 1846 Elo Marginal edge, at twice the token price 1487 Elo
Hard open-ended reasoning Good; trails Opus slightly Anthropic’s recommended choice Insufficient data
Cost predictability Depends heavily on effort Higher per-token, often fewer tokens Reported far fewer output tokens at max
Computer use 80.1% OSWorld 2.1 81.8% Not published

Read the last row of the cost column carefully. A per-token comparison says Sonnet 5.5 and GPT-6 Sol cost the same. A per-task comparison from one secondary source says Sonnet 5.5 can consume six times the output tokens at max effort. Whether that reflects verbosity, more thorough work, or wasted reasoning depends on the outcome quality, which the token count alone cannot tell you. For a wider view of how agent benchmarks are converging and diverging this month, see our roundup of agentic coding benchmarks in September 2026.

Practical Recommendations

If you run Sonnet 5 in production, the case for upgrading is strong on paper: same price, faster output, and fewer tokens per task in several customer reports. The case for doing it carelessly is weak. Budget a sprint for the API changes, another for effort tuning, and keep a rollback.

If you are choosing between Sonnet 5.5 and Opus 5.5, do not decide on a single benchmark. Route by task type: Sonnet 5.5 for well-specified, high-volume work such as bug fixes, document generation, and agent inner loops at medium effort; Opus 5.5 for ambiguous, high-stakes tasks where a failure costs more than the token premium. Then measure cost per successful task on both, because the two-to-one price ratio can be overturned by token counts.

If you are comparing against GPT-6 Sol, compare on your data with token accounting enabled. List price parity hides large differences in verbosity.

Checklist before cutover:

  • Search your code for thinking disabled, forced tool_choice, non-default temperature, top_p, top_k, budget_tokens, and assistant prefill.
  • Update computer use to computer_toolset_20260801 if you use it.
  • Sweep effort levels low through high on your own evaluation set; add xhigh only if high fails.
  • Set per-route token budgets and alerts.
  • Log the model identifier on every response to catch safeguard fallbacks.
  • Re-test structured JSON outputs at your chosen effort.
  • Add execution-based checks, not only schema validation, for infrastructure-as-code and code generation.
  • Re-run any terminal-agent evaluation that rejected Sonnet 5.
  • Review multi-account or history-editing flows for thinking-block constraints.

Frequently Asked Questions

What is Claude Sonnet 5.5 and when was it released?

Claude Sonnet 5.5 is Anthropic’s mid-tier large language model in the Claude 5.5 family, released on September 28, 2026 under the API identifier claude-sonnet-5-5. It is positioned as a faster, lower-cost complement to Claude Opus 5.5 and succeeds Sonnet 5, which shipped June 30, 2026. It is closed-weights, available on the Claude Platform, Amazon Web Services, Google Cloud, and Microsoft Azure, and has a June 2026 knowledge cutoff.

How much does Claude Sonnet 5.5 cost?

List pricing is $2 per million input tokens and $10 per million output tokens, unchanged from Sonnet 5. Cache reads cost $0.20 per million and five-minute cache writes cost $2.50 per million. The Batch API halves input and output prices to $1 and $5. Opus 5.5 costs $4 input and $20 output. Actual spend depends on effort level, because higher effort produces many more output tokens, so cost per task can differ sharply from the list price.

What is the Claude Sonnet 5.5 context window?

The context window is 1 million tokens, with a maximum output of 128K tokens per response. Secondary reporting says the Batch API can extend output to 300K tokens with a beta header, which I could not verify in Anthropic’s primary text. The minimum cacheable prompt was reduced from 1,024 to 512 tokens. A large window is a capacity limit, so test retrieval quality on your own long documents before relying on it.

How does Claude Sonnet 5.5 compare with Sonnet 5?

Prices are identical. Anthropic reports output more than 30% faster and up to 30% lower cost per task, plus large benchmark gains, including Terminal-Bench 4.0 from 10.3% to 70.6% and CursorBench 4.0 from 34.1% to 55.5%. It also adds new safeguards and rejects several API settings that Sonnet 5 accepted, including disabled thinking and forced tool choice. Existing prompts should work unchanged, but client code often needs edits.

Is the 70.6% Terminal-Bench score what I will get?

Probably not at default settings. Anthropic’s headline is a max-effort result. One secondary analysis reports 43.0% at the API’s default high effort, and Artificial Analysis reportedly measured about 64% under its own harness. Benchmark tasks, harnesses, and effort settings differ from production workloads, so treat 70.6% as a ceiling and run a small evaluation on your own tasks before setting expectations or budgets.

Should I use Sonnet 5.5 or Opus 5.5?

Use Sonnet 5.5 for well-scoped, high-volume work like bug fixes, document generation, and agent inner loops, especially at medium effort. Use Opus 5.5 for complex, ambiguous tasks that need sustained judgment, where Anthropic says it remains clearly stronger. On GDPval-AA the two are within two Elo points. If Sonnet 5.5 needs more than about twice the tokens to finish a task, Opus 5.5 can be cheaper per completed task.

Further Reading

Internal:

External primary sources:

By Riju — about

1 Comment

Leave a Reply

Your email address will not be published. Required fields are marked *