Research("Claude Sonnet 5.5 + GPT-6 Sol")
5 sources · 2 companies ·
researched by Claude · reviewed and edited by Mario Belenguer Urpinell
Claude Sonnet 5.5: same price, five breaking changes
What it is
Sonnet 5.5 replaces Sonnet 5 as the “fast and smart” tier. From the official model page: 1M-token context, 128K synchronous output (300K on the Batch API with the output-300k-2026-03-24 beta header), $2 input and $10 output per million tokens, $0.20 cache reads, and a June 2026 knowledge cutoff. It’s available on the Claude API, Bedrock (anthropic.claude-sonnet-5-5), Google Cloud, Microsoft Foundry and Claude Platform on AWS.
Anthropic claims it generates “30%+ faster” than Sonnet 5 and, in its own testing, costs “up to 30% less per task”. On CursorBench 4.0 it reports 55.5% versus 34.1% for Sonnet 5. Those are vendor numbers; worth checking against your own evals.
The tokenizer is unchanged from Sonnet 5, so token counts stay the same. The minimum cacheable prompt drops from 1,024 to 512 tokens.
What breaks
The “what’s new” page lists five breaking changes. The first two are the ones agent code will hit most:
1. thinking: disabled is gone. It returns a 400 pointing you to between_tools, which turns off up-front thinking but keeps the short progress notes between tool calls. It only works at low, medium or high effort; for xhigh or max you need adaptive thinking.
client.messages.create(
- model="claude-sonnet-5",
- thinking={"type": "disabled"},
+ model="claude-sonnet-5-5",
+ thinking={"type": "between_tools"},
...
)
2. No forced tool use. tool_choice set to {"type": "any"} or {"type": "tool", ...} returns a 400, including on the token counting endpoint. The documented replacement is auto plus strict: true, or structured outputs, and telling the model in the prompt when the tool applies.
-tool_choice={"type": "tool", "name": "extract"},
+tool_choice={"type": "auto"},
+# and in the tool definition: "strict": True
3. Thinking blocks are bound. Each block records which model produced it. Sonnet 5.5 can read blocks from Sonnet 5, Opus 4.8 or Haiku 4.5, but not from Opus 5, Opus 5.5 or any Fable model, and no other model can read Sonnet 5.5’s. Switch models mid-conversation and the earlier reasoning is silently dropped. On accounts created on or after August 31, 2026, replaying a block after editing the system prompt, tools or an earlier message returns a 400. The advice is to keep history append-only and use mid-conversation system messages instead of edits. Blocks are also tied to the account that produced them.
4. computer_20251124 is rejected on the Claude API and Google Cloud; you need to move to the computer_toolset_20260801 toolset. Bedrock still accepts the old tool.
5. Advisor tool: Opus 4.8, Opus 4.7 and Sonnet 5 no longer work as advisors for a Sonnet 5.5 executor.
The quiet one
One more change fails no request but can break a UI: text the model writes between tool calls (anything longer than a sentence or two) now comes back as a progress thinking block. With the default display: "omitted", that text is empty, so an agent that shows it to users just goes quiet between tools. Set thinking.display under adaptive thinking, or use between_tools, to get it back.
Also, non-default temperature, top_p or top_k values return a 400. Effort levels have been recalibrated too, so the sweep you ran for Sonnet 5 doesn’t carry over; for agentic coding the docs suggest starting at medium.
In Claude Code
Version 2.1.284 adds Sonnet 5.5 and makes it the default Sonnet on the Anthropic API. According to the announcement, Claude Code defaults to medium effort, while the API defaults to high. If you manage models by policy, note that since 2.1.283 availableModelsMatch: "exact" makes each availableModels entry allow only the version it names, so Sonnet 5.5 stays blocked until you list it.
Same price as the direct competition
On September 22 OpenAI launched GPT-6 Sol (gpt-6-sol) at the exact same price: $2 input, $0.20 cached, $10 output per million, with up to 272K input tokens. Per-token prices are a tie; the real difference will be tokens per task and how many iterations each needs to land a change.
Open questions
- The announcement doesn’t say whether the “up to 30% less per task” figure compares both models at the same effort, which matters now that effort has been recalibrated.
- Binding thinking blocks to the conversation makes life harder for agents that rewrite or summarize their own history. On-demand compaction (beta
compact-2026-09-04) looks like the intended answer, but it’s one more beta header to track.
Sources
WebFetch × 5