
Claude Sonnet 4.5 Deprecation: Migrate to Sonnet 5.5
Claude Sonnet 4.5 (claude-sonnet-4-5-20250929) was deprecated on 30 September 2026 and retires on the Claude API on 30 November 2026. After that, requests to it fail. Anthropic's replacement is claude-sonnet-5-5, which is cheaper per token ($2/$10 vs $3/$15) — but it is not a drop-in swap. I ran the same requests against both models this morning. Seven of them that Sonnet 4.5 happily accepts came back as HTTP 400 on Sonnet 5.5, and one that didn't error cost 29% more. Here's everything that broke, with the exact error text, and the order I'd fix it in.
What exactly is being retired, and when?
One model ID: claude-sonnet-4-5-20250929. Anthropic emailed affected accounts on 30 September and the model deprecations page lists 30 November 2026 as the retirement date. That's 61 days of notice, one day above the 60-day minimum Anthropic commits to.
The date applies to Anthropic-operated surfaces: the Claude API, Claude Platform on AWS and Microsoft Foundry. Amazon Bedrock and Google Cloud run their own retirement schedules, so if you call Sonnet 4.5 through Bedrock you may have longer — or not. Check their model tables instead of assuming.
The part people skip is finding every caller. Sonnet 4.5 was the default Sonnet for most of late 2025, which means it's hard-coded in places you forgot about: an n8n Anthropic node from last year, a .envon a VPS, a Make scenario a client built. The Console has an audit for this — Usage → Export gives you a CSV broken down by API key and model. Do that first, then grep. A quick grep -rn "sonnet-4-5" across my own automation scripts this morning turned up a video pipeline still hard-coding claude-sonnet-4-5-20250929— a script I'd have sworn was on Sonnet 5.
Which requests break on Claude Sonnet 5.5?
I took a small lead-classification request and a 20,000-character extraction prompt and sent variations of each to both models on 1 October 2026. Everything below returned HTTP 400 on claude-sonnet-5-5. The error strings are copied from the responses.
| What you send | What Sonnet 5.5 says |
|---|---|
Assistant prefill (last message is {"role":"assistant","content":"{"}) | This model does not support assistant message prefill. The conversation must end with a user message. |
temperature: 0 | `temperature` is deprecated for this model. |
thinking: {type: "enabled", budget_tokens: 1024} | "thinking.type.enabled" is not supported for this model. Use "thinking.type.adaptive" and "output_config.effort" to control thinking behavior. |
thinking: {type: "disabled"} | To turn thinking off on this model, send "thinking": {"type": "between_tools"}instead… |
tool_choice: {type: "tool", name: "classify_lead"} (or any) | tool_choice: type "tool" and "any" are not supported for this model. |
between_tools + effort: "xhigh" (or max) | output_config.effort 'xhigh' is not supported when thinking is disabled on this model. Use effort 'high' or below, or enable thinking. |
computer_20250124 or computer_20251124 tool (Claude API) | 'claude-sonnet-5-5' does not support tool types: … (per Anthropic's docs; I didn't run computer use) |
The prefill one will hit the most people.Sonnet 4.5 was the last Sonnet that let you start Claude's reply with {to force raw JSON. It's a habit from 2024 and it's everywhere. The same test showed why people did it: asked to “output only JSON”, Sonnet 4.5 still wrapped its answer in a ```json fence. Sonnet 5.5 returned bare JSON without being forced. The right replacement is structured outputs (output_config.format) — I covered the trade-offs in structured outputs vs tool use.
Two notes. temperature: 1is accepted because it's the default value; any other number fails, so delete the parameter rather than tuning it. And forced tool use isn't coming back: use tool_choice: auto, mark the tool strict: true, and say in the prompt when to call it. My classifier called the tool in all three runs with just “Use the classify_lead tool” appended.
Is Sonnet 5.5 actually cheaper than Sonnet 4.5?
Per token, yes. Per request, it depends entirely on one setting you probably never sent. Two things move at once.
1. The tokenizer counts more tokens
Sonnet 5.5 uses the newer tokenizer that arrived with Claude 4.7. I ran the token counting endpoint on the same 20,000-character system prompt (a TSX page, so code-heavy): 6,328 tokens on Sonnet 4.5, 7,823 on Sonnet 5.5 — 23.6% more. Anthropic says “about 30%” and that it varies by content. Even so, the rate cut wins on input: 7,823 × $2 is about 18% cheaper than 6,328 × $3.
2. It thinks by default
This is the one with no error. On Sonnet 4.5, a request with no thinking field ran without thinking. On Sonnet 5.5 the same request runs adaptive thinking at high effort. Here's the extraction task, one run each, same prompt:
| Config | Output tokens | Latency | Cost |
|---|---|---|---|
| Sonnet 4.5, defaults | 149 | 3.0 s | $0.0213 |
| Sonnet 5.5, defaults | 1,172 (957 thinking) | 8.0 s | $0.0274 (+29%) |
| Sonnet 5.5, effort medium | 266 (0 thinking) | 2.2 s | $0.0184 (−14%) |
| Sonnet 5.5, between_tools + low | 346 | 2.5 s | $0.0192 (−10%) |
So a lazy swap of the model ID made this call 29% more expensive and 2.7× slower, on a model that's supposed to be cheaper and faster. One explicit output_config.effort: "medium"flipped it to 14% cheaper and faster than before. It's a single run on a small task, so treat the percentages as direction, not gospel — but the direction is the whole lesson. If you're not sure which level fits which route, I wrote up how I pick an effort level.
Also raise max_tokens. It now covers thinking plus text, and thinking is billed as output. A route that used to fit in 300 tokens can get truncated mid-thought.
What changes without throwing an error?
content[0].textbreaks. With thinking on, the response starts with athinkingblock whose text is empty by default. Read blocks bytype. Every quick script I've ever written doescontent[0].text.- Narration between tool calls moves. Notes longer than a sentence or two now arrive as progress-update
thinkingblocks, empty at the default display. A UI that streamed them goes quiet.between_toolsbrings the text back; with adaptive thinking setdisplayto"summarized". - Thinking blocks are tied to the conversation.Edit an earlier turn, the system prompt or the tools and replay the block, and newer accounts get a 400. Keep history append-only — which also protects your cache, as I found in cache miss diagnostics.
- Images cost more.Sonnet 5.5 reads up to 2576 px on the long edge; a 2000×1500 image costs about 2.5× the tokens it did on 4.5. Downscale first if you don't need the detail.
- Caching gets easier. The minimum cacheable prompt drops from 1,024 to 512 tokens, so short system prompts that never cached now do.
- Beta headers to delete:
interleaved-thinking-2025-05-14, any context-window header, andfine-grained-tool-streaming-2025-05-14(replace witheager_input_streaming: trueper tool). Moveoutput_formattooutput_config.format.
Should you go to Sonnet 4.6 instead?
It's a legitimate stopgap. claude-sonnet-4-6is active until at least 17 February 2027, costs the same $3/$15 as 4.5, keeps the old tokenizer, doesn't think unless you ask, and still accepts temperature and (deprecated) budget_tokens. It does reject prefill. If you have fifty client workflows and two weeks, that's the smallest diff.
My take: don't, unless you're truly out of time.You'd pay 50% more per token for an older model that already has a retirement floor on the calendar, then do this whole migration again next year. Most of the work — prefill, forced tools, parsing by block type — you'd do anyway. I wrote about the same dilemma one tier up in Opus 5.5 vs Opus 5, and the answer was the same: fix it once, properly.
The migration checklist I'm using
- Find callers. Console Usage export, then
grep -rn "sonnet-4-5"across repos, n8n exports and.envfiles. - Swap the ID to
claude-sonnet-5-5(no date suffix;anthropic.claude-sonnet-5-5on Bedrock). - Kill the 400s. Remove prefills,
temperature/top_p/top_k,budget_tokens; replace forcedtool_choicewithauto+strict: true. - Set effort on every route.
mediumfor well-scoped agent steps,loworbetween_toolsfor classification and chat,highonly where quality measurably needs it. - Fix parsing. Read by block type, pass thinking blocks back untouched, keep history append-only, raise
max_tokens. - Re-baseline.Count tokens on real prompts, compare cost and latency per task, not per million. If it's not latency-sensitive, push it through the Batch API for another 50% off.
If you use Claude Code, /claude-api migrate this project to claude-sonnet-5-5does most of steps 2–4 and hands you a checklist. Review its diff like you'd review a junior's; it doesn't know which routes are latency-sensitive. Anthropic's Sonnet 5.5 migration guide has the full before/after per starting model.
One more thing on timing: don't wait for 29 November. Deprecated models are, in Anthropic's own words, likely to be less reliable than active ones. I'd have production off 4.5 by mid-November and leave the last two weeks for the caller you didn't find.
FAQ
When does Claude Sonnet 4.5 stop working?
Anthropic deprecated claude-sonnet-4-5-20250929 on 30 September 2026 and lists 30 November 2026 as its retirement date on the Claude API, Claude Platform on AWS and Microsoft Foundry. Requests after retirement fail. Amazon Bedrock and Google Cloud set their own schedules, so check those model tables separately.
What should I replace Claude Sonnet 4.5 with?
Anthropic recommends claude-sonnet-5-5, which costs $2/$10 per million input/output tokens against Sonnet 4.5's $3/$15. Claude Sonnet 4.6 is a lower-effort stopgap at the same $3/$15 price, active until at least 17 February 2027, with fewer breaking changes because it keeps the old tokenizer and does not think by default.
Is Claude Sonnet 5.5 cheaper than Sonnet 4.5?
Per token yes, per request not always. Sonnet 5.5's tokenizer produced 23.6 percent more tokens for the same prompt in my test, which still left input about 18 percent cheaper. But Sonnet 5.5 thinks by default at high effort, and on a small extraction task that made it 29 percent more expensive than Sonnet 4.5. Setting effort to medium made the same task 14 percent cheaper.
How do I turn off thinking on Claude Sonnet 5.5?
Send thinking type between_tools. thinking type disabled returns a 400 that tells you to use between_tools instead. between_tools only works at low, medium or high effort; with xhigh or max it returns a 400, and it accepts no display, budget_tokens or block_binding fields.
Got Sonnet 4.5 Baked Into Client Work?
The ID swap is a minute. Finding every caller, fixing the 400s and proving the bill didn't go up is the real job. I do these migrations for agencies and SaaS teams — send me your setup and I'll tell you what breaks before November does.
Hire CarlosRetirement dates, prices and error behavior verified 1 October 2026 against Anthropic's pricing documentation and live calls to claude-sonnet-4-5-20250929 and claude-sonnet-5-5. Cost figures are single runs on my prompts at list price; run your own traffic before you budget on them.
Related Posts
AI Models
Claude Opus 5.5 vs Opus 5: What Actually Changes
Claude Opus 5.5 is cheaper than Opus 5 on every line — $4/$20 per MTok and cache reads at $0.20 instead of $0.50 — with the same 1M context and 128K output. But four request shapes now return 400, and the default effort level drops from high to medium with no error attached. The full migration, the real cost math, and why your agent UI went quiet between tool calls.
AI Models
Claude Prompt Cache Miss? What Cache Diagnostics Shows
Pass the previous response id in a diagnostics object and the Claude API names what broke your prompt cache: model, system prompt, tools or message history. A timestamp, reordered tools and one trailing space each threw away 21,000 cached tokens in my test.
AI Models
Fable 5.1 Pricing: Cheaper Only If You Cache
Fable 5.1 kept the $10/$50 per-token price and cut cache reads 75% to $0.25/M. Your bill drops by 0.75 times your cache-read share — and not a point more. The math, plus the three changes that break agent loops.