Claude Opus 5.5 makes agent cost the main question
Anthropic's Claude Opus 5.5 lowers token and cache-read prices, shifting the buyer test toward agent-task cost rather than benchmark rank alone.
Anthropic has released Claude Opus 5.5 with lower API prices and cheaper cache reads, turning its newest frontier model into a budgeting question for teams that run long coding agents.
The company announced Claude Opus 5.5 on September 22 as the first model in its 5.5 family. Anthropic says it performs at roughly the level of Claude Fable 5.1 on most work and costs 40 percent less to run than Opus 5. That is a company claim, but the pricing mechanics are visible in the model documentation: Opus 5.5 is listed at $4 per million input tokens and $20 per million output tokens, down from Opus 5's $5 and $25 rates. Cache reads fall from $0.50 to $0.20 per million tokens.
That cache line matters because agent sessions keep rereading the same context. Anthropic's own cost explainer says a Claude Code task is a loop: the model reads the conversation, calls a tool, reads the result and repeats until the work is done. The explainer uses Opus 5.5 list prices and says cache reads are often the largest input line by volume, while output and thinking tokens carry the higher price per token.
The price cut lands where agents spend
The new rates change the economics in three places developers can actually measure:
- Fresh input costs $4 per million tokens, 20 percent below Opus 5's list rate.
- Output costs $20 per million tokens, also 20 percent below Opus 5.
- Cache reads cost $0.20 per million tokens, a 60 percent reduction from Opus 5.
For a short chat, the headline price may be the input and output rates. For a coding agent that carries repository context, test output and tool logs through many turns, the cache-read discount can matter more. Anthropic says Opus 5.5 also uses fewer tokens per task and generates output more than 30 percent faster than Opus 5. Those efficiency claims still need workload-specific verification, because agent cost depends on turns, cache hit rate, effort setting, task failures and how much thinking the model spends before writing visible output.
The platform docs add another migration detail: adaptive thinking is always on for Opus 5.5. Developers can control thinking depth with the effort parameter, but they cannot disable thinking entirely. The same page lists four breaking changes for code already running on Claude Opus 5, including forced-tool-use behavior and a computer-use toolset change on the Claude API and Google Cloud. Anthropic's what's-new page also says Fast mode is a Claude API research preview and is not available on Amazon Bedrock, Claude Platform on AWS, Google Cloud or Microsoft Foundry.
Benchmarks point to a buyer trade-off
Anthropic frames the release around agentic coding and knowledge work. In its launch post, the company says Opus 5.5 at default effort scores 54.6 percent on FrontierCode v1.1, above GPT-6 Astra's reported top score of 53.3 percent, and says it matches GPT-6 Astra on Terminal-Bench 4.0 at about 40 percent of the cost per task. Anthropic also notes that some benchmark results use its own setup or reported competitor figures, so those comparisons should be read as vendor evidence rather than neutral proof.
Independent benchmark service Artificial Analysis reported Opus 5.5 at 58 on its Intelligence Index, the highest score it had measured, and confirmed the $4/$20 list pricing plus the cache-read drop to $0.20. Its write-up also adds a useful caveat: Opus 5.5 at max effort used about 119,000 output tokens per Intelligence Index task, compared with about 73,000 for Opus 5 and about 27,000 for GPT-6 Astra. In other words, lower per-token prices do not automatically mean every task becomes cheaper.
VentureBeat's launch coverage reached the same basic pricing conclusion: Opus 5.5 undercuts Opus 5 on input, output and cache costs while Anthropic claims stronger results on agentic coding, knowledge work and scientific reasoning benchmarks. That is corroboration of the release and pricing frame, not a substitute for running the model on a real internal backlog.
The practical buying decision is therefore narrower than the launch chart. A team considering Opus 5.5 should compare the same tasks under the same harness, with usage logs turned on, and track at least five fields: turns, fresh input, cache reads, output tokens and pass/fail quality. A model that finishes in fewer turns can save money even if it thinks more. A model that retries or over-explains can erase part of the list-price cut.
The next milestone is adoption data outside early-access testers and benchmark suites. Until then, Opus 5.5 is best read as a frontier-model price reset with a clear target: long-running software agents where cached context and tool loops dominate the bill.
Comments ()