Growth Cab Apply to GC
Blog/AI TOOLS
AI TOOLS · October 1, 2026 · 12 MIN READ

Claude Opus API Pricing: What Opus 5.5 Really Costs

Claude Opus API pricing for Opus 5.5: $4 input, $20 output, $0.20 cache reads. What a real agent session costs, the hidden fees and 7 ways to cut the bill.

Federico DonatoneBy · Founder & CEO, Growth Cab
Claude Opus API Pricing: What Opus 5.5 Really Costs

Claude Opus API pricing for Opus 5.5 is $4 per million input tokens and $20 per million output tokens, 20% below Opus 5. Cache reads cost $0.20 per million, 60% less than on Opus 5. The Batch API halves both rates to $2 and $10. Anthropic says a typical job now costs about 40% less than on Opus 5.

I posted about this today on LinkedIn, because the price cut is only half the story. Cost has been my biggest pain with Claude since Fable came out. For two months I barely used Opus 5, because it talked too much. Opus 5.5 talks less, and on agent work that matters more than the list price. Here is what you will actually pay, with the math.

Federico Donatonein
Federico Donatone
Founder & CEO, Growth Cab · This article started as a LinkedIn post

“Anthropic made Opus 5.5 WAY cheaper. Here’s what’s really going on behind the price.”

Read the original post →

Claude Opus API Pricing in 2026: What Does Each Model Cost?

Opus 5.5 is the cheapest Opus model Anthropic has ever listed: $4 per million input tokens and $20 per million output tokens. Every older Opus from 4.5 to 5 still costs $5 and $25. Opus 4.1 and earlier, now retired on the Claude API, cost three times that, at $15 and $75. The table below puts the current Claude API prices side by side.

Claude API prices per million tokens, from Anthropic's pricing page, checked October 1, 2026 (USD)
ModelInput5-min cache write1-hour cache writeCache readOutputBatch input / output
Claude Opus 5.5$4$5$8$0.20$20$2 / $10
Claude Opus 5$5$6.25$10$0.50$25$2.50 / $12.50
Claude Opus 4.5 to 4.8$5$6.25$10$0.50$25$2.50 / $12.50
Claude Fable 5.1$10$12.50$20$0.25$50$5 / $25
Claude Sonnet 5.5$2$2.50$4$0.20$10$1 / $5
Claude Sonnet 5$2$2.50$4$0.20$10$1 / $5
Claude Haiku 4.5$1$1.25$2$0.10$5$0.50 / $2.50

Two details in that table are easy to miss. First, Opus 5.5 and Sonnet 5.5 now charge the same $0.20 for a cache read. Second, Fable 5.1 reads its cache for $0.25, only 5 cents more than Opus 5.5. The gap between the models is real on fresh input and output. On cached context it almost disappears.

For a simple chatbot with no caching, the math is plain. Ten million input tokens and two million output tokens a month cost $80 on Opus 5.5 and $40 on Sonnet 5.5. Agents are where the table gets interesting, because they reread the same context all day. One token is about four characters, or 0.75 words in English, per Anthropic.

If you are choosing between a subscription and paying per token for your team, my comparison of Claude vs ChatGPT for business covers the seat prices.

Why Is Opus 5.5 40% Cheaper When Tokens Only Dropped 20%?

Anthropic says Opus 5.5 costs about 40% less than Opus 5 on typical workloads at default settings. The 20% cut on input and output explains only half of that. The rest comes from two things. The model writes fewer words to finish a job. And rereading cached context, the part of the bill that grows fastest, got 60% cheaper.

Start with the words. Factory, quoted on Anthropic's launch page, says Opus 5.5 matched Opus 5 at high effort while using 20 to 25% fewer output tokens. Box says it used a third of the tokens Opus 5 did. Fewer output tokens means a smaller bill at $20 per million. A cheaper model wins the price war. A model that writes less wins it twice.

Then the rereading. An agent that works on its own sends the same instructions, files and history back to the model on every step. Anthropic says context per request grew 2.6 times among Claude Code developers between March and September 2026. The ratio of input to output went from 189 to 1 to 324 to 1. Most of that input is cached, so the cache read price now drives the bill.

There is also a supply story behind the price. In April, Anthropic secured up to 5 gigawatts of Amazon Trainium and Graviton capacity, with nearly 1 gigawatt of Trainium capacity coming online by the end of 2026. More chips make a lower price easier to hold. Anthropic also says Opus 5.5 writes output more than 30% faster than Opus 5.

One caveat I would not skip. The default effort on Opus 5.5 is medium, while Opus 5 ran at high. Part of the default-settings saving likely comes from that switch. Anthropic also says Opus 5.5 thinks more per turn at the same effort level. So if you raise effort back to high, rerun your own numbers before you budget for a 40% cut.

What Does a Real Agent Session Cost on Opus 5.5?

In my worked example, one long agent session costs $1.75 on Opus 5.5 against $2.94 on Opus 5, a 40% cut on identical tokens. It is an illustration built on stated assumptions, and I did not measure it. The 40% here comes from prices alone, before Opus 5.5 writes a single word less. I shaped it on the 324 to 1 input to output ratio Anthropic reports for Claude Code sessions.

Picture an account research agent for outbound. It reads the client brief, the ideal customer profile and a list of target companies. Then it runs dozens of searches, reads results and searches again. On September 29 I wrote about Opus 5.5 running searches like these through a data connector. Every step resends the same brief, so caching does most of the work.

One agent session: 3,000,000 cached tokens read, 150,000 written to a 5-minute cache, 50,000 fresh input, 10,000 output (my assumptions, list prices of October 1, 2026)
ModelCache readsCache writesFresh inputOutputTotal
Opus 5$1.50$0.94$0.25$0.25$2.94
Opus 5.5$0.60$0.75$0.20$0.20$1.75
Opus 5.5 without cachingn/an/a$12.80$0.20$13.00
Sonnet 5.5$0.60$0.38$0.10$0.10$1.18
Fable 5.1$0.75$1.88$0.50$0.50$3.63
Haiku 4.5$0.30$0.19$0.05$0.05$0.59

Run 20 of those sessions a day for 22 working days and Opus 5 costs $1,293 a month. Opus 5.5 costs $770. The same work without caching would cost $5,720. That last line is the one I would show a finance team. The model choice saves hundreds. Caching saves thousands.

The table also shows why I would not switch everything to Sonnet 5.5 to save money. On this cached workload Opus 5.5 costs about 1.5 times Sonnet 5.5, while the list price says twice. If Opus finishes the job in fewer steps, that gap shrinks further. Test both on your own task before you decide.

For a simple way to run that comparison on real work, use the scoring sheet in my guide on how to evaluate AI models.

How Does Prompt Caching Change Your Claude API Cost?

Prompt caching stores the start of your prompt so later requests read it at a discount. On Opus 5.5 a cache read costs 5% of the normal input price, half the usual 10% rate. Writing to the cache costs extra: 1.25 times input for a 5-minute cache and 2 times input for a 1-hour cache.

Anthropic's pricing page gives the break-even rule. A 5-minute cache pays for itself after one read. A 1-hour cache pays off after two reads. OpenAI's GPT-6.1 Sol, released at the end of September 2026, charges $0.10 per million cached input tokens, half of Opus 5.5.

On Opus 5.5 the smallest prompt you can cache is 512 tokens. You can turn caching on with one cache_control field at the top of the request, and the API places the breakpoints for you.

Use the 1-hour cache when your agent or your team pauses for more than five minutes between steps. In my example, 1-hour writes would raise the session from $1.75 to $2.20. That is worth it only if the 5-minute cache keeps expiring and forcing full rewrites. Anthropic recommends the 1-hour setting for long sessions.

Is Opus 5.5 Worth It Over Sonnet 5.5 or Fable 5.1?

For most agent work, Opus 5.5 is the best value in Anthropic's lineup right now. It costs 40% of Fable 5.1 on input and output. Anthropic says it performs at the level of Fable 5.1 on most work, and in my first sessions its output beat Fable.

OpenAI's GPT-6.1 Sol lists at $2 and $10, half of Opus 5.5 per token, so test cost per finished job. Above 272,000 input tokens Sol bills $4 and $15 for the whole request, while Opus 5.5 has no long-context surcharge.

Which Claude model I would pick by job (my judgment, October 1, 2026)
JobMy pickWhy
Long agent runs with tools and big contextOpus 5.5Cache reads at $0.20 and fewer steps per job
High-volume classification or extractionHaiku 4.5$1 input, $0.50 on Batch
Everyday drafting at scaleSonnet 5.5Half of Opus on fresh input and output
The hardest reasoning, when cost is secondaryFable 5.1Top tier at $10 and $50
Overnight reports nobody reads until morningOpus 5.5 on Batch$2 and $10, results within 24 hours

What Hidden Costs Show Up on a Claude API Bill?

The list price is not the whole bill. Here are nine things to check on Opus 5.5 before you trust a headline comparison. Effort is the one I would check first. The numbers come from Anthropic's documentation.

  1. Thinking is always on. Opus 5.5 rejects requests that disable it, and thinking tokens are billed as output.
  2. Effort sets the spend. The default is medium. Higher effort means more thinking and more output tokens.
  3. Fast mode costs double: $8 input and $40 output per million on Opus 5.5.
  4. US-only processing adds 10%. Setting inference_geo to us applies a 1.1x multiplier to every token.
  5. Tools add tokens. Any request that includes tools adds a 286-token system prompt on Opus 5.5, and web search costs $10 per 1,000 searches.
  6. The tokenizer changed with Claude 4.7. Anthropic says it produces about 30% more tokens for the same text, so per-token comparisons with older models mislead.
  7. Cloud endpoints vary. On Amazon Bedrock and Google Cloud, regional endpoints cost 10% more than global ones.
  8. Long context has no surcharge. The full 1M-token window bills at standard rates.
  9. Claude Managed Agents adds $0.08 per session-hour of running time on top of tokens.

How Do You Cut Your Claude Opus API Bill?

The fastest savings come from caching and from keeping the cache warm. After that, set effort on purpose and move non-urgent work to the Batch API. This is the checklist I would run on any Claude agent before arguing about which model to buy.

AI FRONTIER
Get one useful AI play every Thursday
The AI changes that matter, Federico's direct read and one practical play, plus the free 10-page Operator Pack.
Free · under five minutes · unsubscribe anytime
  1. Turn on prompt caching with the automatic cache_control field.
  2. Put stable content first: instructions, tools, the client brief. Put changing content last.
  3. Pick the model at the start of a session. Switching models mid-session loses the cache.
  4. Change effort per message (a beta feature) instead of changing the top-level effort value, which resets the cache.
  5. Use the 1-hour cache only for sessions with long pauses.
  6. Send reports, enrichment and scoring jobs to the Batch API at half price. Cache discounts stack on top.
  7. Check cache hits weekly in your usage data. A low hit rate is money leaking.

If you run several agents at once, my notes on managing multiple AI agents show how we split the work so each one keeps a stable prompt.

Should You Use the Claude API or a Claude Subscription?

Use a subscription when people chat with Claude, and the API when software calls it. Pro costs $20 a month, Max from $100, and Team $25 per seat ($20 billed annually). Each gives a person a flat price with usage limits. The API has no monthly fee and bills every token, which suits agents, products and batch jobs. Many teams need both.

The trade-off is predictability. A seat never surprises you on the invoice, but it can stop you mid-task when you hit a limit. The API never stops you, but it can surprise you. With caching on and effort set on purpose, Opus 5.5 is the first Opus I would let an agent run on all day without watching the meter.

If you mostly work in the app, my guide to Claude Max usage limits explains how the five-hour and weekly windows work.

And if you want agents like these running inside a real outbound motion, that is the work our B2B lead generation service does for companies selling contracts of $50,000 or more.

Frequently asked questions

How much does Claude Opus 5.5 cost per million tokens?

Claude Opus 5.5 costs $4 per million input tokens and $20 per million output tokens on the Claude API. Cache reads cost $0.20 per million, 5-minute cache writes $5 and 1-hour cache writes $8. The Batch API halves input and output to $2 and $10. Fast mode costs $8 and $40.

Is Opus 5.5 cheaper than Opus 5?

Yes. Input and output prices are 20% lower and cache reads are 60% lower than on Opus 5. Anthropic says typical workloads cost about 40% less at default settings, partly because the model writes fewer tokens. Note that the default effort also dropped from high to medium, so check your own settings.

How much does prompt caching save on the Claude API?

On Opus 5.5 a cache read costs 5% of the normal input price. In my worked example of a long agent session, caching cut the cost from $13.00 to $1.75. A 5-minute cache pays for itself after one read, and a 1-hour cache after two reads.

Is the Claude API cheaper than a Claude subscription?

It depends on volume and who uses it. For a person chatting a few hours a day, a subscription seat is usually cheaper and predictable. For agents and software that run all day, the API is the only option that scales, and caching keeps the cost under control.

Is there a free tier for the Claude API?

New accounts get a small amount of free credit to test the API. There is no free monthly allowance, so after that every token bills at the rates above. Higher usage tiers raise your rate limits. They do not change the price per token.

Does Opus 5.5 cost more on Amazon Bedrock or Google Cloud?

On Amazon Bedrock and Google Cloud, the cloud provider sets and bills the price, so check its own price list before you commit. Regional endpoints, and multi-region endpoints on Google Cloud, cost 10% more than global endpoints. Fast mode is available only on the Claude API.

Sources

Anthropic, Introducing Claude Opus 5.5, September 22, 2026, read October 1, 2026.

Claude Platform Docs, Pricing, read October 1, 2026.

Claude Platform Docs, What's new in Claude Opus 5.5, read October 1, 2026.

Claude Blog, Claude Opus 5.5 is built for coding sessions that use more context, September 24, 2026, read October 1, 2026.

Anthropic, Anthropic and Amazon expand compute collaboration, April 20, 2026, read October 1, 2026.

Anthropic, Claude pricing, read October 1, 2026.

OpenAI, Introducing GPT-6.1 Sol, read October 1, 2026.

OpenAI, API pricing, read October 1, 2026.

Want a GTM engine that runs like this?

Growth Cab is the #1 GTM & sales advisory in the US & Europe. We build the outbound, LinkedIn, and closing systems behind these playbooks for founders selling high-ACV deals.

Apply to GC ← All articles
AI FRONTIER

Turn this week's AI noise
into one useful move

Every Thursday: the signal, Federico's direct view and one practical play. Join free and get the 10-page AI Frontier Operator Pack.

Free · under five minutes · unsubscribe anytime