Growth Cab Apply to GC
Blog/AI TOOLS
AI TOOLS · October 8, 2026 · 12 MIN READ

Claude Haiku Pricing: What Haiku 5.5 Really Costs

Claude Haiku pricing for Haiku 5.5 and 4.5: $0.10 input and $0.50 output under 100K tokens, the long-prompt catch, real cost examples and when to use it.

Federico DonatoneBy · Founder & CEO, Growth Cab
Claude Haiku Pricing: What Haiku 5.5 Really Costs

Claude Haiku 5.5 costs $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens. Above that, the price rises to $0.50 and $2.50. Haiku 4.5 still costs $1 and $5. Cache reads start at $0.01 per million, and the Batch API halves every price.

I posted about Haiku 5.5 on LinkedIn on October 8, 2026, the day after Anthropic released it, and opened with one line: the price is a joke. It is. It also hides three catches: a second price tier for long prompts, a tokenizer that counts more tokens for the same text, and thinking that is on by default. Here is what you actually pay, with the math.

Federico Donatonein
Federico Donatone
Founder & CEO, Growth Cab · This article started as a LinkedIn post

“Anthropic just dropped Haiku 5.5. The price is a joke.”

Read the original post →

Claude Haiku Pricing in 2026: What Does Each Haiku Cost?

Haiku 5.5 is the cheapest Claude model Anthropic has ever listed. On short prompts it costs a tenth of Haiku 4.5 per token, on every line of the bill. The table puts both Haiku models next to Sonnet 5.5 and Opus 5.5, from Anthropic's pricing page.

Claude API prices per million tokens in USD, from Anthropic's pricing page, read October 8, 2026
ModelInput5-min cache write1-hour cache writeCache readOutputBatch input / output
Claude Haiku 5.5, prompts up to 100K tokens$0.10$0.125$0.20$0.01$0.50$0.05 / $0.25
Claude Haiku 5.5, prompts over 100K tokens$0.50$0.625$1$0.05$2.50$0.25 / $1.25
Claude Haiku 4.5$1$1.25$2$0.10$5$0.50 / $2.50
Claude Sonnet 5.5$2$2.50$4$0.10$10$1 / $5
Claude Opus 5.5$4$5$8$0.20$20$2 / $10

Anthropic says prompts up to 100,000 tokens make up around 90% of the requests that went to Haiku 4.5. So for most teams the first row is the price that matters. The second row is the one that surprises people, and I explain it below.

Three more lines on the bill are worth knowing before you budget.

For every other Claude model, from Sonnet 5.5 to Fable 5.1, see my full guide to Claude API pricing.

Is Haiku 5.5 Really 75% Cheaper Than Haiku 4.5?

On average, yes, by Anthropic's own estimate. Per token it is 90% cheaper on short prompts and 50% cheaper on long ones. Anthropic's 75% is an average across real request sizes and token use. The first thing that pulls it below the list cut is a newer tokenizer that counts the same text as about 30% more tokens.

Here is how that plays out. A text that is 1,000 tokens on Haiku 4.5 becomes about 1,300 on Haiku 5.5. Send a thousand of them on short prompts and you pay $0.13 instead of $1, an 87% cut. On long prompts you pay $0.65 instead of $1, a 35% cut. The headline 90% holds only before the recount.

The second is thinking. Haiku 5.5 thinks by default, at medium effort, and thinking tokens are billed as output. Haiku 4.5 did not think unless you asked. If your old prompts got short answers, set a lower effort for simple jobs and measure the output tokens for a week.

Does Haiku 5.5 Cost More per Task Than Haiku 4.5?

On long agentic jobs it can. Vals AI, an independent evaluator, ran both models on Vibe Code Bench, a test where the model builds a web app from scratch. Haiku 5.5 scored 90.4% at $6.07 per test. Haiku 4.5 with thinking scored 11.4% at $1.31. The new model wins by a mile, and it spends far more tokens to get there.

Vibe Code Bench v1.1 results on the Vals AI leaderboard, OpenHands harness, read October 8, 2026
ModelScoreCost per testList price per million tokens
Claude Sonnet 5.592.4%$31.25$2 / $10
Claude Haiku 5.590.4%$6.07$0.10 / $0.50
Claude Opus 5.590.3%$57.92$4 / $20
GPT-6 Luna81.7%$1.35$0.10 / $0.50
Claude Haiku 4.5 (Thinking)11.4%$1.31$1 / $5

Read that table three ways. Haiku 5.5 edges Opus 5.5, 90.4% to 90.3%, at about a tenth of the cost per test. Against Sonnet 5.5 it gets within two points for a fifth of the cost. Against GPT-6 Luna, it scores nine points higher but costs 4.5 times more per test, at the same list price. The gap is the number of tokens each model burns on the job.

So the 75% saving is real for short, well-defined requests, where there is little room to think. On open-ended agent work, the effort setting decides the bill. Start new workloads at low effort, raise it only where quality drops, and track cost per finished task. Cheap per token. Sometimes expensive per job.

What Is the 100,000-Token Catch?

Haiku 5.5 is the only current Claude model with two prices by prompt length. Anthropic's long-context note says Claude 4.6 and later models bill a 900,000-token request at the same rate as a 9,000-token one, with Haiku 5.5 as the exception. Cross 100,000 tokens and every line of the bill is five times higher.

So one long prompt can cost more than several short ones. Say you need a summary of a 150,000-token contract. In one request the input costs about $0.075. Split into two halves of 75,000 tokens, it costs $0.015. Same words, a fifth of the price, plus a small step at the end to merge the two summaries.

Anthropic prices cache reads by the same prompt-length tier, so treat cached tokens as part of the 100,000. Check which tier your requests land in. The usage data in the Claude Console shows it before you design around it.

This is the checklist I would run before sending real volume to Haiku 5.5.

  1. Count tokens with the model set to claude-haiku-5-5. Counts made on Haiku 4.5 run about 30% low.
  2. Keep the system prompt, tools and examples lean. They count on every request.
  3. Split long documents into chunks under 100,000 tokens and merge the results.
  4. In long chats, replace old turns with a short summary before the history crosses the line.
  5. If a job must read more than 100,000 tokens, Haiku 5.5 at $0.50 still costs a quarter of Sonnet 5.5. Use the Batch API to halve it again.

What Does Haiku 5.5 Cost on Real Work?

In my worked examples, a support classifier that costs $212 a month on Haiku 4.5 costs $28 to $43 on Haiku 5.5. These are illustrations built on stated assumptions, and I did not measure them. Every row applies the 30% tokenizer gap from Anthropic's docs and leaves out cache writes.

Three workloads priced on list rates of October 8, 2026 (my assumptions, token counts as measured on Haiku 5.5, USD)
WorkloadHaiku 4.5Haiku 5.5Sonnet 5.5
100,000 support tickets a month, 2,000 input and 150 output tokens each, low effort with no thinking$212$28$550
The same tickets with 300 extra thinking tokens each$212 (no thinking)$43$850
1,000 subagent tasks, each with 20,000 cached tokens, 1,000 fresh and 500 output$4.23$0.55$9.00

The pattern is clear: the smaller and more repetitive the job, the bigger the gap. On the subagent run, the same 1,000 tasks would cost $18 on Opus 5.5, about 33 times more. That is the setup from my post. Let the expensive model plan the work, and let Haiku run the thousand small tasks underneath it, for cents.

If you are building the planner and the workers yourself, my guide on how to build an AI agent with Claude shows the setup step by step.

Claude Haiku 5.5 vs GPT-6 Luna: Which Is Cheaper?

On list price they are tied. GPT-6 Luna, OpenAI's small model, costs $0.10 per million input tokens, $0.01 for cached input and $0.50 for output. That is exactly Haiku 5.5's price under 100,000 tokens. The differences sit in the fine print, so here they are side by side.

Claude Haiku 5.5 and GPT-6 Luna, from Anthropic's and OpenAI's pricing pages and Anthropic's launch post, read October 8, 2026
QuestionClaude Haiku 5.5GPT-6 Luna
Input / output per million tokens$0.10 / $0.50 under 100K tokens$0.10 / $0.50
Cached input per million tokens$0.01 under 100K tokens$0.01
Where the price changesPrompts over 100,000 tokens pay $0.50 / $2.50Listed rates apply to context under 272,000 tokens
Batch discount50%50%
Computer use, OSWorld 2.1 (Anthropic's test)72.4%48.9%
Agentic coding, Terminal-Bench 4.0 (Anthropic's test)39.2%16.4%
Vibe Code Bench v1.1 score and cost per test (Vals AI)90.4% at $6.0781.7% at $1.35

Tokens are not equal across vendors. Each company counts text with its own tokenizer, so the same prompt can cost a different amount at the same list price. Run 100 of your real requests through both models and put the two bills side by side.

On Anthropic's benchmarks, Haiku 5.5 beats Luna on computer use and agentic coding by a wide margin. Vals AI's independent test agrees on quality, and also shows Haiku 5.5 spending far more per finished task. Pick Haiku 5.5 when the extra quality pays for itself, and Luna when the job is simple enough for either.

When Should You Use Haiku 5.5, and When Should You Skip It?

Use Haiku 5.5 for narrow jobs you run thousands of times. Anthropic points it at summaries, compaction, database queries, classification, live customer support, browser use and subagent work. For complex agentic coding, Anthropic itself says Sonnet 5.5 and Opus 5.5 remain the better choices.

Which model I would pick by job (my judgment, October 8, 2026)
JobMy pickWhy
Classify, tag or route leads and ticketsHaiku 5.5$0.10 input on short prompts
Extract fields from emails, call notes and formsHaiku 5.5 on Batch$0.05 input when results can wait
Subagent tasks under a planner modelHaiku 5.5Cache reads at $0.01
Live chat and support repliesHaiku 5.5Anthropic's fastest model
Terminal work and multi-step codingSonnet 5.5 or Opus 5.5Anthropic says they stay ahead on Terminal-Bench 4.0
Building a small web app from scratchHaiku 5.590.4% on Vibe Code Bench, about a tenth of Opus 5.5 per test
One 300,000-token documentHaiku 5.5, split into chunksFive times cheaper per token under 100K

In outbound, this is the model I would put on the boring layers: cleaning lists, tagging replies, scoring accounts against an ideal customer profile. Those jobs are small, frequent and easy to check. A wrong tag costs a few cents. A wrong strategy costs a quarter, so that one stays with Opus and with people.

AI FRONTIER
Get one useful AI play every Thursday
The AI changes that matter, Federico's direct read and one practical play, plus the free 10-page Operator Pack.
Free · under five minutes · unsubscribe anytime

Can You Run Haiku 5.5 for Free?

Almost, if you already pay for Claude Max or Team. Anthropic is rolling out monthly API credits this week: $100 on Max 5x, $200 on Max 20x, and $20 or $100 per Team seat, pooled up to $500. The credits work on any model on the Claude Platform. They do not cover interactive Claude Code sessions.

At Haiku 5.5 prices, $100 buys one billion input tokens on short prompts. That covers the ticket classifier above two to three and a half times over, every month. Pro plans get no API credit. The credits go to one Console organization, and the offer appears after seven days on an eligible plan.

For what each subscription includes, from Free to Enterprise, see my guide to Claude pricing.

How Do You Switch to Haiku 5.5 Without Surprises?

The model ID is claude-haiku-5-5, and it is live on the Claude API, Amazon Web Services, Google Cloud and Microsoft Azure. Before you move production traffic, check these seven changes from Anthropic's migration guide. Each one breaks old code or an old budget.

  1. Recount tokens with the new model. Old counts run about 30% low, so budgets and max_tokens need updating.
  2. Remove temperature, top_p and top_k. Any non-default value returns a 400 error.
  3. Replace budget_tokens with adaptive thinking and the effort setting. Manual extended thinking now returns an error.
  4. End every request with a user message. Assistant prefill returns an error.
  5. Read content blocks by type. A response can now start with a thinking block.
  6. Handle stop_reason refusal. Safety classifiers can decline a request, and there is no server-side fallback.
  7. Move computer use to computer_toolset_20260801 on the Claude API and Google Cloud.

And if you want agents like these inside a real outbound motion, with cheap models on the routine work and people on the judgment calls, that is what our B2B lead generation service does.

Frequently asked questions

How much does Claude Haiku cost?

Claude Haiku 5.5 costs $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens, and $0.50 and $2.50 above that. Haiku 4.5 costs $1 and $5. The Batch API halves both models' prices, and Haiku 5.5 cache reads start at $0.01 per million tokens.

Is Claude Haiku 5.5 free?

No. The API bills every token, and Pro plans get no API credit. Claude Max and Team subscribers now get monthly API credits, from $100 on Max 5x to a $500 pool on Team, and those credits cover Haiku 5.5 and every other model on the Claude Platform.

What is the context window of Claude Haiku 5.5?

One million tokens, up from 200,000 on Haiku 4.5, with up to 128,000 output tokens per response. You can use the full window, but any prompt over 100,000 tokens pays the higher price tier: $0.50 per million input tokens and $2.50 per million output tokens.

Is Haiku 5.5 cheaper than GPT-6 Luna?

Under 100,000 tokens the list prices are identical: $0.10 input, $0.01 cached input and $0.50 output per million tokens. Above 100,000 tokens Haiku 5.5 costs five times more. Each vendor counts tokens with its own tokenizer, so compare the real bill on the same set of requests.

Should I still use Claude Haiku 4.5?

For new work, no. Haiku 5.5 is cheaper per token on every price line, has a five times larger context window and scores far higher on benchmarks. On long agent jobs it can spend more per task. Keep Haiku 4.5 if you rely on Priority Tier, which Haiku 5.5 does not support, or while you migrate old code.

Sources

Anthropic, Introducing Claude Haiku 5.5, October 7, 2026, read October 8, 2026.

Claude Platform Docs, Pricing, read October 8, 2026.

Claude Platform Docs, Claude Haiku 5.5 overview, read October 8, 2026.

Claude Platform Docs, What's new in Claude Haiku 5.5, read October 8, 2026.

Claude Platform Docs, Claude Haiku 5.5 migration guide, read October 8, 2026.

Claude Help Center, Monthly API credits for Max and Team plans, read October 8, 2026.

OpenAI, API pricing, read October 8, 2026.

Vals AI, Vibe Code Bench leaderboard, read October 8, 2026.

Want a GTM engine that runs like this?

Growth Cab is the #1 GTM & sales advisory in the US & Europe. We build the outbound, LinkedIn, and closing systems behind these playbooks for founders selling high-ACV deals.

Apply to GC ← All articles
AI FRONTIER

Turn this week's AI noise
into one useful move

Every Thursday: the signal, Federico's direct view and one practical play. Join free and get the 10-page AI Frontier Operator Pack.

Free · under five minutes · unsubscribe anytime