Home  ›  Insights  ›  Claude Haiku 5.5 for service teams
Field Note

Claude Haiku 5.5 for Salesforce service teams: the call

Anthropic cut its small model’s price by up to 90% and it is the top AI story on Hacker News today. On short prompts the cut is real. On long case histories a new tokenizer and a 100,000-token price tier take most of it back. Here is the arithmetic, and what to test this week.

The call

PILOT: shadow-test Claude Haiku 5.5 on your high-volume service work now. Priced, generally available on every major cloud, and ahead of GPT-6 Luna on an independent index. The Salesforce service lead and the AI platform owner should run it against their current model this week, measured per case, not per token.

Anthropic released Haiku 5.5 on 7 October (Anthropic). By 8 October it led the Hacker News AI front page with 702 points and 349 comments (independent, HN digest).

What it is, and is not

It is Anthropic’s small model for summaries, classification, database queries, and sub-agent work, now with adjustable effort settings (vendor, Anthropic). It lists $0.10 per million input tokens and $0.50 per million output for requests up to 100,000 tokens, against $1.00 and $5.00 for Haiku 4.5 (vendor). Artificial Analysis scores it 43 on its Intelligence Index (independent), against 38 for GPT-6 Luna (independent). Context is 1 million tokens (independent, Artificial Analysis). HubSpot reports 92.8% on its internal CRM suite, averaged over three runs (vendor page, customer-reported).

It is not a flat 90% saving. Anthropic’s own footnote gives 90% up to 100,000 tokens, 50% above it, and about 75% on average (vendor). It is not cheaper than Luna per token either: both list $0.10 input, $0.50 output, and $0.01 for cache reads (vendor). And it is not a Sonnet replacement for complex agentic coding; Anthropic says so itself.

The number nobody is quoting

Coverage quotes 90%. Two details on Anthropic’s own pages change it. The new tokenizer counts roughly 30% more tokens for the same text, and requests over 100,000 tokens move to a tier priced five times higher. Anthropic’s 50% figure above that line only works if the whole request is billed at the higher rate, so that is how we count it.

FigureValueLabel
Haiku 5.5 input, request up to 100,000 tokens$0.10 per millionvendor
Haiku 5.5 input, request over 100,000 tokens$0.50 per millionvendor
Haiku 4.5 input, any length$1.00 per millionvendor
New tokenizer, same textroughly 30% more input tokens than Haiku 4.5vendor migration guide, via Kingy AI
Prompt of 10,000 Haiku 4.5 tokens4.5: 10,000 × $1.00/M = $0.0100. 5.5: 13,000 × $0.10/M = $0.0013. 87% cheaper, not 90%our maths
Where the cliff starts100,000 ÷ 1.3 = about 76,900 Haiku 4.5 tokensour maths
Prompt of 80,000 Haiku 4.5 tokens4.5: $0.0800. 5.5: 104,000 × $0.50/M = $0.0520. 35% cheaper, not 90%our maths
Either side of the cliff on Haiku 5.599,000 tokens: $0.0099. 101,000 tokens: $0.0505. 5.1x the bill for 2% more textour maths
GPT-6 Luna input, up to 272,000 tokens$0.10 per million, so a 150,000-token request is $0.015 on Luna and $0.075 on Haiku 5.5, 5xvendor price; tier via Kingy AI; ratio is our maths

Most service prompts are short, so most of the saving holds. The ones that are not short are the ones you care about: a full case history, a long email thread, a contract, a knowledge article set. Any prompt that was about 77,000 tokens on Haiku 4.5 now lands in the expensive tier. Price per token is the headline. Price per task is the business case.

Missing:

  • Whether the roughly 30% tokenizer uplift also applies to output tokens.
  • Batch and regional pricing on the launch page.
  • Data residency and retention terms for Haiku 5.5 on each cloud.
  • Rate limits for the new model.
  • A retirement date for Haiku 4.5, so you know how long the old baseline lasts.
  • Independent speed and cost-to-run figures. Artificial Analysis lists both as not yet available.
  • Einstein Studio availability. Salesforce has not announced it.

Who owns what

Price per token is Anthropic’s job. Price per task is ours.

In a Salesforce estate, call it from Apex through a Named Credential to the Claude API or the cloud you already buy from, triggered by a Flow on case create or close. Write the model, token count, and cost back to the case as fields so service ops can see cost per case in a report. Ground it with Data Cloud only once you have capped prompt length. For pure routing questions, OpenAI’s Decisions API went to public beta on 6 October at $0.10 per million input tokens with no output charge (vendor, OpenAI forum); put it in the same shadow run.

Monday morning

  • Step 1. Pull 500 real prompts from your top three AI workloads and count them on Haiku 5.5 with Anthropic’s token-counting endpoint. Owner: AI platform lead.
  • Step 2. Flag every prompt above 95,000 tokens. Summarise first, split, or route those to Luna. Owner: Salesforce architect.
  • Step 3. Shadow-run Haiku 5.5 against your current model on the same cases. Compare accuracy and cost per case. Owner: service operations lead.
  • Step 4. Add the Decisions API to the routing slice of the same run. Owner: AI lead.
  • Step 5. Confirm region and retention terms for your cloud before any customer data goes in. Owner: CISO.

What would change our call: we move this to WATCH if your own token counts erase the saving on your largest workload, or if independent speed and cost figures land well below Anthropic’s tables.

Living page. Last updated 8 Oct 2026. What would change our call: per-task cost on real prompts, or independent speed and cost figures.

Frequently asked questions

How much cheaper is Claude Haiku 5.5 than Haiku 4.5?

Anthropic says 90% cheaper for requests up to 100,000 tokens, 50% above that, and about 75% on average. Its migration guide also says the new tokenizer counts roughly 30% more input tokens for the same text. On our arithmetic that makes the input saving about 87% on short prompts and about 35% once a request crosses 100,000 tokens.

Is Claude Haiku 5.5 cheaper than GPT-6 Luna?

Not per token up to 100,000 tokens: both list $0.10 per million input tokens, $0.50 per million output, and $0.01 for cache reads. Above 100,000 tokens Haiku 5.5 input rises to $0.50, five times Luna’s rate. The tokenizers differ, so compare cost per task on your own prompts. Artificial Analysis scores Haiku 5.5 at 43 on its Intelligence Index against 38 for Luna.

Can we call Claude Haiku 5.5 from Salesforce?

Yes, through an Apex callout to the Claude API, Amazon Bedrock, Google Cloud, or Microsoft Azure, with the endpoint and key held in a Named Credential. Salesforce has not announced Haiku 5.5 in Einstein Studio’s model list, so check before you plan a Prompt Builder or Agentforce rollout on it.