AI

Anthropic cuts Claude Haiku 5.5 API price by 90 percent to match OpenAI GPT 6 Luna

Anthropic priced its newest small language model at $0.10 per million input tokens and $0.50 per million output tokens, matching OpenAI's GPT-6 Luna and cutting most Haiku bills by about 90 percent.

T
By TechQuire Daily Staff TechQuire Daily Staff
October 7, 2026 / 7 min read

Anthropic released Claude Haiku 5.5 on October 7, 2026, pricing its smallest and fastest model at $0.10 per million input tokens and $0.50 per million output tokens for prompts under 100,000 tokens. That is a 90 percent cut against the $1.00 and $5.00 per million tokens that Haiku 4.5, released in October 2025, charged for the same work.

The launch makes Haiku 5.5 the third model in the company's Claude 5.5 generation, arriving two weeks after Opus 5.5 on September 22 and following Sonnet 5.5 earlier in the fall. Anthropic describes it as the cheapest, fastest and most capable small model it has ever shipped.

Haiku models are built for the repetitive work that eats API budgets at scale: classifying support tickets, extracting fields from documents, summarizing long inputs, compacting conversation histories, running database queries and executing quick checks inside a coding agent's loop thousands of times a day. Anthropic positions the new model as a subagent that works alongside Opus 5.5 and Sonnet 5.5 rather than as a replacement for either.

The release lands in a crowded autumn for frontier labs. OpenAI shipped GPT-6 Sol and Luna on September 22 with API prices cut roughly 50 percent from the GPT-5.6 line, and xAI released Grok 4.7 the same week, according to Startup Fortune.

Key Facts

For prompts up to 100,000 tokens, which Anthropic said covered about 90 percent of requests to the previous Haiku model, Haiku 5.5 costs $0.10 per million input tokens and $0.50 per million output tokens. SiliconANGLE reported on October 7 that this is a 90 percent reduction on input and output alike, and that longer prompts receive a 50 percent discount, stepping up to $0.50 per million input tokens and $2.50 per million output tokens. Cache reads fall to $0.01 per million tokens from $0.10.

Those $0.10 and $0.50 rates match OpenAI's GPT-6 Luna exactly. Anthropic still frames the headline saving as roughly 75 percent cheaper on average, a figure that accounts for a new tokenizer which consumes slightly more tokens per task, so the nominal 90 percent price cut does not translate one for one into a 90 percent bill reduction.

Benchmarks favor Haiku 5.5 over Luna on every test where both have published scores. On OSWorld 2.1, which measures agents operating a real computer, Haiku 5.5 scores 72.4 percent against Luna's 48.9 percent and Haiku 4.5's 15.7 percent. On Terminal-Bench 4.0 for agentic coding it scores 39.2 percent where Luna scores 16.4 percent and the previous generation scored zero. On Humanity's Last Exam it reaches 45.9 percent without tools and 57.4 percent with them, up from 10.2 percent and 18.7 percent. Anthropic says Haiku 5.5 is also its fastest model to date, which matters as much as price for pipelines that fire thousands of sequential calls.

Other results put Haiku 5.5 at 1620 on GDPval-AA v2.1 for knowledge work, 1578 on AA-Briefcase v1.1, 46.4 percent on FrontierCode 1.1 (Main) and 46.4 percent on Chartography visual reasoning. Sonnet 5.5 scores higher on all four, at 1840, 1824, 52.1 percent and 61.6 percent respectively.

The model carries a one million token context window and a 128,000 token maximum output. Startup Fortune reported on October 8 that the context window is up from the 200,000 tokens that Haiku 4.5 shipped with a year ago. Haiku 5.5 is Anthropic's first Haiku class model with an adjustable effort setting, letting developers dial the same request toward lower cost or higher intelligence rather than switching model tiers entirely.

Analysis

Matching GPT-6 Luna's price to the cent is not a coincidence, and it is not primarily a capability story. What this really means is that the high volume tier has become a commodity market where the incumbent lab is willing to price at its rival's number and win the argument on benchmarks rather than on discounts. Anthropic did not undercut Luna. It met Luna and published a table showing 72.4 percent against 48.9 percent on computer use and 39.2 percent against 16.4 percent on agentic coding.

The 75 percent average figure is the more honest number for buyers building budgets, because the tokenizer change means a fixed workload may consume more tokens than it did on Haiku 4.5. Even so, the direction of travel is severe. A task that cost $5.00 per million output tokens on the older model now costs $0.50 below the threshold, and the 50 percent cut above it means long context work once reserved for bigger models is now priced within reach of batch pipelines. A 90 percent list price cut that lands as a 75 percent bill reduction is still the largest single price move Anthropic has made on a production model.

Anthropic's own guidance points complex agentic coding users to Sonnet 5.5 and Opus 5.5, which suggests the company is deliberately protecting its premium tiers rather than cannibalizing them. The Tech Portal reported on October 8 that the model is not meant to replace Sonnet or Opus for complex reasoning. The bigger picture here is that Anthropic is selling a ladder rather than a single model: Haiku for the thousands of cheap calls inside an agent loop, Sonnet for the strategic layer and Opus for the hardest reasoning.

The supporting price moves reinforce that reading. Anthropic halved Sonnet 5.5 cache reads from $0.20 to $0.10 per million tokens, and SiliconANGLE reported on October 7 that the change is expected to take about 20 percent off most agentic work on that model. The company also introduced monthly API credits for Claude Max and Team subscribers: $100 per month for Max 5x, double that for Max 20x, and up to $500 shared across Team accounts.

Why It Matters

Small models are where volume lives, and volume is where cloud AI revenue compounds. If 90 percent of requests to the previous Haiku fell under the 100,000 token threshold, then the price change touches nearly every production deployment built on that tier. Customers quoted by Anthropic describe the effect in operational terms rather than cost terms. Asana reported more than 30 percent lower task completion latency and up to 2.5 times faster inference per agent turn, and Aaron Vinh, a staff software engineer there, called it a noticeably snappier experience. HubSpot scored 92.8 percent averaged over three runs on a CRM evaluation suite. Box scored 11 points higher than Haiku 4.5 at about half the latency. AlphaSense measured 0.84 against 0.76 on 400 Ask in Document queries.

Cognition confirmed Haiku 5.5 as a subagent option in Devin Fusion, with the combination holding a FrontierCode score of 66.2 at reduced cost and latency. Devin Fusion pairs a larger model with Haiku as a sidekick, a structure that quietly multiplies the number of small model calls behind every coding task. That is a concrete demonstration of the subagent pattern Anthropic is promoting, and it is the pattern most likely to multiply token consumption across the industry.

Safety and access changed too. Anthropic said alignment testing turned up far fewer misaligned behaviors than Haiku 4.5, and its cybersecurity safeguards allow more defensive security work than Sonnet 5.5 permits, while penetration testing and other attacker techniques remain blocked. Organizations needing wider access can apply to the Cyber Verification Program, which Anthropic expanded on October 6.

Next Up

Haiku 5.5 is available immediately as claude-haiku-5-5 on the Claude Platform and through Amazon Web Services, Google Cloud and Microsoft Azure. The Tech Portal reported on October 8 that the launch completes Anthropic's current model lineup and arrives ahead of the company's planned IPO, which makes the next quarter a test of whether cheap tokens translate into paid volume.

Two questions remain open. Haiku 4.5 scored 73.3 percent on SWE-bench Verified at a third of Sonnet 4's cost, and Haiku 5.5 has no confirmed public SWE-bench number yet. With OpenAI's Luna sitting at the same $0.10 and $0.50 rates, the next move in the small model tier will be decided by benchmarks and latency rather than by list price.

Tagged

Comments (0)

No comments yet. Be the first to share your thoughts.