---
title: "Claude Haiku 5.5 and Mistral Large 4: Prices, APIs and Gotchas"
url: "https://nareshkumar.online/blog/claude-haiku-5-5-mistral-large-4"
author: "Naresh Kumar"
published: "2026-10-10"
updated: "2026-10-10"
category: "Engineering"
tags: ["AI Agents", "Developer Tools", "LLMs", "Python"]
---

# Claude Haiku 5.5 and Mistral Large 4: Prices, APIs and Gotchas

**In short:** Claude Haiku 5.5 costs $0.10 per million input tokens, a tenth of Haiku 4.5, but it is not a drop-in swap: temperature, thinking budgets and prefill now return errors, and the same text uses about 30 percent more tokens. Mistral Large 4 is a preview with weights due by the end of October.

Two models with public APIs came out this week: Anthropic's [Claude Haiku 5.5](https://www.anthropic.com/claude-haiku-5-5) on October 7, 2026, and [Mistral Large 4](https://mistral.ai/news/mistral-large-4/) on October 6, as a preview. A third, older model kept coming up in the discussion: DeepSeek V4.1-Flash, released on September 10, mostly because of its price. They do different jobs. Below is what each costs, how to call it from Python, and what breaks if you only change the model ID. I had no API keys for this post, so I ran each official snippet with the current SDK against a local stub that records the request instead of sending it.

## The three models side by side

| Model | Released | API model ID | USD per million tokens, input / output | Context | Weights |
| --- | --- | --- | --- | --- | --- |
| Claude Haiku 5.5 | October 7, 2026 | `claude-haiku-5-5` | $0.10 / $0.50 for prompts up to 100,000 tokens; $0.50 / $2.50 above | 1M, 128K output | Closed |
| Mistral Large 4 | October 6, 2026 (preview) | `mistral-large-4` | $1.36 / $4.18 list; $0.68 / $2.09 sale price on the docs page | 1M | Promised by the end of October |
| DeepSeek V4.1-Flash | September 10, 2026 | `deepseek-flash` | $0.30 / $1.20 at peak hours, half that off-peak | 1M, 384K output | MIT license |

Prices come from [Anthropic's model page](https://platform.claude.com/docs/en/models/haiku-5-5/overview), [Mistral's model docs](https://docs.mistral.ai/models/mistral-large-4-0) and [DeepSeek's pricing page](https://api-docs.deepseek.com/quick_start/pricing), read on October 10. Mistral gives no end date for its sale price. DeepSeek's input price is for cache misses; peak hours are 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays.

## Claude Haiku 5.5: a tenth of the price, not a drop-in swap

Anthropic positions Haiku 5.5 for high-volume work: classification, extraction, routing and subagents. Per token it costs 90 percent less than Haiku 4.5, which is $1 / $5. Two details change the arithmetic. A prompt over 100,000 tokens pays five times the rate, and that threshold counts cache reads and writes too. And the model uses Anthropic's newer tokenizer, so the same text counts as about 30 percent more tokens, according to the [migration guide](https://platform.claude.com/docs/en/models/haiku-5-5/migration-guide). Simon Willison [measured](https://simonwillison.net/2026/Oct/7/claude-haiku-5-5/) about 1.25 times as many tokens on one long prompt.

### What returns an error after you change the model ID

The migration guide lists the changes. These are the ones existing Haiku 4.5 code is most likely to hit:

- `temperature` and `top_p` return a 400 error at anything but their defaults, and `top_k` at any value. In the Python SDK, version 1.0 and later removed them: with anthropic 1.13.0, my test call stopped with `TypeError: Messages.create() got an unexpected keyword argument 'temperature'` before any request was sent. Code that sets `temperature=0` for repeatable classification is the usual case.
- `thinking: {"type": "enabled", "budget_tokens": N}` returns a 400. Thinking is now adaptive and on by default; you steer it with `output_config.effort` (`low` to `max`, default `medium`).
- A final assistant turn used as a prefill returns a 400, even with thinking off. Use structured outputs or move the instruction into the user turn.
- Responses can start with a `thinking` block. Code that reads `response.content[0].text` breaks; select blocks by `type`.
- Thinking tokens count toward `max_tokens`. A small limit tuned for Haiku 4.5 can end with `stop_reason: "max_tokens"` before any text.
- The `computer_20250124` tool returns a 400, and on Amazon Bedrock structured outputs are not available for this model.

Here is a classification call that follows the new rules, adapted from the example in Anthropic's [effort docs](https://platform.claude.com/docs/en/build-with-claude/effort):

`classify.py`

```python
import anthropic

client = anthropic.Anthropic()

response = client.messages.create(
    model="claude-haiku-5-5",
    max_tokens=4096,
    messages=[
        {
            "role": "user",
            "content": "Classify this log line as bug, config or noise: "
            "OperationalError: FATAL: password authentication failed for user app",
        }
    ],
    output_config={"effort": "low"},
)

for block in response.content:
    if block.type == "text":
        print(block.text)
```

Against my stub, which answered with a thinking block followed by a text block, the loop printed only the text, and the SDK sent `"output_config": {"effort": "low"}` in the request body. For short, high-volume tasks, start at `low`: thinking tokens are billed as output, and the guide says the model can skip thinking on simple requests at lower effort.

> **Tip:** Claude Haiku 4.5 is still listed as active on Anthropic's [deprecations page](https://platform.claude.com/docs/en/about-claude/model-deprecations), and Anthropic promises at least 60 days' notice before it retires a model. You have time to recount your prompts and rerun your evaluations before switching. If you use GitHub Copilot, Haiku 5.5 is [generally available](https://github.blog/changelog/2026-10-07-claude-haiku-5-5-in-github-copilot) there on the Pro, Pro+, Max, Business and Enterprise plans, rolling out gradually.

## Mistral Large 4: a large open-weight model, in preview

Mistral Large 4 is a mixture-of-experts model with about 1.05 trillion parameters, 52 billion of them active per token, plus a 1.6 billion parameter vision encoder; it takes text and images. Today it is a public preview on Mistral's own API. The announcement says the weights follow at the end of October, but neither it nor the docs name a license yet, so do not plan a self-hosted deployment until the license is published. Its predecessor, [Mistral Large 3](https://docs.mistral.ai/models/mistral-large-3-25-12), is Apache 2.0, costs $0.50 / $1.50 and has a 256K context.

On Mistral's own benchmark table, Large 4 scores 61.7 percent on DeepSWE v1.1 and 28.3 percent on Terminal-Bench 4.0. Vendor tables use their own setups, so treat them as a reason to test, not as a ranking. Mistral also says it runs a European deployment end to end, which matters if your data has to stay in the EU. The Python call, verbatim from the docs, ran unchanged with mistralai 3.2.0 against my stub:

`mistral_large.py`

```python
import os
from mistralai.client import Mistral

client = Mistral(api_key=os.environ["MISTRAL_API_KEY"])

response = client.chat.complete(
    model="mistral-large-4",
    messages=[
        {
            "role": "user",
            "content": "Write a Python function to merge two sorted linked lists into one sorted list.",
        }
    ],
    reasoning_effort="high",
)

print(response.choices[0].message.content)
```

Note the import: in mistralai 3.2.0, the version I used, the client comes from `mistralai.client`. Because the model is a preview, name it explicitly and expect changes before general availability.

## DeepSeek V4.1-Flash: the cheap one from September

DeepSeek V4.1-Flash is not new, but on October 8 a Hacker News thread asking why the industry was not more worried about it passed 1,000 points. DeepSeek's [announcement](https://api-docs.deepseek.com/news/news260910) describes a mixture-of-experts model with 552 billion parameters, about 8 billion active for input. The weights are on [Hugging Face](https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash) under the MIT license. The old `deepseek-v4-flash` name is retired and routes to it. Thinking is on by default, and the API is OpenAI-compatible, so the official example uses the `openai` package. It ran unchanged with openai 3.28.0 against my stub:

`deepseek_flash.py`

```python
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ.get('DEEPSEEK_API_KEY'),
    base_url="https://api.deepseek.com")

response = client.chat.completions.create(
    model="deepseek-flash",
    messages=[
        {"role": "system", "content": "You are a helpful assistant"},
        {"role": "user", "content": "Hello"},
    ],
    stream=False,
    reasoning_effort="high",
    extra_body={"thinking": {"type": "enabled"}}
)

print(response.choices[0].message.content)
```

Before you send customer data to any third-party API, check where it is processed and what your contracts allow. With MIT weights, self-hosting is an option, but a 552-billion-parameter model needs serious hardware.

## What one workload costs

Price tables are hard to compare in your head, so here is one workload: 100,000 requests with 2,000 input tokens and a 200-token answer each. The script uses list prices and applies the 30 percent token increase to Haiku 5.5. It treats the other vendors' token counts as equal to Haiku 4.5's, which they are not exactly, and it leaves out thinking tokens, caching and batch discounts (50 percent at Anthropic). Change the numbers to match your own traffic.

`llm_cost.py`

```python
"""Rough cost of one workload on each model, from list prices (USD per million tokens)."""

REQUESTS = 100_000
INPUT_TOKENS = 2_000   # per request, as counted on Claude Haiku 4.5
OUTPUT_TOKENS = 200    # visible answer only; thinking tokens are billed as output too

MODELS = {
    # name: (input price, output price, token multiplier vs Haiku 4.5)
    "Claude Haiku 4.5": (1.00, 5.00, 1.0),
    "Claude Haiku 5.5": (0.10, 0.50, 1.3),          # about 30% more tokens for the same text
    "Mistral Large 4 (list)": (1.36, 4.18, 1.0),
    "Mistral Large 4 (sale)": (0.68, 2.09, 1.0),
    "DeepSeek V4.1-Flash (peak)": (0.30, 1.20, 1.0),
    "DeepSeek V4.1-Flash (off-peak)": (0.15, 0.60, 1.0),
}

for name, (price_in, price_out, factor) in MODELS.items():
    tokens_in = REQUESTS * INPUT_TOKENS * factor / 1e6
    tokens_out = REQUESTS * OUTPUT_TOKENS * factor / 1e6
    cost = tokens_in * price_in + tokens_out * price_out
    print(f"{name:32} ${cost:8.2f}")
```

| Model | Cost of the workload |
| --- | --- |
| Claude Haiku 4.5 | $300.00 |
| Claude Haiku 5.5 | $39.00 |
| Mistral Large 4, list price | $355.60 |
| Mistral Large 4, sale price | $177.80 |
| DeepSeek V4.1-Flash, peak | $84.00 |
| DeepSeek V4.1-Flash, off-peak | $42.00 |

Even with 30 percent more tokens, Haiku 5.5 comes out about 87 percent cheaper than Haiku 4.5 on this input-heavy workload. Mistral Large 4 is in a different weight class, and its price reflects that.

## Which one to use

- **High-volume classification, extraction and subagents in an app already on Claude:** Haiku 5.5 at `low` effort, after you have worked through the migration list above and rerun your evaluations.
- **Weights you can host yourself:** DeepSeek V4.1-Flash today under MIT; Mistral Large 4 once its weights and license are out.
- **Hard agentic coding:** not Haiku 5.5. On Anthropic's own table it scores 39.2 percent on Terminal-Bench 4.0 against 70.6 percent for Sonnet 5.5, and Anthropic recommends Sonnet or Opus for that work. Try Mistral Large 4 on your own tasks during the preview. My [GPT-6 guide](https://nareshkumar.online/blog/gpt-6-intelligent-ui-developers) covers OpenAI's options, and the [coding agent trends](https://nareshkumar.online/blog/ai-coding-agent-trends) post covers the tools around them.
- **Mechanical edits:** cheap models are good enough for rewrites like the ones in a [Python 3.10 to 3.12 upgrade](https://nareshkumar.online/blog/python-3-10-end-of-life), but run pyupgrade and django-upgrade first. They are free and give the same result every time.

## Frequently asked questions

**What is the model ID for Claude Haiku 5.5?**

`claude-haiku-5-5` on the Claude API, Google Cloud and Microsoft Foundry, and `anthropic.claude-haiku-5-5` on Amazon Bedrock. It has no date suffix and no separate alias.

**How much does Claude Haiku 5.5 cost?**

$0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens, and $0.50 and $2.50 above that. Cache reads cost $0.01 per million tokens, and the Batch API is half price.

**Why does my request to Claude Haiku 5.5 return a 400 error?**

The usual causes are a non-default `temperature`, `top_p` or `top_k`, a thinking request with `budget_tokens`, an assistant prefill as the last message, or the old `computer_20250124` tool. Each worked on Haiku 4.5 and is rejected on 5.5.

**Is Mistral Large 4 open source?**

Mistral calls it open-weight and says the weights will be released by the end of October 2026, but it has not named a license yet. Until then you can use it as a preview through Mistral's API.

**Is DeepSeek V4.1-Flash a new model?**

No. DeepSeek released it on September 10, 2026. It drew attention this week because of its price: $0.30 per million input tokens and $1.20 per million output tokens at peak hours, with open weights under the MIT license.
