On this page
- The three models side by side
- Claude Haiku 5.5: a tenth of the price, not a drop-in swap
- What returns an error after you change the model ID
- Mistral Large 4: a large open-weight model, in preview
- DeepSeek V4.1-Flash: the cheap one from September
- What one workload costs
- Which one to use
- Frequently asked questions
Two models with public APIs came out this week: Anthropic's Claude Haiku 5.5 on October 7, 2026, and Mistral Large 4 on October 6, as a preview. A third, older model kept coming up in the discussion: DeepSeek V4.1-Flash, released on September 10, mostly because of its price. They do different jobs. Below is what each costs, how to call it from Python, and what breaks if you only change the model ID. I had no API keys for this post, so I ran each official snippet with the current SDK against a local stub that records the request instead of sending it.
The three models side by side#
| Model | Released | API model ID | USD per million tokens, input / output | Context | Weights |
|---|---|---|---|---|---|
| Claude Haiku 5.5 | October 7, 2026 | claude-haiku-5-5 | $0.10 / $0.50 for prompts up to 100,000 tokens; $0.50 / $2.50 above | 1M, 128K output | Closed |
| Mistral Large 4 | October 6, 2026 (preview) | mistral-large-4 | $1.36 / $4.18 list; $0.68 / $2.09 sale price on the docs page | 1M | Promised by the end of October |
| DeepSeek V4.1-Flash | September 10, 2026 | deepseek-flash | $0.30 / $1.20 at peak hours, half that off-peak | 1M, 384K output | MIT license |
Prices come from Anthropic's model page, Mistral's model docs and DeepSeek's pricing page, read on October 10. Mistral gives no end date for its sale price. DeepSeek's input price is for cache misses; peak hours are 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays.
Claude Haiku 5.5: a tenth of the price, not a drop-in swap#
Anthropic positions Haiku 5.5 for high-volume work: classification, extraction, routing and subagents. Per token it costs 90 percent less than Haiku 4.5, which is $1 / $5. Two details change the arithmetic. A prompt over 100,000 tokens pays five times the rate, and that threshold counts cache reads and writes too. And the model uses Anthropic's newer tokenizer, so the same text counts as about 30 percent more tokens, according to the migration guide. Simon Willison measured about 1.25 times as many tokens on one long prompt.
What returns an error after you change the model ID#
The migration guide lists the changes. These are the ones existing Haiku 4.5 code is most likely to hit:
temperatureandtop_preturn a 400 error at anything but their defaults, andtop_kat any value. In the Python SDK, version 1.0 and later removed them: with anthropic 1.13.0, my test call stopped withTypeError: Messages.create() got an unexpected keyword argument 'temperature'before any request was sent. Code that setstemperature=0for repeatable classification is the usual case.thinking: {"type": "enabled", "budget_tokens": N}returns a 400. Thinking is now adaptive and on by default; you steer it withoutput_config.effort(lowtomax, defaultmedium).- A final assistant turn used as a prefill returns a 400, even with thinking off. Use structured outputs or move the instruction into the user turn.
- Responses can start with a
thinkingblock. Code that readsresponse.content[0].textbreaks; select blocks bytype. - Thinking tokens count toward
max_tokens. A small limit tuned for Haiku 4.5 can end withstop_reason: "max_tokens"before any text. - The
computer_20250124tool returns a 400, and on Amazon Bedrock structured outputs are not available for this model.
Here is a classification call that follows the new rules, adapted from the example in Anthropic's effort docs:
import anthropic
client = anthropic.Anthropic()
response = client.messages.create(
model="claude-haiku-5-5",
max_tokens=4096,
messages=[
{
"role": "user",
"content": "Classify this log line as bug, config or noise: "
"OperationalError: FATAL: password authentication failed for user app",
}
],
output_config={"effort": "low"},
)
for block in response.content:
if block.type == "text":
print(block.text)Against my stub, which answered with a thinking block followed by a text block, the loop printed only the text, and the SDK sent "output_config": {"effort": "low"} in the request body. For short, high-volume tasks, start at low: thinking tokens are billed as output, and the guide says the model can skip thinking on simple requests at lower effort.
Mistral Large 4: a large open-weight model, in preview#
Mistral Large 4 is a mixture-of-experts model with about 1.05 trillion parameters, 52 billion of them active per token, plus a 1.6 billion parameter vision encoder; it takes text and images. Today it is a public preview on Mistral's own API. The announcement says the weights follow at the end of October, but neither it nor the docs name a license yet, so do not plan a self-hosted deployment until the license is published. Its predecessor, Mistral Large 3, is Apache 2.0, costs $0.50 / $1.50 and has a 256K context.
On Mistral's own benchmark table, Large 4 scores 61.7 percent on DeepSWE v1.1 and 28.3 percent on Terminal-Bench 4.0. Vendor tables use their own setups, so treat them as a reason to test, not as a ranking. Mistral also says it runs a European deployment end to end, which matters if your data has to stay in the EU. The Python call, verbatim from the docs, ran unchanged with mistralai 3.2.0 against my stub:
import os
from mistralai.client import Mistral
client = Mistral(api_key=os.environ["MISTRAL_API_KEY"])
response = client.chat.complete(
model="mistral-large-4",
messages=[
{
"role": "user",
"content": "Write a Python function to merge two sorted linked lists into one sorted list.",
}
],
reasoning_effort="high",
)
print(response.choices[0].message.content)Note the import: in mistralai 3.2.0, the version I used, the client comes from mistralai.client. Because the model is a preview, name it explicitly and expect changes before general availability.
DeepSeek V4.1-Flash: the cheap one from September#
DeepSeek V4.1-Flash is not new, but on October 8 a Hacker News thread asking why the industry was not more worried about it passed 1,000 points. DeepSeek's announcement describes a mixture-of-experts model with 552 billion parameters, about 8 billion active for input. The weights are on Hugging Face under the MIT license. The old deepseek-v4-flash name is retired and routes to it. Thinking is on by default, and the API is OpenAI-compatible, so the official example uses the openai package. It ran unchanged with openai 3.28.0 against my stub:
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ.get('DEEPSEEK_API_KEY'),
base_url="https://api.deepseek.com")
response = client.chat.completions.create(
model="deepseek-flash",
messages=[
{"role": "system", "content": "You are a helpful assistant"},
{"role": "user", "content": "Hello"},
],
stream=False,
reasoning_effort="high",
extra_body={"thinking": {"type": "enabled"}}
)
print(response.choices[0].message.content)Before you send customer data to any third-party API, check where it is processed and what your contracts allow. With MIT weights, self-hosting is an option, but a 552-billion-parameter model needs serious hardware.
What one workload costs#
Price tables are hard to compare in your head, so here is one workload: 100,000 requests with 2,000 input tokens and a 200-token answer each. The script uses list prices and applies the 30 percent token increase to Haiku 5.5. It treats the other vendors' token counts as equal to Haiku 4.5's, which they are not exactly, and it leaves out thinking tokens, caching and batch discounts (50 percent at Anthropic). Change the numbers to match your own traffic.
"""Rough cost of one workload on each model, from list prices (USD per million tokens)."""
REQUESTS = 100_000
INPUT_TOKENS = 2_000 # per request, as counted on Claude Haiku 4.5
OUTPUT_TOKENS = 200 # visible answer only; thinking tokens are billed as output too
MODELS = {
# name: (input price, output price, token multiplier vs Haiku 4.5)
"Claude Haiku 4.5": (1.00, 5.00, 1.0),
"Claude Haiku 5.5": (0.10, 0.50, 1.3), # about 30% more tokens for the same text
"Mistral Large 4 (list)": (1.36, 4.18, 1.0),
"Mistral Large 4 (sale)": (0.68, 2.09, 1.0),
"DeepSeek V4.1-Flash (peak)": (0.30, 1.20, 1.0),
"DeepSeek V4.1-Flash (off-peak)": (0.15, 0.60, 1.0),
}
for name, (price_in, price_out, factor) in MODELS.items():
tokens_in = REQUESTS * INPUT_TOKENS * factor / 1e6
tokens_out = REQUESTS * OUTPUT_TOKENS * factor / 1e6
cost = tokens_in * price_in + tokens_out * price_out
print(f"{name:32} ${cost:8.2f}")| Model | Cost of the workload |
|---|---|
| Claude Haiku 4.5 | $300.00 |
| Claude Haiku 5.5 | $39.00 |
| Mistral Large 4, list price | $355.60 |
| Mistral Large 4, sale price | $177.80 |
| DeepSeek V4.1-Flash, peak | $84.00 |
| DeepSeek V4.1-Flash, off-peak | $42.00 |
Even with 30 percent more tokens, Haiku 5.5 comes out about 87 percent cheaper than Haiku 4.5 on this input-heavy workload. Mistral Large 4 is in a different weight class, and its price reflects that.
Which one to use#
- High-volume classification, extraction and subagents in an app already on Claude: Haiku 5.5 at
loweffort, after you have worked through the migration list above and rerun your evaluations. - Weights you can host yourself: DeepSeek V4.1-Flash today under MIT; Mistral Large 4 once its weights and license are out.
- Hard agentic coding: not Haiku 5.5. On Anthropic's own table it scores 39.2 percent on Terminal-Bench 4.0 against 70.6 percent for Sonnet 5.5, and Anthropic recommends Sonnet or Opus for that work. Try Mistral Large 4 on your own tasks during the preview. My GPT-6 guide covers OpenAI's options, and the coding agent trends post covers the tools around them.
- Mechanical edits: cheap models are good enough for rewrites like the ones in a Python 3.10 to 3.12 upgrade, but run pyupgrade and django-upgrade first. They are free and give the same result every time.
Frequently asked questions#
What is the model ID for Claude Haiku 5.5?
claude-haiku-5-5 on the Claude API, Google Cloud and Microsoft Foundry, and anthropic.claude-haiku-5-5 on Amazon Bedrock. It has no date suffix and no separate alias.
How much does Claude Haiku 5.5 cost?
$0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens, and $0.50 and $2.50 above that. Cache reads cost $0.01 per million tokens, and the Batch API is half price.
Why does my request to Claude Haiku 5.5 return a 400 error?
The usual causes are a non-default temperature, top_p or top_k, a thinking request with budget_tokens, an assistant prefill as the last message, or the old computer_20250124 tool. Each worked on Haiku 4.5 and is rejected on 5.5.
Is Mistral Large 4 open source?
Mistral calls it open-weight and says the weights will be released by the end of October 2026, but it has not named a license yet. Until then you can use it as a preview through Mistral's API.
Is DeepSeek V4.1-Flash a new model?
No. DeepSeek released it on September 10, 2026. It drew attention this week because of its price: $0.30 per million input tokens and $1.20 per million output tokens at peak hours, with open weights under the MIT license.
Comments
No comments yet. Be the first.