Claude Haiku 5.5 cuts small-model prices 90 percent, and Mistral previews a trillion-parameter Large 4
Plus: Reflection's open-weight Beam, Google's on-device EmbeddingGemma 2, and Apple tightens Mac access for AI agents.
The week on the beat
1. Mistral released Large 4 in public preview, a 1.05-trillion-parameter model it says will ship open weights at the end of October
It has 52B active parameters and a 1M context on Mistral's API, and it is on sale at $0.68/$2.09 per million, half its $1.36/$4.18 list, with no end date given. Artificial Analysis scores the preview 38.4 on its intelligence index. Mistral is pitching it on cybersecurity work and European data sovereignty; for builders, the open-weight release later this month is the part to wait for.
2. Nvidia-backed Reflection announced Beam, an open-weight model aimed squarely at the Chinese open models many builders use today
It is a 501B mixture-of-experts model with 23B active parameters, and Reflection says it matches GLM-5.2 on reasoning with three to four times less compute. Weights are promised under Apache 2.0 this month; for now access is through a waitlist.
3. Apple is tightening Full Disk Access on macOS because of AI agents
Apple says increasingly capable agents make broad access to users' files, messages, mail and browsing history riskier, and it will make users grant the permission more deliberately. Reuters ties the change to complaints about Meta's Muse reading Mac data. If your agent or desktop app relies on Full Disk Access, expect your users to see a more explicit permission flow.
4. Microsoft AI released MAI-Transcribe-2-Streaming, a real-time speech-to-text model for voice agents
MarkTechPost reports it ranks first of 38 models on Artificial Analysis's streaming word-error-rate leaderboard, and The Decoder says a new text-to-speech model shipped alongside it. Microsoft is aiming both at contact centres, so if live transcription has been the weak link in your voice agent, it is worth a test.
5. Google released EmbeddingGemma 2, an open 740M-parameter embedding model that runs on-device
It maps text, images, video, audio and code into one 768-dimension space, needs about 191 MB of RAM, ships under Apache 2.0, and Google says it beats some rival models twice its size. If your retrieval stack sends every document to a hosted embedding API, this is worth a local benchmark.
6. Anthropic cut small-model prices by 90 percent with Claude Haiku 5.5
Anthropic's pricing page lists it at $0.10/$0.50 per million for prompts up to 100K tokens, against Haiku 4.5's $1/$5, and The Decoder reports its OSWorld computer-use score jumped from 15.7 to 72.4 percent. One catch: a new tokenizer turns the same text into more tokens, so the real saving per task is smaller than 90 percent. If you run subagents or bulk summarization on a small model, this is the first thing to re-test this week.
Model moves
- New: Claude Haiku 5.5 (Anthropic). $0.10/$0.50 per 1M for prompts up to 100K tokens and $0.50/$2.50 above that, with a 1M context and $0.01 cache reads. Batch is half price. Model page →
- New: Mistral Large 4 (Mistral, public preview). On sale at $0.68/$2.09 per 1M from a $1.36/$4.18 list, with a 1M context on Mistral's API. Open weights are due at the end of October. Model page →
- New tier: GPT-6 Astra Ultrafast (OpenAI). $60/$300 per 1M, six times the standard rate, which OpenAI says is up to 6x faster in the API. It launched at DevDay on September 29. Model page →
- Benchmark revisions on Epoch's listings, October 5. APEX moved three models in both directions: Nemotron 3 Ultra from 11.5% to 22.7%, GLM-5.2 from 35.6% to 45.2%, and Muse Spark 1.1 down from 41.9% to 31.8%. Claude Sonnet 5.5's reported WebDev Arena score rose from 1,699 to 1,786. Recent model changes →
- No vendor list-price changes this week.
- Correction to issue #10: Grok 4.7's list price is $2/$6 per 1M. We printed $1.60/$4.80 on September 24. That was the rate OpenRouter was charging in Grok 4.7's first days, which our tracker took for xAI's list price. xAI's own page and launch coverage both put it at $2/$6 from the start, and OpenRouter moved to that rate in early October. Model page →
Personal take
Open-weight models used to be the cheaper option. This week, though, the best deal is a closed-model system. Claude Haiku 5.5 scores 43.4 on Artificial Analysis's index at maximum effort and costs about 21 cents per task. In comparison, Mistral Large 4, the main open-weight model in Europe, scores 38.4 and costs $1.13 per task at its listed price. Neither top open-weight model is currently available for download. Mistral plans to release theirs at the end of October, and Reflection says theirs will be out this month. Open weights still let you control where your data goes and what runs on your own hardware, which is enough for many organizations to make their choice. But if you picked an open-weight model mainly for the price, it's worth checking the numbers again this week.
Until next Thursday, Anmol
Get the next issue
Free, weekly, unsubscribe anytime. That’s the whole pitch.
Free forever. No spam. One-click unsubscribe. See our Privacy Policy.