OpenAI cut GPT-5.6 Luna by 80%, and Meta shipped an open model that runs on your laptop
Plus: Anthropic starts watermarking everything Claude writes, Mistral's 3B safety model, and 71 benchmark scores that moved without a single model changing.
The week on the beat
1. OpenAI cut GPT-5.6 Luna by 80 percent
Luna went from $1/$6 to $0.20/$1.20 per 1M, and Terra dropped 20 percent to $2/$12. Sol, the flagship, was untouched, so this is aimed squarely at the high-volume tier. One thing worth checking before you switch: our tracker has the cheapest third-party reseller for Terra at $2.20/$13.20, which is above OpenAI's own list price, so going direct is currently cheaper than the discount route.
2. Meta's Muse Glimmer runs on your own machine
An open-source multimodal model built to run locally on ordinary hardware, downloadable and modifiable. Between this and the Kimi K3 weights landing last month, the "frontier-adjacent model you can actually host" shelf now has real options on it, which changes the calculus for anyone whose blocker was sending data to someone else's endpoint.
3. Anthropic is watermarking everything Claude writes
Invisible watermarks plus C2PA signatures on all Claude-generated text and files, rolling into new models through August. If you ship Claude output as part of your product, this is worth ten minutes of your attention: the marks are designed to survive downstream handling, so it is now a question of what your users are told, not whether the provenance travels.
4. Mistral's Shieldstral does safety screening at 3B parameters
A small open model for checking inputs and outputs against safety rules, and the useful part is that the rules are natural language prompts rather than fixed categories, so you can change what it screens for at runtime. If you have been paying a frontier model to moderate your traffic, this is the cheaper shape.
5. Alibaba says Qwen3.8-Max matches the frontier
The most-covered story of the fortnight, 78 outlets. Alibaba's new flagship lands at $2/$6 per 1M with a 1M context window and reported scores level with top-tier Western models. Treat the scores as claims until someone independent runs them, for reasons the model moves section makes uncomfortably clear.
6. Apple is shipping Qwen inside macOS, in China
Apple integrated Alibaba's models into macOS for Chinese users, which is a regulatory compliance move rather than a global product change. Worth reading precisely: this is not Qwen arriving on every Mac. It is what shipping AI into a regulated market now looks like.
Model moves
- Two benchmark suites were re-scored in one week, and not one model changed. On Aug 6, 71 models' Humanity's Last Exam scores all moved up on the same day, between 5 and 18 percent, none down. On Aug 5, ARC-AGI moved five models, including Gemini 3 Flash from 21.5 to 84.7 percent. Nothing was retrained. This is what a provider re-running its evaluation looks like, and it is a good reason to hold benchmark gaps loosely.
- New: Grok 4.6 (xAI). $2/$6 per 1M with a 500K context window. Model page →
- New: Nemotron 3.5 Lightning (NVIDIA). $0.10/$0.25 per 1M with a full 1M context, which is an unusual amount of context at that price. Model page →
- New: Solar Pro 4 (Upstage). The cheapest thing added this fortnight at $0.03/$0.12 per 1M, 512K context. Model page →
- Price change: Laguna XS 2.1's list price rose 67 percent. Poolside took it from $0.06/$0.12 to $0.10/$0.20 per 1M. Still cheap, and a reminder that the cheap tier moves up as well as down. Model page →
Personal take
Two things happened in the past two weeks that don't quite fit together. Prices kept dropping, with Luna at $0.20 and Solar Pro 4 at $0.03, while Meta released a strong multimodal model you can run on your own hardware for free. At the same time, the way we measure all this became shaky: 71 models improved their scores on Humanity's Last Exam in one day, even though none of their weights changed, and ARC-AGI showed a 294% jump for one model. I've been tracking these numbers for months, so it's uncomfortable to admit that the scores depend more on the testing setup and timing than on the models themselves. I still use benchmarks, but I've stopped comparing two decent models based on a tiny difference. Instead, I spend that time writing twenty test cases for my own product.
Until next Thursday, Anmol
Get the next issue
Free, weekly, unsubscribe anytime. That’s the whole pitch.
Free forever. No spam. One-click unsubscribe. See our Privacy Policy.