OpenAI cut Sol's price for three months, and a free stealth model is outscoring it on coding
Plus: IBM's Granite 4.2 lands under Apache 2.0, GLM 5.3 Flash arrives with a 1.3M context, and OpenAI's own chip claims to beat Nvidia's.
The week on the beat
1. OpenAI cut GPT-5.6 Sol by 20 percent, and the clock is running
Sol went from $5/$30 to $4/$20 per 1M on August 22. Read the second half of that sentence carefully: this is a promotional rate that runs about three months, at least through November 21, and it covers the API and Codex credits, not the subscription tiers. Output is where the real move is, down 33 percent. If you shifted workloads onto Sol because the arithmetic finally worked, put a reminder in your calendar for mid-November, because the arithmetic has an expiry date on it.
2. A free, anonymous model is beating the paid ones at coding
Ox Alpha appeared on OpenRouter on August 20 with no announced owner, a 1M context window, and a price of zero. Early independent tests put it at 80 percent on the DeepSWE coding benchmark against 65 for Claude Fable 5 and 52 for GPT-5.6 Sol, which are claims worth testing yourself rather than taking on trust. Community fingerprinting points at Z.ai's GLM family, though nobody has confirmed it. The free window was described as roughly a week, so it is closing around now: if you want your own read on it, today is the day.
3. IBM put Granite 4.2 out under Apache 2.0
Three sizes, 3B, 8B and 30B, with a 512K context and a thinking mode you switch on per request rather than picking a separate model. The 8B and 30B were trained with reinforcement learning inside real software engineering, terminal and web search environments. Apache 2.0 means unrestricted commercial use, which is the part that matters if your blocker on open weights has been the licence rather than the quality.
4. The cheap tier got cheaper and much longer
Two models landed on August 26 that reset what the bottom of the market looks like. GLM 5.3 Flash from Z.ai lists at $0.15/$0.50 per 1M with a 1.31M context window, the longest context we track at anywhere near that price. Qwen3.8 Flash arrives at $0.16/$0.47 with 1M. For anything that is mostly reading, long documents, transcripts, big codebases, the cost of a large context window has quietly stopped being the constraint.
5. OpenAI published its Hugging Face report, and Alabama subpoenaed it
The official post-mortem on the agent campaign we covered a fortnight ago says more than 1,000 agents coordinated during the incident and that OpenAI saw warning signs weeks before it and could have reacted sooner. Alabama's attorney general has opened an investigation. If you run agents with any network reach, the useful part is not the enforcement angle, it is that the report describes emergent coordination between instances that nobody designed and nobody caught in testing.
6. OpenAI says its own chip beats Nvidia's
Jalapeño is OpenAI's first custom inference chip, and the company's published results claim it outperforms Nvidia's Blackwell on inference speed and power efficiency. These are the vendor's own benchmarks on the vendor's own silicon, so hold them loosely. It matters anyway, because the cost of serving a token is the thing that eventually decides whether promotional pricing like Sol's becomes permanent or expires.
Model moves
- Correction to issue #5: Nemotron 3.5 Lightning does not have a 1M context. We reported it at $0.10/$0.25 per 1M with a full 1M window. Our tracker now shows 262K, a quarter of that, alongside a price drop to $0.08/$0.20. The spec changed under us on August 20, but we published the earlier figure, so treat the 1M line from that issue as withdrawn. Model page →
- Epoch refreshed its benchmark suite on August 26 and four tracked models moved, in both directions, across four different benchmarks. The largest was GPT-5.1-Codex-Max on METR task horizon, from 161.8 to 223.7. No weights changed; this is the source re-running its evaluations.
- New: Muse Spark 1.2 Contributor (Meta). $0.10/$0.20 per 1M with a 1M context. Model page →
- New: DeepSeek V4 Flash Vision Exp (DeepSeek). $0.22/$0.66 per 1M with a 1M context, and vision on the experimental line. Model page →
Personal take
In the past three weeks, there have been three price announcements, but none of them constitutes a price reduction in the sense that I would have understood the phrase a year ago. Gemini 3.7 Flash is being offered at half price until the end of December, after which the price will double. DeepSeek has replaced its flat rate with peak and off-peak pricing. Sol is now 20% cheaper until about November 21. All of these options are cheaper now, and each one comes with a specific date, which is different to plan for than a gradually falling price. What I keep coming back to is that the figure on the pricing page is beginning to act less like a cost and more like an offer, and offers are designed for people who haven't yet made a commitment. So the question about my own setup isn't what I am currently paying, but what I'd pay if all the promotional rates I am currently benefiting from expired in the same quarter.
Until next Thursday, Anmol
Get the next issue
Free, weekly, unsubscribe anytime. That’s the whole pitch.
Free forever. No spam. One-click unsubscribe. See our Privacy Policy.