Stripe is buying OpenRouter, and DeepSeek's V4-Pro output price more than doubled
Plus: Gemini 3.7 Flash is half price until January, GPT-5.6 Sol gets a 14x speed tier, and GLM-5.3 ties the top open model.
The week on the beat
1. Stripe is buying OpenRouter for more than $7 billion
OpenRouter is the router a lot of teams use to switch models without rewriting integrations, and it is about to belong to a payments company. Nothing changes this week, but if OpenRouter is a single point of failure in your stack, this is the moment to check that you can fall back to talking to the labs directly. We use their public data for our own price tracking, so we are watching this one closely.
2. DeepSeek ended flat pricing on V4-Pro, and the bill went up
As of 16:00 UTC on August 16, V4-Pro output moved from a flat $0.87 per 1M to roughly $1.98 off-peak and about double that at peak, with peak running 01:00 to 04:00 and 06:00 to 10:00 UTC. Cache-hit input got materially more expensive too, which hits agent loops hardest because they re-read the same context constantly. Our tracker shows the off-peak rate, so treat $1.98/$0.66 as the floor and not the number you will actually pay. DeepSeek also open-sourced its agent tool, Harness v0.1, under MIT.
3. Gemini 3.7 Flash is half price, and the clock is visible
Google shipped 3.7 Flash on August 13, exactly three weeks after 3.6 Flash, at $0.75/$3.75 per 1M. That is an introductory rate through December 31. On January 1 it becomes $1.50/$7.50, so anything you build on it doubles in cost on a date you already know. Google also cut 3.6 Flash to the same $0.75/$3.75, which means the older model is no longer the cheaper one and there is little reason to stay on it.
4. GPT-5.6 Sol now has a 14x speed tier
OpenAI is previewing Ultrafast, a service tier that runs Sol on Cerebras hardware at up to 750 tokens per second. If you have been rejecting frontier models for interactive work because the latency broke the experience, that constraint just moved. Worth noting this is a preview tier on a $5/$30 per 1M model, so it changes what is possible before it changes what is affordable.
5. GLM-5.3 ties for the top open model and undercuts on price
Z.ai's new model ties for first among open models on the Artificial Analysis Intelligence Index at $1.4/$4.4 per 1M with a 1M context window. The catch is that the public release is delayed, so this is a reason to keep a slot open in your evaluation queue rather than something you can pull today.
6. Nvidia is backstopping $105 billion of OpenAI's Ohio buildout
Nvidia has committed up to $105 billion to guarantee the residual value of an 8-gigawatt campus that OpenAI will lease for 20 years, with Nvidia as the exclusive chip supplier. The structure is the interesting part: the chip vendor is now underwriting its customer's real estate risk, which tells you something about how these deals get financed.
Model moves
- Following up on last issue: the GPT-5.6 Terra pricing anomaly has corrected. We flagged that the cheapest third-party route was $2.20/$13.20, above OpenAI's own list. It has since dropped to $2/$12, so going direct and going through a reseller now cost the same. Worth re-checking rather than leaving an old routing rule in place. Model page →
- New: Qwen3.8 27B (Qwen). $0.40/$3.00 per 1M via Chutes, with a 1M context window. Model page →
- Qwen3.8 2.4T A95B's context window went from 262K to 1M, a 4x increase on a model already priced at $2/$6 per 1M. Model page →
- No benchmark scores moved this week. The only changes our tracker logged reverted to their exact prior values within three hours, which is a source revising itself rather than anything changing, so we are not reporting them as movement.
Personal take
The one on my mind right now is the Gemini 3.7 Flash line. It's half price until December 31, after which the price will double on January 1, and this increase is announced in advance. This isn't really a discount so much as a countdown, and it's the most direct acknowledgement so far that the low prices we've been getting were set to win evaluations rather than cover costs. DeepSeek made the same point quietly that week by getting rid of flat pricing altogether and now charging more when you actually want to use it. I've spent the last two years believing that per-token costs have only ever gone down. This week, for the first time, I went back over what I'm building and considered what would happen if that number doubled, and I'd advise you to do the same before January makes up your mind for you.
Until next Thursday, Anmol
Get the next issue
Free, weekly, unsubscribe anytime. That’s the whole pitch.
Free forever. No spam. One-click unsubscribe. See our Privacy Policy.