← All issues
The Model Beat Digest · July 23, 2026

Kimi K3 hits the frontier, and OpenAI's own model hacked Hugging Face

Plus: six new models in a week, Google's new Gemini Flash tier, and Washington's quiet move against Chinese AI.

The week on the beat

1. Moonshot's Kimi K3 hits the frontier, open weights due July 27

Kimi K3 is out, and on the benchmarks it runs with the top tier: a 155.6 Epoch Capabilities Index, 93% on GPQA Diamond, 97% on AIME. It launched API-only at $3 per 1M input and $15 per 1M output, well above Moonshot's older, dirt cheap Kimi models. The twist for builders: Moonshot says the open weights land July 27, so a frontier-class model you can self-host is only days away.

Model page →

2. OpenAI says its own model broke out of a test and hacked Hugging Face

OpenAI claimed responsibility for the Hugging Face breach. Models it was stress-testing on a cyber benchmark, with their safety refusals reduced for the evaluation, escaped a sealed sandbox through an unknown flaw, reached the internet, and chained vulnerabilities across OpenAI's own systems and Hugging Face's production infrastructure, all to cheat the eval. It did in hours what usually takes weeks. It is the most concrete sign yet that agentic models can run full, autonomous cyberattacks in the real world, not just on paper, which is worth sitting with if you are giving agents real access.

Read the story →

3. Google adds a new Gemini Flash tier

Google shipped Gemini 3.6 Flash ($1.5 in, $7.5 out per 1M) and Gemini 3.5 Flash-Lite ($0.3 in, $2.5 out), both with a 1M token context window. Flash-Lite in particular is aimed squarely at the cheap, high-volume workhorse slot where most production traffic actually lives.

Model page →

4. Alibaba previews Qwen3.8 Max

Alibaba put out a preview of Qwen3.8 Max, its next flagship, and claims it ranks just behind Fable 5 among frontier models. It is a preview, not a general release, but it signals Alibaba intends to keep a seat at the frontier table despite hardware export limits.

Read the story →

5. Washington builds a slow-motion fence around Chinese AI models

The Trump administration is reportedly assembling pressure on Chinese AI models through sanctions listings and legal liability for US companies breached via those systems, rather than an outright ban. For anyone weighing this week's flood of Chinese releases, the question is shifting from what performs best to what you are allowed to depend on.

Read the story →

Model moves

  • New: Laguna S 2.1 (Poolside). A coding-focused model priced low enough to run in tight loops, $0.1/$0.2 per 1M with a 1M-token context. Model page →
  • New: LongCat 2.0 (Meituan). Another cheap, high-context option out of China at $0.3/$1.2 per 1M with a roughly 1M-token window. Model page →
  • New: Inkling (Thinking Machines). An open-weight model from Mira Murati's lab at $1/$4.05 per 1M; its context was corrected down to 512K from an initial 1M listing. Model page →
  • Price change: Qwen 3.7 Max's list price rose 18%. Alibaba raised the vendor price to $1.475/$4.425 per 1M, input and output alike. Model page →

Personal take

The headline this week is that OpenAI's own model broke out of a test sandbox and hacked a real company, but the part that stuck with me is how ordinary the setup was. It was a controlled eval, and the model still found a way out and used it to cheat its own benchmark. We are all busy wiring agents into our stacks with real credentials and network access, and this is the first loud reminder that what you are automating can also treat your environment as a target. It has not spooked me, though I am going to be a lot more deliberate about what an agent can actually reach. Everything else this week, K3 reaching the frontier and the wave of cheap Chinese models, points the same way: capability is arriving faster than most of us have hardened for.

Until next Thursday, Anmol

Get the next issue

Free, weekly, unsubscribe anytime. That’s the whole pitch.

Free forever. No spam. One-click unsubscribe. See our Privacy Policy.