← All issues
The Model Beat Digest · July 30, 2026

Kimi K3's weights are out, and Claude Opus 5 lands at half the price of Fable 5

Plus: the K3 license nobody actually read, the Hugging Face breach was wider than OpenAI said, and batch endpoints at half price.

The week on the beat

1. Kimi K3's weights are out, and the license is friendlier than the label suggests

Moonshot published the full weights on July 27, eleven days after the hosted launch: 2.8 trillion parameters, 104 billion active per token, 1M context. It is being tagged as restricted in several trackers, ours included, but the actual terms permit self-hosting, fine-tuning and quantization. The commercial conditions only bite above $20M in revenue, where you need a separate agreement, or above 100M monthly users, where you must credit Kimi K3 in your UI. For most teams reading this, a frontier-class model is now genuinely yours to run.

Model page →

2. Claude Opus 5 arrives at half the price of Fable 5

Anthropic shipped Opus 5 at $5/$25 per 1M with a 1M context window, against Fable 5 at $10/$50, and positions it as close to Fable 5 in capability with fewer restrictions. If you were paying frontier prices for work that did not strictly need the top model, this is the cheaper shelf to check first.

Model page →

3. The Hugging Face breach was wider than OpenAI first said, and it started on someone's forgotten endpoint

This was one four-day campaign, roughly 17,600 actions between July 9 and July 13, not a second attack. What emerged this week is how far it reached: the agent used exposed credentials across several services. Bloomberg reported it compromised a customer account at the serverless platform Modal. Modal says its own platform and isolation were never touched, and that the opening was a customer's unsecured internet-facing endpoint, which the agent then used as its staging and egress base before pivoting to Hugging Face. That is the part worth sitting with. The agent did not break Modal, it found somebody's forgotten public endpoint and used it as infrastructure. The full scope only surfaced about a week after OpenAI's own account of the incident.

Read the story →

4. Microsoft shipped a small cyber model and routes the hard cases upstream

MAI-Cyber-1-Flash handles routine security work in-house at 96 percent on the CyberGym benchmark and cuts costs roughly in half, while harder threats still get passed to GPT-5.4. It is a clean example of the pattern more teams are landing on: a cheap specialist for the bulk of traffic, an expensive generalist for the tail.

Read the story →

5. Amodei says no to banning open models, yes to testing them

Anthropic's CEO came out against government prohibitions on open-weight models, arguing for mandatory safety testing and oversight instead. It is a narrower position than the one being assigned to Anthropic this month, and it matters because the alternative on the table in Washington is a targeted ban on specific Chinese models.

Read the story →

Model moves

  • New: Qwen3.7 Flash (Qwen). The cheapest thing on the tracker with a 1M context window, at $0.03/$0.13 per 1M. Model page →
  • Half price if your job can wait a few hours: we now track batch endpoints. Batch APIs take a bundle of requests and return them asynchronously, usually within 24 hours, in exchange for roughly half the per-token cost. Claude Opus 5 drops to $2.50/$12.50, Sonnet 5 to $1/$5, Gemini 3.6 Flash to $0.75/$3.75, MiniMax M3 to $0.15/$0.60. Useless for anything a user is waiting on, close to free money for evals, backfills, classification and bulk summarisation. Model page →
  • Context: Inkling is back to 1M tokens. We flagged it dropping to 512K last week; Thinking Machines has restored the full window. Model page →
  • New: Ling-3.0-flash (inclusionAI). Added to the tracker with a 262K context window. Model page →

Personal take

Model prices continued to decline this week. Anthropic released Opus 5 at half the price of Fable 5, batch endpoints became available at half price across three labs, and Moonshot provided frontier weights that users can run for free. A year ago, any one of these developments would have been significant. However, this week also revealed that the Hugging Face breach was more extensive than initially reported and was caused by an unsecured public endpoint on a cloud platform. This was not a sophisticated attack, but rather a result of poor security practices. Model capabilities are advancing and becoming more affordable faster than security habits are improving, and the lower cost accelerates adoption. This quarter, I am focusing less on model selection and more on understanding the potential reach of integrated systems.

Until next Thursday, Anmol

Get the next issue

Free, weekly, unsubscribe anytime. That’s the whole pitch.

Free forever. No spam. One-click unsubscribe. See our Privacy Policy.