Kimi K3's weights are out, and Claude Opus 5 lands at half the price of Fable 5
Plus: the K3 license nobody actually read, the Hugging Face breach was wider than OpenAI said, and batch endpoints at half price.
The week on the beat
1. Kimi K3's weights are out, and the license is friendlier than the label suggests
Moonshot published the full weights on July 27, eleven days after the hosted launch: 2.8 trillion parameters, 104 billion active per token, 1M context. It is being tagged as restricted in several trackers, ours included, but the actual terms permit self-hosting, fine-tuning and quantization. The commercial conditions only bite above $20M in revenue, where you need a separate agreement, or above 100M monthly users, where you must credit Kimi K3 in your UI. For most teams reading this, a frontier-class model is now genuinely yours to run.
2. Claude Opus 5 arrives at half the price of Fable 5
Anthropic shipped Opus 5 at $5/$25 per 1M with a 1M context window, against Fable 5 at $10/$50, and positions it as close to Fable 5 in capability with fewer restrictions. If you were paying frontier prices for work that did not strictly need the top model, this is the cheaper shelf to check first.
3. The Hugging Face breach was wider than OpenAI first said, and it started on someone's forgotten endpoint
This was one four-day campaign, roughly 17,600 actions between July 9 and July 13, not a second attack. What emerged this week is how far it reached: the agent used exposed credentials across several services. Bloomberg reported it compromised a customer account at the serverless platform Modal. Modal says its own platform and isolation were never touched, and that the opening was a customer's unsecured internet-facing endpoint, which the agent then used as its staging and egress base before pivoting to Hugging Face. That is the part worth sitting with. The agent did not break Modal, it found somebody's forgotten public endpoint and used it as infrastructure. The full scope only surfaced about a week after OpenAI's own account of the incident.
4. Microsoft shipped a small cyber model and routes the hard cases upstream
MAI-Cyber-1-Flash handles routine security work in-house at 96 percent on the CyberGym benchmark and cuts costs roughly in half, while harder threats still get passed to GPT-5.4. It is a clean example of the pattern more teams are landing on: a cheap specialist for the bulk of traffic, an expensive generalist for the tail.
5. Amodei says no to banning open models, yes to testing them
Anthropic's CEO came out against government prohibitions on open-weight models, arguing for mandatory safety testing and oversight instead. It is a narrower position than the one being assigned to Anthropic this month, and it matters because the alternative on the table in Washington is a targeted ban on specific Chinese models.
Model moves
- New: Qwen3.7 Flash (Qwen). The cheapest thing on the tracker with a 1M context window, at $0.03/$0.13 per 1M. Model page →
- Half price if your job can wait a few hours: we now track batch endpoints. Batch APIs take a bundle of requests and return them asynchronously, usually within 24 hours, in exchange for roughly half the per-token cost. Claude Opus 5 drops to $2.50/$12.50, Sonnet 5 to $1/$5, Gemini 3.6 Flash to $0.75/$3.75, MiniMax M3 to $0.15/$0.60. Useless for anything a user is waiting on, close to free money for evals, backfills, classification and bulk summarisation. Model page →
- Context: Inkling is back to 1M tokens. We flagged it dropping to 512K last week; Thinking Machines has restored the full window. Model page →
- New: Ling-3.0-flash (inclusionAI). Added to the tracker with a 262K context window. Model page →
Personal take
Model prices continued to decline this week. Anthropic released Opus 5 at half the price of Fable 5, batch endpoints became available at half price across three labs, and Moonshot provided frontier weights that users can run for free. A year ago, any one of these developments would have been significant. However, this week also revealed that the Hugging Face breach was more extensive than initially reported and was caused by an unsecured public endpoint on a cloud platform. This was not a sophisticated attack, but rather a result of poor security practices. Model capabilities are advancing and becoming more affordable faster than security habits are improving, and the lower cost accelerates adoption. This quarter, I am focusing less on model selection and more on understanding the potential reach of integrated systems.
Until next Thursday, Anmol
Get the next issue
Free, weekly, unsubscribe anytime. That’s the whole pitch.
Free forever. No spam. One-click unsubscribe. See our Privacy Policy.