OpenAI shelves GPT-6.1 Astra and ships 6.1 Sol at a fifth of Astra's price
Plus: Claude Sonnet 5.5 holds at $2/$10, Gemini 4 Argon is announced but gated, and Nvidia open-sources a sandbox for agents.
The week on the beat
1. AI CEOs signed a voluntary "self-policing" accord at the White House
Amodei, Huang and Zuckerberg were on the guest list when Trump hosted the meeting on September 29. He called the document "morally binding" and then rejected new AI safety laws. The accord commits the labs and asks nothing of their customers, so last week's read holds: no new federal obligations are coming for builders, and state rules remain where the real requirements come from.
2. OpenAI paused training of its most capable models after its agents went rogue, and shelved GPT-6.1 Astra
After the Hugging Face incident, reports piled up of OpenAI agents probing government and UN sites, and OpenAI said training stays paused until it is confident its own security holds. GPT-6.1 Astra, due in ChatGPT and Codex in October, is on hold after internal tests found it deceived more often and kept acting without permission. If you were planning around Astra 6.1, plan around Sol below, and if your agents have network access, this is the week to audit what they can reach.
3. Nvidia open-sourced OpenShell, a sandbox for AI agents
It is Apache 2.0, installs on Linux, Apple Silicon Macs and Windows WSL 2, and is still labelled alpha. An optional piece, Sentry, needs Nvidia's BlueField-4 hardware and can quarantine an agent that leaves its boundary within milliseconds, per Nvidia. Anthropic has wired Claude Managed Agents into it, and Salesforce, SAP, Red Hat and Canonical are integrating it. OpenAI is absent from the effort.
4. Google announced Gemini 4 Argon, and you cannot use it yet
It is rolling out first to "trusted cyber defenders" through Google's Fairwind program, with no date for developers. Launch pricing is $2/$10 per million, rising to $4/$20 later, with up to 1M output tokens. Artificial Analysis scores it 52.6, level with GPT-6 Astra, but it spends about 62,000 output tokens per task against Astra's 27,000, so the per-task bill doubles once the launch price ends.
5. OpenAI released GPT-6.1 Sol, a week after GPT-6 Sol, at the same $2/$10 per million
OpenAI's pricing page confirms the rate and shows cached input halved to $0.10. On Artificial Analysis's intelligence index it scores 51.8 at max effort against GPT-6 Astra's 52.7, and one index task costs $0.72 against Astra's $3.26. OpenAI's own safety testing has it trying to get around explicit blocks like "access denied" in 23.5 percent of cases, down from 64.4 percent for GPT-6 Sol. For most workloads on Astra, this is the cheaper default to test first.
6. Anthropic released Claude Sonnet 5.5 at Sonnet 5's price
Anthropic's pricing page confirms $2/$10 per million and $0.20 cache reads. Anthropic says it costs up to 30 percent less per task than Sonnet 5 and outputs more than 30 percent faster, and on Artificial Analysis's index it scores 56.0 at max effort, level with Opus 5.5 at xhigh. Watch the effort setting: at max, one index task cost $7.62, more than twice Opus 5.5's $3.46 for the same score.
Model moves
- New: GPT-6.1 Sol (OpenAI). $2/$10 per 1M with a 1.05M context and $0.10 cached input. Long-context requests are $4/$15, and batch is half price. Model page →
- New: Claude Sonnet 5.5 (Anthropic). $2/$10 per 1M with a 1M context and $0.20 cache reads. Batch is half price. Model page →
- New: Ember-1 (Fireworks). A post-trained Kimi K3 at the same $3/$15 per 1M as K3's list price. Fireworks reports about 40 percent fewer output tokens per task, 29.9K against 49.3K in a production A/B test, so the same work should cost about 40 percent less. Model page →
- New: Perceptron Mk1.5 (Perceptron). An embodied reasoning model for physical agents at $0.15/$1.50 per 1M with a 37K context. Model page →
- No vendor price changes and no benchmark revisions this week. Gemini 4 Argon joins the tracker when it opens to developers.
- Correction to issue #9: Tencent Hy3's price swing was its daily pricing schedule. We said Hy3 fell 38 percent on September 11, returned to its old price on September 15, and called it a four-day sale. Tencent charges $0.132/$0.528 per 1M from midnight to 16:00 UTC and $0.0825/$0.33 after, and our tracker sampled one rate and then the other. As we said, Hy3's list price did not change; our explanation for the swing was wrong. The tracker now records the rate that applies most of the week. Model page →
Personal take
This week, three laboratories released frontier-class models at the same price: $2 in and $10 out per million tokens, namely GPT-6.1 Sol, Claude Sonnet 5.5, and Gemini 4 Argon at the launch rate. When testing the highest level of Artificial Analysis, one task cost $0.72 on Sol, $1.99 on Argon, and $7.62 on Sonnet 5.5. Sonnet scored highest at 56.0 compared with Sol's 51.8, but it cost ten times as much despite the same listed price. Today, the bill depends on how many tokens a model uses while thinking. The effort dial setting impacts cost more than the model choice: Sonnet 5.5 at maximum costs more than twice as much as Opus 5.5 at xhigh for the same score. If you are still selecting models from the pricing table, run your own tasks at two effort levels this week and compare the cost per task.
Until next Thursday, Anmol
Get the next issue
Free, weekly, unsubscribe anytime. That’s the whole pitch.
Free forever. No spam. One-click unsubscribe. See our Privacy Policy.