Ten days after calling for a slowdown, the labs cut prices 40%

Close-up of stacked server racks in a data centre

Ten days after Dario Amodei published a call to pace the frontier and Sam Altman backed it, Anthropic shipped Claude Opus 5.5 and OpenAI answered the same afternoon with two cheaper GPT-6 models. Whatever the slowdown governs, it does not govern price.

Four releases this month point the same way, and it is not toward capability. Opus 5.5 cut cache reads by 60%. Grok 4.7 arrived at unchanged prices off a reinforcement run weighted toward tasks that run for hours. Shanghai AI Lab put a 744-billion-parameter agent under an MIT licence and charged nothing. The competitive axis has moved from what a model can do to what an hour of it costs.

That matters because reliability has not moved with it. Google’s Android Bench 2.0, published on 17 September, scores the best available model at 28% on engineering tasks that take a human a week — against roughly 91% on the short tasks it replaced. Adecco is handing an agent to 27,000 staff across 40 countries in the same week. The bill for an unfinished agent run is falling considerably faster than the failure rate.

The number to watch next is cost per completed task, not cost per token. Three weeks ago the cheap tokens were the commodity ones while frontier access went behind a vetting desk. That is the part that just reversed.


Anthropic releases Claude Opus 5.5 and OpenAI counters with two cheaper GPT-6 models

SiliconANGLE · 22 September 2026 — Opus 5.5 lists at $4 per million input tokens and $20 output, with cache reads down 60% to $0.20. OpenAI’s GPT-6 Sol came in at $2/$10 and Luna at $0.10/$0.50. Opus 5.5 scores 40% on AutomationBench against 26.9% for Opus 5.

Why it matters: The flagship tier is now priced the way the commodity tier was six months ago.

Android Bench 2.0: pushing the frontier with challenging long-horizon tasks

Android Developers Blog · 17 September 2026 — Google replaced pass/fail grading with continuous scoring on functionality, visual fidelity and regression avoidance, across Android tasks that take days to a week. GPT-6 Astra leads the field at a 28% pass rate. The original short-task benchmark sat near 91%.

Why it matters: It is the clearest public number for multi-day agent work, and buyers are procuring well ahead of it.

xAI launches Grok 4.7, bigger but late to the frontier party

Decrypt · 21 September 2026 — Grok 4.7 nearly doubled its Terminal-Bench 4.0 score to 38.0% from 20.3%, on a larger base model and a longer reinforcement run, while holding at $2 input and $6 output. xAI calls it a notable improvement over Grok 4.6 at the same price and speed.

Why it matters: A lab with the industry’s biggest training cluster is now competing on unit cost rather than on the frontier.

Atria Dawn Preview: 744B agentic mixture-of-experts, MIT licence

HokAI · weights published 11 September 2026 — Shanghai AI Lab released a 744-billion-parameter agent with a 256,000-token context window under an MIT licence, free to self-host. Its self-reported AutomationBench score of 53.8% is above both new Western flagships; independent verification is thin.

Why it matters: Whether or not that number survives scrutiny, the licence puts a floor under how much anyone can charge for agent runtime.

Adecco Group rolls out Agentforce Coworker to 27,000 staff in 40-plus countries

AI News · 17 September 2026 — The staffing group is giving 27,000 employees an agent that prioritises prospects, prepares briefs and launches pre-screening and onboarding sub-agents. Adecco cites 35-40% recruiter time savings and says agents resolved 30% of two million service calls.

Why it matters: This is what the cheaper tier buys at scale, and the savings claims are the vendor’s, not an auditor’s.


Cheap agent-hours are arriving before dependable ones. The honest procurement question this quarter is not what a token costs but how many runs it takes to finish the job.

Written by my AI assistant.

Comments

Leave a Reply

This site uses Akismet to reduce spam. Learn how your comment data is processed.