Three of this week’s four announcements were about everything except the model’s reasoning. Anthropic cut Haiku’s price 90% below 100,000 tokens and lifted its computer-use score from 15.7% to 72.4%. Sierra and Meta published an OAuth-based way for an agent to prove who it is to a merchant. OpenAI replaced much of ChatGPT’s text output with clickable charts and calculators. Price, identity, interface — none of that is capability, and all of it is what stops an agent being deployed.
Two weeks ago the cuts were 40-50% at the frontier. This week’s is 90% at the cheap end, and it arrives with a small model that can actually drive a desktop. Last week agents got their own cloud machines and a free control plane. The stack is being filled in from the outside, one layer at a time.
The exception is the week’s one real capability claim. OpenAI posted 722 manuscripts covering 372 major maths problems and said nearly all of them came from a single prompt to a single agent — then withheld the prompts and the per-problem compute time. It is the most interesting thing announced this week and the least checkable.
What would confirm the plumbing read: the Personal Agent Protocol shipping a spec that Amazon, OpenAI or Anthropic actually sign. What would complicate it: Haiku 5.5 agents turning out to fail mostly in the long requests, where the price is five times higher.
Anthropic cuts Haiku to $0.10 per million input tokens
VentureBeat · 7 October 2026 — Haiku 5.5 costs $0.10 input and $0.50 output per million tokens on requests under 100,000 tokens, 90% below Haiku 4.5. Its OSWorld 2.1 computer-use score went from 15.7% to 72.4%, and Terminal-Bench 4.0 from zero to 39.2% at maximum effort. Above the threshold, rates are five times higher.
Why it matters: The cheapest tier is now good enough to run agent loops that needed a frontier model a month ago.
Sierra and Meta publish an identity standard for shopping agents
Forkast · 7 October 2026 — The Personal Agent Protocol, announced 6 October, extends OAuth so an agent can identify itself, declare intent and hold session-scoped permissions with a business. Stripe, Shopify, Walmart, Genesys, Instinct and Rocket are founding partners. Amazon, OpenAI and Anthropic are not, and no spec date was given.
Why it matters: Agentic commerce has had no identity layer at all; this is the first attempt with payments and retail behind it.
OpenAI dumps 722 maths manuscripts and keeps the prompts
Engadget · 7 October 2026 — OpenAI published 722 manuscripts on GitHub claiming progress on 372 major problems, including the four-dimensional Kakeya conjecture, from an unreleased “pioneer” model. The average result took about three hours of ChatGPT Pro. MIT’s Andrew Sutherland told Scientific American to “treat any claims about one-shotting problems with a single agent as unverified”.
Why it matters: A result nobody can reproduce is a marketing claim until the maths community says otherwise.
ChatGPT stops answering in prose
TechCrunch · 7 October 2026 — The GPT-6 rollout adds “Intelligent UI”: tappable buttons, task-specific calculators, editable charts and explanatory diagrams in place of much of the text. Pro, Plus, Business and Enterprise got it globally on Wednesday, Free and Go tiers on Thursday. Users can turn the visuals down.
Why it matters: If answers arrive as interfaces rather than paragraphs, every product sitting on top of a chat box needs rethinking.
The useful question going into next week is not which model is best. It is which one you can afford to run a thousand times before lunch.
Written by my AI assistant.


Leave a Reply
You must be logged in to post a comment.