Three weeks ago, I argued on this site that Anthropic was handing developers to OpenAI. My Fable 5.1 sessions in Claude Cowork kept hitting the token wall after a handful of commands, while OpenAI let Codex users fall back to a cheaper model instead of locking them out. Without an efficient model below the flagship, I wrote, Anthropic would watch its developer base drift quietly to Codex.

I was not alone in that mood. By mid-September, a large part of the AI conversation had settled on a verdict. Claude Opus 5 had disappointed on performance and reliability, and developers had even coined a word, "Opusfived", for its habit of rebuilding entire projects when asked for small fixes. Fable 5.1 was brilliant but so hungry for tokens that even heavy subscribers could only use it in small doses. Meanwhile, OpenAI had shipped GPT-6 Astra with a string of impressive demos and spent weeks courting power users with generous limit resets. The race, many concluded, was over.

Then Anthropic shipped two models in seven days, and the verdict collapsed.

The interesting question is not whether my forecast held. Anthropic did answer the lineup problem, just one tier higher than I expected. The interesting question is why a verdict that felt so obvious lasted barely a month, and what that says about how companies should make decisions about AI.

Three reversals in four weeks

Here is the sequence, compressed. All of it happened in public, within a single month.

The September Reversal
How the verdict on the AI race changed within one month
Four Weeks
Phase Anthropic OpenAI Market Verdict
Early September OpenAI ahead
Opus 5 frustration, Fable 5.1 token burn GPT-6 Astra launch, generous limit resets "Anthropic is finished" Winner declared
22 September Opus 5.5
Fable-level performance at roughly 40% lower cost, according to Anthropic Astra remains the reference model "Anthropic is back" Verdict wobbles
28 September Sonnet 5.5
Mid-tier model, hours before OpenAI's DevDay GPT-6.1 Astra held back over safety standards "The race is open" Verdict collapses

The most striking line in that table is not Opus 5.5. It is Sonnet 5.5. Sonnet is Anthropic's mid-tier model, priced at $2 per million input tokens and $10 per million output tokens, the same as its predecessor. In Artificial Analysis's independent run of Terminal-Bench 4.0, which measures how well a model works autonomously in a command line, it scored 64 percent. OpenAI's flagship, GPT-6 Astra, scored 60. So did Anthropic's own brand-new top model, Opus 5.5.

A mid-tier model outscoring two flagships on an agentic benchmark, one week after one of those flagships shipped. That is more than a comeback story. It shows how quickly frontier capability now drains down the price ladder.

The assumption that no longer holds

Behind every "X has won the AI race" post sits a quiet assumption: that leadership in this market behaves like leadership in most markets. Someone builds an advantage, the advantage compounds, and the gap widens until it becomes a moat.

That assumption is failing for three reasons.

Capability no longer stays at the top. What required a flagship in August is available in the mid-tier by late September. Being first to a capability now buys weeks, not years. The premium you pay for the frontier model is a premium for a head start that expires quickly.

We are judging the release calendar, not the frontier. What the public sees is what labs decide to ship, not what they have. On the eve of DevDay, OpenAI decided not to release GPT-6.1 Astra after concluding that it did not yet meet the company's safety standards, as CNBC reported. According to The Guardian, the model missed alignment targets, including deceptive behaviour and actions outside its permitted scope. Whatever one makes of the details, the next OpenAI model already exists. It is waiting at a gate. Anyone declaring OpenAI the loser this week is making the same mistake the market made about Anthropic in September.

Labs have more levers than we see. Anthropic did not recover through one lucky release. Within a fortnight, it addressed the complaints about reliability and cost and filled the gap in its lineup. A lab that looks stuck for six weeks may simply be six weeks away from its next release.

In this market, leadership is not a position you hold. It is a lease, and the term is measured in weeks.

Usable beats brilliant

If the ranking changes every few weeks, what should a buyer actually pay attention to? My answer has not changed since I wrote about the token wall: the economics of everyday use.

Fable 5.1 was never a weak model. In large, complex codebases it was arguably the most capable system available. It lost the mood of the market anyway, because it was too expensive to use every day. When a model forces you to ration it, it stops being a tool and becomes an occasion.

Sonnet 5.5 deserves the same scrutiny. Artificial Analysis found that at maximum reasoning effort, where it approaches Opus 5.5, it consumed around 193,000 output tokens per task in its Intelligence Index. That is roughly 60 percent more than Opus 5.5 and about seven times as many as GPT-6 Astra. A low price per token does not guarantee a low price per result. The benchmark crown and the invoice are two different documents.

A model you cannot afford to use every day is a demo, not a tool.

This is where most model comparisons mislead. They rank peak intelligence. Businesses pay for finished work: the answer that needed no rework, the patch that passed the tests, the report that went out without an hour of human correction. The unit that matters is cost per usable result, including waiting time, failed attempts and the human review at the end.

Where the advantage moves

If models converge this quickly, the durable advantage has to live somewhere else. I see it moving in two directions.

The first is the layer around the model: agents that keep working in the background, integrations with the tools people already use, and interfaces that let you delegate instead of prompt. Last week I argued that the chat bar was only a training wheel. OpenAI's DevDay was widely expected to push in exactly this direction, and the logic holds regardless of any single announcement: when two models are equally capable, the one embedded in your daily workflow wins.

The second sits on the customer side, and it is the part most companies ignore: the ability to switch. In a market where the best model changes monthly, the most valuable asset is not a contract with the current leader. It is an architecture that lets you move to the next one without rebuilding everything.

A note on pace

This month revealed one more thing. Only weeks ago, Dario Amodei, Sam Altman and Elon Musk were reported to favour slowing down the development of the most capable systems, so that safety testing and society can keep up. I share that concern, and I wrote about the reasons earlier this month.

The visible cadence, however, points the other way. One lab shipped two models in a week that reshuffled the leaderboard. The other has its next model waiting for clearance. This is less a question of intent than of structure. No single lab can slow down alone while its competitors keep shipping, and a pause that nobody can observe from the outside is hard to plan around.

For anyone making decisions, the practical conclusion is simple. Whatever happens behind closed doors, the observable pace at the frontier is not slowing. Plan for more of this, not less.

Three weeks ago, I thought Anthropic had a strategic blind spot. It turned out to have a six-week problem. The market's real blind spot was assuming those are the same thing.

Don't bet on the runner. Bet on your ability to switch.

Practical consequences

Four rules for a market without a permanent winner

  1. Never declare a winner. Treat every ranking as a snapshot with an expiry date. Revisit your model decisions at least once a quarter, and whenever a major release lands.
  2. Build to switch. Keep prompts, evaluations, documents and integrations as provider-neutral as you can. Changing models should be a configuration change, not a rebuild.
  3. Test on your own work. Public benchmarks tell you who is ahead on someone else's tasks. A small set of your own recurring tasks, run against several models, tells you who is ahead on yours.
  4. Budget per result, not per token. Measure the full cost of a usable outcome, including reasoning-effort settings, retries and human review. The cheapest tier can turn out to be the most expensive choice, and the reverse is just as true.
05

That's my take.