On a call with a club this week I mentioned, almost in passing, that a new frontier model had come out the day before. Claude Opus 5.5. Nobody reacted. Not one question. People expect new models almost monthly now. What they wanted to talk about was cost, control and who would own whatever got built. They were right to.
Because under every AI decision a club, venue or rights holder makes, there are questions nobody quite knows the answer to yet. Do we commit now, or wait for the next model and save ourselves a rebuild? Is it expensive? Can we swap models around? Is this stuff actually getting smarter? Are things really moving that fast?
I wanted to dig into the data and see how fast things are actually moving. Is AI improvement speeding up? Or have we hit a wall? And, crucially, what does it mean for rights holders and what they are trying to build today?
I took every model OpenAI and Anthropic have put into ChatGPT and Claude since November 2022, kept the strongest one each vendor launched on each release day, and lined them up against two independent measures: Epoch AI’s Capabilities Index and METR’s task-length results.
What the data shows
Capability rises steadily, in small steps
Epoch’s index rolls more than 50 benchmarks into a single score. It takes account of how hard each benchmark is, so it keeps working as the easier tests max out, and it has no ceiling, because it isn’t a percentage. The scale is pinned so that Claude 3.5 Sonnet (June 2024) scores 130 and GPT-5 (August 2025) scores 150. For reference, GPT-4 sits at about 126 and today’s best models at about 166. So 20 points is roughly the distance between Claude 3.5 Sonnet and GPT-5.
The best available model has gained about 14 to 16 points a year on Epoch’s index since the reasoning models of late 2024. Before them it was about 3 points a year. That rate has held for 18 months, and it isn’t bending in either direction.
There have been two step changes. GPT-4 added 16 points in March 2023. Reasoning models (o1-preview, o1 and o3) added 17 between September 2024 and April 2025. Everything else has arrived in steps of 1 to 4 points. The last 12 months delivered as much as the GPT-4 leap did, just cut into smaller pieces.

Releases are now frequent and interchangeable
OpenAI and Anthropic between them launched 15 new top models in the last 12 months, up from 6 two years earlier. Each now ships one roughly every eight weeks, and the lead has swapped between them again and again. The gap between releases has fallen from about 150 days in 2023 to about 40.

The size of task is compounding
This is the measure that matters most for anyone building with AI. METR gives models real software and research tasks, with tools and no human help, and records how long a task, in skilled-person time, the model completes at least half the time.
GPT-4 managed about 5 minutes in March 2023. Claude Opus 4.6 managed about 12 hours in February 2026. That’s 130 times longer in three years, doubling roughly every four months.

I see it in our own work. This month an agent rebuilt three client decks in our house style in one evening, closely enough that a non-technical co-founder could pick them up and edit them. That used to be days of somebody’s time.
But the headline figure needs reading carefully:
- 12 hours is a 50% pass rate. At 80%, Opus 4.6 handles about 70 minutes, from METR’s own data.
- The tasks are clean. METR found about half of AI code changes that pass the tests would not be accepted into a real codebase as they stand.
- Real work lags further. On the Remote Labor Index, which uses paid freelance projects, the best agent completes about 21% to the client’s standard. A year ago it was 2.5%.
Capability on a benchmark and usefulness inside a commercial team are moving at different speeds. Dwarkesh Patel’s line is the most accurate summary I’ve found: models are getting more impressive at the rate the optimists predict, and more useful at the rate the sceptics predict.
What that means for how you plan
Launch dates are the wrong clock
The next release from either vendor will likely add a couple of points (until we hit the next breakthrough, if we do; see below). If a capability doesn’t work for your commercial team today, that is almost always a design problem, not a model problem. Waiting a quarter buys you one small step and costs you a quarter of learning on your own data. Design problems are what Earl thrives on.
One head of technology I spoke to recently described where most organisations are: an AI button in every system, and leadership using chat to polish presentations and search the web. That isn’t a model shortage. The models are fine. It’s that nothing groundbreaking has been built around them yet. The time for that is always right now.
Treat the model as a replaceable part
You should now plan for a change underneath you every eight weeks, and make it boring. We run Earl on this principle. Every call we have is recorded, transcribed, written up and filed against the right organisation automatically. The engine doing the transcription has changed three times since May. The job it does hasn’t changed once. Our own mail client runs Apple’s on-device model on the phone and Claude on the web, which means, as I told that club this week, “the phone can have no connection, no signal, and it still does stuff.” One capability, several models, each chosen for the job and each swappable.
Test on your own work, not on benchmarks
The gap between impressive and useful is where AI projects in sport fail, and the failures are rarely dramatic. They’re confident.
When our meeting transcriber’s live feed dropped out and a backup model listened to an empty room, it reported someone talking in Russian. Another silent call came back as a German anecdote about how tiring three-year-olds are. The AI in our CRM once linked a mention of Google Cloud to Oracle, with a perfectly reasonable-sounding explanation attached. No benchmark would ever catch any of them. Only our own testing, data and compliance rules do. This stuff is as important as the cool AI parts.
AI sounds exactly as confident when it’s wrong as when it’s right. So the acceptance test should be an exam built from your own questions, questions you already know the answers to because they already happened, scored on day one and rerun every time the model underneath changes. That’s also how you find out whether a new release actually helps you, rather than taking the launch post’s word for it. Earl likes exams.
Scope to the reliable horizon, with a person approving
At 8 in 10 reliability, the best models today handle roughly an hour of skilled work in one go. That’s the right unit to design around: preparing a sponsor renewal pack, drafting a partner performance report, pulling matchday data from four sources into one view. Decision preparation and assisted execution, with a person signing off.
Clubs are getting there on their own. A data engineer at one club, building his first internal tool with an AI coding agent, told us unprompted that he’d need an approval step before any ticket-price recommendation reaches the ticketing system. Nobody had asked. He’d just thought about what happens when it’s wrong. He was right.
If the trend holds, that horizon keeps growing. Design the workflow so AI can be handed bigger pieces as it proves itself on your exam, not all at once on the strength of a benchmark.
Don’t confuse cheap screens with a cheap system
More clubs now tell us they could build this themselves, and some genuinely can. The screens are cheap now. We built the first version of our own CRM in a summer. But as I said to a club recently, a first pass is the easy bit. Enterprise grade and reliable is a different thing. Good design, clever UI, systems thinking: none of these things are cheap.
On cost, the model itself is rarely the expensive part, and it keeps getting cheaper. Anthropic says Opus 5.5 costs about 40% less to run than Opus 5, which came out two months earlier. What costs money is the design, the checking, the audit trail, the security review and the ability to switch models without starting again. That’s the part to be sure of, whoever builds it.
Is a breakthrough coming, or a stall?
The people closest to the work don’t agree. Ilya Sutskever, who co-founded OpenAI, argues the age of simply scaling up is over and new ideas are needed, and in July said his company SSI now has “research that is worthy of scaling up”. Demis Hassabis at Google DeepMind thinks we may need “one or two more breakthroughs”, naming continual learning and memory as the gaps. Dario Amodei at Anthropic describes “a smooth, unyielding increase” underneath the noise.
What they share is the diagnosis. The missing pieces are models that learn on the job and stay reliable on messy work. Those are exactly the pieces a club’s commercial operation depends on.
So I wouldn’t bet a roadmap on either outcome. Nothing in the data has slowed yet, and nobody can date the next breakthrough. Anyone asking you to sign a long agreement on the strength of knowing what the future looks like is claiming something I don’t think anyone knows, including the big AI labs. Build so that a better model is an upgrade you test and switch on, not a project you restart.
The model will change every eight weeks. The questions your commercial team needs answered won’t. Build around those.
Method and sources. Every model available in the ChatGPT and Claude web chat since November 2022, dated from the day it reached users and checked against OpenAI’s and Anthropic’s release notes. For each vendor and release day, the highest-scoring model counts. Capability: Epoch AI Capabilities Index (CC BY 4.0), which combines more than 50 benchmarks. Five scores Epoch has not yet published, including Claude Opus 5.5, are estimated from raw benchmark results using Epoch’s published parameters; tested against the 268 models Epoch has scored, the method lands within about a point. Task length: METR time horizons. Also METR on unmergeable code, the Remote Labor Index, Dwarkesh Patel, Ilya Sutskever and SSI’s July announcement, Demis Hassabis via Fortune and Dario Amodei.
