Claude's strongest model took only 6% of tokens: how much will businesses pay for the final performance edge?
Ramp's enterprise token-management sample shows Fable 5 at only 6% of Anthropic tokens and 11.4% of model spending in its first relatively complete month, while GPT-5.6 Sol represented 25% of OpenAI tokens. The sample is not the full market, but Anthropic then launched half-price Opus 5, described it as close to Fable, and made it the Max default. Enterprise buying is shifting from the strongest model by default toward total cost per successful task.
Conclusion
Ramp’s August enterprise AI-spending data reveals an unusual split. Anthropic’s overall business adoption has passed OpenAI’s, yet Fable 5 — then its strongest and most expensive public model — accounted for only 6% of Anthropic tokens and 11.4% of model spending in its first relatively complete sales month.
This does not prove that Fable 5 “failed to sell,” or that businesses do not want Claude. Ramp’s model-level data comes from Token Spend Management customers that skew more technical, and July was only the first full month after Fable 5 returned to global availability.
It does raise a practical question: when a model costs $10/M input and $50/M output, while a half-price model is already close, how much will businesses pay for the last increment of performance?
My read is that AI procurement is moving from “buy the strongest by default” to “buy enough capability for each job.” Flagship models will remain important, but increasingly as specialist escalation paths for tasks where failure is expensive.
Businesses are buying Claude, but not chasing its most expensive Claude
Start with the context. Ramp reports that 43.5% of eligible US businesses paid Anthropic for subscriptions or tokens in July, ahead of OpenAI at 39.7%.
The story is not that companies abandoned Claude. Anthropic’s business penetration was still rising.
The unusual part appears inside Anthropic’s model mix. Ramp says Fable 5 represented 6% of purchased Anthropic tokens over the prior month, but 11.4% of model spending because of its high price.
GPT-5.6 Sol, by comparison, represented 25% of OpenAI tokens and 23% of spending. Ramp estimates that Fable 5 generated about 75% as much model-attributed spending as Sol in July.
On August 23, the Financial Times pushed the contrast into mainstream technology coverage, with a headline saying Anthropic’s best model was struggling to attract users as cheaper tools thrived.
Fable 5 was not short on capability
When Anthropic launched Fable 5 on June 9, it said the model exceeded the capabilities of anything it had previously made generally available, particularly in software engineering, knowledge work, science and long-running autonomous tasks.
It was also priced outside the everyday tier:
- Input: $10 per million tokens
- Output: $50 per million tokens
Agent work repeatedly reads context, calls tools, writes plans, executes, fails and retries. Doubling the model rate can decide whether a workflow is economically sustainable, rather than adding a few cents to a chat.
The 6% share therefore looks less like “the model is useless” and more like businesses deciding which tasks deserve it.
Anthropic introduced a half-price answer one month later
On July 24, Anthropic launched Opus 5. Its opening description was unusually direct: the model comes close to Fable 5’s frontier intelligence at half the price.
Anthropic made Opus 5 the default model on Claude Max. It also said Opus reached new state-of-the-art results on coding and knowledge-work evaluations including Frontier-Bench and GDPval-AA. At maximum effort on CursorBench, it performed within 0.5% of Fable 5’s peak score at roughly half the cost per task.
That is not an admission that Fable 5 was mispriced. Fable may retain distinct value on certain extreme, scientific or safety-sensitive tasks.
The product decision is still clear: high-volume everyday work needs a cheaper model that is close enough. The strongest model and the default model are becoming different roles.
The performance premium has a visible boundary
Earlier model gaps were large. Weaker models might fail to read a long document, use tools reliably or complete a complex coding task. Businesses had little choice but to buy frontier capability.
The comparison is increasingly closer to 95 versus 92. If the first model costs twice as much, procurement teams ask whether the extra success rate pays for the extra cost.
Customer-message classification, document triage, product-data cleanup and ordinary code changes rarely need the most expensive model at every step. The useful metric is no longer only dollars per million tokens:
How much did it cost to finish one verifiable piece of work?
A cheap model that fails three times may not be cheap. A premium model that moves success from 96% to 97% may not deserve full-volume deployment.
Better agents will tier their models
The likely outcome is not that every business switches to cheap models. It is finer routing.
- Classification, extraction and formatting go to low-cost models;
- Ordinary writing, code changes and research go to the daily workhorse;
- Difficult debugging, critical contracts and high-value research escalate to a Fable-class model;
- Rules, humans or another model review the final output.
Cursor, model gateways and internal enterprise routers are all competing for this layer. A useful router does not always pick the smartest model. It knows when spending more materially reduces rework.
What the data cannot prove
Ramp’s model-share statistics come from its Token Spend Management product, not every Ramp customer or the full AI market. Ramp explicitly says this sample skews more technical than its usual AI Index sample.
Fable 5 also had an abnormal launch. It arrived on June 9, access was suspended on June 12 after US export controls, and global availability returned on July 1. July was its first relatively complete month, while cloud-platform restoration and subscription-credit policies could also affect adoption.
This article therefore does not treat 6% as a market verdict. The narrower conclusion is:
The first enterprise spending data does not show the strongest model automatically absorbing budgets, and Anthropic’s next product move placed cost-effectiveness much closer to the center.
What developers and businesses should do now
Do not begin with “which model is number one?” Run the same 10–20 real tasks through the candidates.
Record four things:
- Whether the job was completed;
- How long a human spent fixing it;
- How many model and tool attempts it took;
- Total cost per accepted result.
If Fable turns a high-value failure into a completed task, its premium may be cheap. If an ordinary model already completes the work reliably, using Fable everywhere turns benchmark leadership into a bill.
For companies, the flagship should usually be an escalation path rather than the default path. Individual users can apply the same rule: switch to the strongest model for hard work, not every message.
Our judgment
Fable 5 does not prove that “the strongest AI cannot sell.” It offers one of the clearest warnings yet that the final points of frontier performance are expensive, and customers will not pay an unlimited premium.
The most important model may not be the top aggregate scorer. It may be the one that minimizes cost per successful result on your work.
That turns competition among Opus, Sol, Kimi, Qwen, DeepSeek and model routers into a procurement problem, not just a benchmark race.
What to watch next
- Whether Fable 5’s token and spending share rises in its second and third complete months;
- Whether adoption changes after large cloud platforms fully restore supply;
- Whether Opus 5 remains Anthropic’s default workhorse;
- Whether enterprise routers publish success, rework and total-cost data by task class.
Frequently asked questions
Does 6% of tokens mean Fable 5 generated 6% of Anthropic revenue?
No. Ramp says Fable represented 6% of Anthropic tokens and 11.4% of model spending in its sample. This is not Anthropic’s complete financial revenue.
Have businesses stopped buying the strongest models?
The data cannot support that claim. A more plausible reading is that businesses reserve frontier models for a smaller set of high-value tasks.
Why not compare token prices directly?
Models can generate different token volumes, take different numbers of tool steps, and have different completion and rework rates. Cost per successful task is more useful.
Sources and evidence boundaries
- Enterprise spending data: Ramp August 2026 AI Index
- Trusted media: Financial Times: Anthropic’s best AI model struggles to attract users
- Primary source: Anthropic: Claude Fable 5 and Mythos 5
- Primary source: Anthropic: Introducing Claude Opus 5
- Primary timeline: Anthropic: Redeploying Fable 5
This article separates model purchasing inside Ramp’s sample from the broader enterprise AI market. The FT article is subscription-gated, so we use only its public headline and publication timing as a media-attention signal; all figures and methodology come from Ramp’s original report.