Fable 5.1 is here, Anthropic flexes its muscles once again.
Fable 5.1 and Mythos 5.1 were suddenly launched in the early hours of the morning.
The previous-generation Fable 5 has almost become synonymous with SOTA models, and before its halo faded, Anthropic had already started iterating at a fast pace on its own products.
The two models were released simultaneously: Fable 5.1 is targeted at general users and enterprises, while Mythos 5.1 opens the same set of cutting-edge capabilities to "verified cybersecurity and life science institutions".
In terms of pricing, the base price remains unchanged: Fable 5.1 still costs $10 per million input Tokens and $50 per million output Tokens. However, the cache read price has been cut by 75%, dropping from $1 per million Tokens to $0.25 per million Tokens.
Performance improvements are clearly concentrated on long-duration, multi-step Agent tasks. The score of Terminal-Bench-Science rose from 24.7% to 52.6%, and the score of AutomationBench increased from 17.1% to 31.4%.
Compared with the model itself, what may be more noteworthy is the timing of the model release — this is very likely the last heavyweight model launch before Anthropic publicly submits its listing documents.
In a few days, it will be September 7, the US Labor Day. According to the plan previously disclosed by The Information, Anthropic, which has secretly submitted its IPO documents, may officially publish its prospectus after Labor Day, hold an investor day in mid-September, and go public as soon as from the end of September to the beginning of October.
The aftermath of Anthropic's previous "nominal increase but actual decrease" adjustment to quota also appeared in the comment section of the new model release. On August 29, Anthropic announced that it would raise the "standard weekly quota" by 25% starting from September 14, but this figure is calculated based on the old standard before May; compared with the temporary quota currently in use by users, the actual quota has decreased by 17% instead.
Shortly after Fable 5.1 went online, some users asked Anthropic whether it could offer "more transparent and more generous" session and weekly quotas at the same time.
Fable 5.1: Pushing Agents to Tackle Longer Tasks
If you only look at the specifications, Fable 5.1 is not a large-scale upgrade.
It inherits the 1 million Token context window and 128K maximum output of Fable 5, and the base price remains the same: $10 per million input Tokens, and $50 per million output Tokens.
In the current lineup of flagship models, Fable 5.1 is still in a significantly more expensive tier: calculated based on 1 million input Tokens plus 1 million output Tokens, the price of Fable 5.1 is 2.5 times that of the current GPT-5.6 Sol, and more than 10 times that of GLM-5.3.
Anthropic officials even still recommend that most tasks start with Opus 5, whose price is only half of Fable 5.1, and turn to Fable 5.1 only when facing high-difficulty reasoning, long-duration Agent tasks, or when Opus 5 is still insufficient under high-effort mode.
The noteworthy change lies in the cost of continuous Agent operation.
The 5-minute and 1-hour cache write prices for Fable 5.1 remain $12.5 and $20 per million Tokens respectively, but the cache read price has dropped from $1 per million Tokens for Fable 5 to $0.25 per million Tokens, which is equivalent to a direct 75% cut.
Cache read can be understood as: after a segment of context has been processed by the model, subsequent requests do not need to be charged at the full input price every time. For an ordinary Q&A, this part may not be important; but for an Agent that runs continuously for several hours, it will repeatedly read system prompts, code bases, tool definitions and previously processed context. The longer the task and the more tool calls, the higher the proportion of cache read in the total cost.
Anthropic has conducted calculations using actual usage data from four weeks in August this year. In scenarios billed by Tokens, switching the same task from Fable 5 to Fable 5.1 is expected to reduce the overall cost of typical workloads by about 25%; for complex coding and highly Agent-automated tasks with heavier context and more tool calls, the maximum cost reduction can reach about 45%.
In other words, although Anthropic did not sell Fable 5.1 itself at a lower price, it specifically reduced the cost of "letting it work for a long time".
However, the lower price is only relative to the past, and the base price is still high. Some developers reminded in the release comment section that even with a very high cache hit rate, costs can still accumulate rapidly, and a real complex coding task may cost $20 to $50.
In addition, this price reduction mainly applies to APIs and other scenarios settled by Tokens. For users who directly subscribe to Claude Code, another set of "quota accounting" has just caused controversy.
On August 29, the official Claude Developers account announced that starting from September 14, for Pro, Max, Team and seat-based Enterprise plans, the "standard weekly quota" of Claude Code will be permanently increased by 25%. But the problem is that Anthropic had previously temporarily increased this set of quotas by 50%, and this temporary quota will remain in effect until September 14.
In other words, if the original quota is recorded as 100, users now actually get 150; after September 14, the quota will become 125. Nominally, it is a "permanent 25% increase" relative to the old standard, but it is about 17% less than the quota that users can actually use now. Anthropic later added a special explanation to confirm this calculation.
In the comment section after the release of Fable 5.1, some users continued to ask Anthropic whether it could provide "more transparent and more generous" session and weekly quotas, and stop playing similar word games.
However, from the perspective of performance changes, Fable 5.1 is indeed increasingly becoming a model specially prepared for long-duration tasks.
Among the benchmarks announced by Anthropic, the most dramatic improvement is in Terminal-Bench-Science 0.1, with the score of Fable 5 rising from 24.7% to 52.6%, which is almost doubled. Terminal-Bench 4.0, which tests the Agent capability in terminal environments, increased from 42.0% to 55.8%, and AutomationBench for complex business workflows increased from 17.1% to 31.4%.
But when it comes to more traditional reasoning and coding benchmarks, the improvement is not that dramatic.
Humanity's Last Exam without tools increased from 57.8% to 60.9%; with tools enabled, the score rose from 63.8% to 65.0%. CursorBench 3.2.0 increased from 70.5% to 73.4%, and the knowledge work benchmark GDPval-AA v2 increased from 1723 to 1853.
Overall, the improvements of Fable 5.1 are obviously concentrated on tasks that require the model to continuously call tools, process multiple steps, check intermediate results, and then decide the next action, rather than being evenly distributed across all capabilities.
Early enterprise tests also reflect this trend.
In its most difficult browser Agent benchmark, Browserbase let Fable 5.1 complete 82% of the tasks, while Opus 5 completed 74% and Fable 5 only completed 57%.
In a test conducted by Ramp, Fable 5.1 ran continuously for 38 hours without human intervention. It first found that the results of a previous machine learning experiment were actually affected by label artifacts, then corrected the problem on its own, launched 6 sets of parallel experiments at the same time and ran them all night, and finally returned with new experimental results and suggestions for the next step. In another open-ended task, Ramp only asked it to find "problems that no one is responsible for but are most worth solving". It found an unhandled alert related to production accidents by itself, retrieved the logs and gave the repair plan.
The feedback from Canva is more focused on generation quality. Its test shows that the text expression of Fable 5.1 is easier to understand and more compliant with writing specifications; in the blind test against Fable 5, Canva prefers the output of 5.1. Moreover, in Canva Code, Fable 5.1 directly created a rhythm game: it not only generated levels and real music, but also matched the gameplay rhythm with the music beats. Canva said that no other model they tested had achieved this.
From benchmarks to enterprise tests, the most obvious improvement of Fable 5.1 this time is that it begins to complete complex tasks more stably: it can continuously judge, correct errors and push forward in tasks lasting dozens of hours, and can also organize text, code, music and interaction more completely in a single generation.
Claude Steps into the Laboratory
Fable 5.1 and Mythos 5.1 look like two different models, but Anthropic made it very clear in its official documentation:
Mythos 5.1 is "identical" to Fable 5.1, the only difference is that it provides "more permissive safeguards" for audited users.
In other words, the only difference between the two lies in security restrictions: Fable 5.