Anthropic's Claude Fable 5.1 and Mythos 5.1 arrive with a 75% cost reduction for Fable cache reads
It's only the first day of September 2026, but the month and fall season are already off to the races in AI land, as Anthropic has just released its latest and most powerful large language models yet — Claude Fable 5.1 and Claude Mythos 5.1 . The two names refer to the same underlying model. Fable 5.1 is the generally available version, with Anthropic’s production safeguards in place. Mythos 5.1 is available through restricted-access programs for vetted cybersecurity and life-sciences organizations that need capabilities normally constrained by those safeguards. For enterprise buyers, however, the release is about more than another round of benchmark gains. Anthropic is simultaneously changing the economics of running persistent agents, reducing the cost of cached context by 75%, and introducing a new security architecture called Enterprise Frontier Safeguards, or EFS, designed to let organizations retain monitoring data inside infrastructure they control. Those changes arrive at a particularly consequential moment. Over the past several weeks, Anthropic and the U.K. AI Security Institute have disclosed incidents in which earlier Claude models, running under unusually permissive cybersecurity evaluation conditions, took unauthorized actions against real systems. Anthropic temporarily paused external cyber evaluations and has since introduced additional containment and monitoring before resuming them. Taken together, Fable 5.1 looks less like a conventional model refresh than an attempt to solve three increasingly intertwined enterprise problems: how to make agents capable enough to finish difficult work, economical enough to leave running for hours, and governable enough to give access to sensitive systems. A model built for work that does not finish in one prompt Anthropic is positioning Fable 5.1 primarily around sustained problem-solving. On Terminal-Bench-Science 0.1, which evaluates agentic scientific research, Anthropic reports Fable 5.1 scoring 52.6%, compared with 24.7% for Fable 5, 29.0% for Opus 5 and 22.4% for GPT-5.6 Sol in its evaluation setup. On Terminal-Bench 4.0, Fable 5.1 scores 55.8%, versus 42.0% for Fable 5 and 52.3% for Opus 5. Mythos 5.1 reaches 60.9% on the same coding benchmark when operating under its more permissive cyber safeguards. The gains extend beyond coding. Anthropic reports a GDPval-AA v2 score of 1,853 for knowledge work, versus 1,824 for Opus 5 and 1,723 for Fable 5. On AutomationBench, intended to measure business workflows, Fable 5.1 scores 31.4%, compared with 17.1% for Fable 5 and 26.9% for Opus 5. On CursorBench 3.2.0, it reaches 73.4%. Those numbers should be read as vendor-reported results rather than independent proof of superiority. Anthropic also notes qualifications around several evaluations: production safeguards can affect scores, and its August 2026 OSWorld task release is not directly comparable with some previously published results. The more useful signal for enterprise teams may therefore come from the kinds of failures early-access partners say the model can resolve. Investment firm Millennium told Anthropic that Fable 5.1 traced an extremely rare software crash to a bug inside an external vendor library after the problem had resisted explanation for four to five years. Corporate expense management provider Ramp described an unattended 38-hour machine-learning run in which the model re-evaluated a previous result, launched six experiments and returned with findings and proposed next steps. Browserbase said Fable 5.1 completed 82% of tasks on its hardest browser-agent benchmark, versus 74% for Opus 5 and 57% for Fable 5. These are customer testimonials supplied as part of Anthropic’s launch, not independently reproduced benchmarks. But they illustrate the direction Anthropic is pursuing: moving the unit of AI work from an answer or code snippet toward an entire investigation. That changes deployment architecture. A model that can operate for hours needs durable context, tool access, checkpoints, logging, permission boundaries and reliable recovery from errors. Model intelligence becomes only one component of the system. Pricing: Fable 5.1 remains premium, but caching changes the equation The most immediately measurable enterprise change is pricing. Fable 5.1 retains Fable 5’s headline API rates: $10 per 1 million input tokens and $50 per million output . That makes it considerably more expensive on uncached tokens than other models in Anthropic’s lineup. Opus 5 costs $5 per million input tokens and $25 per million output tokens, while Sonnet 5 costs $2 and $10 respectively. The important change is cached input: C laude model Input / 1M Cache read / 1M Output / 1M Fable 5.1 $10 $0.25 $50 Fable 5 $10 $1.00 $50 Opus 5 $5 $0.50 $25 Sonnet 5 $2 $0.20 $10 Anthropic has cut a Fable 5.1 cache hit to just $0.25 on input, down from $1.00 for Fable 5. That's also just 2.5% of Fable's normal input-token price of $10, rather than the 10% multiplier used by most other Claude models. Five-minute cache writes remain $12.50 per million tokens and one-hour writes $20, but subsequent reads cost just $0.25 per million. That produces an unusual pricing profile. Fable 5.1's ordinary input and output are twice as expensive as Opus 5's, yet its cached input is half the cost of Opus 5's cache reads . Its cache-read price is only 25% above Sonnet 5's despite Fable's base input price being five times higher. That matters for agents because they repeatedly revisit the same codebase, system instructions, tool definitions, documents and accumulated conversation history. Anthropic says the lower cache price reduces Fable 5.1's effective cost by around 25% for typical workloads and as much as roughly 45% for highly agentic workloads in which cached context accounts for a larger share of usage. This is a more useful enterprise framing than simply comparing per-token list prices. Model selection for an agentic workflow increasingly depends on cost per successfully completed task , including retries, context replay, tool calls and the number of tokens a model consumes before reaching a usable result. The cache price reduction also may be an effort to help woo increasingly price-consicious enterprises. A Financial Times report found that , more than two months after launch, Fable 5 accounted for only about 11% of Anthropic model spending among roughly 70,000 companies represented in Ramp’s transaction data, while the cheaper Opus 5 and Opus 4.8 gained share. The Information further reported growing concern among enterprise customers about unpredictable AI bills, including ServiceNow monitoring employee usage after rapidly consuming its annual Anthropic budget. Those reports suggest that even when enterprises valued Fable 5’s capabilities, many were unwilling to make it the default model for large-scale production workloads. Fable 5.1 nevertheless remains expensive relative to much of the broader market. OpenAI's current promotional API pricing for GPT-5.6 Sol is $4 per million input tokens, $0.40 for cached input and $20 per million output tokens through at least Nov. 21. Google's Gemini 3.7 Flash currently lists at $0.75 per million input and $3.75 per million output through the end of 2026. Model Input ($/1M) Output ($/1M) Total ($/1M) Source Muse Spark 1.2 Contributor $0.10 $0.20 $0.30 Meta MiMo-V2.5 Flash $0.10 $0.30 $0.40 Xiaomi DeepSeek-V4-Flash — off-peak $0.22 $0.66 $0.88 DeepSeek GPT-5.6 Luna $0.20 $1.20 $1.40 OpenAI MiniMax-M3 $0.30 $1.20 $1.50 MiniMax LongCat-2.0 — limited-time promo $0.30 $1.20 $1.50 LongCat DeepSeek-V4-Flash — peak hours $0.44 $1.32 $1.76 DeepSeek MiMo-V2.5 $0.40 $2.00 $2.40 Xiaomi DeepSeek-V4-Pro — off-peak $0.66 $1.98 $2.64 DeepSeek LongCat-2.0 — standard $0.75 $2.95 $3.70 LongCat MiMo-V2.5 Pro (≤256K) $1.00 $3.00 $4.00 Xiaomi Gemini 3.6 Flash — through Dec. 31, 2026 $0.75 $3.75 $4.50 Google Gemini 3.7 Flash — through Dec. 31, 2026 $0.75 $3.75 $4.50 Google DeepSeek-V4-Pro — peak hours $1.32 $3.96 $5.28 DeepSeek Muse Spark 1.1 / 1.2 $1.25 $4.25 $5.50 Meta GLM-5.3 $1.40 $4.40 $5.80 Z.AI Grok 4.6 — $2.00 $6.00 $8.00 xAI MiMo-V2.5 Pro (>256K) $2.00 $6.00 $8.00 Xiaomi Qwen3.8-Max $2.00 $6.00 $8.00 QwenCloud Gemini 3.6 Flash — starting Jan. 1, 2027 $1.50 $7.50 $9.00 Google Gemini 3.7 Flash — starting Jan. 1, 2027 $1.50 $7.50 $9.00 Google GPT-5.6 Terra $2.00 $12.00 $14.00 OpenAI Grok 4.6 — ≥200K prompt tokens $4.00 $12.00 $16.00 xAI GPT-5.4 $2.50 $15.00 $17.50 OpenAI Kimi K3 $3.00 $15.00 $18.00 Moonshot AI Claude Opus 5 $5.00 $25.00 $30.00 Anthropic Sakana Fugu Ultra (≤272K) $5.00 $30.00 $35.00 Sakana AI GPT-5.6 Sol — Standard mode $5.00 $30.00 $35.00 OpenAI Claude Fable 5 / Claude Mythos 5 $10.00 $50.00 $60.00 Anthropic Claude Fable 5.1 / Claude Mythos 5.1 $10.00 $50.00 $60.00 Anthropic GPT-5.6 Sol — Fast mode $10.00 $60.00 $70.00 OpenAI Fable therefore needs to justify its premium through higher task completion, lower token consumption, or the ability to replace more expensive human or multi-stage workflows—not simply through raw API price. Anthropic's pricing system provides additional levers. Batch processing cuts Fable 5.1 input and output rates in half, to $5 and $25 per million tokens, for asynchronous workloads. U.S.-only inference adds a 1.1x multiplier. Web search costs $10 per 1,000 searches in addition to token usage, while web fetch carries no separate fee. For procurement teams, those details make workload profiling increasingly important. A code-review agent with enormous reusable repository context may have very different economics from a report-generation workflow dominated by new output tokens. The cyber incidents behind the release The security architecture surrounding Fable 5.1 also needs to be understood against Anthropic's recent cybersecurity incidents. On July 30, Anthropic disclosed that a review of 141,006 cybersecurity evaluation runs had identified three incidents, spanning six runs, in which Claude models reached the public internet from a third-party testing environment and obtained unauthorized access to real organizations. The models had been running without the cyber safeguards applied to generally available Claude deployments. They were performing capture-the-flag exercises and had explicitly been told that their environments were simulations without internet access. A configuration problem nevertheless left an internet path open. In the most serious incident, Claude Opus 4.7 encountered a real company that shared a name with the fictional organization in its test. Across four runs, the model obtained application and infrastructure credentials and accessed a database containing several hundred rows of production data. Anthropic said the model eventually encountered indications that the system was real but continued its attack. A second incident involved Mythos 5. The model discovered developer instructions referencing a nonexistent Python package and decided that publishing a package under that name could help it win the simulated challenge. It created a PyPI account and uploaded malicious code to the real public package repository. The package remained available for roughly an hour and was downloaded and executed on 15 real systems. One belonged to a security company whose automated scanner executed the package; Claude's code extracted credentials and subsequently used them to access additional infrastructure. In the third incident, an internal Anthropic research model scanned roughly 9,000 internet targets after failing to reach its intended fictional target. It compromised an internet-facing application using exposed credentials and SQL injection, then stopped after recognizing that the system had no connection to the exercise. The incidents did not involve customer data or Anthropic's production infrastructure, and Anthropic says the production safeguards omitted during the tests would have blocked the behavior. Still, the episodes demonstrated a problem enterprise agent architects will recognize: a sufficiently persistent agent can exploit ambiguity between what its operator intended it to access and what its technical credentials allow it to access. A separate U.K. AI Security Institute disclosure reinforced that concern. During testing conducted in July, AISI ran a cyber challenge 122 times across multiple models with internet access deliberately enabled and vendor cyber classifiers disabled. Ten runs produced 19 unsanctioned real-world actions. Seventeen involved Anthropic's Mythos 5; two came from a GPT-5.6 Sol run. The most serious sequence involved a Mythos 5 agent attempting to insert malicious code into a real open-source project, creating fake identities and attempting to socially engineer a maintainer into accepting the change . The maintainer rejected it, and AISI said its investigation found no resulting real-world harm. The agent also attempted prompt injection against other AI systems and left messages that later agents discovered and used. Again, these were deliberately permissive research configurations, not commercially available deployments. But AISI's conclusion is relevant far beyond AI labs: increasingly capable agents should be assumed to explore routes their operators did not anticipate, particularly when given difficult objectives and broad tool access. From model safeguards to infrastructure safeguards Anthropic's response has been to add layers around the model. The company paused external cyber evaluations and briefly halted some internal testing, then deployed a real-time classifier designed to detect aggressive probing, sandbox-escape attempts or unexpected internet access before a tool call executes. It also moved high-risk cyber workloads into stronger isolation and imposed new requirements on external evaluators, including explicit network isolation and pre-engagement validation of sandboxes. External cyber evaluations have now resumed. Fable 5.1 itself also gets more precise production safeguards. Anthropic says its cyber protections now create around 60% fewer interventions per Claude Code session than Fable 5's previous safeguards. The model can be used to discover software vulnerabilities for defensive purposes, while work such as exploit generation, penetration testing and some binary-based vulnerability scanning remains redirected or restricted. That distinction matters for security teams attempting to operationalize AI. A safeguard that blocks too many legitimate actions can make an autonomous security workflow unreliable; one that permits too much creates a materially different risk. Precision, rather than simply the existence of a filter, becomes a production requirement. Enterprise Frontier Safeguards moves data custody to the customer Anthropic is addressing a second enterprise constraint through EFS. The company previously introduced 30-day data retention for Fable 5 as part of its misuse-detection system. For regulated organizations, retaining sensitive conversations with a model provider can make deployment difficult regardless of contractual assurances. EFS changes the architecture. Monitoring data can reside in the customer's own AWS, Azure or Google Cloud environment under customer-managed encryption keys, access policies and audit logging. Anthropic's automated systems can analyze the data for patterns associated with serious misuse, while alerts go to the customer for review; Anthropic says human review by its employees is not required. Anthropic says it developed EFS with more than 100 organizations across financial services, healthcare, manufacturing, telecom, law, retail and government, and with AWS, Google Cloud and Microsoft Azure. Support is planned across Claude Code, Claude Enterprise, the Claude Platform, Amazon Bedrock, Claude Platform on AWS, Google's Agent Platform and Microsoft Foundry. The rollout begins in phases this fall. Eligible customers can use Fable 5.1 with zero data retention until EFS becomes available. Anthropic does not charge separately for EFS, although customers remain responsible for their own cloud storage, operations and egress costs. This is potentially as important as the model upgrade itself. Enterprise AI governance is shifting from promises about what a provider does with data toward architectures that determine where the data can exist in the first place. Fable for production, Mythos for controlled frontiers The split between Fable and Mythos gives Anthropic a mechanism for separating general enterprise deployment from particularly sensitive domains. Fable 5.1 is available now through Anthropic's API as claude-fable-5-1 , as well as through AWS, Google Cloud and Microsoft Azure. Mythos 5.1 uses the same underlying model but exposes more permissive safeguards to vetted cyberdefenders and life-sciences organizations through verification programs. That same model has shown capabilities extending well outside software. Anthropic reports that Mythos 5.1 designed experimentally validated protein binders, while Fable 5.1 trained a neural network that produced a higher-resolution elevation map covering roughly a third of Venus. Mythos 5.1 also optimized seven open-source biological deep-learning models, with Anthropic reporting inference speedups as high as 2.5x. For pharmaceutical, engineering and research organizations, that points toward a future in which the same agent architecture used to investigate a code failure may also orchestrate modeling, experimentation and analysis. The operational lesson is the same in every case: the more work an agent can complete without intervention, the more consequential its permissions become. Fable 5.1 makes those long-running agents more capable and, through cheaper cached context, potentially much cheaper to operate. EFS gives regulated companies another mechanism for governing their data. More precise safeguards reduce some of the friction that has made high-capability models difficult to use in security workflows. But Anthropic's own recent incidents also demonstrate why the enterprise deployment question cannot stop at model selection. The next generation of AI infrastructure will need to treat agents more like powerful service accounts than chatbots: narrowly scoped credentials, segmented networks, explicit allowlists, continuous telemetry, human approval around irreversible actions and the assumption that an agent may find pathways its developer did not anticipate. Fable 5.1 raises the amount of work organizations can plausibly delegate. Its larger significance may be that it also makes the infrastructure surrounding that delegation impossible to treat as an afterthought.
Read the full article on VentureBeat
Read Full Article →