Anthropic has launched Claude Fable 5.1 and Claude Mythos 5.1, twin versions of the same underlying model that the company says are its most advanced yet for coding and knowledge work, with Mythos 5.1 offering reduced safeguards for vetted cybersecurity and life sciences use.

Fable 5.1, generally available now, is around 25% cheaper than its predecessor for typical workloads, and up to roughly 45% cheaper for highly agentic tasks, driven by a 75% cut to cache-read pricing to $0.25 per million tokens. Base pricing otherwise remains at $10 per million input tokens and $50 per million output tokens. Anthropic said the model outperforms Fable 5 across benchmarks it tested, including coding, computer use and multidisciplinary reasoning, and cited early access partners including Jane Street, Cognition, MongoDB and Red Hat reporting stronger performance and cost efficiency in their own testing.

The company also introduced Enterprise Frontier Safeguards, a system letting enterprise customers store data on their own cloud infrastructure while retaining what Anthropic describes as zero-data-retention-equivalent privacy, rolling out in phases from this autumn.

Agentic Exploits- Deterministic gates for a probabilistic problem
David Girvin, CEO and co-founder of Assury, joins Stewart Tinson to dig into what’s actually happening when agentic AI goes wrong, and why he thinks most of the industry is solving the wrong layer of the problem. David explains the difference between prompt-level exploits and execution-level ones, arguing that the real danger starts the moment an agent moves from generating text to calling tools: deleting databases, reading files, sending emails. He walks through real-world incidents, including a Mexican government breach chain that escalated from just over a thousand prompts to over five thousand AI-executed actions across multiple agencies before detection, and the UK AI Security Institute’s recent cyber evaluation, in which agents took unsanctioned action including fabricating identities to socially engineer a real GitHub maintainer. The conversation covers why David is sceptical of “guardrails” language and AI-governing-AI approaches, arguing that only deterministic, architectural controls can reliably constrain agent behaviour, alongside human review reserved for genuinely high-stakes actions rather than blanket approval fatigue. He breaks down credential starvation, session risk accumulation, and why classifier-based tools keep failing inconsistently on identical actions, pointing to a named frontier lab’s own zero trust paper as an example of the industry misjudging what actually works. Elsewhere, David discusses the exposed MCP server problem, the widening trust gap between small specialist security vendors and platform incumbents, and why he believes regulation, not product quality alone, is what finally drives enterprise security spend. He closes with the exploit that concerns him most for the year ahead: session-level, goal-directed deception with no attacker involved at all.

On safety, Anthropic said its testing found Mythos 5.1 has greater biological and cyber capabilities than its predecessor but remains below the next risk tier in its Responsible Scaling Policy, so it is being deployed with the same restrictions applied to Mythos 5. The company said its own alignment evaluations showed Mythos 5.1 was better aligned than Mythos 5 on most measures, including being less likely to attempt reward hacking or use motivated reasoning to justify its actions.

Separately, Anthropic said it has strengthened mechanisms to resist "distillation" attacks, in which rivals extract a model's capabilities at scale, by preventing new API accounts from editing Claude's prior reasoning within a conversation. It also confirmed it has signed the EU AI Act's Code of Practice on Transparency of AI-Generated Content, adding an invisible watermark to outputs from models released after 2 August and opening a private-preview detection API for eligible organisations such as regulators and fact-checkers.

Anthropic also detailed early scientific applications, including Mythos 5.1 designing high-affinity protein binders, speeding up open-source genomics models by up to 2.5 times, and Fable 5.1 producing a higher-resolution elevation map of Venus using decades-old NASA radar data.


Agentic Exploits- Deterministic gates for a probabilistic problem
David Girvin, CEO and co-founder of Assury, joins Stewart Tinson to dig into what’s actually happening when agentic AI goes wrong, and why he thinks most of the industry is solving the wrong layer of the problem. David explains the difference between prompt-level exploits and execution-level ones, arguing that the real danger starts the moment an agent moves from generating text to calling tools: deleting databases, reading files, sending emails. He walks through real-world incidents, including a Mexican government breach chain that escalated from just over a thousand prompts to over five thousand AI-executed actions across multiple agencies before detection, and the UK AI Security Institute’s recent cyber evaluation, in which agents took unsanctioned action including fabricating identities to socially engineer a real GitHub maintainer. The conversation covers why David is sceptical of “guardrails” language and AI-governing-AI approaches, arguing that only deterministic, architectural controls can reliably constrain agent behaviour, alongside human review reserved for genuinely high-stakes actions rather than blanket approval fatigue. He breaks down credential starvation, session risk accumulation, and why classifier-based tools keep failing inconsistently on identical actions, pointing to a named frontier lab’s own zero trust paper as an example of the industry misjudging what actually works. Elsewhere, David discusses the exposed MCP server problem, the widening trust gap between small specialist security vendors and platform incumbents, and why he believes regulation, not product quality alone, is what finally drives enterprise security spend. He closes with the exploit that concerns him most for the year ahead: session-level, goal-directed deception with no attacker involved at all.
Share this post
The link has been copied!