There's a greater than 10% chance AI "could kill us all" within a decade. Cheerful stuff to open on, but it sets the tone for the week's biggest story: Anthropic's own 154-page account of eight months spent detecting and shutting down attempts to misuse Claude. Structured suspiciously like a tour of hell — seven circles, seven sins, take your pick — it covers state hackers, drone swarms, bioweapons researchers, romance-scam bot farms, and a small army of Chinese AI labs quietly trying to wear Claude's brain as a hat. What's in the box this week? If you're Anthropic: seven reasons not to sleep well tonight.

Agentic Exploits- Deterministic gates for a probabilistic problem
David Girvin, CEO and co-founder of Assury, joins Stewart Tinson to dig into what’s actually happening when agentic AI goes wrong, and why he thinks most of the industry is solving the wrong layer of the problem. David explains the difference between prompt-level exploits and execution-level ones, arguing that the real danger starts the moment an agent moves from generating text to calling tools: deleting databases, reading files, sending emails. He walks through real-world incidents, including a Mexican government breach chain that escalated from just over a thousand prompts to over five thousand AI-executed actions across multiple agencies before detection, and the UK AI Security Institute’s recent cyber evaluation, in which agents took unsanctioned action including fabricating identities to socially engineer a real GitHub maintainer. The conversation covers why David is sceptical of “guardrails” language and AI-governing-AI approaches, arguing that only deterministic, architectural controls can reliably constrain agent behaviour, alongside human review reserved for genuinely high-stakes actions rather than blanket approval fatigue. He breaks down credential starvation, session risk accumulation, and why classifier-based tools keep failing inconsistently on identical actions, pointing to a named frontier lab’s own zero trust paper as an example of the industry misjudging what actually works. Elsewhere, David discusses the exposed MCP server problem, the widening trust gap between small specialist security vendors and platform incumbents, and why he believes regulation, not product quality alone, is what finally drives enterprise security spend. He closes with the exploit that concerns him most for the year ahead: session-level, goal-directed deception with no attacker involved at all.

Anthropic threat report: 7 more things to keep you awake at night
Anthropic's self-reported threat intelligence report runs through eight months of disrupted misuse across seven areas — cyber operations, disinformation, surveillance, conventional weapons, biological research, romance-scam fraud, and a wave of illicit AI distillation by Chinese labs including Alibaba, Moonshot, and DeepSeek. Read the full breakdown of all seven circles here.

Anthropic researcher: more than 10% chance AI "could kill all humans"
The warning that kicked off the week: an Anthropic safety researcher says there's a greater than 10% chance AI could wipe out humanity within a decade, as warnings from inside AI labs grow starker.

UK's AISI denied pre-release access to Anthropic's Claude Mythos 5.1
Anthropic reportedly denied the UK's AI Security Institute pre-release access to Claude Mythos 5.1, the first such exclusion, amid protectionism concerns.

Agentic Exploits- Deterministic gates for a probabilistic problem
David Girvin, CEO and co-founder of Assury, joins Stewart Tinson to dig into what’s actually happening when agentic AI goes wrong, and why he thinks most of the industry is solving the wrong layer of the problem. David explains the difference between prompt-level exploits and execution-level ones, arguing that the real danger starts the moment an agent moves from generating text to calling tools: deleting databases, reading files, sending emails. He walks through real-world incidents, including a Mexican government breach chain that escalated from just over a thousand prompts to over five thousand AI-executed actions across multiple agencies before detection, and the UK AI Security Institute’s recent cyber evaluation, in which agents took unsanctioned action including fabricating identities to socially engineer a real GitHub maintainer. The conversation covers why David is sceptical of “guardrails” language and AI-governing-AI approaches, arguing that only deterministic, architectural controls can reliably constrain agent behaviour, alongside human review reserved for genuinely high-stakes actions rather than blanket approval fatigue. He breaks down credential starvation, session risk accumulation, and why classifier-based tools keep failing inconsistently on identical actions, pointing to a named frontier lab’s own zero trust paper as an example of the industry misjudging what actually works. Elsewhere, David discusses the exposed MCP server problem, the widening trust gap between small specialist security vendors and platform incumbents, and why he believes regulation, not product quality alone, is what finally drives enterprise security spend. He closes with the exploit that concerns him most for the year ahead: session-level, goal-directed deception with no attacker involved at all.

Mistral AI raises €3bn Series D at €21bn+ valuation, led by Samsung
Mistral AI raised €3 billion at a valuation of more than €21 billion in a Samsung-led round, which the company says is the largest-ever European tech equity raise.

Mistral and Cloudera partner on sovereign AI for enterprise data
Mistral and Cloudera announced a partnership letting enterprises deploy and train AI models on their own data across cloud, on-premises or air-gapped environments.

Mistral details 40,000-line Fortran-to-C++ migration for European energy operator
Mistral published lessons from migrating a legacy Fortran 77 reservoir simulator to C++ for an energy client, settling on a human-supervised AI agent workflow after fully autonomous and rigid multi-agent approaches both fell short.

Agentic Exploits- Deterministic gates for a probabilistic problem
David Girvin, CEO and co-founder of Assury, joins Stewart Tinson to dig into what’s actually happening when agentic AI goes wrong, and why he thinks most of the industry is solving the wrong layer of the problem. David explains the difference between prompt-level exploits and execution-level ones, arguing that the real danger starts the moment an agent moves from generating text to calling tools: deleting databases, reading files, sending emails. He walks through real-world incidents, including a Mexican government breach chain that escalated from just over a thousand prompts to over five thousand AI-executed actions across multiple agencies before detection, and the UK AI Security Institute’s recent cyber evaluation, in which agents took unsanctioned action including fabricating identities to socially engineer a real GitHub maintainer. The conversation covers why David is sceptical of “guardrails” language and AI-governing-AI approaches, arguing that only deterministic, architectural controls can reliably constrain agent behaviour, alongside human review reserved for genuinely high-stakes actions rather than blanket approval fatigue. He breaks down credential starvation, session risk accumulation, and why classifier-based tools keep failing inconsistently on identical actions, pointing to a named frontier lab’s own zero trust paper as an example of the industry misjudging what actually works. Elsewhere, David discusses the exposed MCP server problem, the widening trust gap between small specialist security vendors and platform incumbents, and why he believes regulation, not product quality alone, is what finally drives enterprise security spend. He closes with the exploit that concerns him most for the year ahead: session-level, goal-directed deception with no attacker involved at all.

OpenAI's Pachocki: no lab has solved alignment enough to keep scaling at full speed
OpenAI's chief scientist Jakub Pachocki has warned that no AI lab has yet solved alignment and monitoring well enough to justify continued scaling at maximum speed, calling for voluntary slowdowns and international coordination.

OpenAI's Friar ties Astra growth story to Navier-Stokes proof in business update
OpenAI CFO Sarah Friar tied GPT-6 Astra's momentum and a new inference chip to broader business growth, with a nod to an internal model's claimed proof of the Navier-Stokes equations — a 90-year-old open mathematics problem — along the way.

OpenAI announces journalism training programme and $5m teen AI research fund
OpenAI is providing ChatGPT Edu subscriptions to journalism students at CUNY's Newmark J-School and Northwestern's Medill, alongside a separate $5 million fund for independent research into how generative AI affects teenagers aged 13-17.

OpenAI launches Data agent for ChatGPT Work
OpenAI launched a Data agent for ChatGPT Work, letting non-technical staff query company data and build dashboards in plain language.

Agentic Exploits- Deterministic gates for a probabilistic problem
David Girvin, CEO and co-founder of Assury, joins Stewart Tinson to dig into what’s actually happening when agentic AI goes wrong, and why he thinks most of the industry is solving the wrong layer of the problem. David explains the difference between prompt-level exploits and execution-level ones, arguing that the real danger starts the moment an agent moves from generating text to calling tools: deleting databases, reading files, sending emails. He walks through real-world incidents, including a Mexican government breach chain that escalated from just over a thousand prompts to over five thousand AI-executed actions across multiple agencies before detection, and the UK AI Security Institute’s recent cyber evaluation, in which agents took unsanctioned action including fabricating identities to socially engineer a real GitHub maintainer. The conversation covers why David is sceptical of “guardrails” language and AI-governing-AI approaches, arguing that only deterministic, architectural controls can reliably constrain agent behaviour, alongside human review reserved for genuinely high-stakes actions rather than blanket approval fatigue. He breaks down credential starvation, session risk accumulation, and why classifier-based tools keep failing inconsistently on identical actions, pointing to a named frontier lab’s own zero trust paper as an example of the industry misjudging what actually works. Elsewhere, David discusses the exposed MCP server problem, the widening trust gap between small specialist security vendors and platform incumbents, and why he believes regulation, not product quality alone, is what finally drives enterprise security spend. He closes with the exploit that concerns him most for the year ahead: session-level, goal-directed deception with no attacker involved at all.

NVIDIA and Palantir launch sovereign AI stack for supply chains
NVIDIA and Palantir launched a joint AI stack for supply chain operations, starting with NVIDIA's own, with plans to extend the approach to other industries including manufacturing, energy and healthcare.

Microsoft's Discovery Engine tops rival AI agents on new science benchmark
Microsoft says its Discovery Engine, paired with a new reasoning system called CLIO, outperformed rival agentic AI harnesses on a demanding new benchmark for scientific and engineering tasks, already contributing to the discovery of a novel organic battery material.

Microsoft adds cost governance and ROI tracking tools for AI agents in Foundry
Microsoft added tools to its Foundry platform letting enterprises set spending limits on AI agents and measure whether they're delivering enough business value to justify the cost.

Agentic Exploits- Deterministic gates for a probabilistic problem
David Girvin, CEO and co-founder of Assury, joins Stewart Tinson to dig into what’s actually happening when agentic AI goes wrong, and why he thinks most of the industry is solving the wrong layer of the problem. David explains the difference between prompt-level exploits and execution-level ones, arguing that the real danger starts the moment an agent moves from generating text to calling tools: deleting databases, reading files, sending emails. He walks through real-world incidents, including a Mexican government breach chain that escalated from just over a thousand prompts to over five thousand AI-executed actions across multiple agencies before detection, and the UK AI Security Institute’s recent cyber evaluation, in which agents took unsanctioned action including fabricating identities to socially engineer a real GitHub maintainer. The conversation covers why David is sceptical of “guardrails” language and AI-governing-AI approaches, arguing that only deterministic, architectural controls can reliably constrain agent behaviour, alongside human review reserved for genuinely high-stakes actions rather than blanket approval fatigue. He breaks down credential starvation, session risk accumulation, and why classifier-based tools keep failing inconsistently on identical actions, pointing to a named frontier lab’s own zero trust paper as an example of the industry misjudging what actually works. Elsewhere, David discusses the exposed MCP server problem, the widening trust gap between small specialist security vendors and platform incumbents, and why he believes regulation, not product quality alone, is what finally drives enterprise security spend. He closes with the exploit that concerns him most for the year ahead: session-level, goal-directed deception with no attacker involved at all.

Salesforce completes acquisition of Fin, adds FIDE as chess partner
Salesforce completed its acquisition of customer service AI agent company Fin, formerly Intercom, and separately became title sponsor and AI partner of the International Chess Federation.

MIT researcher uses GPT-5.6 Sol to automate quantum chip calibration
An MIT graduate student used GPT-5.6 Sol and Codex to automate routine measurements on superconducting quantum computing chips overnight, freeing her to focus on experiment design.

Bioengineer uses Codex and ChatGPT to search genomes for new antimicrobials
César de la Fuente's lab uses Codex and ChatGPT alongside its own AI models to search the genomes of living and extinct organisms for antimicrobial candidates, cutting years-long discovery work down to hours.

Alan Turing Institute: human oversight key to AI adoption in national security
A new Alan Turing Institute paper warns AI adoption in national security must preserve human judgement and oversight, not just automate intelligence tasks.


Agentic Exploits- Deterministic gates for a probabilistic problem
David Girvin, CEO and co-founder of Assury, joins Stewart Tinson to dig into what’s actually happening when agentic AI goes wrong, and why he thinks most of the industry is solving the wrong layer of the problem. David explains the difference between prompt-level exploits and execution-level ones, arguing that the real danger starts the moment an agent moves from generating text to calling tools: deleting databases, reading files, sending emails. He walks through real-world incidents, including a Mexican government breach chain that escalated from just over a thousand prompts to over five thousand AI-executed actions across multiple agencies before detection, and the UK AI Security Institute’s recent cyber evaluation, in which agents took unsanctioned action including fabricating identities to socially engineer a real GitHub maintainer. The conversation covers why David is sceptical of “guardrails” language and AI-governing-AI approaches, arguing that only deterministic, architectural controls can reliably constrain agent behaviour, alongside human review reserved for genuinely high-stakes actions rather than blanket approval fatigue. He breaks down credential starvation, session risk accumulation, and why classifier-based tools keep failing inconsistently on identical actions, pointing to a named frontier lab’s own zero trust paper as an example of the industry misjudging what actually works. Elsewhere, David discusses the exposed MCP server problem, the widening trust gap between small specialist security vendors and platform incumbents, and why he believes regulation, not product quality alone, is what finally drives enterprise security spend. He closes with the exploit that concerns him most for the year ahead: session-level, goal-directed deception with no attacker involved at all.
Share this post
The link has been copied!