xAI's Grok 4.6 has topped an independent biosecurity benchmark measuring how reliably AI models refuse disguised or hazardous biological requests while still completing legitimate research tasks.

The analysis, published on 1 September by AI evaluation firm LatchBio, tested Grok 4.6 against other frontier models on two benchmark suites. On BioSecBench-Refusal, which hides biosecurity hazards inside routine-looking research tasks, Grok 4.6 was the only model to score above 50% on both refusing dangerous red-team tasks and completing routine biological work, averaging 62.1% across different testing setups. It refused 59.2% of red-team tasks while still completing 64.8% of legitimate ones.

On a separate biosurveillance benchmark, testing pathogen genomic monitoring workflows, Grok 4.6 scored 53.5%, placing behind Anthropic's Opus 5 but ahead of OpenAI's GPT-5.6 Sol.

Agentic Exploits- Deterministic gates for a probabilistic problem
David Girvin, CEO and co-founder of Assury, joins Stewart Tinson to dig into what’s actually happening when agentic AI goes wrong, and why he thinks most of the industry is solving the wrong layer of the problem. David explains the difference between prompt-level exploits and execution-level ones, arguing that the real danger starts the moment an agent moves from generating text to calling tools: deleting databases, reading files, sending emails. He walks through real-world incidents, including a Mexican government breach chain that escalated from just over a thousand prompts to over five thousand AI-executed actions across multiple agencies before detection, and the UK AI Security Institute’s recent cyber evaluation, in which agents took unsanctioned action including fabricating identities to socially engineer a real GitHub maintainer. The conversation covers why David is sceptical of “guardrails” language and AI-governing-AI approaches, arguing that only deterministic, architectural controls can reliably constrain agent behaviour, alongside human review reserved for genuinely high-stakes actions rather than blanket approval fatigue. He breaks down credential starvation, session risk accumulation, and why classifier-based tools keep failing inconsistently on identical actions, pointing to a named frontier lab’s own zero trust paper as an example of the industry misjudging what actually works. Elsewhere, David discusses the exposed MCP server problem, the widening trust gap between small specialist security vendors and platform incumbents, and why he believes regulation, not product quality alone, is what finally drives enterprise security spend. He closes with the exploit that concerns him most for the year ahead: session-level, goal-directed deception with no attacker involved at all.

xAI said evaluation traces show Grok 4.6 reasoning over task context and environment to assess intent before deciding whether to proceed, rather than reacting to specific keywords. The company said this represents a marked improvement over its earlier Grok 4.5 and Grok 4.3 models.

xAI described its safeguards for Grok as layered and defence-in-depth, combining refusal training on how and when to decline tasks, inference-time filters designed to reject harmful requests before they reach the model, behavioural controls applied during deployment, and post-deployment monitoring to detect adversarial usage patterns at the session and user level.

Looking ahead, xAI said it plans wider and more ambitious testing of future models, including broader pre-deployment evaluation suites, additional third-party assessments, improved post-deployment monitoring, and closer work with biology-focused institutions. The company added that it treats overrefusal of legitimate biological and public-health work as an equally serious risk to permitting malicious use.


Agentic Exploits- Deterministic gates for a probabilistic problem
David Girvin, CEO and co-founder of Assury, joins Stewart Tinson to dig into what’s actually happening when agentic AI goes wrong, and why he thinks most of the industry is solving the wrong layer of the problem. David explains the difference between prompt-level exploits and execution-level ones, arguing that the real danger starts the moment an agent moves from generating text to calling tools: deleting databases, reading files, sending emails. He walks through real-world incidents, including a Mexican government breach chain that escalated from just over a thousand prompts to over five thousand AI-executed actions across multiple agencies before detection, and the UK AI Security Institute’s recent cyber evaluation, in which agents took unsanctioned action including fabricating identities to socially engineer a real GitHub maintainer. The conversation covers why David is sceptical of “guardrails” language and AI-governing-AI approaches, arguing that only deterministic, architectural controls can reliably constrain agent behaviour, alongside human review reserved for genuinely high-stakes actions rather than blanket approval fatigue. He breaks down credential starvation, session risk accumulation, and why classifier-based tools keep failing inconsistently on identical actions, pointing to a named frontier lab’s own zero trust paper as an example of the industry misjudging what actually works. Elsewhere, David discusses the exposed MCP server problem, the widening trust gap between small specialist security vendors and platform incumbents, and why he believes regulation, not product quality alone, is what finally drives enterprise security spend. He closes with the exploit that concerns him most for the year ahead: session-level, goal-directed deception with no attacker involved at all.
Share this post
The link has been copied!