xAI's Grok 4.6 has topped an independent biosecurity benchmark measuring how reliably AI models refuse disguised or hazardous biological requests while still completing legitimate research tasks.
The analysis, published on 1 September by AI evaluation firm LatchBio, tested Grok 4.6 against other frontier models on two benchmark suites. On BioSecBench-Refusal, which hides biosecurity hazards inside routine-looking research tasks, Grok 4.6 was the only model to score above 50% on both refusing dangerous red-team tasks and completing routine biological work, averaging 62.1% across different testing setups. It refused 59.2% of red-team tasks while still completing 64.8% of legitimate ones.
On a separate biosurveillance benchmark, testing pathogen genomic monitoring workflows, Grok 4.6 scored 53.5%, placing behind Anthropic's Opus 5 but ahead of OpenAI's GPT-5.6 Sol.

xAI said evaluation traces show Grok 4.6 reasoning over task context and environment to assess intent before deciding whether to proceed, rather than reacting to specific keywords. The company said this represents a marked improvement over its earlier Grok 4.5 and Grok 4.3 models.
xAI described its safeguards for Grok as layered and defence-in-depth, combining refusal training on how and when to decline tasks, inference-time filters designed to reject harmful requests before they reach the model, behavioural controls applied during deployment, and post-deployment monitoring to detect adversarial usage patterns at the session and user level.
Looking ahead, xAI said it plans wider and more ambitious testing of future models, including broader pre-deployment evaluation suites, additional third-party assessments, improved post-deployment monitoring, and closer work with biology-focused institutions. The company added that it treats overrefusal of legitimate biological and public-health work as an equally serious risk to permitting malicious use.
