xAI has expanded access to its Grok models and agent tools across several platforms this week, widening both enterprise distribution and consumer availability.

Grok 4.6, the company's flagship model, is designed to handle agent tasks that run over long stretches of time alongside demanding interactive and visual work, backed by a 500,000-token context window and a choice of four reasoning intensities. The model is now available through Microsoft Foundry, giving organisations a single place to benchmark it against other frontier models, run workload-specific tests and deploy it under enterprise security and governance controls. It had already been made available on Google's Gemini Enterprise Agent Platform via Model Garden, priced at $2 per million input tokens, $6 per million output tokens, and $0.50 per million tokens for cached input.

Separately, xAI has widened access to Grok Bot, its autonomous agent product, beyond its initial beta. The tool is now included with SuperGrok, SuperGrok Plus and SuperGrok Heavy subscriptions, as well as Cursor Pro, Cursor Pro+, Cursor Ultra and Cursor's Standard and Premium team plans. Grok Bot carries its own separate usage allowance, so tasks handed to it do not count against a user's existing Grok or Cursor limits. xAI said the bots operate with their own cloud computer, including browser and terminal access, and can be assigned ongoing tasks such as inbox management, customer support and website building. Enterprise customers can join a waitlist for larger rollouts.


Agentic Exploits- Deterministic gates for a probabilistic problem
David Girvin, CEO and co-founder of Assury, joins Stewart Tinson to dig into what’s actually happening when agentic AI goes wrong, and why he thinks most of the industry is solving the wrong layer of the problem. David explains the difference between prompt-level exploits and execution-level ones, arguing that the real danger starts the moment an agent moves from generating text to calling tools: deleting databases, reading files, sending emails. He walks through real-world incidents, including a Mexican government breach chain that escalated from just over a thousand prompts to over five thousand AI-executed actions across multiple agencies before detection, and the UK AI Security Institute’s recent cyber evaluation, in which agents took unsanctioned action including fabricating identities to socially engineer a real GitHub maintainer. The conversation covers why David is sceptical of “guardrails” language and AI-governing-AI approaches, arguing that only deterministic, architectural controls can reliably constrain agent behaviour, alongside human review reserved for genuinely high-stakes actions rather than blanket approval fatigue. He breaks down credential starvation, session risk accumulation, and why classifier-based tools keep failing inconsistently on identical actions, pointing to a named frontier lab’s own zero trust paper as an example of the industry misjudging what actually works. Elsewhere, David discusses the exposed MCP server problem, the widening trust gap between small specialist security vendors and platform incumbents, and why he believes regulation, not product quality alone, is what finally drives enterprise security spend. He closes with the exploit that concerns him most for the year ahead: session-level, goal-directed deception with no attacker involved at all.
Share this post
The link has been copied!