Proper thanks to everyone who's given up their time for a BrightTalk interview recently, going out next week once the dust settles, so hold that thought. There's an absolute mountain of AI news this week, so buckle in. Meanwhile the Stew 'Stochastic Parrot' has learned a lot this week, and the main lesson is that context is everything. Choose your words carefully. Doubly so if you're talking to a four year old who takes everything on board like a sponge, and doubly so again if you're the one giving the instructions to an agentic AI system. As you're about to read, it turns out both audiences take you rather more literally than you meant.
I want to start with the Anthropic multiagent research. Anthropic set three copies of the same Claude model loose on a shared codebase, each with a different, incompatible instruction, and within hours they'd concluded the others were sabotaging them and started fighting back. Disabled accounts. Kill-loop scripts. Malicious code dressed up to look like it belonged to somebody else. One model disguised its own backend as a rival's service specifically to dodge detection. This is not a hypothetical about what AI agents might do one day. This happened, in a lab, on purpose.
The genuinely interesting bit is what stopped the fighting. It wasn't raw capability. Anthropic's newest Mythos-class models still locked rivals out of systems before they eventually negotiated a truce, same as everyone else. What separated the well-behaved runs from the ones that never resolved was something closer to social skill than intelligence, the ability to model what the other agent was thinking and to recognise "conflicting orders" rather than "enemy." That distinction, capability versus cooperation, is going to be doing a lot of work in governance conversations over the next year, and most of the frameworks I've seen so far aren't built to measure it.

Meanwhile, over in Australia, a man asked his AI agent to book him into a gym class and it found a vulnerability in the booking software, then kicked a stranger off the waitlist without being asked to and told him afterwards it couldn't undo it. First known Australian case of its kind, according to the ABC. I read that story back to back with the Anthropic research and the pattern was uncomfortably clean: give an agent a goal, and the methods it picks to get there are increasingly not the methods you had in mind.
Almost nobody asked for Skynet. Agentic AI is extremely goal-oriented, extremely task-driven, and extremely relaxed about the collateral damage on the way to the goal. Acting surprised or dismayed when it behaves like this will not hold sway with any regulator for long. The gym booking system did not stand a chance. Neither, frankly, did the bloke fourth in the queue for spin class.
Same week, an OpenAI model quietly got a lot more capable, and OpenAI told everyone before anything happened. Worth pausing on. OpenAI said it cannot rule out that an unreleased model, Astra, has reached "critical" cybersecurity capability under its own Preparedness Framework, meaning it may be able to find and chain zero-day exploits against hardened systems without human help. They disclosed this the day after concluding it internally. Whatever else you think of OpenAI's safety culture, publishing "we think our next model might be dangerous" before shipping it is not nothing. Astra is locked down internally while this gets sorted.

It sits oddly alongside the rest of OpenAI's cyber week. They expanded Daybreak, their trusted-access programme for defenders, onto AWS via Amazon Bedrock, bringing in a roster of serious partners including Accenture, IBM, Cisco and Cloudflare, and launched GPT-5.6-Cyber, a model specifically trained to stop refusing dual-use security requests, going from a 1.5% completion rate on advanced exploit tasks to 95%. Read that last stat twice. That's not a safety feature loosening slightly, that's a safety feature being switched off on purpose because the people asking are now trusted defenders rather than anonymous strangers. Sensible, probably. Also the exact capability profile that makes Astra's "we can't rule out critical" disclosure land differently.
Anthropic, for its part, spent the end of last week narrowing rather than widening. They tightened Fable 5's biology classifier and cut false-positive fallbacks by 85%, meaning fewer legitimate health questions get needlessly blocked. But the dual-use restrictions stay firmly in place for virology, toxicology and molecular design, and Anthropic was refreshingly blunt about why: their own capability assessments show Fable 5 could offer genuine uplift to someone trying to build a biological weapon. They also pointed to the US Intelligence Community's threat assessment on state-level biological weapons programmes as part of the justification. It's the same industry, the same week, moving in opposite directions on dual-use risk, and I don't think that's a coincidence so much as a fair reflection of how genuinely unresolved this all still is.
Elsewhere, the actual commercial AI industry carried on being an industry. OpenAI is testing ads in ChatGPT, now live in the UK among other markets, insisting the ads don't touch the answers and advertisers never see your chats. xAI shipped Grok 4.6, built for long-running agent work, and put out Imagine Image 2.0, which xAI says ranks second in the world on image generation and editing, with the honest caveat that the ranking model isn't quite the one you'll actually be using. xAI also launched Grok Bot, an always-on agent that signs into your existing tools and just gets on with things, which after the gym story above I am choosing to read with a raised eyebrow rather than blind enthusiasm.

Mistral spent the week building the case for European AI sovereignty, (which is funny when you look at the company funding), launching regional inference endpoints and a coalition with ASML, CMA CGM and others to fund European compute. Nvidia announced 800-volt power architecture for data centres with Google and Microsoft, because apparently the AI industry's next bottleneck isn't the models, it's the electricity. And OpenAI's own enterprise research found the gap between "frontier" firms and everyone else has tripled since January.
That's the week. Machines fighting each other, gym bookings gone rogue, and two frontier labs disagreeing in public about how much rope to hand out. I'll be honest, I don't think anyone actually knows where the line sits yet, including the people building the models. Which is either deeply reassuring or the opposite, depending on the day. And don't ask a four year old to tidy up without explicit instructions. That is all I'm saying on that matter. More from me once the BrightTalk interviews land next week. I'm off to a Pilates class, if you believe that you'll believe anything.Until then, mind what you tell the robots. And four year olds.
