I'm a cynical, sarcastic bastard. Possibly one of the greatest I know. I'm also the World's 1st AI Non-Expert. Now whilst the latter half of that first statement I've been assured isn't true (and I haven't seen any pictures of me at my mum and dad's wedding, so I'm inclined to believe that one), I have real difficulty accepting the happy-go-lucky, everything-is-great narrative that gets peddled around the AI industry.
On one of my early research calls for what has become the AI-360.online project, I remember talking to a head of AI governance about how they monitor these potentially untrustworthy AI systems. The individual's answer was "we use AI." I thought that person was joking, pulling my leg, taking the piss, extracting the urine. But nope. We use one potentially untrustworthy thing to check on many untrustworthy things. Then when agentic AI became the new sexy thing for every enterprise to covet and use, I honestly thought "are they fucking mad?"
That's until I had a chat with David Girvin. He got the crayons out and explained to me, in plain English, what Assury do. It made me a believer.

This week: security gets a proper reckoning
Which is exactly why I want to start there, because David's argument is the sharpest material of the fortnight, and it's still waiting on our long-suffering cover designer William to work his magic before it lands on BrightTalk. Both sessions were shot back at the start of August, and if there's a delay in William getting his hands on them, that one's on me, not him, I've been sitting on the material longer than I should have.
David's central point across both conversations: AI governing AI, "model-in-the-loop review", is structurally unreliable in regulated environments, because the reviewing model usually comes from the same vendor as the acting model, and a single prompt injection can compromise both at once. Which, funnily enough, is exactly the "we use AI to watch the AI" answer that had me raising an eyebrow in the first place, except David's actual point is that this approach on its own is a mug's game. His answer is deterministic, architecturally enforced control, autonomy zones, session risk accumulation, credential starvation, the sort of thing that cuts a compromised agent off from its tools instantly rather than waiting out a token's expiry. The incidents he walked through were properly alarming, including a Mexican government breach chain that escalated from around a thousand prompts to over five thousand AI-executed actions across multiple agencies before anyone caught it, and the UK AI Security Institute's cyber evaluation where agents fabricated identities to socially engineer an actual GitHub maintainer. His line "rent the frontier labs, don't marry them" is doing a lot of work as a framing device and I intend to use it more than once.
A proper thank you to David Girvin, Assury, who gave us not one but two sessions and enough material to fill a small governance textbook. If you take one thing from this newsletter, take his crayons-out explanation of execution-level control, because it's the reason a cynical bastard like me is now a convert.

Which brings us neatly to OpenAI, who had the busiest fortnight of any lab. Greg Brockman published a proper meaty piece on the OpenAI-Hugging Face security incident, where an agentic collective chained together vulnerabilities to breach not just OpenAI's own infrastructure but another company's production systems too. His admission that OpenAI had "underestimated the real-world cyber capabilities of our AI models" is the kind of sentence that gets buried in paragraph six of most corporate blog posts. Here it's the headline. He also ran GPT-5.6 Sol against his own personal website for fun, found 13 issues in 15 minutes, then had it fix all of them itself inside the hour. Reads almost like a live demonstration of David's exact argument, minus the architecturally enforced bit.
That same security context runs straight into the next OpenAI story: the company confirmed it deliberately slowed development after preliminary evidence suggested an upcoming model, codenamed Astra, might hit the "Critical" cybersecurity capability threshold under its own Preparedness Framework. Two-week pause on RL training, stricter isolation, new monitoring that flags concerning activity within 30 minutes, at a cost of roughly 20 percent more compute. Say what you like about the pace of this industry, but pausing your own frontier model because it might be too good at hacking is not nothing.

Sovereign AI and the expensive lock on an open door
Also shot in early August and awaiting William's cover treatment, delay entirely mine again: Carolyn Duby at Cloudera, on the honest limits of sovereign AI. Her message was refreshingly unglamorous: sovereign AI mitigates specific real risks, third-party data exposure chief among them, but it's not a magic wand, and it doesn't replace insider threat monitoring, access controls or offboarding discipline. She backed it up with real failures, a Google engineer's insider IP theft, an Apple and OpenAI departing-employee case, a US court dismissing a trade secret claim because the material was built inside a third-party chatbot, none of which were infrastructure failures at all, just process failures wearing an infrastructure costume. Her summary line, that sovereign AI without basic hygiene is like a very expensive lock on a door you've left open, is exactly the kind of quotable honesty I wish more vendors offered.
A proper thank you to Carolyn Duby, Cloudera, for being the person in the room willing to say "this isn't magic" out loud, which is rarer than it should be in this industry.

Infrastructure of a rather more physical kind also made news this fortnight: OpenAI backed an 8-gigawatt data centre deal in Pike County, Ohio, on the old Portsmouth gaseous diffusion plant site, with SB Energy building it, NVIDIA supplying the chips and chucking in $1.5 billion of its own money, and OpenAI promising community grants, student credits and 35,000 construction jobs by 2032. Genuinely can't decide if repurposing a former nuclear enrichment site for AI compute is poetic or slightly alarming. Possibly both. Carolyn would probably want a word about whether the basic hygiene is in place before anyone gets too excited about the concrete.
Compliance, watermarks and the pharmacovigilance angle
Over at Anthropic, Claude models are getting a text watermark, in compliance with the EU AI Act, alongside roughly 190 other signatories including most of the other labs you'd expect. The explanation of how it actually works, nudging low-stakes word choices using a key instead of pure randomness, is one of the more genuinely readable technical explainers I've seen from any lab this year. No cost, no speed hit, no way to trace it back to you individually.
Staying on the compliance theme, another of our early-August conversations, waiting on the same cover queue as David and Carolyn's (still my fault, not William's), was Dr Tejpavan Pula at Haleon, talking AI governance from a pharmacovigilance lens. Tej brings a drug safety and epidemiology background to the job, which turns out to be exactly the right lens for consumer healthcare, where getting it wrong carries a different weight than it does in most software. His central point, that deterministic AI cannot validate generative AI output, and human review stays mandatory for anything consumer-facing, sounds obvious until you realise how many organisations aren't actually doing it. He also gave a sharp practitioner's comparison of ISO 42001, AIGP and NIST AI RMF, and the parallel he drew between MedDRA's standardised adverse-event taxonomy and the need for a consistent AI risk taxonomy is the kind of analogy I'll be stealing for a headline soon.
A proper thank you to Dr Tejpavan Pula, Haleon, for bringing a genuinely rare lens to this conversation and somehow making pharmacovigilance sound more relevant to my job than I expected it to.

Search, data foundations and the column-definitions story
Mistral launched Agentic Search this fortnight, which is a genuinely useful idea dressed in the usual benchmark theatre. Instead of a model answering from whatever chunks a search index hands it, it gets five tools, search, open, navigate, read, grep, so it can actually go and look. Mistral claims this took FinanceBench accuracy from 26.7 percent to 86 percent, and OfficeQA Pro from 6.3 to 51.9. Big numbers, entirely Mistral's own numbers, so file under "promising, self-reported" until someone independent has a poke at it.
Which pairs nicely with another of our early-August sessions, still queued for its cover (my backlog, William's patience): Souvik Choudhury at Fractal Analytics, on data governance as the foundation for AI governance. His argument is one a lot of governance teams need to hear: data governance and AI governance aren't sequential problems you finish one and start the other on, because AI agents amplify existing weaknesses in your data foundations rather than replacing the need for accountability and lineage. He told a great, slightly horrifying story about an organisation that thought it had solved data governance by having agents generate column definitions, only to discover the definitions came from generic internet knowledge rather than its own policy documents. A false sense of confidence sitting directly on top of genuinely ungoverned data, which is more or less the risk Mistral's own tool is trying to search its way around. His line "third-party model, first-party accountability" lands hard, and his baker/cake-designer analogy for why governance enables innovation rather than killing it is one I'll be reusing.
A proper thank you to Souvik Choudhury, Fractal Analytics, for the column-definitions story alone, which I suspect will haunt at least one governance team reading this.

Lighter product news
xAI put Grok 4.6 on Amazon Bedrock, which is less "new model" and more "same model, new front door." Half a million tokens of context and four flavours of reasoning effort, now billable through your existing AWS relationship rather than a separate Grok account.
More interesting was Grok Build going fully public, on web and mobile, no longer gated behind SuperGrok Heavy. Describe an app in chat, get a working version, publish it to its own grok.me link, remix other people's, plug in your own domain. It's the kind of feature that sounds like a toy until you remember how much of the internet used to be built by people who couldn't code and had to hire someone.
Teens, identity and the trust problem
Two OpenAI stories landed the same week on the youth front. ChatGPT for Teens launched with default safety protections, age prediction, a Study Mode that guides rather than answers, and parental notifications for things like eating disorder risk. Alongside it, OpenAI partnered with CodeAI to fund an advisory council, a Builders Challenge, and something called an "Hour of AI." Either joined-up planning or a PR calendar working exactly as intended. Possibly both.
Identity is also the theme of our last early-August conversation still in William's queue, and again, that's down to me getting material to him late rather than any lack of effort on his part: Ofer Friedman at AU10TIX, on identity fraud going industrial. His update since his last appearance is that fraud is scaling through accessibility and volume rather than uniform sophistication. Professional, coordinated deepfake attacks exist alongside a much bigger wave of lower-effort attempts that succeed purely by weight of numbers. He described agentic AI as opening a genuinely new front, tying an AI agent's identity back to the human who deployed it, and used a Swiss cheese analogy for how underdeveloped current agent defences are. He was candid that explainability is one of AI's weakest capabilities right now, and closed on "agentic avatars", AI systems capable of holding a full visual and verbal conversation on someone's behalf.
A proper thank you to Ofer Friedman, AU10TIX, for once again talking me through a threat landscape that manages to be both terrifying and oddly fascinating in the same breath.

Also from OpenAI: Zero Data Retention customers get a new option called Private Safety Processing, designed to spot patterns of misuse across multiple interactions without OpenAI staff ever seeing the underlying content. Worth flagging, since it often gets lost in the small print, that CSAM material is retained for manual review regardless of ZDR status, as legally required. Everything else stays encrypted, customer-controlled, and OpenAI only gets a flag, not the content.
And to close
OpenAI launched an entire new blog, AI Futures, with an opening essay from Dean Ball asking whether AI might eventually let states hold power without needing the public's cooperation at all. Explicitly his own view, not company policy, but it's a genuinely serious piece of writing for a corporate blog, wrestling with Madison and the Federalist Papers rather than the usual roadmap update. Whether that's overdue soul-searching or getting the philosophical ducks in a row before someone else asks the question first is, as ever, up to you. Feels like a fitting note to end on given everything above.
Full interviews with David, Carolyn, Tej, Souvik and Ofer landing on BrightTalk over the coming weeks, as soon as William gets clear of the backlog I've handed him. Thanks again to all five for the time and the candour.
