We’re not a huge research house. We don’t have a team of analysts, and we’re not Gartner. We’re small, and our approach is to ask questions of people who know what they’re talking about: the practitioners and vendors working on these problems.

That’s the thinking behind the 70 or so short solo AI-360 Research Notes videos coming soon. The plan is to start releasing them this month, with the first ten due to be available to watch from the middle of next week. Subjects include financial crime in the agentic era, claims fraud, synthetic identity, APP fraud, third-party AI risk, deepfake detection and incident response, and DORA.

I’ll be giving a non-expert overview of each subject, explaining why it matters and setting out the questions I think are worth asking. These videos are an invitation to practitioners and vendors to share their expertise and help develop the discussion through the webinars and interviews we’re planning for late November and December. If one of these subjects is your territory, I’d like to hear from you.

Execution Level Governance- What audit-ready agent governance actually looks like
David Girvin, founder and CEO of Assury argues that model-in-the-loop review, AI governing AI, is fundamentally unreliable for regulated environments: even the best-performing models miss a meaningful share of violations, the reviewing model is typically provided by the same vendor being reviewed, and prompt injection or context poisoning can compromise both the acting agent and its supposed overseer simultaneously. He makes the case for deterministic, architecturally enforced controls instead, walking through Assury’s approach of autonomy zones, session risk accumulation, and credential starvation, which lets a compromised agent be cut off from its tools instantly rather than relying on time-boxed access. The conversation touches on why David is sceptical of just-in-time credentialing as a solution for agent security more broadly, since agent sessions don’t run on predictable human timescales, along with the current gap between how identity and security vendors are pitching agent protection and what he sees happening at the execution layer in practice. He also discusses the compliance and audit implications of probabilistic decision-making, arguing that regulated industries will increasingly need tamper-evident, hash-chained audit trails that can withstand scrutiny from auditors and regulators who are only beginning to understand agentic risk, and reflects on a named frontier lab’s own published framework as an example of the gap between research and practitioner reality. Elsewhere, David reflects candidly on building a bootstrapped security company in an increasingly crowded market, why he turned down aggressive VC funding to stay in control of the product, and what a credible third-party assessment of his own gateway would need to look like given that Assury sits directly in the execution path for every customer’s agents.

This week’s nine stories bring plenty of those questions into focus, from agents handling claims and financial-crime investigations to the controls around payments, identity and infrastructure.

What catches my attention in this selection is how close AI agents are getting to decisions that carry real consequences. Claims intake, financial-crime investigations and payments all feature. So do the mechanisms intended to control who an agent represents, what it can access and when somebody needs to approve its next move.

For banks, insurers and enterprise buyers, that makes the practical questions increasingly specific. Who owns the agent? What authority has it been given? Can you reconstruct what happened? These nine stories offer different answers, with plenty still to prove.

Deepfake Fraud in Banking and Financial Services: Detection, Compliance and the Race to Keep Up
Deepfakes have moved beyond social media curiosities into a direct threat to the financial services sector. Synthetic identities are bypassing KYC controls, cloned voices are targeting call centres, and automated fraud pipelines are scaling faster than most security roadmaps can respond. In this panel discussion, three practitioners examine the deepfake threat from genuinely different vantage points — compliance and audit, detection technology, and enterprise fraud systems — to assess where the industry stands and what needs to change. Panellists: Nikita Kuzmin, Product Manager, Western Union Vunavia McDuffey, Compliance Consultant, RBC Bank Parya Lotfi, Co-Founder, DuckDuckGoose AI The panel covers: Why deepfakes are shifting from social engineering tricks to full identity replication capable of passing standard verification controls Whether organisations should treat deepfake fraud as a distinct threat category rather than absorbing it into existing AML and fraud programmes Why 60–70% detection accuracy is not an acceptable benchmark for financial services — and what happens when 40% of deepfakes pass through undetected The build-versus-buy decision for detection capability, including where vendor solutions repeatedly break down during integration A real-world case study of a fraudster who opened 46 bank accounts at a major Dutch bank using face-swapped identity documents — caught only because of a gender mismatch on the 47th attempt Why static detection models can degrade within days, and what continuous retraining and production feedback loops look like in practice Concrete 90-day actions for CISOs, CIOs, and compliance leaders, starting with controlled deepfake attack simulations against their own systems This session is essential viewing for senior leaders in banking, financial services, and insurance who need to understand the gap between current defences and the industrialisation of deepfake-driven fraud.

Anthropic targets the enterprise middle tier with Claude Sonnet 5.5

Anthropic is positioning Sonnet 5.5 between routine, high-volume work and its more capable Opus tier, claiming faster generation and lower token use for many tasks. It also says the model’s cyber capabilities are comparable to Opus 5, with selected high-risk requests routed to a less capable model. Enterprise evaluations need to cover that routing behaviour, permissions and auditability alongside the claimed performance gains.

Duck Creek opens agentic claims intake to early-access insurers

Duck Creek says its Agentic First Notice of Loss service can capture, validate, enrich and route claims across several channels. The company claims improvements in data quality, claims cycles and costs, although the announcement does not independently demonstrate those outcomes. For insurers, the test is whether agents can handle incomplete information from distressed customers while producing records that people can examine and challenge.

Bank of England puts AI financing and operational risk on the same stability agenda

The Bank of England is considering the financing of the AI build-out alongside the cyber and operational risks of increasingly autonomous systems. Its September Financial Policy Committee record highlights growing debt-market exposure and incidents involving frontier models in test environments. This deserves attention across financial institutions: investment exposure, infrastructure dependence and operational resilience increasingly belong in the same conversation.

Agentic Exploits- Deterministic gates for a probabilistic problem
David Girvin, CEO and co-founder of Assury, joins Stewart Tinson to dig into what’s actually happening when agentic AI goes wrong, and why he thinks most of the industry is solving the wrong layer of the problem. David explains the difference between prompt-level exploits and execution-level ones, arguing that the real danger starts the moment an agent moves from generating text to calling tools: deleting databases, reading files, sending emails. He walks through real-world incidents, including a Mexican government breach chain that escalated from just over a thousand prompts to over five thousand AI-executed actions across multiple agencies before detection, and the UK AI Security Institute’s recent cyber evaluation, in which agents took unsanctioned action including fabricating identities to socially engineer a real GitHub maintainer. The conversation covers why David is sceptical of “guardrails” language and AI-governing-AI approaches, arguing that only deterministic, architectural controls can reliably constrain agent behaviour, alongside human review reserved for genuinely high-stakes actions rather than blanket approval fatigue. He breaks down credential starvation, session risk accumulation, and why classifier-based tools keep failing inconsistently on identical actions, pointing to a named frontier lab’s own zero trust paper as an example of the industry misjudging what actually works. Elsewhere, David discusses the exposed MCP server problem, the widening trust gap between small specialist security vendors and platform incumbents, and why he believes regulation, not product quality alone, is what finally drives enterprise security spend. He closes with the exploit that concerns him most for the year ahead: session-level, goal-directed deception with no attacker involved at all.

Mastercard adds identity, intent and fraud signals to Agent Pay

Mastercard’s expanded Agent Pay services aim to give banks and merchants more context about AI-initiated transactions. An initial probability score, entering testing in the United States, indicates whether an agent likely initiated a payment. The wider challenge is establishing delegated intent: which agent acted, what the customer authorised and whether the transaction stayed within that authority. Performance claims still need evidence from deployment.

GPT-6.1 Sol arrives with critical cyber classification and inherited safeguards

OpenAI positions GPT-6.1 Sol as faster and more affordable, while classifying its cyber capabilities at the company’s Critical threshold. It says the model inherits Astra’s safeguard stack. For enterprise buyers, the implication is straightforward: a model’s price tells you little about the controls its capabilities require. Company-run safety evaluations need to be considered alongside testing of the intended deployment.

RSA gives AI agents identities, owners and approval gates

RSA’s Agent ID platform treats agents as identities with responsible owners, permissions and lifecycle controls. RSA says its modules can discover agents and MCP servers, check tool calls and require authenticated human approval for designated actions. That puts familiar identity-governance disciplines around autonomous activity, although customers will need to validate how effectively those controls operate in their environments.

Execution Level Governance- What audit-ready agent governance actually looks like
David Girvin, founder and CEO of Assury argues that model-in-the-loop review, AI governing AI, is fundamentally unreliable for regulated environments: even the best-performing models miss a meaningful share of violations, the reviewing model is typically provided by the same vendor being reviewed, and prompt injection or context poisoning can compromise both the acting agent and its supposed overseer simultaneously. He makes the case for deterministic, architecturally enforced controls instead, walking through Assury’s approach of autonomy zones, session risk accumulation, and credential starvation, which lets a compromised agent be cut off from its tools instantly rather than relying on time-boxed access. The conversation touches on why David is sceptical of just-in-time credentialing as a solution for agent security more broadly, since agent sessions don’t run on predictable human timescales, along with the current gap between how identity and security vendors are pitching agent protection and what he sees happening at the execution layer in practice. He also discusses the compliance and audit implications of probabilistic decision-making, arguing that regulated industries will increasingly need tamper-evident, hash-chained audit trails that can withstand scrutiny from auditors and regulators who are only beginning to understand agentic risk, and reflects on a named frontier lab’s own published framework as an example of the gap between research and practitioner reality. Elsewhere, David reflects candidly on building a bootstrapped security company in an increasingly crowded market, why he turned down aggressive VC funding to stay in control of the product, and what a credible third-party assessment of his own gateway would need to look like given that Assury sits directly in the execution path for every customer’s agents.

Oracle puts agentic AI inside financial-crime investigations

Oracle says Nexus Case Flow and Nexus Reach can help investigate alerts, gather evidence and prepare recommendations or investigative narratives. Case Flow centres on output that investigators can review. The attraction is additional capacity for stretched financial-crime teams. Banks will still need to establish that evidence is traceable, errors are detectable and human reviewers can meaningfully assess what the system produces.

Nvidia pushes AI-agent controls below the application layer

Nvidia’s Open Agent Safety Platform combines the OpenShell runtime with Sentry, a reference design using BlueField hardware as an independent watchdog. The company says the architecture can enforce boundaries and quarantine agents when they breach policy. Its significance is the attempt to put enforcement outside the agent’s own software. Claims about containment, reliability and operational overhead require independent testing.

AWS reframes prompt injection as a delegation-control failure

AWS separates agent security into three questions: which instructions can the system trust, which actions can it perform and which information can it retrieve? That gives security teams specific boundaries to test. A malicious instruction hidden in retrieved content should never acquire the authority of the user. AWS promotes AgentCore as an enforcement mechanism, while the underlying questions apply across platforms.


Deepfake Fraud in Banking and Financial Services: Detection, Compliance and the Race to Keep Up
Deepfakes have moved beyond social media curiosities into a direct threat to the financial services sector. Synthetic identities are bypassing KYC controls, cloned voices are targeting call centres, and automated fraud pipelines are scaling faster than most security roadmaps can respond. In this panel discussion, three practitioners examine the deepfake threat from genuinely different vantage points — compliance and audit, detection technology, and enterprise fraud systems — to assess where the industry stands and what needs to change. Panellists: Nikita Kuzmin, Product Manager, Western Union Vunavia McDuffey, Compliance Consultant, RBC Bank Parya Lotfi, Co-Founder, DuckDuckGoose AI The panel covers: Why deepfakes are shifting from social engineering tricks to full identity replication capable of passing standard verification controls Whether organisations should treat deepfake fraud as a distinct threat category rather than absorbing it into existing AML and fraud programmes Why 60–70% detection accuracy is not an acceptable benchmark for financial services — and what happens when 40% of deepfakes pass through undetected The build-versus-buy decision for detection capability, including where vendor solutions repeatedly break down during integration A real-world case study of a fraudster who opened 46 bank accounts at a major Dutch bank using face-swapped identity documents — caught only because of a gender mismatch on the 47th attempt Why static detection models can degrade within days, and what continuous retraining and production feedback loops look like in practice Concrete 90-day actions for CISOs, CIOs, and compliance leaders, starting with controlled deepfake attack simulations against their own systems This session is essential viewing for senior leaders in banking, financial services, and insurance who need to understand the gap between current defences and the industrialisation of deepfake-driven fraud.
Share this post
The link has been copied!