We’re not a huge research house. We don’t have a team of analysts, and we’re not Gartner. We’re small, and our approach is to ask questions of people who know what they’re talking about: the practitioners and vendors working on these problems.
That’s the thinking behind the 70 or so short solo AI-360 Research Notes videos coming soon. The plan is to start releasing them this month, with the first ten due to be available to watch from the middle of next week. Subjects include financial crime in the agentic era, claims fraud, synthetic identity, APP fraud, third-party AI risk, deepfake detection and incident response, and DORA.
I’ll be giving a non-expert overview of each subject, explaining why it matters and setting out the questions I think are worth asking. These videos are an invitation to practitioners and vendors to share their expertise and help develop the discussion through the webinars and interviews we’re planning for late November and December. If one of these subjects is your territory, I’d like to hear from you.
This week’s nine stories bring plenty of those questions into focus, from agents handling claims and financial-crime investigations to the controls around payments, identity and infrastructure.
What catches my attention in this selection is how close AI agents are getting to decisions that carry real consequences. Claims intake, financial-crime investigations and payments all feature. So do the mechanisms intended to control who an agent represents, what it can access and when somebody needs to approve its next move.
For banks, insurers and enterprise buyers, that makes the practical questions increasingly specific. Who owns the agent? What authority has it been given? Can you reconstruct what happened? These nine stories offer different answers, with plenty still to prove.
Anthropic is positioning Sonnet 5.5 between routine, high-volume work and its more capable Opus tier, claiming faster generation and lower token use for many tasks. It also says the model’s cyber capabilities are comparable to Opus 5, with selected high-risk requests routed to a less capable model. Enterprise evaluations need to cover that routing behaviour, permissions and auditability alongside the claimed performance gains.
Duck Creek says its Agentic First Notice of Loss service can capture, validate, enrich and route claims across several channels. The company claims improvements in data quality, claims cycles and costs, although the announcement does not independently demonstrate those outcomes. For insurers, the test is whether agents can handle incomplete information from distressed customers while producing records that people can examine and challenge.
The Bank of England is considering the financing of the AI build-out alongside the cyber and operational risks of increasingly autonomous systems. Its September Financial Policy Committee record highlights growing debt-market exposure and incidents involving frontier models in test environments. This deserves attention across financial institutions: investment exposure, infrastructure dependence and operational resilience increasingly belong in the same conversation.
Mastercard’s expanded Agent Pay services aim to give banks and merchants more context about AI-initiated transactions. An initial probability score, entering testing in the United States, indicates whether an agent likely initiated a payment. The wider challenge is establishing delegated intent: which agent acted, what the customer authorised and whether the transaction stayed within that authority. Performance claims still need evidence from deployment.
OpenAI positions GPT-6.1 Sol as faster and more affordable, while classifying its cyber capabilities at the company’s Critical threshold. It says the model inherits Astra’s safeguard stack. For enterprise buyers, the implication is straightforward: a model’s price tells you little about the controls its capabilities require. Company-run safety evaluations need to be considered alongside testing of the intended deployment.
RSA’s Agent ID platform treats agents as identities with responsible owners, permissions and lifecycle controls. RSA says its modules can discover agents and MCP servers, check tool calls and require authenticated human approval for designated actions. That puts familiar identity-governance disciplines around autonomous activity, although customers will need to validate how effectively those controls operate in their environments.
Oracle says Nexus Case Flow and Nexus Reach can help investigate alerts, gather evidence and prepare recommendations or investigative narratives. Case Flow centres on output that investigators can review. The attraction is additional capacity for stretched financial-crime teams. Banks will still need to establish that evidence is traceable, errors are detectable and human reviewers can meaningfully assess what the system produces.
Nvidia’s Open Agent Safety Platform combines the OpenShell runtime with Sentry, a reference design using BlueField hardware as an independent watchdog. The company says the architecture can enforce boundaries and quarantine agents when they breach policy. Its significance is the attempt to put enforcement outside the agent’s own software. Claims about containment, reliability and operational overhead require independent testing.
AWS separates agent security into three questions: which instructions can the system trust, which actions can it perform and which information can it retrieve? That gives security teams specific boundaries to test. A malicious instruction hidden in retrieved content should never acquire the authority of the user. AWS promotes AgentCore as an enforcement mechanism, while the underlying questions apply across platforms.