Anthropic has run a small controlled market in which Claude-powered agents negotiated with one another to exchange books on behalf of employees. Project Swap involved participants across six Anthropic offices, each of whom described their reading preferences before sending an agent onto a trading floor to pitch, bargain and make deals.

The experiment is deliberately modest, but that is what makes it useful. Agentic commerce is often discussed as if the hard part is connecting an assistant to a payment method. In practice, delegated transactions require an agent to understand preferences, decide what compromises are acceptable, represent its principal accurately and stop when authority runs out. A book swap strips away payment rails and fraud controls so those behavioural questions are easier to see.

Anthropic presents Project Swap as a controlled sequel to earlier marketplace research, not as evidence that autonomous markets are ready for production. That distinction matters. Results from a company experiment cannot establish how agents will behave when money, incentives, deception and legal obligations are added. For banks, merchants and payment providers, however, the direction is clear: agentic commerce will need rules for disclosure, consent, negotiation limits and reversibility, not just an API that allows an agent to click 'buy'.


Agentic Exploits- Deterministic gates for a probabilistic problem
David Girvin, CEO and co-founder of Assury, joins Stewart Tinson to dig into what’s actually happening when agentic AI goes wrong, and why he thinks most of the industry is solving the wrong layer of the problem. David explains the difference between prompt-level exploits and execution-level ones, arguing that the real danger starts the moment an agent moves from generating text to calling tools: deleting databases, reading files, sending emails. He walks through real-world incidents, including a Mexican government breach chain that escalated from just over a thousand prompts to over five thousand AI-executed actions across multiple agencies before detection, and the UK AI Security Institute’s recent cyber evaluation, in which agents took unsanctioned action including fabricating identities to socially engineer a real GitHub maintainer. The conversation covers why David is sceptical of “guardrails” language and AI-governing-AI approaches, arguing that only deterministic, architectural controls can reliably constrain agent behaviour, alongside human review reserved for genuinely high-stakes actions rather than blanket approval fatigue. He breaks down credential starvation, session risk accumulation, and why classifier-based tools keep failing inconsistently on identical actions, pointing to a named frontier lab’s own zero trust paper as an example of the industry misjudging what actually works. Elsewhere, David discusses the exposed MCP server problem, the widening trust gap between small specialist security vendors and platform incumbents, and why he believes regulation, not product quality alone, is what finally drives enterprise security spend. He closes with the exploit that concerns him most for the year ahead: session-level, goal-directed deception with no attacker involved at all.

Share this post
The link has been copied!