Anthropic has released Claude Sonnet 5.5 across its own platform, Amazon Web Services, Google Cloud and Microsoft Azure. The company positions the model between high-volume routine work and its more capable Opus tier, claiming generation speeds more than 30 per cent faster than Sonnet 5 and lower token use for many tasks.

Anthropic’s benchmark results should be treated as vendor-reported evidence rather than universal performance. It says Sonnet 5.5 scored 70.6 per cent on Terminal-Bench 4.0 and improved materially on coding and long-horizon knowledge-work tests. More significant for regulated buyers is the safety note: Anthropic says the model’s cyber capabilities are comparable to Opus 5, so selected high-risk requests will fall back to a less capable model. It is also the first Sonnet release with classifiers intended to resist industrial-scale reasoning extraction.

This is the commercial shape of the frontier-model market in 2026: capability gains are being packaged with cost controls, effort settings and increasingly explicit security tiers. Banks, insurers and other large enterprises should evaluate the entire operating envelope — model routing, fallback behaviour, data retention, tool permissions and auditability — rather than comparing headline benchmark scores. Sonnet 5.5 may improve the economics of everyday agent workflows, but stronger capability also enlarges the control surface.


Execution Level Governance- What audit-ready agent governance actually looks like
David Girvin, founder and CEO of Assury argues that model-in-the-loop review, AI governing AI, is fundamentally unreliable for regulated environments: even the best-performing models miss a meaningful share of violations, the reviewing model is typically provided by the same vendor being reviewed, and prompt injection or context poisoning can compromise both the acting agent and its supposed overseer simultaneously. He makes the case for deterministic, architecturally enforced controls instead, walking through Assury’s approach of autonomy zones, session risk accumulation, and credential starvation, which lets a compromised agent be cut off from its tools instantly rather than relying on time-boxed access. The conversation touches on why David is sceptical of just-in-time credentialing as a solution for agent security more broadly, since agent sessions don’t run on predictable human timescales, along with the current gap between how identity and security vendors are pitching agent protection and what he sees happening at the execution layer in practice. He also discusses the compliance and audit implications of probabilistic decision-making, arguing that regulated industries will increasingly need tamper-evident, hash-chained audit trails that can withstand scrutiny from auditors and regulators who are only beginning to understand agentic risk, and reflects on a named frontier lab’s own published framework as an example of the gap between research and practitioner reality. Elsewhere, David reflects candidly on building a bootstrapped security company in an increasingly crowded market, why he turned down aggressive VC funding to stay in control of the product, and what a credible third-party assessment of his own gateway would need to look like given that Assury sits directly in the execution path for every customer’s agents.
Share this post
The link has been copied!