Microsoft has said its Discovery Engine, paired with a new reasoning system called CLIO, outperformed rival agentic AI harnesses on a demanding new benchmark for scientific and engineering tasks.

In a blog post published on 8 September, Aseem Datar, Corporate Vice President of Product Innovation at Microsoft Discovery & Quantum, said the combination scored 61.6% in health and medicine, 75.2% in physical sciences and 64.6% in life sciences on Agent's Last Exam, a test of long-running, tool-using professional work. CLIO, which stands for Cognitive Loop via In-Situ Optimization, works by running several reasoning attempts at once, each tackling the problem differently, before pooling what they find and settling on whichever line of reasoning is best backed by evidence. Rather than following a fixed process throughout, it is designed to change course mid-task, whether that means continuing down the same path, trying a different one, calling on a different underlying model, or pulling a human specialist into the process.

Microsoft said the approach has already contributed to the discovery of a novel organic redox flow battery, and pointed to further potential uses in chip design simulation, manufacturing and consumer goods formulation, materials and drug discovery, and lab automation.

Datar was careful to frame the technology as a complement to human researchers rather than a replacement, describing the tools as a way to help scientists explore more hypotheses and learn faster from evidence, while leaving validation to experts.


Swiss Cheese Defences- Identity fraud goes industrial and off-the-shelf
Ofer distinguishes between two current attack patterns: highly sophisticated, professionally coordinated deepfake and injection attacks designed to beat detection outright, and a much larger volume of lower-effort, high-scale attempts that rely on bombarding systems rather than disguising themselves well. He argues the real story right now is industrialisation of scale rather than uniform improvement in quality, though both are accelerating in parallel. The conversation covers why agentic AI is opening a new front in identity fraud, particularly the unresolved problem of tying an AI agent’s identity back to the human who deployed it and the permissions it holds. Ofer is candid about the current state of agent defences, describing them as underdeveloped and easy to hijack or poison, comparing the current state of play to Swiss cheese. He also discusses cross-industry fraud detection, including how signals, rather than raw data, are now being shared across platforms including AU10TIX and Reality Defender to surface fraud rings invisible to any single organisation. The conversation also covers explainability as one of AI’s most underdeveloped capabilities, with Ofer arguing that flagging a session as fraudulent without a credible, defensible reason will increasingly run into regulatory and practical limits. Elsewhere, the discussion covers credential laundering and the manufactured construction of fake digital histories and footprints, and the shift toward digital ID wallets that Ofer believes will make physical document fraud increasingly rare. He closes on what he calls “agentic avatars,” AI systems capable of holding a full visual and verbal conversation on someone’s behalf, and why he expects identity verification to have to become genuinely immersive across every form of media as a result.
Share this post
The link has been copied!