Microsoft has said its Discovery Engine, paired with a new reasoning system called CLIO, outperformed rival agentic AI harnesses on a demanding new benchmark for scientific and engineering tasks.
In a blog post published on 8 September, Aseem Datar, Corporate Vice President of Product Innovation at Microsoft Discovery & Quantum, said the combination scored 61.6% in health and medicine, 75.2% in physical sciences and 64.6% in life sciences on Agent's Last Exam, a test of long-running, tool-using professional work. CLIO, which stands for Cognitive Loop via In-Situ Optimization, works by running several reasoning attempts at once, each tackling the problem differently, before pooling what they find and settling on whichever line of reasoning is best backed by evidence. Rather than following a fixed process throughout, it is designed to change course mid-task, whether that means continuing down the same path, trying a different one, calling on a different underlying model, or pulling a human specialist into the process.
Microsoft said the approach has already contributed to the discovery of a novel organic redox flow battery, and pointed to further potential uses in chip design simulation, manufacturing and consumer goods formulation, materials and drug discovery, and lab automation.
Datar was careful to frame the technology as a complement to human researchers rather than a replacement, describing the tools as a way to help scientists explore more hypotheses and learn faster from evidence, while leaving validation to experts.
