The causality test that humbled six AI agents
A brand new benchmark out of the University of Michigan simply handed six frontier AI agents a causal reasoning examination. The scores expose a niche that uncooked functionality metrics are inclined to paper over. The benchmark is named CausalDS, revealed on arXiv on July 9, 2026, by Andrej Leban and Yuekai Sun. It exams one…
