Security

Abstract illustration of red dots and thin vertical lines rising from a perspective grid against a dark navy background.

The tests that grade AI may be getting it wrong

Q&A

Benchmarks – the standardized tests that rank AI models on safety, bias, and reasoning – drive markets and shape regulation. New Stanford research finds they often don’t measure what they claim to.