The Measurement Gap: Why AI Evaluation Needs Psychometrics
Artificial Intelligence systems are being deployed at scale into consequential domains — healthcare decisions, hiring, credit, criminal justice — before the field has developed a principled way to evaluate them. The dominant evaluation method, benchmarking, was borrowed from software engineering and was designed for a fundamentally different class of system. It cannot detect the problems that matter most: subtle bias, construct mismeasurement, and evaluation instruments that produce misleading results precisely when accuracy is most critical.
