What we measure before anyone relies on it: which model wins which prompt, where a model fails, and what it cannot read.