How good are frontier models at physics?
A paper led by Ali Ansari and many coauthors re-evaluates frontier language models on six physics benchmarks using expert re-grading of text-only, verifiable problems, and finds existing evaluations are flawed.