When LLM judges agree, should we believe them?
The discussion proposes dependence-aware label aggregation using Ising models to model correlations among LLM judge panels, distinguishing independent evidence from shared mistakes; this improves accuracy by 9–14% in tests and recommends reporting confidence adjusted for judge correlation.