arXiv paper: WorldCup Arena: Prospective, Leakage-Free Evaluation of Frontier LLMs on a Live Tournament
A new arXiv AI paper by Zhenran Wang, Zhonghan Bian, and Jinsong Li, and 1 more studies WorldCup Arena: Prospective, Leakage-Free Evaluation of Frontier LLMs on a Live Tournament.
Follow arXiv AI/ML to make it a durable For You signal.
Researchers evaluated six frontier LLMs—all using extended thinking and native web search—on a live 2026 World Cup forecasting task where no answer existed at prediction time, eliminating data leakage by design. Across 4,494 scored predictions, models averaged 63.9% accuracy on match outcomes, essentially matching the bookmaker’s favorite, which they mostly mirrored. They disproportionately avoided predicting draws and goals, clustered scoreline picks onto a single result, and performed accurately only on lopsided fixtures while failing on close matches where information was richest. A majority vote across models did not improve accuracy, and no model substantially outperformed peers. The paper releases all dossiers, fixtures, results, predictions, and scoring code as a new benchmark for prospective forecasting evaluation.