Diversity is the Strength of the AI Crowd

June 29, 2026 Β· Grace Period Β· πŸ› the ICML 2026 Workshop on Forecasting as a New Frontier of Intelligence

⏳ Grace Period
This paper is less than 90 days old. We give authors time to release their code before passing judgment.
Authors Matthew Aitchison, Scott Jeen, Toby Shevlane, Ben Day arXiv ID 2606.29661 Category cs.AI: Artificial Intelligence Citations 0 Venue the ICML 2026 Workshop on Forecasting as a New Frontier of Intelligence
Abstract
Top AI forecasting systems are approaching superforecaster-level accuracy on future world events, but still rely primarily on off-the-shelf LLMs combined with forecasting-specific context gathering and scaffolding. We study how to improve this recipe through ensembling: given a fixed number of samples, which off-the-shelf model forecasts should be combined to maximize accuracy? On binary questions from the Metaculus AI Benchmark, we find that individual accuracy is not enough: many frontier LLMs make highly correlated predictions, limiting the value of additional forecasts from the same or similar models. Instead, the strongest ensembles combine accurate but diverse forecasters, with models such as \model{Grok 4} contributing disproportionately because their predictions are less correlated with other frontier LLMs. These results suggest that the strength of the AI crowd comes not from sampling more forecasts indiscriminately, but from combining forecasts across models with complementary errors, motivating forecasting systems that explicitly optimize for both model quality and diversity.
Community shame:
Not yet rated
Community Contributions

Found the code? Know the venue? Think something is wrong? Let us know!

πŸ“œ Similar Papers

In the same crypt β€” Artificial Intelligence