Research

AI Forecasters Are Closing the Gap With Human Experts

A dynamic benchmark shows large language models improving steadily at predicting future events, while central banks test transformer-based models for macroeconomic forecasting.

Editorial·1 Sep 2026
AI Forecasters Are Closing the Gap With Human Experts

Large language models are getting measurably better at predicting future events, but they still cannot match the judgment of trained human forecasters on complex questions. That is the central finding from ForecastBench, a dynamic benchmark that evaluates AI forecasting accuracy on 1,000 automatically generated questions about real-world events, refreshed every two weeks to prevent models from simply memorizing answers from training data.

The benchmark, introduced in September 2024 by researchers including Ezra Karger, Houtan Bastani, and Philip E. Tetlock under the Forecasting Research Institute, represents the most systematic attempt yet to answer a question with profound implications for business strategy, policy design, and risk assessment: can machines reliably tell us what happens next? The institute, funded by Open Philanthropy for at least three years, has turned ForecastBench into a public resource with live leaderboards, allowing researchers and practitioners to track progress across different models and question categories.

The short answer, according to the data, is nuanced. In a 200-question test subset, expert human forecasters outperformed the top-performing LLM with statistical significance — a p-value below 0.001. Yet the same research shows that LLMs have improved consistently over time and now forecast better than non-expert humans. A linear trend based on current rates of improvement projects that AI models could reach parity with elite “superforecasters” by late 2027. That trajectory, while speculative, is already reshaping how institutions think about prediction.

Why forecasting matters beyond the benchmark

Forecasting is not an academic exercise. It underpins decisions about capital allocation, supply chains, geopolitical risk, and monetary policy. If AI systems can generate reliable probabilistic estimates about future events — inflation trajectories, election outcomes, regulatory shifts — they become tools for scenario planning that go far beyond conventional data analytics. The stakes are particularly high in economics, where a single missed inflection point can translate into billions in losses or misguided policy interventions.

The Bank for International Settlements (BIS) has already demonstrated this in the macroeconomic domain. In March 2026, BIS researchers published BISTRO, a transformer-based model designed specifically for macroeconomic time series forecasting. Unlike traditional econometric models that assumed inflation would revert to its historical mean after the 2021 surge, BISTRO correctly anticipated that elevated inflation would persist. That single result, according to the BIS, showed that LLM-style architectures can capture complex, non-linear patterns in economic data that conventional models miss. The model was trained on a broad set of macroeconomic time series, including inflation, output, and labor market indicators across multiple countries.

Key contributors to BISTRO include Batuhan Koyuncu, Byeungchun Kwon, and Hyun Song Shin. Their work signals a shift: central banks and financial institutions are no longer treating LLMs solely as language tools. They are exploring whether the same architectures can serve as general-purpose engines for numerical prediction, with implications for how monetary policy is formulated and how financial stability risks are assessed.

How ForecastBench works and what it shows

ForecastBench addresses a persistent problem in AI evaluation: contamination. If a model has seen the answer to a question during training, its performance on that question says little about its ability to reason about genuinely new situations. By generating fresh questions every two weeks and updating the benchmark continuously, the Forecasting Research Institute ensures that models are tested on events that have not yet occurred when the questions are posed. The 1,000-question format covers a wide range of domains, from geopolitics to economics to technology, with clear criteria for what counts as a correct forecast and a defined timeframe for resolution.

The results so far reveal a consistent pattern. LLMs perform well on questions where relevant information is widely available and the underlying dynamics are relatively stable. They struggle more with questions that require integrating ambiguous signals, weighing conflicting expert opinions, or anticipating rare but consequential events. Expert human forecasters, particularly those trained in probabilistic reasoning and base-rate calibration, still hold an edge in those harder cases. The 200-question subset that produced the statistically significant gap between experts and the top LLM was designed to stress-test exactly these capabilities, including questions with sparse data and high uncertainty.

But the gap is narrowing. The linear improvement trend observed in the benchmark data is notable because it suggests that advances in model architecture, training data quality, and fine-tuning techniques are translating directly into better forecasting performance. If that trend continues, the late-2027 projection for parity with superforecasters becomes a plausible milestone rather than a distant aspiration. The public leaderboard makes this progress visible, allowing organizations to compare models on a like-for-like basis and to identify which systems are most reliable for specific types of questions.

BISTRO and the macroeconomic frontier

The BIS work adds a different dimension to the forecasting story. While ForecastBench tests general knowledge and reasoning, BISTRO applies transformer architectures specifically to structured economic data. What makes BISTRO significant is not just its accuracy on the 2021 inflation episode. It is the demonstration that a single architecture can handle the kind of regime shifts and structural breaks that have historically confounded econometric models. Traditional models often rely on assumptions about mean reversion or stable relationships between variables. When those assumptions break down — as they did during the pandemic and subsequent inflation surge — the models fail. BISTRO’s ability to learn patterns directly from data, without strong prior assumptions, gives it an advantage in such environments.

The model’s performance on the 2021 inflation surge is a concrete example of this advantage. While conventional forecasting approaches projected a rapid return to pre-pandemic inflation levels, BISTRO identified persistent upward pressure in the data. That divergence was not a matter of marginal accuracy; it reflected a fundamentally different reading of the underlying economic dynamics. For central banks, which rely on inflation forecasts to set interest rates and communicate policy intentions, such differences have direct consequences for credibility and effectiveness.

However, the BIS researchers are careful not to overclaim. BISTRO’s performance on inflation does not automatically generalize to all macroeconomic variables or all time periods. The model is a proof of concept, not a replacement for central bank forecasting teams. Human judgment, institutional knowledge, and the ability to incorporate qualitative information remain essential components of policy analysis. The BIS publication frames BISTRO as a complement to existing tools, one that can surface patterns human analysts might miss but that still requires expert oversight to interpret and act upon.

Uncertainties and open questions

Several caveats temper the optimistic projections. First, the linear trend toward superforecaster parity is an extrapolation, not a guarantee. Progress in AI has historically been uneven, with periods of rapid advance followed by plateaus. The specific challenges of forecasting — including the need to reason about counterfactuals, handle ambiguous evidence, and calibrate uncertainty — may prove harder to solve than benchmark improvements suggest. A model that performs well on a 200-question subset may still fail on the long tail of high-stakes, low-probability events that matter most in practice.

Second, the mechanisms by which LLMs generate forecasts are not fully understood. Researchers at the Vector Institute, presenting at ICML 2026, have challenged the assumption that LLMs use unique internal “circuits” for specific tasks. Their work on mechanistic interpretability suggests that forecasting abilities may be distributed across the network rather than localized in identifiable components. This has practical implications: if forecasting emerges from diffuse interactions rather than dedicated modules, it may be harder to improve through targeted fine-tuning or to audit for reliability. A model that cannot explain why it made a particular prediction is harder to trust in high-stakes settings.

Third, generalization across domains remains unproven. A model that excels at macroeconomic forecasting may not perform equally well on questions about technological disruption, political instability, or climate-related risks. Each domain has its own data characteristics, uncertainty profiles, and expert communities. The fact that BISTRO works for inflation does not mean a similar model would work for predicting election outcomes or supply chain disruptions. The ForecastBench data itself shows significant variation in LLM performance across question categories, with some domains showing much faster improvement than others.

For executives and specialists, the practical takeaway is clear. LLMs are becoming reliable tools for baseline forecasting and scenario analysis in domains where data is abundant and patterns are relatively stable. They can generate useful probability estimates, identify relevant variables, and update predictions as new information arrives. But for high-stakes decisions — where the cost of being wrong is large and the situation is genuinely novel — human expertise still matters. The most effective approach, according to the research, is likely to be a hybrid one: using AI to generate and update forecasts, while relying on trained human forecasters to challenge assumptions, interpret ambiguous signals, and make final judgments.

The convergence of AI and forecasting is accelerating. Whether the late-2027 projection proves accurate or not, the direction of travel is unmistakable. Organizations that invest in understanding these tools now — and in building the human expertise to use them well — will be better positioned to navigate an increasingly uncertain world. The benchmark data, the BIS results, and the open questions from interpretability research all point to the same conclusion: forecasting is no longer a purely human domain, but it is not yet a purely machine domain either. The institutions that thrive will be those that treat AI as a partner in prediction, not a substitute for judgment.

#AI forecasting #LLMs #macroeconomics #benchmarks

Newsletter

Get the AI news that matters

One short brief with the day's most important AI stories — written for professionals.

We send a confirmation link. No spam. Unsubscribe anytime.

WhatsApp