BIS Primer Shows Economists How to Use Large Language Models
A new Bank for International Settlements guide walks researchers through the full lifecycle of LLM projects, from organising text to evaluating outputs, while warning that the tools require discipline and human judgment.
The Bank for International Settlements has issued a practical guide for economists seeking to deploy large language models in their research, arguing that the technology can transform how central banks and financial institutions extract signals from the vast, unstructured text that shapes modern economies. Published on 10 December 2024 as part of the BIS Quarterly Review, the primer walks readers through the full lifecycle of an LLM project—from organising raw text to evaluating model outputs—while cautioning that the tools are not a substitute for careful research design and human judgment.
The report, titled Large language models: a primer for economists, arrives at a moment when economic institutions are awash in textual data: policy statements, earnings calls, regulatory filings, news wires and social media. Much of this information has historically been too costly or cumbersome to analyse at scale. The BIS authors argue that LLMs offer a way to convert that flood of words into structured, quantifiable inputs for forecasting, nowcasting, sentiment analysis and financial surveillance. For a global audience of economists, data scientists and policy professionals, the primer is a signal that one of the world’s most influential financial institutions sees LLMs as a core analytical tool rather than a passing experiment.
Who wrote it and what it covers
The primer was authored by four members of the BIS Monetary and Economic Department: Byeungchun Kwon, a senior financial market analyst; Taejin Park, head of financial markets research support; Fernando Perez-Cruz, a senior adviser on innovation; and Phurichai Rungcharoenkitkul, a principal economist. Perez-Cruz, who holds a PhD in electrical engineering and serves as an adjunct professor at ETH Zurich, brings a depth of technical expertise that is often absent from central bank publications on machine learning. His background in signal processing and statistical learning underpins the primer’s emphasis on rigorous evaluation rather than anecdotal success stories.
The document is explicitly designed as a step-by-step guide. It covers how to organise economic text into usable datasets, how to choose between different model architectures, and how to evaluate whether an LLM is actually producing reliable outputs. The authors emphasise modular workflows, informed model selection and sufficient training data as best practices. At the same time, they warn against two common failure modes: expecting too much from the technology and deploying it without adequate human oversight. The primer does not present LLMs as magic. It presents them as tools that require discipline, clear objectives and a willingness to test assumptions at every stage.
A key technical distinction: decoder versus encoder models
One of the primer’s most useful contributions is its clarification of a technical divide that many economists overlook. The authors distinguish between decoder-based models, such as GPT, and encoder-based models, such as BERT. Decoder models generate text sequentially, predicting one word after another. They are the engines behind chatbots and other generative applications. Encoder models, by contrast, process entire sequences of text at once and build a contextual representation of meaning. For many economic research tasks—classifying documents, extracting sentiment, identifying topics—encoder models may be more appropriate, the authors note, because they are designed to understand text rather than produce it.
This distinction matters because economists often default to the most famous LLMs, which are overwhelmingly decoder-based. The primer encourages a more deliberate choice. A researcher trying to measure the tone of central bank communications, for example, may get better results from a fine-tuned encoder model than from prompting a general-purpose chatbot. The BIS authors do not dismiss generative models; they simply insist that the model should match the task. Choosing the wrong architecture can lead to wasted resources, misleading outputs and a false sense of analytical precision.
Testing the framework on 60,000 news articles
To demonstrate their approach, the authors applied their framework to a concrete research question: what drives US equity markets? They assembled a corpus of more than 60,000 news articles published between 2021 and 2023 and used LLMs to extract structured signals from that text. The analysis found that macroeconomic and monetary policy news are important drivers of equity market movements, but that market sentiment—the collective mood embedded in news coverage—also exerts substantial influence. The finding is not wholly surprising, but the method is notable: it shows how LLMs can turn a mountain of unstructured news into a dataset suitable for rigorous econometric analysis.
The BIS has also released sample code on GitHub to support replication. That is a significant step for an institution whose research is often consumed but rarely reproduced. By making the code public, the authors invite other researchers to test their methods, adapt them to different datasets, and identify weaknesses. It is a move that aligns with broader trends in empirical economics toward transparency and reproducibility, and it lowers the barrier for smaller institutions or academic teams that lack the resources to build LLM pipelines from scratch.
Implications for central banks and financial institutions
The primer’s publication is not merely an academic exercise. It reflects a growing recognition among central banks and financial regulators that unstructured text is a critical input for policy. The European Central Bank has used text mining to analyse speeches and press conferences. The US Federal Reserve has explored natural language processing for measuring economic uncertainty. The BIS itself has long argued that better data—including data extracted from text—can improve financial stability monitoring. What this primer adds is a practical, institutionally grounded methodology that other organisations can follow, adapt and refine.
The authors are careful to temper expectations. They note that LLMs are not a replacement for economic theory or domain expertise. A model that identifies sentiment in news articles does not, by itself, explain why sentiment moves markets. Nor does it account for selection bias in news coverage, changes in journalistic norms, or the complex feedback loops between media and markets. The primer repeatedly returns to the need for human oversight: economists must validate outputs, interrogate assumptions, and remain sceptical of results that look too clean. In high-stakes settings—monetary policy decisions, financial stability assessments—an unchecked LLM pipeline could amplify errors rather than reduce them.
The report also raises questions about data quality and access. LLMs require large volumes of text, but not all text is equally useful. Regulatory filings differ from news articles, which differ from social media posts. The primer’s emphasis on data organisation reflects a hard truth: the most sophisticated model will fail if it is fed messy, biased or incomplete data. For institutions in smaller economies or with fewer resources, building such datasets may be a significant challenge. The BIS does not solve that problem, but it does make the requirements explicit, helping institutions plan realistically rather than stumble into avoidable pitfalls.
Looking ahead, the primer suggests that LLMs will become a standard part of the economist’s toolkit, much as econometric software did decades ago. The question is not whether to use them, but how to use them well. The BIS authors offer a framework that is at once technically informed and institutionally cautious. They see enormous potential in using LLMs to read the economy’s text, but they insist that the final judgment must remain human. As central banks and financial institutions around the world grapple with the same challenge, this primer is likely to become a reference point—not because it offers definitive answers, but because it asks the right questions and provides a disciplined path toward answering them.
Sources
Written by an AI editorial process from the sources above. Errors may occur.
Newsletter
Get the AI news that matters
One short brief with the day's most important AI stories — written for professionals.
We send a confirmation link. No spam. Unsubscribe anytime.
Read next
AI Accelerates Drug Discovery from Concept to Clinic
Artificial intelligence is slashing development timelines and costs in pharmaceutical R&D, with AI-designed drugs now entering clinical trials in record time.
27 Sep 2026
Google Moves Gemini Team Under DeepMind Leadership
Google integrates its consumer AI app team into DeepMind to accelerate generative AI development and streamline research-to-product pipelines.
25 Sep 2026
AI in Drug Discovery: From Target ID to Clinical Trials
Artificial intelligence is accelerating drug discovery, but clinical validation remains the final frontier.
24 Sep 2026