Words That Unite The World: A Unified Framework for Deciphering Central Bank Communications

Siddhartha Somani (Georgia Institute of Technology) · Agam Shah (Georgia Institute of Technology) · Siddhant Sukhani (Stanford University) · Huzaifa Pardawala (Georgia Institute of Technology) · Saketh Budideti (Georgia Institute of Technology) · Riya Bhadani (Georgia Institute of Technology) · Rudra Gopal (Georgia Institute of Technology (MS QCF)) · Rutwik Routu (Duke University) · Michael Galarnyk (Georgia Institute of Technology) · Soungmin Lee (Georgia Institute of Technology) · Arnav Hiray (Georgia Institute of Technology) · Akshar Ravichandran (Georgia Institute of Technology) · Eric Kim (Georgia Institute of Technology) · Pranav Aluru (Georgia Institute of Technology) · Joshua Zhang (Georgia Institute of Technology) · Sebastian Jaskowski (Microsoft) · Veer Guda (Georgia Institute of Technology) · Meghaj Tarte (Georgia Institute of Technology) · Liqin Ye (Georgia Institute of Technology) · Spencer Gosden (Georgia Institute of Technology) · Rachel Yuh (Georgia Institute of Technology) · Sloka Chava (Fulton Science Academy ) · Sahasra Chava (Georgia Institute of Technology) · Dylan Patrick Kelly (Georgia Institute of Technology) · Aiden Chiang (Georgia Institute of Technology) · Harsit Mittal (Georgia Institute of Technology) · Sudheer Chava (Georgia Institute of Technology)
benchmarking experimentsdisagreement resolutionsdual annotatorseconomic utilityerror analysesfew-shot learningmonetary policy corpuspredictive taskspretrained language modelsstance detectiontemporal classificationuncertainty estimationworld central banks datasetzero-shot learning

Central banks around the world play a crucial role in maintaining economic stability. Deciphering policy implications in their communications is essential, especially as misinterpretations can disproportionately impact vulnerable populations. To address this, we introduce the World Central Banks (WCB) dataset, the most comprehensive monetary policy corpus to date, comprising over 380k sentences from 25 central banks across diverse geographic regions, spanning 28 years of historical data. After uniformly sampling 1k sentences per bank (25k total) across all available years, we annotate and review each sentence using dual annotators, disagreement resolutions, and secondary expert reviews. We define three tasks: Stance Detection, Temporal Classification, and UncertaintyEstimation, with each sentence annotated for all three. We benchmark seven Pretrained Language Models (PLMs) and nine Large Language Models (LLMs) (Zero-Shot, Few-Shot, and with annotation guide) on these tasks, running 15,075 benchmarking experiments. We find that a model trained on aggregated data across banks significantly surpasses a model trained on an individual bank's data, confirming the principle *"the whole is greater than the sum of its parts."* Additionally, rigorous human evaluations, error analyses, and predictive tasks validate our framework's economic utility. Our artifacts are accessible through the HuggingFace and GitHub under the CC-BY-NC-SA 4.0 license.