Abstract
This study investigates the role of textual information in forecasting realized volatility infinancial markets. Using state-of-the-art large language models (LLMs) such as LLaMA,
OPT, BERT, and RoBERTa, we develop novel textual regression models to predict oneday-,
one-week-, two-week-, and one-month-ahead volatility for corn, soybean, and wheat
markets based on news articles published by media agencies. Our results show that these
models significantly outperform autoregressive models like HAR, as well as sentiment-based
approaches, particularly in long-term forecasting. The findings also reveal relationships
between forecasting accuracy and both the size and type of the language model. Shapley
value analysis demonstrates that market-related terms enhance short-term forecasts, while
production- and weather-related terms drive long-term predictions. Our findings underscore
the utility of textual data for forecasting volatility and provide a foundation for further
applications in financial markets and risk management.