Forecasting Cybersecurity Incidents Using Geopolitical Data and Large Language Models
Predicting security incidents is a profound task critical for informing proactive defensive measures and cyber-insurance policies. Prior work tackling this problem mainly utilized structured, manually defined features based on network measurements (e.g., protocol misconfigurations). Still, despite leading to promising performance, the network-based features may fail to capture aspects related to adversaries' motives. To fill this gap, our work leverages geopolitical data mined from public sources--which may help capture attacker motives--to forecast security incidents. Specifically, our approach relies on news articles and transcribed podcasts that are fed to large language models to automatically produce rich representations. The representations are then fed to a classifier trained to forecast future incidents based on historical ones. Our evaluation with a large incidents dataset (>15,700 records) demonstrates substantial accuracy (71.4% ROC AUC) with geopolitical data alone. Notably, combining geopolitical data and network measurements outperforms the state-of-the-art technique based on network features alone (81.3% vs.77.5% ROC AUC). Our analysis also helps shed light on when geopolitical data is most helpful and the data sources that are most useful for accurate forecasting.