Anonymization and Information Loss
Anonymizing financial texts prevents large language models (LLMs) from exploiting look-ahead bias, but inadvertently weakens the extracted signal. We propose a framework disentangling this information loss from bias removal. Predicting S&P credit downgrades, raw texts significantly outperform anonymized texts. Look-ahead bias does not explain this difference; rather, anonymization degrades predictive performance by masking informative numerical and agent entities, fundamentally altering how LLMs interpret the remaining context. This degradation spans various LLMs, tasks, and text types. To quantify it, we introduce the "anonymization gap," a metric requiring no outcome data.