SearcharxivSearch

arXiv subjects

Dexin Zhou

Publications and source records attributed to Dexin Zhou.

4 recordsLinked to original sources

Anonymization and Information Loss

Anonymizing financial texts prevents large language models (LLMs) from exploiting look-ahead bias, but inadvertently weakens the extracted signal. We propose a framework disentangling this information loss from bias removal. Predicting S&P credit downgrades, raw texts significantly outperform anonymized texts. Look-ahead bias does not explain this difference; rather, anonymization degrades predictive performance by masking informative numerical and agent entities, fundamentally altering how LLMs interpret the remaining context. This degradation spans various LLMs, tasks, and text types. To quantify it, we introduce the "anonymization gap," a metric requiring no outcome data.

q-fin.GN

The Promise and Peril of Generative AI: Evidence from GPT as Sell-Side Analysts

Large language models (LLMs) promise to democratize financial analysis by reducing information-processing costs. Yet equal access does not ensure equal outcomes, as the locus of friction may shift from processing information to evaluating model outputs. We study GPT's earnings forecasts following corporate earnings releases and document two patterns. First, GPT's narrative attention is consistent and human-like but not always associated with higher forecast accuracy. Second, its quantitative reasoning varies substantially across contexts, challenging the view that LLMs are uniformly weak at numerical tasks. Building on these insights, we propose a diagnostic framework that links forecast accuracy to observable processing features (i.e., narrative focus, numerical reasoning, and self-assessed confidence). These indicators serve as proxies for this new form of information friction and alert investors when to exercise caution. Our study has implications for information frictions, regulatory oversight, and the economics of AI-mediated financial markets.

q-fin.GN

What Does ChatGPT Make of Historical Stock Returns? Extrapolation and Miscalibration in LLM Stock Return Forecasts

We examine how large language models (LLMs) interpret historical stock returns and compare their forecasts with estimates from a crowd-sourced platform for ranking stocks. While stock returns exhibit short-term reversals, LLM forecasts over-extrapolate, placing excessive weight on recent performance similar to humans. LLM forecasts appear optimistic relative to historical and future realized returns. When prompted for 80% confidence interval predictions, LLM responses are better calibrated than survey evidence but are pessimistic about outliers, leading to skewed forecast distributions. The findings suggest LLMs manifest common behavioral biases when forecasting expected returns but are better at gauging risks than humans.

q-fin.GN

Data Clustering via Principal Direction Gap Partitioning

We explore the geometrical interpretation of the PCA based clustering algorithm Principal Direction Divisive Partitioning (PDDP). We give several examples where this algorithm breaks down, and suggest a new method, gap partitioning, which takes into account natural gaps in the data between clusters. Geometric features of the PCA space are derived and illustrated and experimental results are given which show our method is comparable on the datasets used in the original paper on PDDP.

stat.ML