SearcharxivSearch

arXiv subjects

Linying Lv

Publications and source records attributed to Linying Lv.

4 recordsLinked to original sources

Instruction Tuning Chronologically Consistent Language Models

We introduce a family of chronologically consistent, instruction-tuned large language models to eliminate lookahead bias. Each model is trained only on data available before a clearly defined knowledge-cutoff date, ensuring strict temporal separation from any post-cutoff data. The resulting framework offers (i) a simple, conversational chat interface, (ii) fully open, fixed model weights that guarantee replicability, and (iii) a conservative lower bound on forecast accuracy, isolating the share of predictability that survives once training leakage is removed. Together, these features provide researchers with an easy-to-use generative AI tool useful for a wide range of prediction tasks that is free of lookahead bias.

cs.LG

Chronologically Consistent Large Language Models

Large language models are increasingly used in social sciences, but their training data can introduce lookahead bias and training leakage. A good chronologically consistent language model requires efficient use of training data to maintain accuracy despite time-restricted data. Here, we overcome this challenge by training a suite of chronologically consistent large language models, ChronoBERT and ChronoGPT, which incorporate only the text data that would have been available at each point in time. Despite this strict temporal constraint, our models achieve strong performance on natural language processing benchmarks, outperforming or matching widely used models (e.g., BERT), and remain competitive with larger open-weight models. Lookahead bias is model and application-specific because even if a chronologically consistent language model has poorer language comprehension, a regression or prediction model applied on top of the language model can compensate. In an asset pricing application predicting next-day stock returns from financial news, we find that ChronoBERT and ChronoGPT's real-time outputs achieve Sharpe ratios comparable to a much larger Llama model, indicating that lookahead bias is modest. Our results demonstrate a scalable, practical framework to mitigate training leakage, ensuring more credible backtests and predictions across finance and other social science domains.

q-fin.GN

Do Sell-side Analyst Reports Have Investment Value?

This paper documents novel investment value in analyst report text. Using 1.2 million reports from 2000-2023, I embed narratives with large language models (LLMs) and fit machine learning (ML) forecasts of future long-term returns. Portfolios formed on the report narrative forecasts earn sizable and significant performance that is incremental to analysts' numerical outputs and to a broad set of established factors and characteristic-based predictors. The effect is stronger after adverse news and is amplified for growth stocks with aggressive investment. To open the black box, I apply a Shapley decomposition that attributes portfolio performance to distinct topics. Analysts' strategic outlook contributes the most to portfolio performance, especially forward-looking fundamental assessments. Beyond providing direct evidence that analyst narratives contain value-relevant assessments that diffuse into price over time, this study illustrates how interpretable LLM-plus-ML pipelines can scale and augment human judgment in investment decisions.

q-fin.PR

The Value of Information from Sell-side Analysts

I examine the value of information from sell-side analysts by analyzing a large corpus of their written reports. Using embeddings from state-of-the-art large language models, I show that qualitative information in analyst reports explains above 10% of contemporaneous stock returns out-of-sample, a value that is more economically significant than quantitative forecasts. I then perform a Shapley value decomposition to assess how much each topic within the reports contributes to explaining stock returns. The results show that analysts' income statement analyses account for more than half of the reports' explanatory power. Expressing these findings in economic terms, I estimate that early acquisition of analyst reports can yield significant profits. Analyst information value peaks in the first week following earnings announcements, highlighting their vital role in interpreting new financial data.

q-fin.GN