SearcharxivSearch

arXiv · 2503.16974

Assessing Consistency and Reproducibility in the Outputs of Large Language Models: Evidence Across Diverse Finance and Accounting Tasks

Abstract

This study provides the first comprehensive assessment of consistency and reproducibility in Large Language Model (LLM) outputs in finance and accounting research. We evaluate how consistently LLMs produce outputs given identical inputs through extensive experimentation with 50 independent runs across five common tasks: classification, sentiment analysis, summarization, text generation, and prediction. Using three OpenAI models (GPT-3.5-turbo, GPT-4o-mini, and GPT-4o), we generate over 3.4 million outputs from diverse financial source texts and data, covering MD&As, FOMC statements, finance news articles, earnings call transcripts, and financial statements. Our findings reveal substantial but task-dependent consistency, with binary classification and sentiment analysis achieving near-perfect reproducibility, while complex tasks show greater variability. More advanced models do not consistently demonstrate better consistency and reproducibility, with task-specific patterns emerging. LLMs significantly outperform expert human annotators in consistency and maintain high agreement even where human experts significantly disagree. We further find that simple aggregation strategies across 3-5 runs dramatically improve consistency. We also find that aggregation may come with an additional benefit of improved accuracy for sentiment analysis when using newer models. Simulation analysis reveals that despite measurable inconsistency in LLM outputs, downstream statistical inferences remain remarkably robust. These findings address concerns about what we term "G-hacking," the selective reporting of favorable outcomes from multiple generative AI runs, by demonstrating that such risks are relatively low for finance and accounting tasks.

Explore related subjects

Keep this discovery

BibTeXRIS

Julian Junyan Wang, Victor Xiaoqi Wang. 2025-03-21. Assessing Consistency and Reproducibility in the Outputs of Large Language Models: Evidence Across Diverse Finance and Accounting Tasks. https://arxiv.org/abs/2503.16974

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

AI for AI: Optimizing Additional Infrastructure Build-out to Power Artificial Intelligence Data Centers

The twenty-first century's transformative technology, artificial intelligence, is increasingly constrained by the twentieth century's transformative technology, the electricity grid. Rapid growth in electricity demand from data centers is leading to higher electricity prices, without a compensating supply-side response. We develop a framework linking data-center load growth, available generation capacity, and market-clearing prices to understand this phenomenon. We first analyze a deterministic model to show how differing estimates of demand and supply growth rates affect prices. We then model the expansion of new data centers and their associated electricity demand, together with build-outs of new electricity supply, as stochastic processes,resulting in probabilistic distributions of supply, demand, and prices rather than a single forecast. Finally, we formulate generation expansion as a stochastic control problem in which a revenue-maximizing investor dynamically chooses the intensity of supply-side investments. The analysis highlights a central challenge of the data-center build-out: even when rapid demand growth increases the need for new generation, the uncertainties related to load forecasts, development execution risks, and value cannibalization from overbuilding capacity may weaken incentives to invest at the pace required to keep electricity prices stable.

q-fin.GN

Measuring DeFi Risk

Decentralized finance (DeFi) lending has grown from nonexistent in 2017 to nearly 40 billion US Dollars in deposited funds in May 2022. Using cryptocurrency as collateral, the platforms match speculative margin trading with yield-seeking depositors lending coins pegged to the dollar (stable coins). Depositors receive claims guaranteed by a basket of collateral, akin to new stable coins. We develop a framework requiring only knowledge of aggregate deposits and borrowings to measure overall system risks to lenders and borrowers. Using evidence from major protocols, the measures identify an increase in system fragility beyond prudent levels around mid 2021, with a potential loss of peg for extreme variations in coin prices. Overall, the model offers an easily implementable aggregate risk metric capturing the perspectives of synthetic investors and offers early warning signals as the industry is moving from deposits guaranteed by collateral to fiat money.

q-fin.GN

Historical Reflections on Interest Rates and the Emergence of the Yield Curve

This text grew out of a historical introduction initially written for a study of interest rates in cryptocurrency markets. The difficulty of defining a term structure for a currency without a conventional bond market led naturally to a more fundamental question: under what historical conditions does a yield curve become observable at all? Credit existed long before modern money, and interest-bearing loans are documented as early as ancient Mesopotamia. For much of history, the surviving evidence lacks the institutional features that facilitate reliable comparisons of interest rates by maturity: standardised debt instruments, sufficiently homogeneous borrowers, regular issuance over a range of maturities, observable market prices, and liquid secondary markets. We trace the gradual emergence of these conditions from ancient Mesopotamia, Greece, and Rome, through medieval and early modern Europe, to the development of modern sovereign debt markets in the nineteenth and twentieth centuries.

q-fin.GN