arXiv · 2610.08675
Same-Number Citation Swaps: Stress-Testing Jev as a Financial Evidence Judge
Abstract
Financial reports repeat values across periods, metrics and accounting lines, allowing an LLM-generated calculation to be numerically correct while citing the wrong financial role. We evaluate what probabilistic evidence verification adds beyond number matching using Jev as a source-support verifier for GPT-4.1-mini calculation traces. A signed-number-at-pointer baseline explains most recovery over exact quotation checks. To isolate the remaining role-recognition problem, we hold operands and arithmetic fixed, move citations between same-number cells, and retain controls that express equivalent facts. These contrasts reveal both wrong-role citations that pass and valid alternative citations that are withheld. Explicit column labels improve selected wrong-role decisions while also lowering support for some equivalent evidence. A constructed follow-up on 36 new source pages, labeled by a non-author reviewer, extends this evaluation and exposes the same tradeoff between detecting role errors and retaining valid citations. The contribution is a controlled evaluation that identifies what a probabilistic financial verifier distinguishes when numerical matching is held fixed. For LLM-based financial assistants, it makes numerical correctness, cited-role support and acceptance outcomes separately assessable.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Chuhong Xu, Bo Su, Ziyao Chen, Ruiyang Xu, Shimeng Dai, Xinyu Qiu. 2026-10-06. Same-Number Citation Swaps: Stress-Testing Jev as a Financial Evidence Judge. https://arxiv.org/abs/2610.08675
Cite the original work for its findings. Save a collection to share your selection of sources.