SearcharxivSearch

arXiv subjects

Lingzhi Chen

Publications and source records attributed to Lingzhi Chen.

5 recordsLinked to original sources

FinToolBench: Evaluating LLM Agents for Real-World Financial Tool Use

The integration of Large Language Models (LLMs) into the financial domain is driving a paradigm shift from passive information retrieval to dynamic, agentic interaction. While general-purpose tool learning has witnessed a surge in benchmarks, the financial sector, characterized by high stakes, strict compliance, and rapid data volatility, remains critically underserved. Existing financial evaluations predominantly focus on static textual analysis or document-based QA, ignoring the complex reality of tool execution. Conversely, general tool benchmarks lack the domain-specific rigor required for finance, often relying on toy environments or a negligible number of financial APIs. To bridge this gap, we introduce FinToolBench, the first real-world, runnable benchmark dedicated to evaluating financial tool learning agents. Unlike prior works limited to a handful of mock tools, FinToolBench establishes a realistic ecosystem coupling 760 executable financial tools with 295 rigorous, tool-required queries. We propose a novel evaluation framework that goes beyond binary execution success, assessing agents on finance-critical dimensions: timeliness, intent type, and regulatory domain alignment. Furthermore, we present FATR, a finance-aware tool retrieval and reasoning baseline that enhances stability and compliance. By providing the first testbed for auditable, agentic financial execution, FinToolBench sets a new standard for trustworthy AI in finance. The tool manifest, execution environment, and evaluation code will be open-sourced to facilitate future research.

cs.AI

Multi-modal Vision Pre-training for Medical Image Analysis

Self-supervised learning has greatly facilitated medical image analysis by suppressing the training data requirement for real-world applications. Current paradigms predominantly rely on self-supervision within uni-modal image data, thereby neglecting the inter-modal correlations essential for effective learning of cross-modal image representations. This limitation is particularly significant for naturally grouped multi-modal data, e.g., multi-parametric MRI scans for a patient undergoing various functional imaging protocols in the same study. To bridge this gap, we conduct a novel multi-modal image pre-training with three proxy tasks to facilitate the learning of cross-modality representations and correlations using multi-modal brain MRI scans (over 2.4 million images in 16,022 scans of 3,755 patients), i.e., cross-modal image reconstruction, modality-aware contrastive learning, and modality template distillation. To demonstrate the generalizability of our pre-trained model, we conduct extensive experiments on various benchmarks with ten downstream tasks. The superior performance of our method is reported in comparison to state-of-the-art pre-training methods, with Dice Score improvement of 0.28\%-14.47\% across six segmentation benchmarks and a consistent accuracy boost of 0.65\%-18.07\% in four individual image classification tasks.

cs.CV

Estimating the index of increase via balancing deterministic and random data

We introduce and explore an empirical index of increase that works in both deterministic and random environments, thus allowing to assess monotonicity of functions that are prone to random measurement-errors. We prove consistency of the index and show how its rate of convergence is influenced by deterministic and random parts of the data. In particular, the obtained results suggest a frequency at which observations should be taken in order to reach any pre-specified level of estimation precision. We illustrate the index using data arising from purely deterministic and error-contaminated functions, which may or may not be monotonic.

math.ST

Influence of tensor interactions on masses and decay widths of dibaryons

The influence of gluon and Goldstone boson induced tensor interactions on the dibaryon masses and D-wave decay widths has been studied in the quark delocalization, color screening model. The effective S-D wave transition interactions induced by gluon and Goldstone boson exchanges decrease rapidly with increasing strangeness of the channel. The tensor contribution of K and $η$ mesons is negligible in this model. There is no six-quark state in the light flavor world studied so far that can become bound by means of these tensor interactions besides the deuteron. The partial D-wave decay widths of the $IJ^p={1/2}2^+$ N$Ω$ state to spin 0 and 1 $ΛΞ$ final states are 12.0 keV and 21.9 keV respectively. This is a very narrow dibaryon resonance that might be detectable in relativistic heavy ion reactions by existing RHIC detectors through the reconstruction of the vertex mass of the decay product $ΛΞ$ and by the COMPAS detector at CERN or at JHF in Japan and the FAIR project in Germany in the future.

hep-ph