Searcharxiv⌕ Search

arXiv subjects

Yi Han

Publications and source records attributed to Yi Han.

At least 37 records · Page 2Linked to original sources

Conv-FinRe: A Conversational and Longitudinal Benchmark for Utility-Grounded Financial Recommendation

Most recommendation benchmarks evaluate how well a model imitates user behavior. In financial advisory, however, observed actions can be noisy or short-sighted under market volatility and may conflict with a user's long-term goals. Treating what users chose as the sole ground truth, therefore, conflates behavioral imitation with decision quality. We introduce Conv-FinRe, a conversational and longitudinal benchmark for stock recommendation that evaluates LLMs beyond behavior matching. Given an onboarding interview, step-wise market context, and advisory dialogues, models must generate rankings over a fixed investment horizon. Crucially, Conv-FinRe provides multi-view references that distinguish descriptive behavior from normative utility grounded in investor-specific risk preferences, enabling diagnosis of whether an LLM follows rational analysis, mimics user noise, or is driven by market momentum. We build the benchmark from real market data and human decision trajectories, instantiate controlled advisory conversations, and evaluate a suite of state-of-the-art LLMs. Results reveal a persistent tension between rational decision quality and behavioral alignment: models that perform well on utility-based ranking often fail to match user choices, whereas behaviorally aligned models can overfit short-term noise. The dataset is publicly released on Hugging Face, and the codebase is available on GitHub.

cs.AI↗

Can LLM Agents Be CFOs? Benchmarking Long-Horizon Resource Allocation in an Uncertain Enterprise Environment

Large language model (LLM) agents are increasingly tested on complex tasks, but their ability to allocate scarce resources over long horizons remains unclear. Unlike reactive tasks with immediate feedback, this setting requires agents to make binding commitments under partial observability, delayed consequences, hard resource budgets, and shifting dynamics. We introduce EnterpriseArena, a 132-month CFO simulator that evaluates long-horizon resource allocation under uncertainty in a FinTech lending firm. Agents must manage liquidity, close books, gather costly signals, and request equity or debt financing across changing macroeconomic regimes. The simulator is built from transformed firm-level financial data, anonymized business documents, decade-scale macroeconomic and industry signals, and expert-validated operating rules. Experiments across 23 LLMs and four agent frameworks show that current agents remain far from robust: only 15.4% of trials survive the full horizon, larger models do not reliably outperform smaller ones, and failures cascade across observation, action timing, and capital sizing. These findings establish long-horizon resource allocation under uncertainty as a distinct capability gap for LLM agents.

cs.AI↗

On the rank of a random symmetric matrix in the large deviation regime

Let $A$ be an $n\times n$ random symmetric matrix with independent identically distributed subgaussian entries of unit variance. We prove the following large deviation inequality for the rank of $A$: for all $1\leq k\leq c\sqrt{n}$, $$\mathbb{P}(\operatorname{Rank}(A)\geq n-k)\geq 1-\exp(-c'kn),$$ for some fixed constants $c,c'>0$. A similar large deviation inequality is proven for the rank of the adjacency matrix of dense Erdos-Renyi graphs. This corank estimate enhances the recent breakthrough of Campos, Jensen, Michelen and Sahasrabudhe that the singularity probability of a random symmetric matrix is exponentially small, and echoes a large deviation inequality of Mark Rudelson for the rank of a random matrix with independent entries.

math.PR↗

Glauber dynamics for random field Ising models on bounded degree graphs and MLSI

We study the ferromagnetic random field Ising model (RFIM) on a graph $G=(V,E)$ having maximal degree $Δ$, where the external field at each vertex is an i.i.d. random variable. When the random field distribution is sufficiently anti-concentrated, we prove that with high probability over the quenched randomness of the external field, the Glauber dynamics of this RFIM mixes in polynomial time as a consequence of a Poincaré inequality. This model is relevant to the Griffiths phase where the correlations decay exponentially fast in expectation over the quenched random field, but contraction does not hold point-wise due to the existence of weak fields that lead to low-temperature behavior. Previously, fast mixing of Glauber dynamics under large disorder was only proven on the integer lattice, and for RFIM on general graphs, only a sampling algorithm based on self-avoiding walks was known. Under a further technical condition that the random fields are bounded, we prove a modified log-Sobolev inequality for the Glauber dynamics. When the random field is weaker but still satisfies weak spatial mixing (exponential decay of correlations from boundary to bulk) in expectation, and the graph has at most $α$-stretched exponential growth for some $α<1$, then we prove a weak Poincaré inequality holds, which gives rise to a polynomial time sampling algorithm based on Glauber dynamics with warm start. The latter result was previously proven for the integer lattice, and we extend its scope to graphs with only a volume growth condition without assuming a local geometry.

math.PR↗

FinCriticalED: A Visual Benchmark for Financial Fact-Level OCR

Recent progress in multimodal large language models (MLLMs) has substantially improved document understanding, yet strong optical character recognition (OCR) performance on surface metrics does not guarantee faithful preservation of decision-critical evidence. This limitation is especially consequential in financial documents, where small visual errors can induce discrete shifts in meaning. To study this gap, we introduce FinCriticalED (Financial Critical Error Detection), a fact-centric visual benchmark for evaluating whether OCR and vision-language systems preserve financially critical evidence beyond lexical similarity. FinCriticalED contains 859 real-world financial document pages with 9,481 expert-annotated facts spanning five critical field types: numeric, temporal, monetary unit, reporting entity, and financial concept. We formulate the task as structured OCR with fact-level verification, and develop a Deterministic-Rule-Guided LLM-as-Judge protocol to assess whether model outputs preserve annotated facts in context. We benchmark 13 systems spanning OCR pipelines, specialized OCR VLMs, open-source MLLMs, and proprietary MLLMs. Results reveal a clear gap between lexical accuracy and factual reliability, with numerical values and monetary units emerging as the most vulnerable fact types, and critical errors concentrating in visually complex, mixed-layout documents with distinct failure patterns across model families. Overall, FinCriticalED provides a rigorous benchmark for trustworthy financial OCR and a practical testbed for evidence fidelity in high-stakes multimodal document understanding. Benchmark and dataset details available at https://the-finai.github.io/FinCriticalED/

cs.CV↗

Quadratic shift-and-stack for Ground-Based Optical Detection of Faint Cislunar Objects

Detecting faint objects in cislunar space using ground-based optical telescopes is difficult because of their low brightness, strong lunar background, and complex, nonlinear apparent motion. Traditional shift-and-stack techniques based on linear motion assumption suffer signal trailing loss due to significant nonlinear motion during long integrations, thus producing a degraded signal-to-noise ratio (SNR). In this paper, we first derive a theoretical criterion based on the point spread function to determine the maximum applicable integration time for linear-motion stacking. We then propose a quadratic shift-and-stack (QSS) method to correct for the first-order nonlinear motion, namely the angular acceleration of cislunar targets. Simulations of typical cislunar orbits verify this theoretical criterion and show that the QSS method significantly improves SNR from stacking and can enhance the detection limit by up to 1 stellar magnitude compared with the linear-motion stacking method. Furthermore, tests using observational data of the cislunar object Tiandu-1 confirm that while linear stacking degrades after a 29-minute integration due to trajectory curvature, the QSS method achieves continuous SNR improvement over a 46-minute integration, outperforming the peak SNR of the linear method by 31%.

astro-ph.IM↗

More scaling limits for 1d random Schrödinger operators with critically decaying and vanishing potentials

Consider the random Schrödinger operator $H_n$ defined on $\{0,1,\cdots,n\}\subset\mathbb{Z}$ $$ (H_nψ)_\ell=ψ_{\ell-1,n}+ψ_{\ell+1,n}+σ\frac{ω_\ell}{a_{\ell,n}}ψ_{\ell,n},\quad ψ_0=ψ_{n+1}=0, $$ where $σ>0$, $ω_\ell$ are i.i.d. random variables and $a_{\ell,n}$ typically has order $\sqrt{n}$ for $\ell\in[εn,(1-ε)n]$ and any $ε>0$. Two important cases: (a) the vanishing case $a_{\ell,n}=\sqrt{n}$ and (b) the decaying case $a_{\ell,n}=\sqrt{\ell}$, were studied before in \cite{kritchevski2011scaling}. In this paper we consider more general decaying profiles that lie in between these two extreme cases. We characterize the scaling limit of transfer matrices and determine the point process limit of eigenvalues near a fixed energy in the bulk, in terms of solutions to coupled SDEs. We obtain new point processes that share similar properties to the $\text{Sch}_τ$ process. We determine the shape profile of eigenfunctions after a suitable rescaling, that corresponds to a uniformly chosen eigenvalue of $H_n$. We also give a more detailed description of the newly defined point processes, including the probability of small and large gaps and a variance estimate.

math.PR↗

SaPaVe: Towards Active Perception and Manipulation in Vision-Language-Action Models for Robotics

Active perception and manipulation are crucial for robots to interact with complex scenes. Existing methods struggle to unify semantic-driven active perception with robust, viewpoint-invariant execution. We propose SaPaVe, an end-to-end framework that jointly learns these capabilities in a data-efficient manner. Our approach decouples camera and manipulation actions rather than placing them in a shared action space, and follows a bottom-up training strategy: we first train semantic camera control on a large-scale dataset, then jointly optimize both action types using hybrid data. To support this framework, we introduce ActiveViewPose-200K, a dataset of 200k image-language-camera movement pairs for semantic camera movement learning, and a 3D geometry-aware module that improves execution robustness under dynamic viewpoints. We also present ActiveManip-Bench, the first benchmark for evaluating active manipulation beyond fixed-view settings. Extensive experiments in both simulation and real-world environments show that SaPaVe outperforms recent vision-language-action models such as GR00T N1 and \(π_0\), achieving up to 31.25\% higher success rates in real-world tasks. These results show that tightly coupled perception and execution, when trained with decoupled yet coordinated strategies, enable efficient and generalizable active manipulation. Project page: https://lmzpai.github.io/SaPaVe

cs.RO↗

TIGeR: Tool-Integrated Geometric Reasoning in Vision-Language Models for Robotics

Vision-Language Models (VLMs) have shown remarkable capabilities in spatial reasoning, yet they remain fundamentally limited to qualitative precision and lack the computational precision required for real-world robotics. Current approaches fail to leverage metric cues from depth sensors and camera calibration, instead reducing geometric problems to pattern recognition tasks that cannot deliver the centimeter-level accuracy essential for robotic manipulation. We present TIGeR (Tool-Integrated Geometric Reasoning), a novel framework that transforms VLMs from perceptual estimators to geometric computers by enabling them to generate and execute precise geometric computations through external tools. Rather than attempting to internalize complex geometric operations within neural networks, TIGeR empowers models to recognize geometric reasoning requirements, synthesize appropriate computational code, and invoke specialized libraries for exact calculations. To support this paradigm, we introduce TIGeR-300K, a comprehensive tool-invocation-oriented dataset covering point transformations, pose estimation, and spatial compatibility verification, complete with tool invocation sequences and intermediate computations. Through a two-stage training pipeline combining supervised fine-tuning (SFT) and reinforcement fine-tuning (RFT) with our proposed hierarchical reward design, TIGeR achieves SOTA performance on geometric reasoning benchmarks while demonstrating centimeter-level precision in real-world robotic manipulation tasks.

cs.RO↗

The USTC-NERCSLIP Systems for the CHiME-9 MCoRec Challenge

This report details our submission to the CHiME-9 MCoRec Challenge on recognizing and clustering multiple concurrent natural conversations within indoor social settings. Unlike conventional meetings centered on a single shared topic, this scenario contains multiple parallel dialogues--up to eight speakers across up to four simultaneous conversations--with a speech overlap rate exceeding 90%. To tackle this, we propose a multimodal cascaded system that leverages per-speaker visual streams extracted from synchronized 360 degree video together with single-channel audio. Our system improves three components of the pipeline by leveraging enhanced audio-visual pretrained models: Active Speaker Detection (ASD), Audio-Visual Target Speech Extraction (AVTSE), and Audio-Visual Speech Recognition (AVSR). The AVSR module further incorporates Whisper and LLM techniques to boost transcription accuracy. Our best single cascaded system achieves a Speaker Word Error Rate (WER) of 32.44% on the development set. By further applying ROVER to fuse outputs from diverse front-end and back-end variants, we reduce Speaker WER to 31.40%. Notably, our LLM-based zero-shot conversational clustering achieves a speaker clustering F1 score of 1.0, yielding a final Joint ASR-Clustering Error Rate (JACER) of 15.70%.

eess.AS↗

Invertibility for non-Hermitian and symmetric random band matrices with sublinear bandwidth and discrete entries

A well-known result in random matrix theory, proven by Kahn, Komlós and Szemerédi in 1995, states that a square random matrix with i.i.d. uniform $\{\pm 1\}$ entries is invertible with probability $1-\exp(-Ω(n))$. As a natural generalization of the model, we consider the invertibility of a class of random band matrices with independent entries where the bandwidth $d_n$ scales like $n^α$, for some $α\in(0,1)$. The band matrix model we consider is sufficiently general and covers existing models such as the block band matrix and periodic band matrix, allowing great flexibility in the variance profile. As the bandwidth is sublinear in the dimension, estimating the invertibility and least singular values of these matrices is a well-known open problem. We make progress towards the invertibility problem by showing that, when $α>\frac{2}{3}$ and when the random variables are i.i.d. uniformly distributed on $\{\pm 1+c\}$ for any fixed integer $c$, then the band matrix is invertible with probability $1-\exp(-Ω(n^{α/2}))$. Previously, even invertibility with probability $1-o(1)$ was not known for these band matrix models except in the very special case of block band matrices. We then extend the invertibility result to symmetric random band matrices with integer entries, and prove the same non-singularity probability estimate whenever $α>\frac{2}{3}$.

math.PR↗

MIRROR: Manifold Ideal Reference ReconstructOR for Generalizable AI-Generated Image Detection

High-fidelity generative models have narrowed the perceptual gap between synthetic and real images, posing serious threats to media security. Most existing AI-generated image (AIGI) detectors rely on artifact-based classification and struggle to generalize to evolving generative traces. In contrast, human judgment relies on stable real-world regularities, with deviations from the human cognitive manifold serving as a more generalizable signal of forgery. Motivated by this insight, we reformulate AIGI detection as a Reference-Comparison problem that verifies consistency with the real-image manifold rather than fitting specific forgery cues. We propose MIRROR (Manifold Ideal Reference ReconstructOR), a framework that explicitly encodes reality priors using a learnable discrete memory bank. MIRROR projects an input into a manifold-consistent ideal reference via sparse linear combination, and uses the resulting residuals as robust detection signals. To evaluate whether detectors reach the "superhuman crossover" required to replace human experts, we introduce the Human-AIGI benchmark, featuring a psychophysically curated human-imperceptible subset. Across 14 benchmarks, MIRROR consistently outperforms prior methods, achieving gains of 2.1% on six standard benchmarks and 8.1% on seven in-the-wild benchmarks. On Human-AIGI, MIRROR reaches 89.6% accuracy across 27 generators, surpassing both lay users and visual experts, and further approaching the human perceptual limit as pretrained backbones scale. The code is publicly available at: https://github.com/349793927/MIRROR

cs.CV↗

RoboBrain 2.5: Depth in Sight, Time in Mind

We introduce RoboBrain 2.5, a next-generation embodied AI foundation model that advances general perception, spatial reasoning, and temporal modeling through extensive training on high-quality spatiotemporal supervision. Building upon its predecessor, RoboBrain 2.5 introduces two major capability upgrades. Specifically, it unlocks Precise 3D Spatial Reasoning by shifting from 2D pixel-relative grounding to depth-aware coordinate prediction and absolute metric constraint comprehension, generating complete 3D manipulation traces as ordered keypoint sequences under physical constraints. Complementing this spatial precision, the model establishes Dense Temporal Value Estimation that provides dense, step-aware progress prediction and execution state understanding across varying viewpoints, producing stable feedback signals for downstream learning. Together, these upgrades extend the framework toward more physically grounded and execution-aware embodied intelligence for complex, fine-grained manipulation. The code and checkpoints are available at project website: https://superrobobrain.github.io

cs.RO↗

CARLA-Round: A Multi-Factor Simulation Dataset for Roundabout Trajectory Prediction

Accurate trajectory prediction of vehicles at roundabouts is critical for reducing traffic accidents, yet it remains highly challenging due to their circular road geometry, continuous merging and yielding interactions, and absence of traffic signals. Developing accurate prediction algorithms relies on reliable, multimodal, and realistic datasets; however, such datasets for roundabout scenarios are scarce, as real-world data collection is often limited by incomplete observations and entangled factors that are difficult to isolate. We present CARLA-Round, a systematically designed simulation dataset for roundabout trajectory prediction. The dataset varies weather conditions (five types) and traffic density levels (spanning Level-of-Service A-E) in a structured manner, resulting in 25 controlled scenarios. Each scenario incorporates realistic mixtures of driving behaviors and provides explicit annotations that are largely absent from existing datasets. Unlike randomly sampled simulation data, this structured design enables precise analysis of how different conditions influence trajectory prediction performance. Validation experiments using standard baselines (LSTM, GCN, GRU+GCN) reveal traffic density dominates prediction difficulty with strong monotonic effects, while weather shows non-linear impacts. The best model achieves 0.312m ADE on real-world rounD dataset, demonstrating effective sim-to-real transfer. This systematic approach quantifies factor impacts impossible to isolate in confounded real-world datasets. Our CARLA-Round dataset is available at https://github.com/Rebecca689/CARLA-Round.

cs.CV↗

RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for Robotics

Spatial referring is a fundamental capability of embodied robots to interact with the 3D physical world. However, even with the powerful pretrained vision language models (VLMs), recent approaches are still not qualified to accurately understand the complex 3D scenes and dynamically reason about the instruction-indicated locations for interaction. To this end, we propose RoboRefer, a 3D-aware VLM that can first achieve precise spatial understanding by integrating a disentangled but dedicated depth encoder via supervised fine-tuning (SFT). Moreover, RoboRefer advances generalized multi-step spatial reasoning via reinforcement fine-tuning (RFT), with metric-sensitive process reward functions tailored for spatial referring tasks. To support SFT and RFT training, we introduce RefSpatial, a large-scale dataset of 20M QA pairs (2x prior), covering 31 spatial relations (vs. 15 prior) and supporting complex reasoning processes (up to 5 steps). In addition, we introduce RefSpatial-Bench, a challenging benchmark filling the gap in evaluating spatial referring with multi-step reasoning. Experiments show that SFT-trained RoboRefer achieves state-of-the-art spatial understanding, with an average success rate of 89.6%. RFT-trained RoboRefer further outperforms all other baselines by a large margin, even surpassing Gemini-2.5-Pro by 17.4% in average accuracy on RefSpatial-Bench. Notably, RoboRefer can be integrated with various control policies to execute long-horizon, dynamic tasks across diverse robots (e,g., UR5, G1 humanoid) in cluttered real-world scenes.

cs.RO↗

Adaptive Causal Coordination Detection for Social Media: A Memory-Guided Framework with Semi-Supervised Learning

Detecting coordinated inauthentic behavior on social media remains a critical and persistent challenge, as most existing approaches rely on superficial correlation analysis, employ static parameter settings, and demand extensive and labor-intensive manual annotation. To address these limitations systematically, we propose the Adaptive Causal Coordination Detection (ACCD) framework. ACCD adopts a three-stage, progressive architecture that leverages a memory-guided adaptive mechanism to dynamically learn and retain optimal detection configurations for diverse coordination scenarios. Specifically, in the first stage, ACCD introduces an adaptive Convergent Cross Mapping (CCM) technique to deeply identify genuine causal relationships between accounts. The second stage integrates active learning with uncertainty sampling within a semi-supervised classification scheme, significantly reducing the burden of manual labeling. The third stage deploys an automated validation module driven by historical detection experience, enabling self-verification and optimization of the detection outcomes. We conduct a comprehensive evaluation using real-world datasets, including the Twitter IRA dataset, Reddit coordination traces, and several widely-adopted bot detection benchmarks. Experimental results demonstrate that ACCD achieves an F1-score of 87.3\% in coordinated attack detection, representing a 15.2\% improvement over the strongest existing baseline. Furthermore, the system reduces manual annotation requirements by 68\% and achieves a 2.8x speedup in processing through hierarchical clustering optimization. In summary, ACCD provides a more accurate, efficient, and highly automated end-to-end solution for identifying coordinated behavior on social platforms, offering substantial practical value and promising potential for broad application.

cs.AI↗

Beyond Artifacts: Real-Centric Envelope Modeling for Reliable AI-Generated Image Detection

The rapid progress of generative models has intensified the need for reliable and robust detection under real-world conditions. However, existing detectors often overfit to generator-specific artifacts and remain highly sensitive to real-world degradations. As generative architectures evolve and images undergo multi-round cross-platform sharing and post-processing (chain degradations), these artifact cues become obsolete and harder to detect. To address this, we propose Real-centric Envelope Modeling (REM), a new paradigm that shifts detection from learning generator artifacts to modeling the robust distribution of real images. REM introduces feature-level perturbations in self-reconstruction to generate near-real samples, and employs an envelope estimator with cross-domain consistency to learn a boundary enclosing the real image manifold. We further build RealChain, a comprehensive benchmark covering both open-source and commercial generators with simulated real-world degradation. Across eight benchmark evaluations, REM achieves an average improvement of 7.5% over state-of-the-art methods, and notably maintains exceptional generalization on the severely degraded RealChain benchmark, establishing a solid foundation for synthetic image detection under real-world conditions. The code and the RealChain benchmark will be made publicly available upon acceptance of the paper.

cs.CV↗

OceanForecastBench: A Benchmark Dataset for Data-Driven Global Ocean Forecasting

Global ocean forecasting aims to predict key ocean variables such as temperature, salinity, and currents, which is essential for understanding and describing oceanic phenomena. In recent years, data-driven deep learning-based ocean forecast models, such as XiHe, WenHai, LangYa and AI-GOMS, have demonstrated significant potential in capturing complex ocean dynamics and improving forecasting efficiency. Despite these advancements, the absence of open-source, standardized benchmarks has led to inconsistent data usage and evaluation methods. This gap hinders efficient model development, impedes fair performance comparison, and constrains interdisciplinary collaboration. To address this challenge, we propose OceanForecastBench, a benchmark offering three core contributions: (1) A high-quality global ocean reanalysis data over 28 years for model training, including 4 ocean variables across 23 depth levels and 4 sea surface variables. (2) A high-reliability satellite and in-situ observations for model evaluation, covering approximately 100 million locations in the global ocean. (3) An evaluation pipeline and a comprehensive benchmark with 6 typical baseline models, leveraging observations to evaluate model performance from multiple perspectives. OceanForecastBench represents the most comprehensive benchmarking framework currently available for data-driven ocean forecasting, offering an open-source platform for model development, evaluation, and comparison. The dataset and code are publicly available at: https://github.com/Ocean-Intelligent-Forecasting/OceanForecastBench.

cs.LG↗