SearcharxivSearch

arXiv subjects

Yi Han

Publications and source records attributed to Yi Han.

At least 19 recordsLinked to original sources

Right Frame, Wrong Rule: Cultural Cues Expose the Financial Knowledge Gap They Were Meant to Close

When a question has valid answers under different normative frameworks, a language model must decide which framework to use and whether it can answer correctly within it. We call this setting normative pluralism and study it in Islamic finance using a four-choice taxonomy that separates framework selection from within-framework correctness. This separation reveals the stereotype trap: a cultural cue steers a model toward one framework, but the model selects an incorrect answer within that framework. Across twelve models, two languages, and fifty demographic signals, cultural cues change framework selection and reveal substantial differences in accuracy, especially among non-frontier models. Under the strongest signal, large open-weight models select the Islamic framework 97% of the time. A two-choice evaluation would report near-perfect alignment, although 57--66% of those selections are incorrect. These findings motivate, but do not directly test, the competence-conditioned routing hypothesis: models may favor frameworks where they are more accurate, while cultural cues may expose framework-specific competence gaps.

cs.CL

The circular law for non-Hermitian random band matrices: optimal bandwidth, periodic profile and discrete law

We consider non-Hermitian random band matrices with growing bandwidth and study convergence of their empirical spectral distributions to the circular law. Let $N$ denote the matrix size and $W_N$ the bandwidth. From a universality perspective, it is conjectured that the circular law holds whenever $W_N\to\infty$. Previous results have mainly required $W_N\gg N^{1/2}$, with the principal exception of \cite{Han2511}, which proves the $W_N\to\infty$ circular law for an open-boundary block-tridiagonal model. Here we prove the circular law at this optimal threshold for several periodic models and under the near-optimal condition $W_N\gg\log N$ for a genuinely discrete model. For the periodic hard-indicator profile and its uniform and polynomially tapered generalizations, we prove the circular law under bounded-density and finite-third-moment assumptions whenever $W_N\to\infty$. At the same threshold, we prove the circular law for continuous, integrable profiles locally bounded below on every finite interval, including exponentially and Gaussian decaying profiles, with circular complex Gaussian entries. For the periodic full-block model, we prove the circular law under bounded-density and finite-third-moment assumptions whenever $W_N\to\infty$. The finite-third-moment assumption in the bounded-density results can be weakened to a finite $(2+\alpha)$-moment assumption. For the same full-block model without a density assumption, we prove the circular law for centered variance-one real subgaussian atoms when $W_N\gg\log N$. The proof uses compositions of random transfer operators. An established high-band circular-law input on a short auxiliary ring calibrates the full exterior coefficient norm, and a boundary-uniform local comparison lifts this calibration to target rings even when $N/W_N$ is arbitrarily large.

math.PR

Agent Safety Should Be a Runtime Contract

The dominant paradigm treats AI safety as a property to be instilled during model training via RLHF, DPO, or Constitutional AI. We argue this is structurally insufficient for autonomous agents that execute code, mutate files, send messages, and modify databases. Agent safety should be a runtime contract enforced by the harness, and the contract has two complementary faces. The preventive face blocks dangerous actions before they happen via sandboxes, permission gates, output filters, and trajectory monitors. The evidential face requires verifiable proof that good actions actually happened, gating task submission on hard evidence such as test runs, log captures, file diffs, and citation grounding. We ground the position in four lines of public evidence, with row-level protocols and data released in the supplementary JSON files: a survey of 52 documented AI-agent and LLM safety incidents, a false-completion audit with 31 non-contested core cases plus one disputed illustrative case, a trajectory-schema audit of 12 public agent systems and harnesses, and a title-level audit of all 28,560 papers accepted at NeurIPS, ICML, and ICLR 2023-2025 showing a pooled 8-12x imbalance between training-time and deployment-time publication. Two prior communities that needed to enforce safety, computer security and the experimental sciences, converged on runtime contracts with both preventive and evidential elements; agentic AI is now under the same pressure. We formalize an Agent Trajectory Schema and Evidence Chain, state a compositional gating proposition based on standard monitor composition, and outline a research agenda. The right unit of safety in agentic AI is the trajectory-with-checkable-evidence, not the model.

cs.CR

Stochastic heat equation with nondegenerate H\"older diffusion coefficient: uniqueness below the three-fourth threshold

We consider stochastic heat equation (SHE) defined on 1-d torus $\mathbb{T}$ of the form $$\partial_t u=\Delta u+g(u)\dot{W},$$where $\dot{W}$ is a space-time white noise and $g$ is a real-valued function which is uniformly elliptic (i.e., $|g|$ is uniformly bounded away from 0), and is globally $\beta$-Holder continuous for some $\beta\in(0,1)$. We prove that weak uniqueness holds as long as $\beta>\frac{2}{3}$. The same uniqueness holds for vector-valued solutions where the coefficient $G$ has the same dimension as the white noise. Previously, uniqueness of solutions to the SHE with Holder diffusion coefficient was only established for $\beta>\frac{3}{4}$ via a Yamada Watanabe argument by Mytnik and Perkins (arxiv:0809.0248) without assuming $g$ is nonzero. And when $\beta<\frac{3}{4}$, Mueller, Mytnik and Perkins (arXiv:1201.2767) constructed a non-unique SPDE example satisfying $g(0)=0$. A later generalized coupling argument for nondegenerate $g$ also stopped at the same threshold $\frac{3}{4}$. Our result shows that uniform ellipticity of $g$ restores uniqueness to SHEs in the Holder regime where the same SHE with non-elliptic $g$ and the same Holder regularity are often non-unique in law. This constitutes the first general class of SHE weak uniqueness results in the $\beta\in(\frac{2}{3},\frac{3}{4}]$ regime.

math.PR

Are Large Language Models Suitable for Graph Computation? Progress and Prospects

Large language models (LLMs) have been increasingly explored for graph computation, where tasks require reasoning over structured relationships and algorithmic operations. Yet, it remains unclear when LLMs can reliably support such computation and how they should be incorporated into graph-solving pipelines. Existing surveys at the intersection of LLMs and graphs primarily focus on graph learning, text-attributed graphs, or graph-language modeling. To bridge this gap, we provide a comprehensive review of LLMs for graph computation through a role-based taxonomy. Specifically, we identify two major paradigms: i) LLMs as executors, where models directly solve graph tasks from graph descriptions and instructions; and ii) LLMs as planners, where models formulate problems, decompose reasoning steps, and invoke external tools or agents for execution. Based on this taxonomy, we analyze the strengths and limitations of current methods. Our review indicates that LLMs are promising for simple, small-scale tasks, but remain unreliable for large-scale and exactness-demanding tasks. Finally, we summarize available datasets and suggest four future directions.

cs.CL

Brown measure convergence for the spectrum of polynomials in Ginibre matrices

Fix a multivariate polynomial $\mathfrak{p}$ in $n$ non-commuting variables of arbitrary degree, and consider $n$ independent $N\times N$ complex Ginibre matrices $X_1^N,\cdots,X_n^N$. We prove that the empirical spectral distribution of $P^N=\mathfrak{p}(X_1^N,\cdots,X_n^N)$ converges as $N$ tends to infinity to the so-called Brown measure of $\mathfrak{p}$ evaluated at free circular variables. For polynomials of degree at most 2, the convergence was proven by Cook, Guionnet, and Husson \cite{cook2022spectrum}, and we prove that the convergence in fact holds for polynomials $\mathfrak{p}$ of any degree. The main step in the proof is a least singular value lower bound for $P^N-z$ for almost all complex shifts $z$, and we prove this via a least singular value lower bound for a wide class of tensorized Ginibre matrices of finite type with a deterministic shift, which is of independent interest. We further show that the Brown measure convergence holds beyond Gaussians: the same convergence holds when the entry law has mean 0, variance 1, bounded density on $\mathbb{C}$ and finite moments of all orders.

math.PR

Herculean: An Agentic Benchmark for Financial Intelligence

As AI agents improve, the central question is no longer whether they can solve isolated well-defined financial tasks, but whether they can reliably carry out financial professional work. Existing financial benchmarks offer only a partial view of this ability, as they primarily evaluate static competencies such as question answering, retrieval, summarization, and classification. We introduce Herculean, the first skilled benchmark for agentic financial intelligence spanning four representative workflows, including Trading, Hedging, Market Insights, and Auditing. Each workflow is instantiated as a standardized MCP-based skill environment with its own tools, interaction dynamics, constraints, and success criteria, enabling consistent end-to-end assessment of heterogeneous agent systems. Across frontier agents, we find agents perform relatively well on Trading and Market Insights, but struggle substantially on Hedging and Auditing, where long-horizon coordination, state consistency, and structured verification are critical. Overall, our results point to a key gap in current agents in turning financial reasoning into dependable workflow execution in high-stakes financial workflows.

cs.AI

Glauber dynamics for random field Ising models on bounded degree graphs and MLSI

We study the ferromagnetic random field Ising model (RFIM) on a graph $G=(V,E)$ having maximal degree $\Delta$, where the external field at each vertex is an i.i.d. random variable. When the random field distribution is sufficiently anti-concentrated, we prove that with high probability over the quenched randomness of the external field, the Glauber dynamics of this RFIM mixes in polynomial time as a consequence of a Poincar\'e inequality. This model is relevant to the Griffiths phase where the correlations decay exponentially fast in expectation over the quenched random field, but contraction does not hold point-wise due to the existence of weak fields that lead to low-temperature behavior. Previously, fast mixing of Glauber dynamics under large disorder was only proven on the integer lattice, and for RFIM on general graphs, only a sampling algorithm based on self-avoiding walks was known. Under a further technical condition that the random fields are bounded, we prove a modified log-Sobolev inequality for the Glauber dynamics. When the random field is weaker but still satisfies weak spatial mixing (exponential decay of correlations from boundary to bulk) in expectation, and the graph has at most $\alpha$-stretched exponential growth for some $\alpha<1$, then we prove a weak Poincar\'e inequality holds, which gives rise to a polynomial time sampling algorithm based on Glauber dynamics with warm start. The latter result was previously proven for the integer lattice, and we extend its scope to graphs with only a volume growth condition without assuming a local geometry.

math.PR

Deployment of AI-Assisted Interventions: Capacity Constraints and Noisy Compliance

AI tools increasingly guide targeted interventions in healthcare, education, and recruiting. Algorithms score individuals, trigger outreach to those above a threshold (e.g., high-risk or high-value), and encourage them to request service; then providers deliver service to those who request. Standard practice sets the threshold and selects the algorithm to maximize predictive accuracy, assuming that better predictions yield better outcomes. We show that this approach is suboptimal when limited service capacity and probabilistic behavioral responses influence who receives service. In such settings, the optimal score threshold must balance two effects: ensuring all capacity is filled (utilization) and ensuring high-value individuals are served despite competition between requests (cannibalization). We characterize the optimal threshold and prove that policies based solely on predictive accuracy are generally suboptimal. Further, because optimal thresholds vary with service capacity, algorithm selection metrics like AUC, which weight all thresholds equally, are misaligned with operational performance. We introduce a new metric--Operational AUC (OpAUC)--and show it leads to optimal algorithm selection. Finally, we conduct a case study on sepsis early warning data and illustrate the magnitude of improvement that can be achieved from improved threshold and algorithm selection.

stat.ME

Quadratic shift-and-stack for Ground-Based Optical Detection of Faint Cislunar Objects

Detecting faint objects in cislunar space using ground-based optical telescopes is difficult because of their low brightness, strong lunar background, and complex, nonlinear apparent motion. Traditional shift-and-stack techniques based on linear motion assumption suffer signal trailing loss due to significant nonlinear motion during long integrations, thus producing a degraded signal-to-noise ratio (SNR). In this paper, we first derive a theoretical criterion based on the point spread function to determine the maximum applicable integration time for linear-motion stacking. We then propose a quadratic shift-and-stack (QSS) method to correct for the first-order nonlinear motion, namely the angular acceleration of cislunar targets. Simulations of typical cislunar orbits verify this theoretical criterion and show that the QSS method significantly improves SNR from stacking and can enhance the detection limit by up to 1 stellar magnitude compared with the linear-motion stacking method. Furthermore, tests using observational data of the cislunar object Tiandu-1 confirm that while linear stacking degrades after a 29-minute integration due to trajectory curvature, the QSS method achieves continuous SNR improvement over a 46-minute integration, outperforming the peak SNR of the linear method by 31%.

astro-ph.IM

Can LLM Agents Be CFOs? Benchmarking Long-Horizon Resource Allocation in an Uncertain Enterprise Environment

Large language model (LLM) agents are increasingly tested on complex tasks, but their ability to allocate scarce resources over long horizons remains unclear. Unlike reactive tasks with immediate feedback, this setting requires agents to make binding commitments under partial observability, delayed consequences, hard resource budgets, and shifting dynamics. We introduce EnterpriseArena, a 132-month CFO simulator that evaluates long-horizon resource allocation under uncertainty in a FinTech lending firm. Agents must manage liquidity, close books, gather costly signals, and request equity or debt financing across changing macroeconomic regimes. The simulator is built from transformed firm-level financial data, anonymized business documents, decade-scale macroeconomic and industry signals, and expert-validated operating rules. Experiments across 23 LLMs and four agent frameworks show that current agents remain far from robust: only 15.4% of trials survive the full horizon, larger models do not reliably outperform smaller ones, and failures cascade across observation, action timing, and capital sizing. These findings establish long-horizon resource allocation under uncertainty as a distinct capability gap for LLM agents.

cs.AI

SaPaVe: Towards Active Perception and Manipulation in Vision-Language-Action Models for Robotics

Active perception and manipulation are crucial for robots to interact with complex scenes. Existing methods struggle to unify semantic-driven active perception with robust, viewpoint-invariant execution. We propose SaPaVe, an end-to-end framework that jointly learns these capabilities in a data-efficient manner. Our approach decouples camera and manipulation actions rather than placing them in a shared action space, and follows a bottom-up training strategy: we first train semantic camera control on a large-scale dataset, then jointly optimize both action types using hybrid data. To support this framework, we introduce ActiveViewPose-200K, a dataset of 200k image-language-camera movement pairs for semantic camera movement learning, and a 3D geometry-aware module that improves execution robustness under dynamic viewpoints. We also present ActiveManip-Bench, the first benchmark for evaluating active manipulation beyond fixed-view settings. Extensive experiments in both simulation and real-world environments show that SaPaVe outperforms recent vision-language-action models such as GR00T N1 and \(\pi_0\), achieving up to 31.25\% higher success rates in real-world tasks. These results show that tightly coupled perception and execution, when trained with decoupled yet coordinated strategies, enable efficient and generalizable active manipulation. Project page: https://lmzpai.github.io/SaPaVe

cs.RO

The USTC-NERCSLIP Systems for the CHiME-9 MCoRec Challenge

This report details our submission to the CHiME-9 MCoRec Challenge on recognizing and clustering multiple concurrent natural conversations within indoor social settings. Unlike conventional meetings centered on a single shared topic, this scenario contains multiple parallel dialogues--up to eight speakers across up to four simultaneous conversations--with a speech overlap rate exceeding 90%. To tackle this, we propose a multimodal cascaded system that leverages per-speaker visual streams extracted from synchronized 360 degree video together with single-channel audio. Our system improves three components of the pipeline by leveraging enhanced audio-visual pretrained models: Active Speaker Detection (ASD), Audio-Visual Target Speech Extraction (AVTSE), and Audio-Visual Speech Recognition (AVSR). The AVSR module further incorporates Whisper and LLM techniques to boost transcription accuracy. Our best single cascaded system achieves a Speaker Word Error Rate (WER) of 32.44% on the development set. By further applying ROVER to fuse outputs from diverse front-end and back-end variants, we reduce Speaker WER to 31.40%. Notably, our LLM-based zero-shot conversational clustering achieves a speaker clustering F1 score of 1.0, yielding a final Joint ASR-Clustering Error Rate (JACER) of 15.70%.

eess.AS

Conv-FinRe: A Conversational and Longitudinal Benchmark for Utility-Grounded Financial Recommendation

Most recommendation benchmarks evaluate how well a model imitates user behavior. In financial advisory, however, observed actions can be noisy or short-sighted under market volatility and may conflict with a user's long-term goals. Treating what users chose as the sole ground truth, therefore, conflates behavioral imitation with decision quality. We introduce Conv-FinRe, a conversational and longitudinal benchmark for stock recommendation that evaluates LLMs beyond behavior matching. Given an onboarding interview, step-wise market context, and advisory dialogues, models must generate rankings over a fixed investment horizon. Crucially, Conv-FinRe provides multi-view references that distinguish descriptive behavior from normative utility grounded in investor-specific risk preferences, enabling diagnosis of whether an LLM follows rational analysis, mimics user noise, or is driven by market momentum. We build the benchmark from real market data and human decision trajectories, instantiate controlled advisory conversations, and evaluate a suite of state-of-the-art LLMs. Results reveal a persistent tension between rational decision quality and behavioral alignment: models that perform well on utility-based ranking often fail to match user choices, whereas behaviorally aligned models can overfit short-term noise. The dataset is publicly released on Hugging Face, and the codebase is available on GitHub.

cs.AI

MIRROR: Manifold Ideal Reference ReconstructOR for Generalizable AI-Generated Image Detection

High-fidelity generative models have narrowed the perceptual gap between synthetic and real images, posing serious threats to media security. Most existing AI-generated image (AIGI) detectors rely on artifact-based classification and struggle to generalize to evolving generative traces. In contrast, human judgment relies on stable real-world regularities, with deviations from the human cognitive manifold serving as a more generalizable signal of forgery. Motivated by this insight, we reformulate AIGI detection as a Reference-Comparison problem that verifies consistency with the real-image manifold rather than fitting specific forgery cues. We propose MIRROR (Manifold Ideal Reference ReconstructOR), a framework that explicitly encodes reality priors using a learnable discrete memory bank. MIRROR projects an input into a manifold-consistent ideal reference via sparse linear combination, and uses the resulting residuals as robust detection signals. To evaluate whether detectors reach the "superhuman crossover" required to replace human experts, we introduce the Human-AIGI benchmark, featuring a psychophysically curated human-imperceptible subset. Across 14 benchmarks, MIRROR consistently outperforms prior methods, achieving gains of 2.1% on six standard benchmarks and 8.1% on seven in-the-wild benchmarks. On Human-AIGI, MIRROR reaches 89.6% accuracy across 27 generators, surpassing both lay users and visual experts, and further approaching the human perceptual limit as pretrained backbones scale. The code is publicly available at: https://github.com/349793927/MIRROR

cs.CV

RoboBrain 2.5: Depth in Sight, Time in Mind

We introduce RoboBrain 2.5, a next-generation embodied AI foundation model that advances general perception, spatial reasoning, and temporal modeling through extensive training on high-quality spatiotemporal supervision. Building upon its predecessor, RoboBrain 2.5 introduces two major capability upgrades. Specifically, it unlocks Precise 3D Spatial Reasoning by shifting from 2D pixel-relative grounding to depth-aware coordinate prediction and absolute metric constraint comprehension, generating complete 3D manipulation traces as ordered keypoint sequences under physical constraints. Complementing this spatial precision, the model establishes Dense Temporal Value Estimation that provides dense, step-aware progress prediction and execution state understanding across varying viewpoints, producing stable feedback signals for downstream learning. Together, these upgrades extend the framework toward more physically grounded and execution-aware embodied intelligence for complex, fine-grained manipulation. The code and checkpoints are available at project website: https://superrobobrain.github.io

cs.RO

CARLA-Round: A Multi-Factor Simulation Dataset for Roundabout Trajectory Prediction

Accurate trajectory prediction of vehicles at roundabouts is critical for reducing traffic accidents, yet it remains highly challenging due to their circular road geometry, continuous merging and yielding interactions, and absence of traffic signals. Developing accurate prediction algorithms relies on reliable, multimodal, and realistic datasets; however, such datasets for roundabout scenarios are scarce, as real-world data collection is often limited by incomplete observations and entangled factors that are difficult to isolate. We present CARLA-Round, a systematically designed simulation dataset for roundabout trajectory prediction. The dataset varies weather conditions (five types) and traffic density levels (spanning Level-of-Service A-E) in a structured manner, resulting in 25 controlled scenarios. Each scenario incorporates realistic mixtures of driving behaviors and provides explicit annotations that are largely absent from existing datasets. Unlike randomly sampled simulation data, this structured design enables precise analysis of how different conditions influence trajectory prediction performance. Validation experiments using standard baselines (LSTM, GCN, GRU+GCN) reveal traffic density dominates prediction difficulty with strong monotonic effects, while weather shows non-linear impacts. The best model achieves 0.312m ADE on real-world rounD dataset, demonstrating effective sim-to-real transfer. This systematic approach quantifies factor impacts impossible to isolate in confounded real-world datasets. Our CARLA-Round dataset is available at https://github.com/Rebecca689/CARLA-Round.

cs.CV

Adaptive Causal Coordination Detection for Social Media: A Memory-Guided Framework with Semi-Supervised Learning

Detecting coordinated inauthentic behavior on social media remains a critical and persistent challenge, as most existing approaches rely on superficial correlation analysis, employ static parameter settings, and demand extensive and labor-intensive manual annotation. To address these limitations systematically, we propose the Adaptive Causal Coordination Detection (ACCD) framework. ACCD adopts a three-stage, progressive architecture that leverages a memory-guided adaptive mechanism to dynamically learn and retain optimal detection configurations for diverse coordination scenarios. Specifically, in the first stage, ACCD introduces an adaptive Convergent Cross Mapping (CCM) technique to deeply identify genuine causal relationships between accounts. The second stage integrates active learning with uncertainty sampling within a semi-supervised classification scheme, significantly reducing the burden of manual labeling. The third stage deploys an automated validation module driven by historical detection experience, enabling self-verification and optimization of the detection outcomes. We conduct a comprehensive evaluation using real-world datasets, including the Twitter IRA dataset, Reddit coordination traces, and several widely-adopted bot detection benchmarks. Experimental results demonstrate that ACCD achieves an F1-score of 87.3\% in coordinated attack detection, representing a 15.2\% improvement over the strongest existing baseline. Furthermore, the system reduces manual annotation requirements by 68\% and achieves a 2.8x speedup in processing through hierarchical clustering optimization. In summary, ACCD provides a more accurate, efficient, and highly automated end-to-end solution for identifying coordinated behavior on social platforms, offering substantial practical value and promising potential for broad application.

cs.AI