SearcharxivSearch

arXiv subjects

Xuan Yao

Publications and source records attributed to Xuan Yao.

At least 19 recordsLinked to original sources

DelistBench: Evaluating Search-Enabled LLMs for Auditable Corporate-Event Database Completion

Financial institutions need an independent way to detect missing, stale, and misclassified corporate-event records in vendor databases. We introduce Search-to-Record, a database-assurance task in which search-enabled large language models reconstruct institution-defined event records from public sources for a known security universe and historical cutoff, and DelistBench, a 1,200-record benchmark for security-level delisting announcements. We evaluate five models in paired closed-book and web-enabled conditions. Web access raises announcement-date accuracy within seven days by 34.0 to 48.0 percentage points and event-status accuracy by approximately 2.8 to 21.7 points; the best system achieves 81.5% overall joint accuracy within seven days. Economy web systems achieve 75.9-78.3% overall joint accuracy within seven days at 4.5-6.6% of the API cost of the most expensive web system. Risk-based triage identifies low-error subsets, although the highest-coverage operating point still sends 27.3% of the balanced test set to review. The evaluation identifies web retrieval as the main source of timing gains and shows that low-cost systems can approach the best system's accuracy. Together, Search-to-Record, DelistBench, and the evaluation provide concrete deployment guidance: calibrate triage to local event prevalence and market mix, preserve positive-event recall, and route positive and ambiguous cases to targeted review.

cs.CL

SC$^{2}$-WM: A Self-Correcting World Model with Closed-Loop Feedback for Vision-and-Language Navigation in Continuous Environments

Vision-and-Language Navigation in Continuous Environments (VLN-CE) requires agents to make fine-grained navigation decisions under partial observability. However, most existing methods rely on open-loop execution, lacking mechanisms to detect and correct internal state drift during inference. We propose SC$^{2}$-WM, a self-correcting world model framework that introduces internal feedback for closed-loop decision making in VLN-CE. Our method derives feedback from world-model foresight to perform state-level plan refinement before action execution. To handle challenging scenarios, we further introduce conditional world-aware adaptation, which enables model-level correction by selectively updating the world model at test time when feedback indicates model capacity insufficiency. Experiments on standard VLN-CE benchmarks demonstrate improved navigation robustness and generalization. Our code is available at https://github.com/sunrise-ikun/SC2_WM.

cs.RO

Who Wins Where? Conformal Model Comparison for Local Superiority

Standard model comparison is global, aggregating losses across the covariate space to declare a single winner. This can obscure heterogeneous performance, where different models are preferable in different regions. We introduce conformalized local model comparison, a split-sample framework for constructing calibrated local best-model maps. Given a model comparison score, such as the difference between two squared losses, the method uses three disjoint splits to fit competing models, estimate local centers and scales from out-of-sample scores, and conformally calibrate residual uncertainty. At a target point, the procedure declares a local winner only when a one-sided conformal bound excludes a tie, with the score's sign determining the favored model. We prove finite-sample marginal control for one-sided erroneous declarations on the realized future comparison score, establish pointwise consistency of the localized mean-score estimator away from tie boundaries, show that aggregate comparison can disagree sharply with the prevalence of local superiority, and derive a squared-loss bias--variance decomposition that clarifies how model structure affects local wins. Synthetic and real-data experiments show that the method recovers heterogeneous winner regions, abstains under uncertainty, and yields higher conditional gain than global selection.

cs.LG

A Positive Mass Theorem for Continuous Metrics

Let $g$ be a continuous metric on $\mathbb R^3$ which is asymptotically flat in the sense that $\vert g_{ij}(x) - \delta_{ij}\vert = O(\vert x\vert^{-\tau})$ for some $\tau > \frac{1}{2}$. Further assume that $g$ can be uniformly approximated on compact sets by smooth metrics with almost non-negative scalar curvature. For such a metric $g$, we define a synthetic ADM mass $m(g)$ using harmonic functions. The harmonic mass $m(g)$ coincides with the usual ADM mass whenever $g$ is smooth and decays rapidly enough that the latter is defined. The harmonic mass can also be computed as a limit of the $C^0$ local mass introduced by Burkhardt-Guim. Our main result is a positive mass theorem: the harmonic mass satisfies $m(g)\geq 0$ and if $m(g) = 0$ then $g$ is flat.

math.DG

Rigidity in the Positive Mass Theorem with $C^0$ Decay

Let $g$ be a smooth metric on $\mathbb R^3$ with non-negative scalar curvature. We show that if $g$ satisfies $\vert g(x)-g_{\text{euc}}(x)\vert = O(\vert x\vert^{-1-\tau})$ for some $\tau > 0$ then $g$ must be flat.

math.DG

Scalar curvature under weak limits of manifolds

We show that scalar curvature lower bounds are preserved under certain weak convergence of smooth three manifolds to a smooth limit. More precisely, suppose that $M_k$ and $M$ are smooth, closed, Riemannian three manifolds. Assume that there are smooth, surjective, $\lambda_k$-Lipschitz maps $f_k\colon M_k \to M$ and that $\text{Vol}(M_k)\to \text{Vol}(M)$ and $\lambda_k\to 1$. Then if each $M_k$ has scalar curvature bounded below by $\kappa$ so does $M$. This result answers questions of Gromov, Sormani, Allen, and others. The proof relies on a delicate comparison between $\mu$-bubbles in $M_k$ and $\mu$-bubbles in $M$.

math.DG

Quantification of $C^0$ Convergence in Dimension Three

We address Gromov's Quantification of $C^0$ Convergence Conjecture in dimension three. Let $B$ be the unit ball in $\mathbb R^3$. Let $g$ and $g_0$ be smooth metrics on $B$. We prove there are constants $C$ and $\epsilon_0$ depending only on $g_0$ so that \[ \inf_{x\in B} R_g(x) \leq R_{g_0}(0) + C \|g-g_0\|_{C^0}^{1/2} \] provided $\|g-g_0\|_{C^0}\leq \epsilon_0$. We also construct examples to show that the exponent $1/2$ is sharp. This explicitly quantifies the fact that scalar curvature lower bounds are preserved under $C^0$ convergence of metrics. When $g_0$ is merely $C^2$ we prove a related estimate with a slightly weaker rate, and when $g_0$ has rotational symmetry we prove a related estimate with a stronger linear rate. To prove these results, we use harmonic functions to define a local quantity that detects the scalar curvature. Then we use classical elliptic PDE estimates to show that this quantity is stable under $C^0$ perturbations of the metric. As a further application of this method, we give a partial answer to a question of Gromov on the preservation of scalar curvature lower bounds for metrics that are converging in measure.

math.DG

Capillary minimal slicing and scalar curvature rigidity

We develop minimal slicing via capillary hypersurfaces to understand positive scalar curvature metric on manifolds with boundary. The method provides rigidity statements once the regularity of minimizers of capillary area functional holds. In particular, in dimension $4$, we prove following comparison and rigidity statement: given a compact Riemannian $4$-manifold $(M^4,g)$ with a mean convex boundary whose boundary is diffeomorphic to boundary of a connected convex domain in $\mathbb R^4$, if the scalar curvature is non-negative and the scaled mean curvature comparison holds along the boundary, then $M$ is isometric to the Euclidean domain.

math.DG

FinDeepForecast: A Live Multi-Agent System for Benchmarking Deep Research Agents in Financial Forecasting

Deep Research (DR) Agents powered by advanced Large Language Models (LLMs) have fundamentally shifted the paradigm for completing complex research tasks. Yet, a comprehensive and live evaluation of their forecasting performance on real-world, research-oriented tasks in high-stakes domains (e.g., finance) remains underexplored. We introduce FinDeepForecast, the first live, end-to-end multi-agent system for automatically evaluating DR agents by continuously generating research-oriented financial forecasting tasks. This system is equipped with a dual-track taxonomy, enabling the dynamic generation of recurrent and non-recurrent forecasting tasks at both corporate and macro levels. With this system, we generate FinDeepForecastBench, a weekly evaluation benchmark over a ten-week horizon, encompassing 8 global economies and 1,314 listed companies, and evaluate 13 representative methods. Extensive experiments show that, while DR agents consistently outperform strong baselines, their performance still falls short of genuine forward-looking financial reasoning. We expect the proposed FinDeepForecast system to consistently facilitate future advancements of DR agents in research-oriented financial forecasting tasks. The benchmark and leaderboard are publicly available on the OpenFinArena Platform.

cs.MA

Stable Bernstein Problem in certain positively curved manifolds

We formulate stable Bernstein type theorems in certain positively curved ambient manifolds. In all dimensions, we prove that for any complete Riemannian manifold $(X^{n+1},g)$, if the Ricci curvature is non-negative and it positive BiRic curvature with $\alpha$-decay, then any complete, two-sided, stable minimal immersion must be totally geodesic and $\text{Ric}(\nu,\nu)$ vanish along the minimal immersion. For $4\leq n+1\leq 6$, we prove that the result still holds if $(X^{n+1},g)$ has uniform positive $3$-intermediate curvature and non-negative $(n-1)$-Ricci curvature, which generalize Chodosh-Li-Stryker's result \cite{chodosh2024complete} for $n+1=4$ to higher dimensions. As an immediate corollary, we show that, in all dimensions, for a complete Riemannian manifold $(X^{n+1},g)$, if it has uniform positive Ricci curvature and non-negative $(n-1)$-Ricci curvature then there is no (not necessarily) complete, two-sided, stable minimal immersion in $(X^{n+1},g)$.

math.DG

FinDeepResearch: Evaluating Deep Research Agents in Rigorous Financial Analysis

Deep Research (DR) agents, powered by advanced Large Language Models (LLMs), have recently garnered increasing attention for their capability in conducting complex research tasks. However, existing literature lacks a rigorous and systematic evaluation of DR Agent's capabilities in critical research analysis. To address this gap, we first propose HisRubric, a novel evaluation framework with a hierarchical analytical structure and a fine-grained grading rubric for rigorously assessing DR agents' capabilities in corporate financial analysis. This framework mirrors the professional analyst's workflow, progressing from data recognition to metric calculation, and finally to strategic summarization and interpretation. Built on this framework, we construct a FinDeepResearch benchmark that comprises 64 listed companies from 8 financial markets across 4 languages, encompassing a total of 15,808 grading items. We further conduct extensive experiments on the FinDeepResearch using 16 representative methods, including 6 DR agents, 5 LLMs equipped with both deep reasoning and search capabilities, and 5 LLMs with deep reasoning capabilities only. The results reveal the strengths and limitations of these approaches across diverse capabilities, financial markets, and languages, offering valuable insights for future research and development. The benchmark and evaluation code is publicly available at https://OpenFinArena.com/.

cs.CL

Evaluating Large Language Models for Financial Reasoning: A CFA-Based Benchmark Study

The rapid advancement of large language models presents significant opportunities for financial applications, yet systematic evaluation in specialized financial contexts remains limited. This study presents the first comprehensive evaluation of state-of-the-art LLMs using 1,560 multiple-choice questions from official mock exams across Levels I-III of CFA, most rigorous professional certifications globally that mirror real-world financial analysis complexity. We compare models distinguished by core design priorities: multi-modal and computationally powerful, reasoning-specialized and highly accurate, and lightweight efficiency-optimized. We assess models under zero-shot prompting and through a novel Retrieval-Augmented Generation pipeline that integrates official CFA curriculum content. The RAG system achieves precise domain-specific knowledge retrieval through hierarchical knowledge organization and structured query generation, significantly enhancing reasoning accuracy in professional financial certification evaluation. Results reveal that reasoning-oriented models consistently outperform others in zero-shot settings, while the RAG pipeline provides substantial improvements particularly for complex scenarios. Comprehensive error analysis identifies knowledge gaps as the primary failure mode, with minimal impact from text readability. These findings provide actionable insights for LLM deployment in finance, offering practitioners evidence-based guidance for model selection and cost-performance optimization.

cs.CL

NavMorph: A Self-Evolving World Model for Vision-and-Language Navigation in Continuous Environments

Vision-and-Language Navigation in Continuous Environments (VLN-CE) requires agents to execute sequential navigation actions in complex environments guided by natural language instructions. Current approaches often struggle with generalizing to novel environments and adapting to ongoing changes during navigation. Inspired by human cognition, we present NavMorph, a self-evolving world model framework that enhances environmental understanding and decision-making in VLN-CE tasks. NavMorph employs compact latent representations to model environmental dynamics, equipping agents with foresight for adaptive planning and policy refinement. By integrating a novel Contextual Evolution Memory, NavMorph leverages scene-contextual information to support effective navigation while maintaining online adaptability. Extensive experiments demonstrate that our method achieves notable performance improvements on popular VLN-CE benchmarks. Code is available at https://github.com/Feliciaxyao/NavMorph.

cs.CV

On the topology of manifolds with positive intermediate curvature

We formulate a conjecture relating the topology of a manifold's universal cover with the existence of metrics with positive $m$-intermediate curvature. We prove the result for manifolds of dimension $n\in\{3,4,5\}$ and for most choices of $m$ when $n=6$. As a corollary, we show that a closed, aspherical 6-manifold cannot admit a metric with positive $4$-intermediate curvature.

math.DG

Bottom-Up Reputation Promotes Cooperation with Multi-Agent Reinforcement Learning

Reputation serves as a powerful mechanism for promoting cooperation in multi-agent systems, as agents are more inclined to cooperate with those of good social standing. While existing multi-agent reinforcement learning methods typically rely on predefined social norms to assign reputations, the question of how a population reaches a consensus on judgement when agents hold private, independent views remains unresolved. In this paper, we propose a novel bottom-up reputation learning method, Learning with Reputation Reward (LR2), designed to promote cooperative behaviour through rewards shaping based on assigned reputation. Our agent architecture includes a dilemma policy that determines cooperation by considering the impact on neighbours, and an evaluation policy that assigns reputations to affect the actions of neighbours while optimizing self-objectives. It operates using local observations and interaction-based rewards, without relying on centralized modules or predefined norms. Our findings demonstrate the effectiveness and adaptability of LR2 across various spatial social dilemma scenarios. Interestingly, we find that LR2 stabilizes and enhances cooperation not only with reward reshaping from bottom-up reputation but also by fostering strategy clustering in structured populations, thereby creating environments conducive to sustained cooperation.

cs.MA

Euclidean Domains with Nearly Maximal Yamabe Quotient

Let $\Omega$ be a smooth, bounded domain in $\mathbb R^3$ with connected boundary. It follows from work of Escobar that the Yamabe quotient of $\Omega$ is at most the Yamabe quotient of a ball, and equality holds if and only if $\Omega$ is a ball. We show that if equality almost holds then the following things are true: (i)$\Omega$ is diffeomorphic to a ball; (ii) There is a small number $\epsilon > 0$ such that $B(x,r) \subset \Omega \subset B(x,r(1+\epsilon))$; (iii) After suitable scaling, $\Omega$ is Gromov-Hausdorff close to the unit ball when considered as a metric space with its induced length metric. We also give a qualitative comparison between $Q$ and the coefficient of quasi-conformality studied in the theory of quasi-conformal maps.

math.DG

A Note on Scalar curvature comparison rigidity for compact domains

We prove a generalization of Gromov's conjecture on scalar curvature rigidity of convex polytopes to arbitrary convex Riemannian polytope type domains via harmonic spinors on convex domians with boundary condition constructed by Brendle. In particular, we prove a rigidity results on comparison of scalar curvature and scaled mean curvature on the boundary for any convex domain in Euclidean space, which is a parallel of Shi-Tam's results.

math.DG

Scalar curvature comparison and rigidity of $3$-dimensional weakly convex domains

For a compact Riemannian $3$-manifold $(M^{3}, g)$ with mean convex boundary which is diffeomorphic to a weakly convex compact domain in $\mathbb{R}^{3}$, we prove that if scalar curvature is nonnegative and the scaled mean curvature comparison $H^{2}g \ge H_{0}^{2} g_{Eucl}$ holds, then $(M,g)$ is flat. Our result is a smooth analog of Gromov's dihedral rigidity conjecture and an effective version of extremity results on weakly convex balls in $\mathbb R^3$. More generally, we prove the comparison and rigidity theorem for several classes of manifold with corners. Our proof uses capillary minimal surfaces with prescribed contact angle together with the construction of foliation with nonnegative mean curvature and with prescribed contact angles.

math.DG