SearcharxivSearch

arXiv subjects

Hui Xu

Publications and source records attributed to Hui Xu.

At least 19 recordsLinked to original sources

Cayley-tree pseudo-orbit tracing: period subgroups and relative geometry

We introduce Cayley-tree POTP, obtained by imposing the pseudo-orbit equations of a finitely generated group action only along a spanning tree of a Cayley graph. For zero-dimensional actions, we characterize this property by equicontinuity along normalized replacement paths; for subshifts, the criterion is expressed in the right-coset space of the common left-period subgroup. These criteria characterize virtual freeness and, for commensurated subgroup pairs, identifies Cayley-tree POTP of the coset full shift with relative quasi-tree geometry and a finite Bass--Serre decomposition. For infinite-index VFP pairs, Cayley-tree POTP of the coset full shift is equivalent to virtual cohomological codimension one, although ordinary POTP holds for every such shift. Finally, we prove that the strong topological Rokhlin property passes to finite-index overgroups. Consequently every finitely generated virtually free group has this property, answering the virtually cyclic case posed by Doucha, and Cayley-tree POTP is generic for its Cantor actions.

math.DS

Training Documents Reranker with Search Rubrics for Deep Research Agent

Retrieval systems help deep research agents generate high-quality answers by providing relevant documents. However, existing retrievers typically select documents through relevance matching, while individually well-matched top-$k$ documents may not form a \textit{set} that satisfies the complex information needs of an agent query (\eg, diverse, concise and authoritative documents). In this paper, we propose search-oriented rubrics that \textit{explicitly} define the requirements that high-quality document sets should satisfy for each agent query. Our search rubrics are organized into a hierarchical structure and synthesized using a powerful LLM. Based on these search rubrics, we further train a document reranker \textbf{RubricRanker} to select a high-quality subset from retrieved documents. We design a two-stage training framework that consists of rubrics-guided supervised fine-tuning and rubric-based reinforcement learning. Extensive experiments demonstrate that RubricRanker outperforms the strongest baseline by 2.6 points on four deep research benchmarks and generalizes well to five RAG benchmarks.

cs.IR

SEER: A Self-Grounded Evidence Interface for Controlled Spatial Relation Classification

Spatial relation questions require a model to identify the queried subject and object before comparing their layout. Yet a VLM can recognize both entities and still answer from the wrong instance or an ambiguous global view. We ask whether making query-specific evidence explicit can mitigate this failure and propose SEER (Self-grounded Evidence for Entity-Relation Reasoning), a training-free inference-time evidence interface for frozen VLMs. SEER hides candidate relations during pair localization, constructs a query-specific view with explicit subject/object roles, and retains the full image and sparse box geometry as complementary evidence. For relation-choice protocols with exact inverse support, an optional refinement swaps the entity roles and changes the forward decision only when exactly one visual state obeys the corresponding inverse relation. On an image-disjoint GQA-Train900 test frozen before model scoring, SEER pools to +3.94 [2.17,5.72] over Full; the gain remains positive under label-independent grounding-order counterbalancing and on the 535 rows whose entity names are unique. The unchanged protocol yields +4.35 to +11.79 on all 2,434 filtered EmbSpatial pair-relation questions across three models. Matched controls separate local refocus from role-explicit conditioning. These results establish query-specific evidence construction as the principal intervention, with reciprocal consistency as a smaller protocol-specific refinement.

cs.CV

Rigid Functions, IP-Systems, and Topological Mild Mixing

We study uniform rigidity and topological mild mixing through continuous observables. For a fixed sequence of times, the observables rigid along that sequence form a closed unital $T^{\pm1}$-invariant algebra and determine the maximal factor uniformly rigid along the prescribed sequence. We then give functional forms of the classical ${\rm SIP}^{*}$- and ${\rm IP}^{*}$-return-time criteria: a topological dynamical system is mildly mixing exactly when it has no nonconstant locally SIP-rigid observable, and in the minimal category the same property is equivalent to the absence of nonconstant locally IP-rigid observables. Finally, a locally IP-rigid observable yields a canonical orbit-name factor carrying marked local data. For fixed local data, the $T^{\pm 1}$-invariant core of the local rigidity algebra determines a uniformly rigid factor whenever the core is nontrivial.

math.DS

Adherence Semigroups and Density Finite-Sums Configurations

In this paper, we use the adherence semigroup to describe finite-sums configurations in topological dynamical systems. We establish a correspondence between exact finite sumsets and powers of adherence elements. As applications, we give dynamical formulations of density finite-sums results, characterize total minimality through the density of adherence-power orbits, and give a positive answer to the ultrafilter question in [Question~8.9, 6]. We also characterize the sets that belong to the common sum of two commuting nonprincipal ultrafilters.

math.DS

Beyond Relevance-Centric Retrieval: Rubric-Oriented Document Set Selection and Ranking

As large language models and AI agents become the primary consumers of search results, document set quality determines the upper bound of downstream generation. Yet existing evaluation systems remain confined to scoring documents independently and aggregating via nDCG, ignoring inter-document interactions (redundancy, conflict, complementarity) and unable to answer what makes one document set better than another. To address these issues, we propose a complete evaluate-diagnose-optimize framework. We design SetwiseEvalKit, a three-level, nine-dimension document set evaluation benchmark covering both short-form and long-form scenarios, comprising approximately 28K high-quality evaluation rubrics. We systematically evaluate 12 rerankers: even the best method achieves no more than 45% coverage, cross-document coordination dimensions are universally weak, and no single method maintains top performance across both settings. Building on this, we propose Rubric4Setwise, a training-free method that converts rubric-based evaluation criteria into document set selection signals, achieving the best downstream generation performance with fewer documents and search rounds. It is the only method that maintains state-of-the-art results across both scenarios, validating the effectiveness of closing the loop from evaluation to optimization.

cs.CL

EgoAERO: Learning Dexterous Manipulation from a Single Egocentric Video without Object Assets

Egocentric RGB-D videos offer a natural source of human dexterous manipulation demonstrations, but existing data is difficult to use for robot learning because object pose, geometry, and contact information are often missing or require pre-scanned object assets. We present EgoAERO, the first framework that learns dexterous manipulation from a single egocentric RGB-D human demonstration without object assets. EgoAERO reconstructs contact-consistent hand-object trajectories through asset-free object tracking and reconstruction, ego motion compensation, and adaptive contact optimization, then converts them into robot policies using two-stage residual learning. We further introduce an online quality assessment mechanism and construct EgoDex-R, a large-scale egocentric dataset with 4.3M RGB-D frames for dexterous policy learning. Simulation and real-world experiments show that EgoAERO enables single-demonstration dexterous manipulation and achieves downstream performance close to CAD-based reconstructions on HOI4D.

cs.RO

Identifying sensitivity-dominant parameters via active subspaces in reduced-order modeling of fluid dynamics

Reduced-order models (ROMs) are widely employed to describe complex system dynamics when simulations with full-order models (FOMs) are computationally prohibitive. This study presents POD-AS-PRS, a novel model-reduction framework based on the active subspaces (AS) technique, which performs dimensionality reduction in both the state and parameter spaces, enabling efficient and high-fidelity approximations of quantities of interest (QoI). The approach employs proper orthogonal decomposition (POD) to extract low-dimensional coefficients from CFD snapshots, which are inputs to a residual neural network (ResNet) with linear layers to learn their nonlinear mapping to QoI. Reverse-mode automatic differentiation (AD) is utilized to compute gradients with respect to the coefficients, enabling AS analysis to identify influential modes by shifting the analysis to the POD coefficient space, thereby achieving a dual-stage dimensionality reduction driven by QoI sensitivity rather than modal energy. A surrogate model is subsequently constructed using a polynomial response surface (PRS) based on AS-derived active variables, retaining only the highly influential POD coefficients to ensure accurate and efficient QoI reconstruction. The framework is validated on periodic and chaotic bluff-body flows, demonstrating high accuracy with few influential parameters, while AD-based gradients achieve a two-order-of-magnitude speed-up over finite-difference approximations. Sensitivity analysis further reveals that the influential coefficients are not necessarily proportional to modal energy, highlighting the critical flow structures. Consequently, POD-AS-PRS identifies a low-dimensional manifold of sensitivity-dominant parameters that govern the QoI, elucidating the essential flow structures and their coupling with control parameters, thereby enabling efficient and accurate QoI reconstruction.

physics.flu-dyn

Quasi-disjointness in topological dynamics

Motivated by Berg's notion of quasi-disjointness for ergodic systems, we introduce and investigate the concept of quasi-disjointness for minimal systems. Several equivalent characterizations are provided. We prove that quasi-disjointness is preserved under taking factors, proximal extensions, and group extensions. As a consequence, we establish that every minimal {\bf PI} system is quasi-disjoint from all minimal systems. In addition, some variant of quasi-disjointness, namely strong quasi-disjointness is also introduced and examined. Particularly, we prove that each {\bf AI} system is strongly quasi-disjoint from all minimal systems.

math.DS

Unique continuation inequalities for the Dunkl-Schr\"odinger equation via uncertainty principles

In this paper, we establish unique continuation inequalities at two time points for the Dunkl--Schr\"odinger equation. The proof is based on quantitative uncertainty principles for the Dunkl transform. In particular, we prove that pairs of (\varepsilon,k)-thin sets form strong annihilating pairs for the Dunkl transform, which yields quantitative unique continuation properties for solutions to the Dunkl--Schr\"odinger equation.

math.AP

Ada-MK: Adaptive MegaKernel Optimization via Automated DAG-based Search for LLM Inference

When large language models (LLMs) serve real-time inference in commercial online advertising systems, end-to-end latency must be strictly bounded to the millisecond range. Yet every token generated during the decode phase triggers thousands of kernel launches, and kernel launch overhead alone can account for 14.6% of end-to-end inference time. MegaKernel eliminates launch overhead and inter-operator HBM round-trips by fusing multiple operators into a single persistent kernel. However, existing MegaKernel implementations face a fundamental tension between portability and efficiency on resource-constrained GPUs such as NVIDIA Ada: hand-tuned solutions are tightly coupled to specific architectures and lack portability, while auto-compiled approaches introduce runtime dynamic scheduling whose branch penalties are unacceptable in latency-critical settings. We observe that under a fixed deployment configuration, the optimal execution path of a MegaKernel is uniquely determined, and runtime dynamic decision-making can be entirely hoisted to compile time. Building on this insight, we propose Ada-MK: (1) a three-dimensional shared-memory constraint model combined with K-dimension splitting that reduces peak shared memory usage by 50%; (2) MLIR-based fine-grained DAG offline search that solidifies the optimal execution path, completely eliminating runtime branching; and (3) a heterogeneous hybrid inference engine that embeds MegaKernel as a plugin into TensorRT-LLM, combining high-throughput Prefill with low-latency Decode. On an NVIDIA L20, Ada-MK improves single-batch throughput by up to 23.6% over vanilla TensorRT-LLM and 50.2% over vLLM, achieving positive gains across all tested scenarios--the first industrial deployment of MegaKernel in a commercial online advertising system.

cs.CL

Efficient LLM-based Advertising via Model Compression and Parallel Verification

Large language models (LLMs) have shown remarkable potential in advertising scenarios such as ad creative generation and targeted advertising. However, deploying LLMs in real-time advertising systems poses significant challenges due to their high inference latency and computational cost. In this paper, we propose an Efficient Generative Targeting framework that integrates adaptive group quantization, layer-adaptive hierarchical sparsification, and prefix-tree parallel verification to accelerate LLM inference while preserving generation quality. Extensive experiments on two real-world advertising scenarios demonstrate that our framework achieves significant speedup with acceptable quality degradation, making it operationally viable for practical deployments.

cs.CL

DVM: A Bytecode Virtual Machine Approach for Dynamic Tensor Computation

Dynamism is common in AI computation, e.g., the dynamic tensor shapes and the dynamic control flows in models. Due to the long compilation time, existing runtime compilation damages the model efficiency, while the offline compilers either suffer from the long compilation time and device memory footprint to cover all the possible execution instances of a dynamic model, or sacrifice optimization opportunities for usability. In this paper, we rethink the feasibility of runtime compilation for dynamic models and identify that the key for it to work is to speed up the compilation or hide the compilation overhead. To do this, we propose a real-time compiler, DVM. In DVM, we design a runtime operator compiler based on a bytecode virtual machine to perform effective and efficient compilation for each dynamic operator instance given its input. Specifically, instead of compiling programs into machine code, we encode the operator program into bytecode on the CPU and decode the bytecode into virtual instructions for direct execution on the NPU. Based on the runtime operator compiler, we further propose an operator fuser, which performs symbol-deduction-based fusion on static graphs and runtime fusion on dynamic graphs. Both pattern- and stacking-based fusion are supported to increase fusion opportunities. Evaluation on operators, subgraphs, and models shows that, compared with TorchInductor, PyTorch-eager and MindSpore-graph-O0, we are up to 11.77$\times$ better in terms of the operator/model efficiency and up to 5 orders of magnitude faster in terms of the maximum compilation time.

cs.PL

Asymptotics for a nonstandard risk model with multivariate subexponential claims and constant interest force

In this paper, the asymptotic behavior of the entrance probability of discounted aggregate claims of a certain family of rare sets is studied, considering the finite and infinite time horizons. This multivariate risk model, driven by a common counting process, has a constant interest rate and allows the interdependence of claim vectors. For the finite time horizon, the multivariate subexponential distribution of the common claim vector and the weak dependence structure of regression dependence are used. For the infinite time horizon, the claim vector is taken from a smaller distribution class, and the weak dependence structure is more general. Both results are derived under some additional assumptions on the moments of the counting process, which is fulfilled by all inhomogeneous renewal processes and many quasi-renewal processes, respectively. Moreover, the results are specialized to the multivariate regularly varying case, where more explicit results on the asymptotic behavior of the entrance probability of the discounted aggregate claims are derived. At the end of the paper, the results obtained are used to study the finite and infinite time horizon ruin problems of a risk model with eventual Brownian perturbations.

math.PR

From Complex Dynamics to DynFormer: Rethinking Transformers for PDEs

Partial differential equations (PDEs) are fundamental for modeling complex physical systems, yet classical numerical solvers face prohibitive computational costs in high-dimensional and multi-scale regimes. While Transformer-based neural operators have emerged as powerful data-driven alternatives, they conventionally treat all discretized spatial points as uniform, independent tokens. This monolithic approach ignores the intrinsic scale separation of physical fields, applying computationally prohibitive global attention that redundantly mixes smooth large-scale dynamics with high-frequency fluctuations. Rethinking Transformers through the lens of complex dynamics, we propose DynFormer, a novel dynamics-informed neural operator. Rather than applying a uniform attention mechanism across all scales, DynFormer explicitly assigns specialized network modules to distinct physical scales. It leverages a Spectral Embedding to isolate low-frequency modes, enabling a Kronecker-structured attention mechanism to efficiently capture large-scale global interactions with reduced complexity. Concurrently, we introduce a Local-Global-Mixing transformation. This module utilizes nonlinear multiplicative frequency mixing to implicitly reconstruct the small-scale, fast-varying turbulent cascades that are slaved to the macroscopic state, without incurring the cost of global attention. Integrating these modules into a hybrid evolutionary architecture ensures robust long-term temporal stability. Extensive memory-aligned evaluations across four PDE benchmarks demonstrate that DynFormer achieves up to a 95% reduction in relative error compared to state-of-the-art baselines, while significantly reducing GPU memory consumption. Our results establish that embedding first-principles physical dynamics into Transformer architectures yields a highly scalable, theoretically grounded blueprint for PDE surrogate modeling.

cs.LG

TraderBench: How Robust Are AI Agents in Adversarial Capital Markets?

Evaluating AI agents in finance faces two key challenges: static benchmarks require costly expert annotation yet miss the dynamic decision-making central to real-world trading, while LLM-based judges introduce uncontrolled variance on domain-specific tasks. We introduce TraderBench, a benchmark that addresses both issues. It combines expert-verified static tasks (knowledge retrieval, analytical reasoning) with adversarial trading simulations scored purely on realized performance-Sharpe ratio, returns, and drawdown-eliminating judge variance entirely. The framework features two novel tracks: crypto trading with four progressive market-manipulation transforms, and options derivatives scoring across P&L accuracy, Greeks, and risk management. Trading scenarios can be refreshed with new market data to prevent benchmark contamination. Evaluating 13 models (8B open-source to frontier) on ~50 tasks, we find: (1) 8 of 13 models score ~33 on crypto with <1-point variation across adversarial conditions, exposing fixed non-adaptive strategies; (2) extended thinking helps retrieval (+26 points) but has zero impact on trading (+0.3 crypto, -0.1 options). These findings reveal that current agents lack genuine market adaptation, underscoring the need for performance-grounded evaluation in finance.

cs.AI

Asymptotics of randomly weighted sums without moment conditions of random weights

In the paper, we investigate the asymptotics of randomly weighted sums with upper tail asymptotically independent and quasi-upper tail asymptotically independent primary random variables without requiring moment assumptions on random weights. For the case of primary random variables with regularly varying tails, we obtain more explicit results via an extension of Breiman's theorem. Then an application of the obtained results is established to asymptotically estimate for the finite-time and infinite-time ruin probabilities in a discrete-time risk model.

math.PR

Seeing the Goal, Missing the Truth: Human Accountability for AI Bias

This research explores how human-defined goals influence the behavior of Large Language Models (LLMs) through purpose-conditioned cognition. Using financial prediction tasks, we show that revealing the downstream use (e.g., predicting stock returns or earnings) of LLM outputs leads the LLM to generate biased sentiment and competition measures, even though these measures are intended to be downstream task-independent. Goal-aware prompting shifts these intermediate measures toward the disclosed downstream objective, producing in-sample overfitting. Specifically, purpose leakage improves performance on data prior to the LLM's knowledge cutoff, but provides no advantage after the cutoff. This bias is strong enough that regularization of prompt instructions cannot fully address this form of overfitting. We further show that the bias can arise from users' unintentional conversational context that hints at the purpose. Overall, we document that AI bias due to "seeing the goal" is not an algorithmic flaw, but stems from human accountability in research design.

q-fin.GN