SearcharxivSearch

arXiv subjects

Xu Gao

Publications and source records attributed to Xu Gao.

At least 19 recordsLinked to original sources

LANTERN: A Closed-Loop Benchmark for VLM-Based Cooperative Driving with Temporally Grounded Warnings

We present LANTERN, a closed-loop benchmark for temporally grounded cooperative warnings. LANTERN separates warning onset, hazard onset, warning termination, and post-hazard recovery, and evaluates each physical event under matched warning and no-warning executions so that the warning's contribution is measured in isolation rather than confounded with onboard vision. The benchmark spans six safety-critical scenario families and provides 3,272 sequences with 236,309 frames for training, together with 120 matched route pairs for closed-loop evaluation. Each hazard route is evaluated under the warning and no-warning conditions, while its no-hazard control penalizes unconditional braking. We further introduce the Cooperative Unified Score (CUS), a safety-gated metric that jointly rewards route progress, anticipation, clearance, and recovery. Fine-tuning a representative VLM driving model raises CUS from 34.6 without warnings to 75.5 with them, demonstrating both the value of cooperative warnings and the discriminative power of the paired protocol. All resources will be made publicly available.

cs.RO

On strong identities of almost-canonically seminormed rings

We investigate the strong identity condition (SIC) for almost-canonically seminormed rings, a class of topological graded rings that includes enveloping algebras of vertex operator algebras. This condition was introduced in the algebro-geometric theory of conformal blocks, where it governs the smoothing of nodal curves. To understand the representation-theoretic meaning of SIC, we develop the representation theory of almost-canonically seminormed rings, including Zhu-type algebras, induced modules, rationality conditions, tensor product compatibility, and an end formula for the mode transition algebra. Our main result characterizes the strong identity condition in terms of orthogonal expansions, projectivity of canonical modules, and Morita-type equivalences induced by Zhu-type algebras. As an application, we show that for vertex operator algebras of CFT type, the smoothing property is equivalent to the Zhu algebra inducing a Morita-type equivalence with the category of admissible modules. Consequently, the strong identity condition identifies the precise representation-theoretic obstruction to extending algebraic smoothing beyond the semisimple setting. We further illustrate the theory through explicit examples, including the Weyl algebra and several irrational vertex operator algebras where the strong identity condition fails.

math.QA

Search for invisible decays of light mesons via $J/\psi \to VP$ $(V=\omega/\phi,P=\eta/\eta')$ decays at STCF

We present a preliminary feasibility study of searches for invisible decays of light mesons via $J/\psi \to VP$ $(V=\omega/\phi,P=\eta/\eta')$ using a traditional analytical method at the proposed Super $\tau$-Charm facility (STCF) which is expected to accumulate $3.4\times10^{12}$ $J/\psi$ events per year, based on an inclusive Monte Carlo sample of $1.3 \times 10^{9}$ $J/\psi$ events. The upper limits on the invisible decay branching fractions at the 90\% confidence level are set as $\mathcal{B}(\omega \to invisible) < 3.7 \times 10^{-7}$, $\mathcal{B}(\phi \to invisible) < 8.9 \times 10^{-7}$, $\mathcal{B}(\eta \to invisible) < 1.8 \times 10^{-7}$ and $\mathcal{B}(\eta' \to invisible) < 4.1 \times 10^{-7}$, respectively, using a projected toy data corresponding to the expected STCF statistics. By using the machine learning technique such as Deep Learning, the upper limit may be further improved to approach theoretical predictions for light dark matter.

hep-ex

From Local Indices to Global Identifiers: Generative Reranking for Recommender Systems via Global Action Space

In modern recommender systems, list-wise reranking serves as a critical phase within the multi-stage pipeline, finalizing the exposed item sequence and directly impacting user satisfaction by modeling complex intra-list item dependencies. Existing methods typically formulate this task as selecting indices from the local input list. However, this approach suffers from a semantically inconsistent action space: the same output neuron (logits) represents different items across different samples, preventing the model from establishing a stable, intrinsic understanding of the items. To address this, we propose GloRank (Global Action Space Ranker), a generative framework that shifts reranking from selecting local indices to generating global identifiers. Specifically, we represent items as sequences of discrete tokens and reformulate reranking as a token generation task. This design effectively decouples the scoring mechanism from the variable input order, ensuring that items are evaluated against a consistent global standard. We further enhance this with a two-stage optimization pipeline: a supervised pre-training phase to initialize the model with high-quality demonstrations, followed by a reinforcement learning-based post-training phase to directly maximize list-wise utility. Extensive experiments on two public benchmarks and a large-scale industrial dataset, coupled with online A/B tests, demonstrate that GloRank consistently outperforms state-of-the-art baselines and achieves superior robustness in cold-start scenarios.

cs.IR

Precision $YN$ and $\bar{n}N$ measurements with an LH$_2$/LD$_2$ target in the BESIII detector

Located at the BEPCII $e^{+}e^{-}$ collider, the BESIII experiment provides a robust platform for investigating (anti)hyperon-nucleon ($YN$) and antineutron-nucleon ($\bar{n}N$) interactions. This is made possible by the high production cross-sections of $J/\psi$ and $\psi(3686)$ resonances and their substantial decay branches into these baryons. Although previous studies using the beam pipe as a target demonstrated feasibility, statistical precision remains constrained by the limited material budget. To address this, we propose installing a dedicated liquid hydrogen or liquid deuterium target between the beam pipe and the Cylindrical Gas Electron Multiplier Inner Tracker. Monte Carlo simulations confirm that the added material has a negligible effect on charged particle tracking. This upgrade is expected to enhance the effective luminosity for scattering on free protons by a factor of 10--30 for $\Lambda$, $\Sigma^{+}$, $\Xi$, and $\bar{n}$ beams, enabling high-precision measurements of $YN$ and $\bar{n}N$ interactions that will significantly advance our knowledge of non-perturbative strong interactions.

hep-ex

Denoising Neural Reranker for Recommender Systems

For multi-stage recommenders in industry, a user request would first trigger a simple and efficient retriever module that selects and ranks a list of relevant items, then the recommender calls a slower but more sophisticated reranking model that refines the item list exposure to the user. To consistently optimize the two-stage retrieval reranking framework, most efforts have focused on learning reranker-aware retrievers. In contrast, there has been limited work on how to achieve a retriever-aware reranker. In this work, we provide evidence that the retriever scores from the previous stage are informative signals that have been underexplored. Specifically, we first empirically show that the reranking task under the two-stage framework is naturally a noise reduction problem on the retriever scores, and theoretically show the limitations of naive utilization techniques of the retriever scores. Following this notion, we derive an adversarial framework DNR that associates the denoising reranker with a carefully designed noise generation module. The resulting DNR solution extends the conventional score error minimization loss with three augmented objectives, including: 1) a denoising objective that aims to denoise the noisy retriever scores to align with the user feedback; 2) an adversarial retriever score generation objective that improves the exploration in the retriever score space; and 3) a distribution regularization term that aims to align the distribution of generated noisy retriever scores with the real ones. We conduct extensive experiments on three public datasets and an industrial recommender system, together with analytical support, to validate the effectiveness of the proposed DNR.

cs.IR

A basis theorem for Genus-One Conformal Blocks and modular invariance of intertwining operators

We prove that trace functions associated to intertwining operators over a strongly rational vertex operator algebra form a global frame of the conformal block bundle $\mathscr{C}_{\mathbb{H}}(W)$ over $\mathbb{H}$. Consequently, for each $\tau\in\mathbb{H}$, these trace functions, evaluated at $\tau$, form a basis of the fiber $\mathscr{C}(E_\tau,\mathsf{p},z,W)$, and the natural $\mathrm{SL}(2,\mathbb{Z})$-action on the fiber is represented in this basis. This result is both a generalization and a refinement of Zhu's and Dong-Li-Mason's modular invariance theorems for trace functions associated to vertex operators and twisted vertex operators, and a specialization and refinement of Huang's and Miyamoto's modular invariance theorems for (logarithmic) intertwining operators for $C_2$-cofinite vertex operator algebras. The proof combines a new construction of a connection on the bundle $\mathscr{C}_{\mathbb{H}}(W)$, Zhu's recursive formulas for trace functions, Frenkel-Zhu's fusion rules theorem, and recent theorems of Damiolini-Gibney-Krashen-Tarasca on the geometry of sheaves of vertex operator algebra conformal blocks over the moduli spaces $\overline{\mathscr{M}}_{g,n}$.

math.QA

Attention Distillation: A Unified Approach to Visual Characteristics Transfer

Recent advances in generative diffusion models have shown a notable inherent understanding of image style and semantics. In this paper, we leverage the self-attention features from pretrained diffusion networks to transfer the visual characteristics from a reference to generated images. Unlike previous work that uses these features as plug-and-play attributes, we propose a novel attention distillation loss calculated between the ideal and current stylization results, based on which we optimize the synthesized image via backpropagation in latent space. Next, we propose an improved Classifier Guidance that integrates attention distillation loss into the denoising sampling process, further accelerating the synthesis and enabling a broad range of image generation applications. Extensive experiments have demonstrated the extraordinary performance of our approach in transferring the examples' style, appearance, and texture to new images in synthesis. Code is available at https://github.com/xugao97/AttentionDistillation.

cs.CV

Highly coherent grain boundaries induced by local pseudo-mirror symmetry in $\beta$-Ga2O3

Grain boundaries have extensive influence on the performance of crystal materials. However, the atomic-scale structure and its relation with local and crystallographic symmetries remain elusive in low-symmetry crystals. Herein, we find that the local pseudo-mirror-symmetric atomic layer is the common physical origin of a series of highly coherent grain boundaries in the low-symmetry $\beta$-Ga2O3 crystal. These include the (100) twin boundary and an emerging series of $(h-1'0'2)/(h+1'0'\bar{2})$ coherent asymmetric grain boundaries (CAGBs). Owing to the local pseudo-mirror symmetry and the special geometric relation of the $\beta$-Ga2O3 conventional cell, these CAGBs place 80% of the boundary atoms in pseudo-coincident sites, exhibiting high coherence under the coincident-site lattice model. With a combination of density functional theory calculations, Czochralski growth experiment, and atomic-scale characterizations, the structure and stability of the $(002)/(20\bar{2})$-A CAGB are confirmed, with a boundary energy density as low as 0.36 J/m2. This CAGB is responsible for the spontaneous formation of a twinned defect facet at the surface steps during the epitaxy growth of $\beta$-Ga2O3, warranting a substrate orientation selection rule for $\beta$-Ga2O3. Through this study, we provide insights into the grain boundary physics in the low-symmetry $\beta$-Ga2O3 crystal while emphasizing the importance of the local pseudo-symmetries in the low-symmetry crystals.

cond-mat.mtrl-sci

Hopping Transfer Optimizes Avalanche Multiplication in Molybdenum Disulfide

Recently, avalanche multiplication has been observed in TMDC-based FETs, enhancing sensor performance with high sensitivity. However, the high voltage required for operation can damage the FETs, making it crucial to reduce the breakdown voltage for effective sensing applications. Here, we demonstrate that the utilization of hopping transfer induced by high-density defects can effectively reduce the breakdown voltage in TMDCs FETs. By substituting oxygen atoms for sulfur atoms in a monolayer of MoS2, we create MoS2-xOx, with x carefully adjusted within the range of 0 to 0.51. Oxygen doping reduces the bandgap of TMDCs and enhances ion collision rates. Moreover, higher levels of oxygen doping (x > 0.41) in MoS2-xOx exhibit nearest-neighbor hopping behavior, leading to a significant enhancement in electron mobility. These improvements result in a decrease in the breakdown voltage of avalanche multiplication from 26.2 V to 12.6 V. Additionally, we propose avalanche multiplication in MoS2-xOx as an efficient sensing mechanism to overcome the limitations of gas sensing. The MoS2-xOx sensors display an ultra-high response to NO2 gas in the air, with a response of 5.8x103 % to NO2 gas of 50 ppb at room temperature, which is nearly two orders of magnitude higher than resistance-type gas detectors based on TMDCs. This work demonstrates that hopping transfer induced by high-density oxygen defects can effectively decrease the breakdown voltage of MoS2-xOx FETs, enhancing avalanche multiplication and serving as a promising mechanism for ultrasensitive gas detection.

physics.app-ph

EMHI: A Multimodal Egocentric Human Motion Dataset with HMD and Body-Worn IMUs

Egocentric human pose estimation (HPE) using wearable sensors is essential for VR/AR applications. Most methods rely solely on either egocentric-view images or sparse Inertial Measurement Unit (IMU) signals, leading to inaccuracies due to self-occlusion in images or the sparseness and drift of inertial sensors. Most importantly, the lack of real-world datasets containing both modalities is a major obstacle to progress in this field. To overcome the barrier, we propose EMHI, a multimodal \textbf{E}gocentric human \textbf{M}otion dataset with \textbf{H}ead-Mounted Display (HMD) and body-worn \textbf{I}MUs, with all data collected under the real VR product suite. Specifically, EMHI provides synchronized stereo images from downward-sloping cameras on the headset and IMU data from body-worn sensors, along with pose annotations in SMPL format. This dataset consists of 885 sequences captured by 58 subjects performing 39 actions, totaling about 28.5 hours of recording. We evaluate the annotations by comparing them with optical marker-based SMPL fitting results. To substantiate the reliability of our dataset, we introduce MEPoser, a new baseline method for multimodal egocentric HPE, which employs a multimodal fusion encoder, temporal feature encoder, and MLP-based regression heads. The experiments on EMHI show that MEPoser outperforms existing single-modal methods and demonstrates the value of our dataset in solving the problem of egocentric HPE. We believe the release of EMHI and the method could advance the research of egocentric HPE and expedite the practical implementation of this technology in VR/AR products.

cs.CV

Explore the LiDAR-Camera Dynamic Adjustment Fusion for 3D Object Detection

Camera and LiDAR serve as informative sensors for accurate and robust autonomous driving systems. However, these sensors often exhibit heterogeneous natures, resulting in distributional modality gaps that present significant challenges for fusion. To address this, a robust fusion technique is crucial, particularly for enhancing 3D object detection. In this paper, we introduce a dynamic adjustment technology aimed at aligning modal distributions and learning effective modality representations to enhance the fusion process. Specifically, we propose a triphase domain aligning module. This module adjusts the feature distributions from both the camera and LiDAR, bringing them closer to the ground truth domain and minimizing differences. Additionally, we explore improved representation acquisition methods for dynamic fusion, which includes modal interaction and specialty enhancement. Finally, an adaptive learning technique that merges the semantics and geometry information for dynamical instance optimization. Extensive experiments in the nuScenes dataset present competitive performance with state-of-the-art approaches. Our code will be released in the future.

cs.CV

Twisted restricted conformal blocks of vertex operator algebras II: twisted restricted conformal blocks on totally ramified orbicurves

In this paper, we introduce a notion of twisted restricted conformal blocks on totally ramified orbicurves and establish an isomorphism between the space of twisted restricted conformal blocks and the space of twisted conformal blocks. The relationships among twisted (restricted) conformal blocks, $g$-twisted (restricted) correlation functions, and twisted intertwining operators are explored. Furthermore, by introducing a geometric generalization of Zhu's algebra and its modules, we obtain a description of the space of coinvariants by modules over associative algebras and show it is finite-dimensional under some conditions. In particular, a more conceptual proof of the $g$-twisted fusion rules theorem in vertex operator algebra theory is provided.

math.AG

Twisted restricted conformal blocks of vertex operator algebras I: $g$-twisted correlation functions and fusion rules

In this paper, we introduce a notion of $g$-twisted restricted conformal block on the three-pointed twisted projective line $\mathfrak{x}\colon\overline{C}\to\mathbb{P^1}$ associated with an untwisted module $M^1$ and the bottom levels of two $g$-twisted modules $M^2$ and $M^3$ over a vertex operator algebra $V$. We show that the space of twisted restricted conformal blocks is isomorphic to the space of $g$-twisted (restricted) correlation functions defined by the same datum and to the space of intertwining operators among these twisted modules. As an application, we derive a twisted version of the Fusion Rules Theorem.

math.QA

The stable Picard group of finite Adams Hopf algebroids with an application to the $\mathbb{R}$-motivic Steenrod subalgebra $\mathcal{A}(1)^{\mathbb{R}}$

In this paper, we investigate the rigidity of the stable comodule category of a specific class of Hopf algebroids known as finite Adams, shedding light on its Picard group. Then we establish a reduction process through base changes, enabling us to effectively compute the Picard group of the $\mathbb{R}$-motivic mod $2$ Steenrod subalgebra $\mathcal{A}(1)^{\mathbb{R}}$. Our computation shows that $\operatorname{Pic}(\mathcal{A}(1)^{\mathbb{R}})$ is isomorphic to $\mathbb{Z}^4$, where two ranks come from the motivic grading, one from the algebraic loop functor, and the last is generated by the $\mathbb{R}$-motivic joker $J$.

math.AT

V2X-Seq: A Large-Scale Sequential Dataset for Vehicle-Infrastructure Cooperative Perception and Forecasting

Utilizing infrastructure and vehicle-side information to track and forecast the behaviors of surrounding traffic participants can significantly improve decision-making and safety in autonomous driving. However, the lack of real-world sequential datasets limits research in this area. To address this issue, we introduce V2X-Seq, the first large-scale sequential V2X dataset, which includes data frames, trajectories, vector maps, and traffic lights captured from natural scenery. V2X-Seq comprises two parts: the sequential perception dataset, which includes more than 15,000 frames captured from 95 scenarios, and the trajectory forecasting dataset, which contains about 80,000 infrastructure-view scenarios, 80,000 vehicle-view scenarios, and 50,000 cooperative-view scenarios captured from 28 intersections' areas, covering 672 hours of data. Based on V2X-Seq, we introduce three new tasks for vehicle-infrastructure cooperative (VIC) autonomous driving: VIC3D Tracking, Online-VIC Forecasting, and Offline-VIC Forecasting. We also provide benchmarks for the introduced tasks. Find data, code, and more up-to-date information at \href{https://github.com/AIR-THU/DAIR-V2X-Seq}{https://github.com/AIR-THU/DAIR-V2X-Seq}.

cs.CV

Simplicial volumes in Bruhat-Tits buildings of split classical type

In a Bruhat-Tits building of split classical type (that is, of type $A_n$, $B_n$, $C_n$, $D_n$, and any combination of them) over a local field, the simplicial volume counts the vertices within the given simplicial distance from a special vertex. This paper aims to study the asymptotic growth of the simplicial volume. A formula of the simplicial volume is deduced from the theory of concave functions. Then the dominant term in its asymptotic growth is found using the theory of $q$-exponential polynomials developed in this paper.

math.NT

Developing Dipole-scheme Heterojunction Photocatalysts

The high recombination rate of photogenerated carriers is the bottleneck of photocatalysis, severely limiting the photocatalytic efficiency. Here, we develop a dipole-scheme (D-scheme for short) photocatalytic model and materials realization. The D-scheme heterojunction not only can effectively separate electrons and holes by a large polarization field, but also boosts photocatalytic redox reactions with large driving photovoltages and without any carrier loss. By means of first-principles and GW calculations, we propose a D-scheme heterojunction prototype with two real polar materials, PtSeTe/LiGaS2. This D-scheme photocatalyst exhibits a high capability of the photogenerated carrier separation and near-infrared light absorption. Moreover, our calculations of the Gibbs free energy imply a high ability of the hydrogen and oxygen evolution reaction by a large driving force. The proposed D-scheme photocatalytic model is generalized and paves a valuable route of significantly improving the photocatalytic efficiency.

cond-mat.mtrl-sci