SearcharxivSearch

arXiv subjects

Hao Xu

Publications and source records attributed to Hao Xu.

At least 19 recordsLinked to original sources

Morphological Decoupling-Based Skeletal Classification for Clinical Assessment of Malocclusion

Malocclusion skeletal grading is a fundamental task in orthodontics, critical for diagnosis and treatment planning. Traditionally, cone-beam computed tomography (CBCT) is used for visual measurement, and the reconstructed lateral cephalograms are handed over to expert dentists for diagnosis. However, manual review is time-consuming, labor-intensive, and subject to inter-operator variability. Therefore, an automatic CBCT-based system is needed for reliable malocclusion skeletal grading. In this case, we develop TeethGNN, a novel graph-based framework designed to combine CBCT image features with morphological information for accurate and efficient malocclusion grading. TeethGNN utilizes a decoupled learnable decoder to directly predict key morphological indicators from CBCT images, eliminating the need for manual measurements. These morphological features are then fused with image features using a graph neural network (GNN), which effectively models the relationships between the modalities. To further enhance robustness and calibration, we introduce a collaborative calibration strategy. This strategy combines multi-scale graph adversarial perturbation for explicit calibration and nonlinear topological graph calibration for implicit confidence adjustment. Extensive experiments and ablation studies on our collected clinical dataset demonstrate that our malocclusion measurement system achieves 77.08\% in accuracy and 89.61\% in AUC, outperforming the compared state-of-the-art methods. These results validate the effectiveness of graph-based multimodal fusion and collaborative calibration in improving malocclusion grading performance. Our system shows strong potential for advancing computer-aided orthodontic diagnosis, providing an accurate and reliable solution for vision-based clinical measurement and diagnosis.

eess.IV

CODE: Cross-Modal Calibration and Dynamic Suppression for Open World Object Detection

Open World Object Detection (OWOD) built on multimodal foundation models often suffers from semantic ambiguity caused by unidirectional text-to-vision matching, while rigid outlier penalties may over-suppress unknown objects near known-class decision boundaries. We propose CODE (Cross-Modal Calibration and Dynamic Suppression), a unified inference-time framework with three complementary components. Cross-Modal Joint Confidence Calibration injects global visual prototypes to calibrate text-driven known-class predictions. Uncertainty-Guided Universal Objectness Enhancement measures classification hesitation from local visual responses to strengthen potential unknown objects. Dynamic Outlier Suppression via Confidence Margin replaces rigid suppression with a margin-aware adjustment that preserves ambiguous out-of-distribution instances. Experiments on the Real-World Detection benchmark demonstrate that, with the OWL-ViT L/14 backbone, CODE achieves 21.7 U-mAP and 40.8 K-mAP in Task 1, surpassing the previous state of the art by 2.6 and 2.3 points, respectively.

cs.CV

SSE-Bio: A Structured Self-Evolving Agent with Agentic Retrieval Policy for Multi-Hop Biomedical Reasoning

Biomedical multi-hop question answering (QA) requires models to connect evidence across intermediate entities such as diseases, drugs, proteins, and phenotypes. Existing agents typically rely on static retrieval workflows or coarse-grained prompt rewriting, which can lead to instruction drift when reasoning procedures need to be updated. We propose SSE-Bio, a structured self-evolving agent with an agentic retrieval policy for multi-hop biomedical reasoning. Instead of globally rewriting agent instructions, SSE-Bio maintains a structured state, selectively retrieves knowledge triplets and prior templates through a trainable proxy policy, and improves its reasoning memory through fine-grained template editing. To optimise retrieval decisions, we introduce a proxy-training strategy based on group relative policy optimization, where the proxy is improved through decision-contrastive groups over alternative retrieval choices. Experiments on three biomedical multi-hop QA benchmarks show that SSE-Bio consistently outperforms existing baselines, achieving an improvement of 6.56 absolute points over the strongest self-evolving baseline on BioHopR.

cs.CL

The Equality Cases of the Weak Simplex Conjecture

Among $n+1$ equiprobable equal-energy signals in $\R^n$ under additive white Gaussian noise with maximum-likelihood decoding, which arrangement maximizes the probability of correct decoding? The question is Shannon's, recorded by Rice in 1950. Mulgund proved in 2026 that the regular-simplex value bounds the correct-decoding probability of every signal set at every signal-to-noise ratio, leaving open whether the simplex is the only maximizer. This paper determines the equality cases in a form stronger than uniqueness. A signal set other than a regular simplex falls strictly below the bound at every positive signal-to-noise ratio. Hence a code meeting the bound at one positive operating point is already a regular simplex, up to vertex relabeling and an orthogonal map. In probabilistic form, among the correlation matrices that signal sets induce, any matrix other than the identity gives a lower-orthant probability strictly above its independent counterpart at every finite threshold, leaving no room for a nontrivial equality. No code of ambient dimension below $n$ attains the bound. Under an energy budget $E$ with unrestricted blocklength the optimal codebook is uniquely the regular simplex of circumradius $\sqrt{E}$. Every optimal codeword therefore exhausts its allowance. Equality in the Simplex Mean Width Conjecture likewise occurs only at the regular simplex. The proof strengthens the first self-convolution step of Mulgund's argument with Royen's correlation theorem. The single-parameter rigidity is machine-checked in Lean 4.

cs.IT

Spectral Localization Principle for Entanglement Harvesting

We propose a unified physical principle for entanglement harvesting: the entanglement that two localized detectors can extract from a quantum field is determined solely by how localized the field's effective spectral density is. We demonstrate this in an analytically solvable model of two qubits coupled to a leaky single-mode cavity, which in turn couples to a continuous electromagnetic bath, and derive the maximum harvestable concurrence in closed form, $\mathcal{C}_{\max}(Q)=2e^{-\pi/(2Q)}(1+e^{-\pi/(2Q)})/(1+3e^{-\pi/Q})$, where $Q\equiv|\Delta|/\kappa$ is the ratio of the qubit-cavity detuning $\Delta$ to the cavity linewidth $\kappa$. In the high-$Q$ limit, $\mathcal{C}_{\max}\simeq1-\pi^{2}/(16Q^{2})$, so the entanglement is robust against cavity loss; in the low-$Q$ limit it decays exponentially to zero, consistent with the irreversible-reservoir character of a continuous field, where maximal entanglement is unattainable. Since $Q$ is proportional to the inverse participation ratio (IPR) of the effective spectral density, it is the single dimensionless parameter governing the crossover from deterministic gate-based entanglement ($Q\to\infty$) to vacuum harvesting ($Q\to0$). Our framework operationalizes the Reeh-Schlieder theorem by quantifying the fraction of vacuum correlations accessible to localized detectors. It also reveals a formal correspondence of the maximal concurrence with the IPR, analogous to the conductivity-participation-ratio relation in Anderson localization. The predicted $\mathcal{C}_{\max}(Q)$ curve is, in principle, directly observable in superconducting circuit QED experiments.

quant-ph

ATOM: Geometry-Aware Microgesture towards Object-Agnostic Tangible Interaction

This paper presents ATOM, an integrated framework towards agnostic and tangible object interactions with microgestures. Our goal is to support microgesture interactions across different everyday objects, with the capability to automatically leverage the geometric affordance of each object. We formulate a fingertip-aware detection pipeline to leverage generative 2D and 3D models for geometry enhancement and refinement. We then introduce a usability-based method to prioritize the detected elements based on their ergonomic suitability for interactions. Building on this foundation, we further develop an AR system to transform everyday handheld objects into tangible user interfaces with 0D, 1D, and 2D microgesture interactions. Across transitions among everyday cooking objects of varying shapes and sizes, ATOM outperformed ablation baselines in task completion, usability (SUS), and workload (NASA-TLX). A further study with 10 objects demonstrates ATOM's generalizability across objects and grasps, highlighting its potential towards fluid, object-agnostic tangible interaction in real-world AR scenarios.

cs.HC

Secure Cooperative THz ISAC via Mamba Empowered Graph Neural Network Precoding

The terahertz (THz) band offers abundant spectrum resources for high-throughput communication and ultra high-precision localization. This paper investigates secure communication in cooperative THz orthogonal frequency-division multiplexing (OFDM) bistatic integrated sensing and communications (ISAC) systems, where multiple base stations (BSs) equipped with extremely large-scale antenna arrays (ELAAs) collaboratively serve downlink users while concurrently locating multiple targets. Malicious targets are assumed to act as potential eavesdroppers attempting to intercept confidential information intended for legitimate users. To mitigate these threats, we formulate a joint optimization problem for analog beamforming, digital precoding, true-time delayers (TTDs), and sensing signal covariance matrix design. The objective is to maximize the minimum secrecy rate subject to Cramer-Rao bound (CRB) constraints that ensure localization accuracy. This problem is highly challenging due to the non-convex CRB constraint, strongly coupled variables, high computational complexity from ELAA, and near-field channel modeling. To address these challenges, we propose a novel data-driven framework that integrates graph neural networks (GNNs) with the Mamba architecture. Our proposed framework first encodes the interactions among users, targets, and BSs into a heterogeneous graph and then employs message passing to optimize vertex features. The Mamba blocks further enhance this process through their selection mechanism and state space modeling capabilities, enabling dynamic and context-aware optimization of beamforming, TTD configurations, and sensing parameters. Numerical simulations validate that the proposed method outperforms both conventional and learning-based baselines, while offering high computational efficiency and strong generalization across different network conditions.

cs.IT

Intelligent Wiretap Code Design: Exploiting Wireless Endogenous Security via Information Theory and Deep Learning Integration

Recent advancements in wireless endogenous security have explored leveraging the inherent randomness of wireless channels to enhance communication security, providing an effective alternative to traditional encryption methods. This paper proposes a wiretap coding scheme within the semantic communication framework, which leverages discrete semantic representations compatible with conventional digital modulation to jointly enhance communication security and reliability. We investigate two eavesdropping scenarios: (i) the eavesdropper employs a maximum a posteriori (MAP) decoder, and (ii) the eavesdropper has access to a decoder identical to that of the legitimate receiver. In the first scenario, we exploit mutual information as a metric to guide the design of an optimized coding strategy, minimizing information leakage while enhancing communication reliability. In the second scenario, considering the limitations of the eavesdropper's decoding capability, we employ generalized mutual information (GMI) to characterize recoverability under the prescribed decoding rule and guide reliability-aware code optimization.

cs.IT

CSI Reconstruction in Fluid Antenna Systems Without Spatial Covariance Priors

Fluid antenna systems (FASs) exploit many candidate ports for spatial diversity, but hardware constraints allow channel observations at only a few active ports. Whether full-port CSI can be recovered without pre-acquired channel statistics remains open. Under the Clarke isotropic scattering model, we show that the channel lies in a low-dimensional spatial modal subspace determined by the scattering environment rather than the total port count. Consequently, recovery becomes feasible when the number of observed ports reaches the modal dimension (i.e., $M\geq r$), even when $M\ll N$. We further establish a sharp feasibility threshold: reliable recovery is impossible below this dimension regardless of SNR, whereas accuracy improves with additional observations above it. By decomposing the recovery error into modal truncation, estimation, and learning components, we derive explicit tradeoffs among RF chains, pilot overhead, transmit power, and training data. These results enable scalable prior-free full-port CSI recovery with few active ports.

cs.IT

WirelessOpsAgent: A Benchmark and Agent Design for Action Assurance in Wireless Networks

Large language model (LLM) agents are emerging as planners for autonomous wireless network operations. Yet a task answer that is correct at proposal time can still be unsafe at execution time if supporting telemetry is stale or inconsistent. Existing benchmarks mainly evaluate task solving from fixed observations and leave support checking at execution time untested. We introduce WirelessOptBench, a benchmark for action assurance in wireless operations. It turns wireless tasks into execution state decision episodes with controlled telemetry faults and action constraints. We further develop WirelessOpsAgent, which grounds candidate actions in current evidence and repairs recoverable support failures before execution. Across three backbone evaluations with 600 episodes each, WirelessOpsAgent achieves up to 0.983 Exact Action Accuracy. On Claude Sonnet 4.6, the Unsafe APPLY Rate decreases from 82.2% to 10.3% relative to the safest baseline. We make WirelessOptBench available at https://anonymous.4open.science/r/wirelessopsbench-artifact-D969/.

cs.NI

Classification of symmetric fusion categories over $\mathbb{R}$

We show that every symmetric fusion category over $\mathbb{R}$ is equivalent to the category of finite-dimensional semi-linear representations of a $\mathbb{Z}_2$-graded finite super group. The proof uses Galois descent for tensor categories over $\mathbb{C}/\mathbb{R}$, reducing the classification to semi-linear $\mathbb{Z}_2$-actions on symmetric fusion categories over $\mathbb{C}$. As a further structural result, we establish a Tannaka-Krein type correspondence between symmetric fusion categories over $\mathbb{R}$ and finite groupoids with a $\mathbb{Z}_2 \times \mathrm{B} \mathbb{Z}_2$-action. This gives a complete real analogue of Deligne's classification result.

math.QA

Projection-Based Outlier Detection in Interval-Valued Functional Data

Outlier detection is a fundamental task for ensuring reliable statistical modeling and inference. Interval-valued functional data (IVFD), in which each observation is represented by an interval-valued curve that preserves the variability and uncertainty within the observation, have attracted increasing attention in statistics and related applications. Developing effective outlier detection procedures for IVFD is therefore an important methodological problem. To address this issue, we develop a robust projection-based outlier detection framework. We first represent each interval-valued functional observation through its center and log-radius functions and apply interval-valued functional principal component analysis (IFPCA) to obtain a joint low-dimensional representation. We then introduce the interval-valued least trimmed functional scores (ILTFS) method, which identifies a robust reference subset by minimizing a trimmed aggregate of standardized IFPCA score distances. Finally, we proposed the ILTFS-FDR outlier detection procedure by converting the resulting projection distances into empirical $p$-values and adjusting using the Benjamini--Hochberg procedure at a prespecified target false discovery rate level. Theoretically, we derive the finite-sample breakdown point of the ILTFS mean estimator and establish the descent property of the concentration-step algorithm. Simulation studies and an empirical application to high-frequency ETF data demonstrate the effectiveness and robustness of the proposed ILTFS-FDR procedure in detecting abnormal interval-valued functional observations.

stat.ME

LingShu: A Large-Scale Symptom-Centric Contextualized Knowledge Graph Bridging Traditional Chinese Medicine and Modern Biomedicine

Biomedical knowledge graphs (KGs) are pivotal for knowledge organization, yet traditional binary relations often struggle to represent the conditional nature of biomedical knowledge. Symptoms provide a shared phenotypic layer for linking Traditional Chinese Medicine (TCM), which relies on symptom patterns for syndrome differentiation and treatment selection, with modern biomedicine, which connects clinical manifestations to diseases and molecular mechanisms. We present LingShu, a large-scale symptom-centric contextualized knowledge graph designed to bridge TCM and modern biomedicine. The exported version of LingShu analyzed in this study comprises 17.33 million atom-level entity records and 39.47 million relation records, including 17.19 million semantic triples and 22.29 million contextualized quadruples. LingShu integrates multi-source data, including clinical electronic medical records, authoritative TCM texts, biomedical ontologies, and curated knowledge bases, through a pipeline combining natural language processing, terminology normalization, and human-in-the-loop verification. A key innovation of LingShu is its hybrid data model: it maintains 64 typed triple relation patterns to ensure broad connectivity, while incorporating 35 contextual quadruple relation patterns to capture conditional medical associations. This dual-structure approach explicitly encodes conditional knowledge, providing a granular representation of the contexts associated with medical relations. These contextualized relations cover syndrome-dependent herb efficacy, disease-contextualized drug effects, population-specific clinical associations, and mechanism-related therapeutic responses. Furthermore, we developed a web platform (http://www.tcmkg.com/) that integrates graph visualization, graph-based reasoning, and an evidence-grounded knowledge question-answering agent.

cs.CL

SONG: A Photorealistic 3D Gaussian Simulation Platform for Benchmarking Social Navigation

Social navigation has progressed from simplified 2D environments toward a more general vision-based setting, in which a robot needs to achieve socially compliant behavior purely from onboard visual observations. Yet supporting simulation platforms have not kept pace: existing options either lack visual observations, lack moving human avatars, or fall short of real-world fidelity in appearance and pedestrian behavior, offering limited support for advancing vision-based social navigation. We introduce SONG, a SOcial Navigation platform powered by 3D Gaussian splatting (3DGS). It leverages 3DGS for both scene and avatar representations, drives pedestrians using semantically grounded trajectories generated by a large language model, and synthesizes their full-body motion with a trajectory-conditioned generator to produce continuous, natural movement. On top of the platform, we curate SONG-Bench, a set of evaluation episodes stratified by difficulty, and propose a multi-dimensional metric suite covering effectiveness, safety, and social compliance. A systematic evaluation of representative navigation baselines reveals three findings: (a) vision-based social navigation is far from solved; (b) a critical safety deficit precedes social etiquette; (c) real-world data matters more than model scale. Crucially, we demonstrate that fine-tuning on our curated data effectively improves the success rate in real-world environments. We hope our platform provides a faithful and rigorous testbed for the next generation of vision-based social navigation research.

cs.RO

Rethinking Layer-Wise Information Allocation for Vision Foundation Model Adaptation

Vision foundation models are increasingly reused as frozen backbones for downstream visual recognition, making parameter-efficient adaptation a central problem. Prompt-based adaptation, including Visual Prompt Tuning (VPT), provides a lightweight way to specialize these models, but its layer-wise behavior remains poorly understood: performance is sensitive to prompt depth, placement, and task distribution, and gains on standard in-domain benchmarks do not always translate into robust generalization. We argue that this limitation is not solely an optimization issue, but a layer-wise information allocation issue: existing prompt-based methods lack principled control over what prompt-conditioned representations should preserve, suppress, and propagate across depth. Inspired by the Information Bottleneck principle, we introduce Prompted Information Bottlenecks (PIB), a framework that regularizes layer-wise compression-sufficiency trade-offs and promotes a more coherent cross-layer information path. The key idea is that effective adaptation should be minimal yet sufficient, retaining task-relevant local evidence in earlier layers while progressively discarding nuisance factors and redundant details in deeper layers. Extensive experiments show that PIB achieves strong performance across 34 datasets, reaching 92.1% on FGVC, 93.01% on HTA, and 77.33% on VTAB-1k, while tuning only 0.35% parameters on average across the main settings. Beyond benchmark accuracy, PIB helps explain the non-monotonic behavior of prompt capacity scaling, reduces shortcut reliance, and improves robustness under distribution shift and fine-grained recognition settings. These results position PIB as both a practical method and an information-allocation perspective for adapting frozen vision foundation models. Our code is available at https://github.com/itsnotacie/MM-26-PIB

cs.CV

HyWorldVLA: A Vision-Language-Action Model with Hybrid World Modeling for Autonomous Driving

Vision-Language-Action (VLA) models augmented with world modeling represent a promising paradigm for end-to-end autonomous driving. While pixel-level future prediction enables fine-grained spatiotemporal reasoning, it compromises robustness in noisy driving scenarios. Conversely, latent-based world models alleviate this sensitivity but often incur limited interpretability and representational degradation due to absent pixel-level grounding. To reconcile this trade-off, we propose HyWorldVLA, a hybrid world-VLA framework that unifies pixel-level supervision and latent representation learning. In the pre-training stage, HyWorldVLA predicts video latents encoded by a pre-trained video VAE, while simultaneously reconstructing video frames to provide precise pixel-level grounding. During the subsequent co-fine-tuning phase, the model exclusively predicts latent features, which are fed into an action expert to generate trajectories. Extensive experiments on NAVSIM v1 and v2 benchmarks demonstrate that HyWorldVLA significantly outperforms both pixel-based and latent-based world model baselines. Notably, we present the first comprehensive qualitative and quantitative analysis of world model noise robustness in autonomous driving, establishing a new benchmark for evaluating future architectures.

cs.CV

Sizable Ligand-Mediated Bond-Dependent Interactions in a Spin-1 Triangular Antiferromagnet NiI$_2$

The bond-dependent anisotropic Kitaev interactions are the key for the Kitaev model, which has attracted intense interest for its potential to host quantum-spin-liquid states and fractional excitations. However, experimental realizations of such interactions remain scarce. Here, we investigate the magnetic excitations of NiI$_2$, a van der Waals magnet with spin $S=1$. By combining inelastic neutron scattering, magnetization measurements, magnetic structure analysis, first-principles calculations, and linear-spin-wave simulations, we identify a minimal model that features substantial Kitaev and off-diagonal $\Gamma$ interactions, which together stabilize the canted magnetic ground state and open a gap in the spin-wave spectrum. Notably, these interactions arise from strong spin-orbit coupling on the ligand ions, despite the quenched orbital moment of the magnetic Ni$^{2+}$ ions. Our results provide compelling experimental evidence for the ligand-driven Kitaev mechanism. This demonstrates a concrete pathway to generating strong bond-dependent anisotropy in systems where the magnetic ions themselves have weak spin-orbit coupling, thereby substantially broadening the range of potential Kitaev materials.

cond-mat.str-el

SeamGen: Artist-Aligned UV Seam Generation via Graph Flow Matching

UV seam placement is a critical yet labor-intensive step in 3D content creation, requiring artists to balance chart shape, seam concealment, and alignment with semantic and geometric features. Existing automatic methods are primarily based on per-object optimization, relying on handcrafted objectives to avoid distortion or on proxies from pretrained models to inject semantic information. However, these strategies are not always well aligned with seams used in industrial production pipelines, often resulting in layouts that deviate from artist-preferred seam patterns and practical production requirements. To address these limitations, we propose SeamGen, a generative model for UV seam generation that aligns with artist preferences and production requirements. Instead of depending on manually designed objectives and constraints, SeamGen learns the distribution of per-edge seam labels from a large corpus of existing seam layouts using a flow-matching generative model. A key challenge is that typical Transformer architectures used in flow matching models are designed for sequential representations, such as point clouds, and cannot naturally account for mesh topology. To enable mesh-native learning, we design a Mesh Transformer backbone that interleaves local graph attention over mesh edges with global self-attention across vertices, capturing both fine-grained geometric cues and long-range topological coherence. To further improve inference-time controllability and quality, we exploit the training-free inpainting capability of flow models for both localized seam refinement and constraint-guided seam generation. Extensive experiments show that by learning priors from professional seam layout data, SeamGen produces UV layouts that better align with artist-authored preferences and achieve superior perceptual quality compared with distortion-based and semantic-proxy baselines.

cs.CV