SearcharxivSearch

arXiv subjects

Xiaodong Wang

Publications and source records attributed to Xiaodong Wang.

At least 19 recordsLinked to original sources

Aha-Flow Distillation: Flow Markers Matter in LLM Reasoning

We identify the Flow Moment, a reasoning pattern characterized by sustained, process-confirming verbalizations such as I'm doing, in contrast to the revision- and backtracking-oriented Aha Moment. We refer to their corresponding linguistic expressions as Flow Markers and Aha Markers, respectively. Based on this observation, we construct Flow-CoT by rewriting the discourse markers of original reasoning traces while preserving their underlying reasoning content, and use it as auxiliary supervision for on-policy self-distillation (OPSD). We further propose \textbf{Aha-Flow Distillation (AFD)}, a dual-mode extension of OPSD that pairs different forms of privileged information with corresponding reasoning instructions. The Aha branch retains concise solution-based supervision, while the Flow branch introduces rewritten Flow-CoT under a direct and confident reasoning instruction. At inference time, the model uses only the standard reflective instruction, so Flow-style reasoning serves purely as a training signal. Experiments on AIME25 and HMMT25 show consistent improvements across Qwen3-8B and Qwen3-4B: AFD improves Avg@12 from 60.8 to 61.3 on Qwen3-8B and from 57.5 to 58.6 on Qwen3-4B over our reproduced OPSD baselines. Controlled ablations further show that, with the same Flow-CoT/Aha-CoT composition, dual-mode training improves Avg@12 from 59.5 to 60.1, indicating that the benefit comes not only from introducing heterogeneous reasoning supervision, but also from how it is organized during self-distillation. The code is available at https://github.com/Wang-Xiaodong1899/Aha-Flow-Distillation.

cs.CL

Quantum Advantage in Multiple Access Wiretap Channels with Entangled Transmitters

We investigate secure communication over a classical multiple-access wiretap channel (MAC-WTC), specifically exploring the benefit of shared entanglement between the transmitters. Under the strict semantic security criterion, we derive an achievable rate region and a regularized expression for the secrecy capacity. We further establish a single-letter upper bound on the secure capacity of the MAC-WTC with entangled transmitters. Our results demonstrate that entanglement strictly enlarges the secrecy capacity of MAC-WTC compared to sharing only classical correlated randomness, a finding we illustrate using a pseudo-telepathy game example. Finally, we establish new strong soft-covering lemmas for the output statistics of multipleaccess channels (MACs) with entangled transmitters. Our results generalize existing results for non-entangled systems.

cs.IT

Spectral Improvements of Geometric Inequalities on Closed Kähler Manifolds

Let $(M,g,J)$ be a closed Kähler manifold satisfying $\operatorname{Ric}\geqslant g$. We establish improved Liouville theorems for the Euler--Lagrange equations associated with the Beckner--Sobolev inequalities by incorporating the first positive eigenvalue of the $\bar\partial$-Laplacian into a differential-identity argument. As a consequence, we obtain improved Sobolev and Beckner inequalities that refine the known Riemannian and Kähler estimates when the first eigenvalue is sufficiently large. We also derive new upper bounds for the diameter of $(M,g)$.

math.DG

Preference Flow Matching with Spectral Factorization for Micro-video Recommendation

Micro-video recommendation aims to infer user preferences from historical interactions and multimodal video content, thereby identifying the next video of interest. However, prevailing methods compress frame sequences into a single holistic representation, entangling the stable visual semantics and the evolving dynamics that jointly shape user preferences. Meanwhile, diffusion- and flow matching-based recommenders condition their generation process solely on coarse behavioral context, leaving its internal temporal structure outside preference formation. We therefore propose PrismRec, a Preference Flow Matching framework with Spectral Factorization for Micro-video Recommendation. Analogous to a prism that disperses white light into its constituent spectrum, PrismRec devises Spectral Semantic Factorization (SSF) to derive complementary static semantic and dynamic factors from frame-level representations via a prior-guided learnable frequency mask in the temporal frequency domain. Then, it proposes Context-Calibrated Preference Matching (CPM) to weigh them with each user's specific sensitivity and inject the calibrated context as a structured condition to steer the matching trajectory toward the target representation, making video content as an intrinsic driver of preference formation rather than auxiliary side information. Experiments on four datasets from two platforms show that PrismRec surpasses the SOTA baseline by up to 22.65%, with the lowest inference cost and peak memory among the compared methods.

cs.IR

S-EMBER: A Large-Scale Benchmark for Streaming Egocentric Memory Retrieval

As wearable devices enable continuous first-person recording, AI assistants must reason across long time horizons to recall past experiences-a capability known as episodic memory. Current benchmarks often rely on offline evaluation with access to entire video files, failing to simulate the streaming reality of wearable intelligence. We introduce S-EMBER (Streaming Egocentric Memory Benchmark for Episodic Retrieval), a large-scale benchmark comprising 3,141 videos totaling 388 hours of organic activity captured via Ray-Ban Meta smart glasses. S-EMBER formalizes grounded streaming episodic retrieval, a paradigm shift from global offline search to causal, active recall triggered by visual events in a continuous stream. We provide 9,448 QA pairs requiring manual visual proof through precise temporal localization and supporting flexible response lengths to simulate natural human-AI interaction. Our extensive benchmarking of frontier models reveals a grounded recall gap: models answer and localize with moderate competence in isolation, yet fall furthest short of human performance when both must hold for the same query, the strongest reaching less than half the human rate. S-EMBER establishes a hardware-authentic foundation for developing grounded, reliable episodic memory in the next generation of wearable AI agents.

cs.CV

Learning-to-Transition for Large-scale and High-Order MIMO Detection

High-order multiple-input multiple-output (MIMO) detection requires efficient search over a large discrete symbol space while producing reliable soft information for channel decoding. This paper develops a learning-to-transition (L2T) framework that formulates MIMO detection as a stochastic sequence of complete-vector transitions. At each transition, a channel-coupled Transformer updates both the instance embedding and the sampling policy, while a blockwise autoregressive factorization captures inter-stream dependence with moderate sequential complexity. For hard-output detection, a transition network is applied recursively and trained through a residual-to-BER curriculum, which first learns the MIMO search geometry from the exact residual metric and then aligns the policy with transmitted-bit accuracy. For soft-output reception, the well-trained hard policy is cloned at the parameter level into every layer of an untied soft-input soft-output iterative detection and decoding (IDD) receiver. This tied-to-untied transfer preserves the learned zero-prior search dynamics while enabling layer- and round-specific specialization under decoder feedback. Within each IDD round, decoder priors tilt candidate generation according to Bayes' rule, and likelihood-weighted terminal hypotheses produce posterior and extrinsic log-likelihood ratios for LDPC decoding. A multi-stage training strategy further stabilizes the hard-to-soft transfer by progressively exposing the receiver to synthetic and in-loop decoder-generated priors.

cs.IT

Reconstruction of Torsion-Free Abelian Groups from Rational Group Fields

For a torsion-free abelian group $G$, let \[ K_{\mathbb{Q}}(G)=\operatorname{Frac} \mathbb{Q}[G] \] be the fraction field of its rational group algebra. We prove that this field determines the group up to isomorphism: \[ K_{\mathbb{Q}}(G) \cong K_{\mathbb{Q}}(H) \quad \Longleftrightarrow \quad G \cong H. \] The main structural input is that, over every field $k$ of characteristic zero, the monomial defect group \[ Δ_k(G)=K_k(G)^\times/(k^\times X^G) \] is free abelian. We give a self-contained proof. It first treats the one-variable Puiseux field $F(t^{\mathbb{Q}})$: factorization from $F(t^{1/n!})$ to $F(t^{1/(n+1)!})$ gives split inclusions because the substituted irreducibles are square-free in characteristic zero. A transfinite decomposition of the divisible hull $\mathbb{Q} \otimes G$ then proves the general case. Given a field isomorphism, we compare the two monomial subgroups inside the common multiplicative group. Their intersection produces isomorphic subgroups $M\leq G$ and $N\leq H$, while the quotients $G/M$ and $H/N$ embed in free defect groups and are therefore free. Relative transcendence degree shows that these two free quotients have the same rank, completing the reconstruction. As consequences, rational-function stabilization of group fields exactly records free stabilization of groups, and Rickard's bounded sequence group yields a field $F$ with $F\cong F(x,y)$ but $F\not\cong F(x)$.

math.RA

RSstitcher -- Merging 2D diffraction frames for Wide Range Reciprocal Space Maps with absorption correction and integration functions

Wide Range Reciprocal Space Mapping (WRRSM) is a technique that allows visualisation of the geometric relationships among multiple hkl spots in a whole reciprocal space map. However, commercial softwares for WRRSMs generation are associated with several issues or limitations, which are overcome by the currently reporting open-source python program RSstitcher (Reciprocal Space Stitcher). RSstitcher merges 2D scan frames formats supported by FabIO and enables WRRSM function on most laboratory X-ray diffractometers equipped with a goniometer cradle and a 2D detector of any sensor size. It is so far the only WRRSMs generation tool that applies diffraction intensity correction due to sample self-absorption, which enables quantitative analyses for WRRSMs including texture measurement and 1D data integration. The conversion equations used in the python program is explained geometrically, including novel WRRSM measurements in ω-ϕ compensated Side Inclination Grazing Incident Diffraction mode for thin film samples. The applications of RSstitcher for bulk and thin film samples are demonstrated using two common 2D X-ray diffraction systems.

physics.ins-det

LogiDroid: Individual Functional Test Generation via Business Logic Extraction and Adaptation

Functional testing is essential for verifying that the business logic of mobile applications aligns with user requirements. Despite its importance, functional testing remains heavily dependent on manual effort due to two core challenges. First, acquiring and reusing business logic from unstructured requirements remains difficult, which hinders the understanding of specific functionalities. Second, a significant semantic gap exists when adapting business logic to the diverse GUI environments, which hinders the generation of test cases for specific mobile applications. To address the preceding challenges, we propose LogiDroid, a two-stage approach that generates individual functional test cases by extracting business logic and adapting it to target applications. First, in the Knowledge Retrieval and Fusion stage, two LLM-based agents are employed to construct a functional test dataset, retrieve relevant test cases, and extract structured business logic for the target functionality. Second, in the Context-Aware Test Generation stage, two other LLM-based agents jointly analyze the extracted business logic and the real time GUI environment to incrementally generate context adaptive functional test cases. This design allows LogiDroid to accurately understand application semantics and use domain expertise to generate complete test cases with verification assertions. We assess the effectiveness of LogiDroid using two widely-used datasets that cover 28 real-world applications and 190 functional requirements. Experimental results show that LogiDroid successfully tested 40% of functional requirements on the FrUITeR dataset (an improvement of over 25% compared to the state-of-the-art approaches) and 65% on the Lin dataset (an improvement of over 55% compared to the state-of-the-art approaches). These results demonstrate the significant effectiveness of LogiDroid in functional test generation.

cs.SE

Self-Evolving In-Context Learning for Direct Pilot-to-Beamformer Design in MU-MISO Systems

We develop an enhanced in-context learning (ICL) framework to improve the performance of pilot-based beamforming in multi-user multiple-input single-output (MU-MISO) systems. The proposed scheme integrates the ICL-Transformer backbone with the pilot encoder-decoder network (EDN) and the beamformer EDN. A crucial feature of our ICL network is that it can handle multiple channel models without retraining, enabled by the construction of model-specific context datasets. To improve convergence and robustness, we introduce three key innovations: (a) a curriculum learning (CL) strategy that smoothly transitions from supervised LMMSE-labeled imitation to unsupervised sum-rate maximization, (b) a self-evolving mechanism that dynamically expands and refines the context datasets for all channel models during CL-based training, and (c) a mismatch-aware extension that incorporates several mismatches into the general ICL framework and bypasses explicit channel calibrations. Ablation studies validate the effectiveness of the in-context architecture and enhanced training strategies. Simulation results over diverse communication environments show that the proposed scheme is able to rapidly adapt to both seen and unseen channel models without gradient-based parameter updates, and can mitigate the mismatch issues via intelligent context constructions. Furthermore, our scheme consistently outperforms the existing beamforming schemes under pilot-based settings, including the WMMSE benchmark and the recent Transformer-based methods.

cs.LG

SPID-Chain: Verifiable Polar-Coded State Validation for Cross-Chain DAG Settlement

Cross-chain settlement must preserve safety across heterogeneous ledgers while tolerating delayed computation, Byzantine participants, and adversarial transaction issuance. This paper presents SPID-Chain, an adapter-compatible settlement architecture for escrow-backed fungible transfers across programmable blockchains. SPID-Chain maintains settlement state through persistent Polar-coded fragments, validates candidate state transitions using hidden linear verification checks, and records certified transfers in a weighted directed acyclic graph (DAG). The design separates native-chain finality from cross-chain settlement: source-chain finality establishes an immutable reservation, whereas weighted DAG confirmation determines when the corresponding destination credit becomes executable. We derive an exact recovery-time distribution for heterogeneous coded workers, a verification-soundness bound for Byzantine responses, and an exact weighted-quorum condition for conflicting-block safety. These components are coupled in a cross-layer stability theorem showing how the coded-validation completion probability determines the effective honest issuance rate and, consequently, the stable adversarial-load region of the settlement DAG. We further establish an end-to-end settlement guarantee covering balance non-negativity, asset conservation, conflict exclusion, replay protection, coded-state consistency, and finite expected lock-to-release latency under the stated liveness conditions. Prototype-assisted simulations indicate that coded validation reduces sensitivity to stragglers, improves validation and confirmation throughput under heterogeneous delays, and produces the predicted transition between stable and unstable DAG operation. The resulting framework provides a verifiable and analytically grounded settlement layer without modifying the native consensus protocol of participating chains.

cs.DC

CMSL: Constructive Multi-Sequence Learning for Recommendation Systems

Sequence learning has emerged as the promising paradigm in recommendation systems, surpassing traditional Deep Learning Recommendation Models (DLRM) by capturing the temporal nuances of user behavior. However, current state-of-the-art architectures operate under a limiting analogy: they treat user history as a monolithic chronological sequence like a sentence in a Large Language Model (LLM). We observe a fundamental divergence between natural language and recommendation data: unlike the linear, logical flow of text, user history is inherently multi-faceted. A user's journey is a fragmented reflection of diverse interests, resulting in much weaker coherence between items than is found in LLM training data. This lack of structural unity leads to context pollution. In single-sequence modeling, unrelated behaviors compete for the same attention budget. This "noisy" signal dilutes the model's focus, effectively capping its ability to discern high-intent patterns from background activity. To address this, we propose Constructive Multi-Sequence Learning (CMSL), a paradigm shift from passive sequence ingestion to active "context engineering" that constructs multiple coherent sequences in latent space. CMSL leverages a learnable Sequence Construction Module to disentangle user history into "pure" thematic strands, followed by a linear attention mechanism to efficiently model these strands at scale. CMSL has been deployed across ranking and retrieval tasks and across four major surfaces at Meta.

cs.IR

Secure Decentralized Federated Learning via Gossip and Virtual Voting

Decentralized federated learning (DFL) removes the central server by letting nodes exchange model updates through peer-to-peer gossip, but existing gossip-based methods often lack provenance finality and resilience to Byzantine or lazy participants. Ledger-assisted federated learning (FL) improves auditability, yet blockchains, shards, or settlement committees can reintroduce global coordination costs that conflict with DFL locality. This paper proposes \emph{gspDAG-FL}, a secure DFL framework that derives consensus from the same gossip history used to disseminate models. Nodes exchange model payloads only with neighbors, while full nodes collect event certificates and receiver-endorsed accepted gossip proofs, reconstruct a compact Topology directed acyclic graph (DAG), and run Hashgraph-style virtual voting followed by compact full-node certificates. Finality is over unique model-origin tuples, not identical local parameter states. To improve resilience, gspDAG-FL combines payload validation, accepted-proof validation, and private semantic audit before aggregation. We formalize the adversarial setting, prove safety and conditional liveness of the control plane, and give a convergence guarantee for certified perturbed gossip under time-varying effective mixing. Experiments on MNIST classification and Penn Treebank language modeling, using fair held-out validation/audit data and networks up to \(N=100\), show that gspDAG-FL achieves learning quality close to validation-based ledger FL while reducing coordination bottlenecks, improving throughput, and maintaining high invalid-origin detection under mixed Byzantine and lazy participation.

cs.LG

Keyless Covert Communication Over Quantum MACs with General Message Sets

We study covert classical communication over quantum multiple-access channels (MACs) with general message sets. Specifically, we consider a fully quantum MAC with arbitrary message sets and an arbitrary number of transmitters. We demonstrate the feasibility of achieving a positive covert rate over this channel and establish general one-shot and asymptotic achievable rate regions. For classical-quantum MACs with general message sets, we establish the covert capacity, when the transmitters are restricted to deterministic encoding. Our result recovers, as a special case, known results for classical communication over classical MACs with general message sets, covert communication of a classical message over a classical channel with two transmitters, and classical communication over quantum MACs. We provide three examples of MACs to which our results can be applied, either directly or indirectly, to achieve positive covert rates. Specifically, we first study covert communication over a finite-dimensional MAC with a helper. We then analyze a classical Gaussian MAC with a helper and derive its covert capacity. Finally, we extend the analysis to a single-mode bosonic MAC with a helper and show that positive covert rates can also be achieved in this setting. To the best of our knowledge, this is the first work to achieve positive-rate covert communication over both classical and quantum MACs.

cs.IT

Covert Communication Over a Quantum MAC with a Helper

We study covert classical communication over a quantum multiple-access channel (MAC) with a helper. Specifically, we consider three transmitters, where one transmitter helps the other two transmitters communicate covertly with a receiver. We demonstrate the feasibility of achieving a positive covert rate over this channel and establish an achievable rate region. Our result recovers as a special case known results for classical communication over classical MACs with a degraded message set, classical communication over quantum MACs, and classical communication over MACs with a helper. To the best of our knowledge, our result is the first to achieve covert communication with positive rates over both classical and quantum MACs.

cs.IT

Ontology-Guided Evidence Path Inference for Multi-hop Knowledge Graph Question Answering

Knowledge graph question answering (KGQA) aims to answer natural-language questions by reasoning over structured facts. Existing multi-hop KGQA methods mainly rely on topic-centered expansion, which faces two key challenges: the search space rapidly grows with noisy mixed-type paths, and retrieved paths may fail to satisfy the semantic constraints of complex questions. To address these challenges, we propose OPI, an ontology-guided evidence path inference framework for multi-hop KGQA. OPI introduces a relation-centric ontology graph to capture the head-tail type constraints of relations, providing a compact interface for answer-side constraints. Based on this ontology graph, OPI first introduces a bidirectional retrieval mechanism by mapping the predicted answer type to compatible final-hop relations and combining topic-side prefix expansion with answer-side final-hop matching, thereby suppressing noisy mixed-type expansion. OPI further adopts an iterative refinement strategy to reassess retrieved paths and candidate answers under the question context, filtering type-compatible but question-irrelevant evidence for more reliable answer prediction. Experiments on WebQSP, CWQ, and MetaQA show that OPI substantially reduces the search space, improves Hit@1/F1 by 4.6/5.0 points on WebQSP and 8.9/3.3 points on CWQ over the strongest prior results, and achieves near-saturated Hit@1 on MetaQA with the retrieval module alone.

cs.AI

Frustration from Localized Zhang-Rice States: A Unified Theory of Doping-Driven Magnetic Transitions in Cuprates

The microscopic mechanism by which doped holes disrupt the antiferromagnetic order is one of the fundamental questions in cuprates. In this work, we propose a unified microscopic theory in which doped holes form spatially localized Zhang-Rice singlets which actively mediate emergent spin exchange. Rather than acting as simple non-magnetic vacancies, these localized states introduce emergent next-nearest $J_2$ and third-nearest $J_3$ neighbor superexchanges. This dopant-induced exchange pathway generates significant magnetic frustration, naturally explaining the rapid collapse of the Néel AFM order and the emergence of a spin-glass phase on the hole-doped side. Our findings provide a comprehensive framework for understanding the complex doping-driven magnetic phase transitions and magnetic electron-hole asymmetry in lightly doped cuprates.

cond-mat.str-el