SearcharxivSearch

arXiv subjects

Xingyu Zhou

Publications and source records attributed to Xingyu Zhou.

At least 19 recordsLinked to original sources

Optimal Low-Rank Quantum State Tomography with Bounded-Sample Joint Measurements

We determine the optimal sample complexity of low-rank quantum state tomography when each measurement may act jointly on at most $t$ samples. For sufficiently small $\varepsilon$, estimating an unknown state on $\mathbb{C}^d$ of rank at most $r$ to trace norm error $\varepsilon$ with constant success probability requires, and is achievable with, $$ \Theta\left( \frac{dr}{\varepsilon^2} \max\left\{1,\frac r{\sqrt t}\right\} \right)$$ samples. The lower bound allows the protocol to choose each joint measurement adaptively using all previous classical outcomes; the matching upper bound is nonadaptive. Thus joint measurements on at most $t$ samples improve the complexity of algorithms making single-sample measurements by at most a factor $\sqrt t$. Further, measuring order $r^2$ samples jointly is necessary and sufficient to attain the unrestricted collective rate. For the lower bound, we vary the support of a state with fixed uniform spectrum and bound the Fisher information trace of every joint measurement on $t$ samples. The adaptive Fisher chain rule and the van Trees inequality then give the trace norm lower bound. For the upper bound, we construct and analyze a nonadaptive tomography protocol based on a Gaussian joint measurement. An explicit second moment identity and a conditional Gaussian law outside the state's support give a rank-dependent error analysis, yielding the matching rate.

quant-ph

A Graph Foundation Model for Large-Scale MIMO Detection

Large-scale multiple-input multiple-output (MIMO) detection is fundamental to modern wireless networks but constrained by performance-complexity trade-offs. Existing detectors, whether classical or learning-based, often fall short in either scalability or generalizability across heterogeneous scenarios. To overcome these limitations, we introduce a wireless-native graph foundation model (GFM) tailored for large-scale MIMO detection. The proposed GFM employs a physics-informed hybrid architecture, integrating the local correlation extraction of message passing neural networks with the global attention of graph Transformers, encoding the physical interference patterns from the expectation propagation algorithm. Via extensive pre-training, this synergy enables the learning of a general-purpose detection mapping scalable across antenna dimensions and channel conditions. For rapid downstream deployment, parameter-efficient fine-tuning is leveraged to adapt the GFM to specific non-ideal system regimes with minimal overhead. To enhance inference efficiency, a mixture-of-experts mechanism is embedded at downstream deployment to dynamically activate only the necessary sub-modules. Evaluations show that the proposed GFM consistently outperforms classical detectors and advanced data-driven baselines in accuracy, configuration generality, and cross-scenario transferability across various challenging zero-shot and few-shot conditions.

cs.IT

Advanced Pixel Diffusion Model with Guided Sparse Global Refinement

Pixel-space diffusion has recently emerged as a promising direction for high-fidelity image generation by modeling images directly in the original pixel domain. However, pixel-space diffusion is computationally demanding due to the extremely high dimensionality of natural images. For efficiency, existing pixel diffusion models either compromise fine details with large-patch tokenization or confine subsequent refinement within individual patches. Such intra-patch refinement inevitably restricts structural continuity across patch boundaries and long-range token interactions, limiting refinement quality. To address these issues, we propose PixSGR, a novel Pixel diffusion framework with Sparse Global Refinement tailored for modeling the distribution of natural images directly in pixel space. PixSGR starts from a supervised low-channel bottleneck to efficiently capture the low-dimensional manifold of natural images. It then progressively expands the channel dimensionality and spatial resolution to recover increasingly fine-grained structures. At the spatial refinement stage, coarse-scale attention maps preselect globally relevant interactions to pre-sparsify fine-scale attention, enabling non-local refinement beyond isolated patches without the quadratic cost of dense attention. Extensive experiments on ImageNet validate the effectiveness of PixSGR. It achieves an FID of 1.51 at 256$\times$256 and maintains performance when scaled to 512$\times$512, attaining an FID of 1.60.

cs.CV

A superflare of BP Tau simultaneously caught by EP X-ray and TESS optical observations

Multiwavelength observations of stellar flares trace the activity of different components of the stars' outer atmosphere, providing insight into their interactions. In the present paper, we report a superflare from BP Tau, simultaneously observed with the Wide-field X-ray Telescope (WXT) on board the Einstein Probe (EP) satellite and TESS. While we attribute the X-ray flux increase to a magnetically powered flare, the optical light curve likely results from the superposition of the flare and an accretion burst. The X-ray flare has a mean flux of $(1.5^{+0.3}_{-0.4})\times10^{-11}$ erg cm$^{-2}$ s$^{-1}$ in the WXT energy band (0.5-4.0 keV), with e-folding times of $1.7\pm1.0$ ks and $14\pm5$ ks for the rise and decay phase, respectively. The corresponding time-integrated flare energy is $(1.0\pm 0.2)\times 10^{36}$ erg. The optical flare has an e-folding time of $0.33\pm0.04$ ks for the rise phase, but the data do not constrain the decay timescale. Assuming a decay phase equal to the rise phase, the resulting optical flare energy is $(2.8\pm0.4)\times10^{34}$ erg in the TESS band ($\sim6,000$-$\sim10,000$ \AA), corresponding to a bolometric energy of $(1.9\pm0.3)\times10^{35}$ erg (assuming a blackbody at 11000 K). The Follow-up X-ray Telescope (FXT) on EP triggered an observation $\sim1.5$ day after the flare, with a flux of $(4.6^{+0.2}_{-0.5})\times10^{-13}$ erg cm$^{-2}$ s$^{-1}$ (0.5-10.0 keV), indicating that BP Tau had returned to quiescence. This work demonstrates the potential of jointly analyzing EP and TESS data for superflares. WXT is expected to detect $\sim800$ superflares per year, with FXT capable of slewing to the flaring star within $\sim3$-5 minutes. The large field of view of both missions offers us the opportunity to study multiwavelength variability during energetic flares.

astro-ph.SR

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination

Embodied cognition requires agents to connect high-level task reasoning with the physical states to be achieved. We introduce Hy-Embodied-RxBrain, an embodied cognition foundation model with joint language-visual reasoning and imagination. Unlike vision-language models that emphasize scene understanding and textual decision making, or generative world models that mainly predict future visual states, RxBrain represents embodied plans in a single planning sequence where language and visual imagination play complementary roles. Language provides the abstract structure of a plan, including task decomposition, planning primitives, constraints, temporal order, and decision logic, while visual imagination grounds this structure through world state prediction and joint subgoal planning, associating each planning step with intermediate and final physical states. RxBrain adopts a unified multimodal Mixture-of-Transformers architecture that supports language, image, and video understanding and generation within one model. To train this capability, we build an automatic pipeline that converts embodied videos into joint text-visual planning supervision by decomposing videos into planning steps and aligning them with visual state transitions. We further introduce RxBrain-Bench to evaluate whether models can represent embodied plans through joint textual and visual components rather than separate understanding or generation. Experiments show that RxBrain maintains embodied understanding and generation abilities, and produces plans with coupled textual reasoning, world state prediction, and joint subgoal planning. We also extend RxBrain to continuous robot action generation, where it shows promising real-robot performance without large-scale action-data pretraining. These results provide an initial step toward foundation models for embodied cognition.

cs.AI

Exoplanetary Tour Design with Solar Sails: TheAntipodes Results in the GTOC13 Problem

Solar sails present an attractive but challenging propulsion method for large-scale, long-duration trajectory design problems. In 2025, the 13th Global Trajectory Optimization Competition (GTOC13) presented a trajectory design problem involving an exoplanetary solar sailing spacecraft in the fictional Altaira system, where the goal is to collect scientific return from flybys of planets, comets, and asteroids. High-scoring solutions combine combinatorial gravity assist tour design with continuous solar sail trajectory optimization. This paper presents the solution approach developed by the team `TheAntipodes' during GTOC13. The approach combines several search and optimization stages: (1) trade studies to identify competitive entry opportunities, (2) large-scale beam search over ballistic gravity assist tours to identify beneficial planetary structures, (3) resonant targeting strategies for Vulcan flyby sequences, and (4) multi-leg solar sail trajectory refinement using sequential convex programming (SCP). A key component of the refinement process is the use of a lossless control-convex solar sail formulation, which allows for large portions of the trajectory, including all gravity assist geometry and flyby timing, to be optimized simultaneously to maximize score. The resulting trajectory placed third, with a score of 337.878 from 133 scoring flybys, and exhibited a structure broadly similar to those of the other high-scoring solutions. This demonstrates the scalability of methods such as SCP for very large trajectory design problems.

astro-ph.IM

Low-Overhead Receiver Design for Data-Dependent Superimposed Training via Deep Learning

Superimposed pilot (SIP) transmission improves spectral efficiency by eliminating the dedicated pilot overhead required in orthogonal pilot (OP)-based schemes. However, SIP suffers from severe pilot-data coupling, which leads to a critical performance-complexity bottleneck at the receiver. To address this issue, this paper proposes a low-overhead transmission framework that revitalizes data-dependent superimposed training (DDST) with enhanced interference mitigation strategies. First, for quasi-static block-fading channels, an enhanced DDST receiver is developed to achieve non-iterative pilot-data decoupling by exploiting data-dependent algebraic structures. Second, to overcome the sensitivity of conventional DDST to channel variations and symbol misidentification in fast time-varying environments, a mix transmission scheme is developed. By strategically applying DDST to a subset of resource elements, the proposed scheme combines the interference-free transmission property of OP with the zero-pilot-overhead advantage of SIP, thereby improving demapping reliability and interference suppression. Furthermore, under the proposed mix scheme, a Vision Transformer-based neural receiver is designed to capture the orthogonal structure between pilots and perturbation-bearing data, as well as the underlying channel correlations, thereby relaxing the stringent quasi-static assumption required for interference disentanglement. Simulation results demonstrate that the proposed framework achieves significant performance gains in the low-to-medium SNR regime under time-varying channels while providing superior computational efficiency compared with state-of-the-art SIP receivers.

cs.IT

When Determinants Are Not Enough: Private Rare Switching

In this note, I would like to share a small research moment where Codex helped me find the right way to adapt rare switching to the private setting. The standard determinant-based update rule in linear bandits and RL works beautifully because the design matrix grows monotonically. But once Gaussian noise is added for privacy, this monotonicity can fail, and the usual analysis no longer goes through. The key reason is that determinant growth controls volume, while regret analysis needs control of the worst direction. To address this, Codex comes up with a different rare-switching rule based on the generalized Rayleigh quotient, which restores logarithmic policy updates and the desired confidence-width comparison up to a constant factor. I present my manually clean-up version of the proof here as well as some personal reflection on this example.

cs.LG

Towards Differentially Private Reinforcement Learning with General Function Approximation

We present the first theoretical guarantees for differentially private online reinforcement learning (RL) with general function approximation, extending beyond prior work restricted to tabular and linear settings. Our approach combines a batched policy update scheme with the exponential mechanism, together with a novel regret analysis. We show that, even under general function approximation, the regret in the model-free setting under differential privacy matches the state of the art for the linear case, scaling as $\widetilde{O}(K^{3/5})$, where $K$ denotes the number of episodes. As an important by-product, we also establish the first regret bound for online RL with batch update that depends on the standard complexity measure of coverability, complementing existing results based on a newly introduced Eluder-Condition class. In addition, we uncover fundamental gaps in recent results for private RL with linear function approximation, thereby clarifying its landscape.

cs.LG

Robust Angles-Only Initial Relative Orbit Determination Using Polynomial Optimization

This paper develops a robust angles-only IROD method based on polynomial optimization for arbitrary nonlinear dynamics. First, the relative motion is approximated by high-order Taylor polynomials within the differential algebra framework, and the resulting cross-product-residual minimization problem is solved through a recursive polynomial optimization procedure. Second, a reduced-order weighting strategy is introduced by projecting the residual onto the two-dimensional tangent subspace of the line of sight, thereby structurally removing the intrinsic singularity of conventional three-dimensional weighting. Third, a zero-solution-avoidance constraint together with an adaptive threshold-selection mechanism is developed to improve robustness against poor initialization, strong measurement noise, and unfavorable observation geometries. Numerical simulations show that the proposed method improves IROD accuracy by about three orders of magnitude relative to the baseline methods, while also reducing the downstream orbit-refinement burden. The reduced-order weighting strategy further improves accuracy by about 43% in the nominal case and remains stable under large-noise conditions, outperforming the conventional three-dimensional weighting by about 81%.

astro-ph.IM

State Transfer Reveals Reuse in Controlled Routing

Prompt-based interventions can change model behavior, but trained success alone does not identify where the behaviorally relevant state is represented. We study this question in controlled routing tasks using interfaces chosen on support data, held-out query evaluation, and matched necessity, sufficiency, and wrong-interface controls. On GPT-2 triop, an early interface supports exact transfer under these tests. On GPT-2 add/sub, zero-retrain compiled transfer at the fixed interface recovers most of donor routing accuracy, while trainable prompt slots can relearn the same behavior at several other positions only after additional support examples and optimization. These results distinguish fixed-interface reuse from prompt relocation in a setting where the two can be tested directly. Qwen routing provides a cross-architecture consistency check for the same matched-interface pattern at the operator token, although donor-specific identity on the local V-path remains unresolved. Generation and reasoning branches are used to map scope: they show broader transport or weaker controller identifiability once control depends on longer trajectories or harder selection. In controlled routing, fixed-interface transfer is therefore stronger evidence of reuse than trained prompt success alone.

cs.AI

A Control Architecture for Training-Free Memory Use

Prompt-injected memory can improve reasoning without updating model weights, but it also creates a control problem: retrieved content helps only when it is applied in the right state. We study this problem in a strict training-free setting and formulate it as applicability control: when to trigger a memory-assisted second pass, when to trust it, and how to maintain the memory bank over time. Our method combines uncertainty-based routing, confidence-based selective acceptance, bank selection across rule and exemplar memory, and evidence-based governance of the memory bank over time. Under a locked training-free protocol with compute-matched controls, it improves two core arithmetic benchmarks by +7.0 points on SVAMP and +7.67 points on ASDiv over baseline. The same architecture also transfers to QA and agent benchmarks with smaller positive effects and shows the same positive direction on a second checkpoint for the main arithmetic tasks. On arithmetic, the main empirical pattern is that the control architecture, rather than raw memory exposure, drives the improvements on SVAMP and ASDiv. Mechanistically, confidence separates helpful from harmful rule-bank interventions, and under fixed retrieval the repair-versus-corrupt difference localizes to rows whose retrieved set actually contains the edited entries.

cs.AI

Taming Sampling Perturbations with Variance Expansion Loss for Latent Diffusion Models

Latent diffusion models have emerged as the dominant framework for high-fidelity and efficient image generation, owing to their ability to learn diffusion processes in compact latent spaces. However, while previous research has focused primarily on reconstruction accuracy and semantic alignment of the latent space, we observe that another critical factor, robustness to sampling perturbations, also plays a crucial role in determining generation quality. Through empirical and theoretical analyses, we show that the commonly used $\beta$-VAE-based tokenizers in latent diffusion models, tend to produce overly compact latent manifolds that are highly sensitive to stochastic perturbations during diffusion sampling, leading to visual degradation. To address this issue, we propose a simple yet effective solution that constructs a latent space robust to sampling perturbations while maintaining strong reconstruction fidelity. This is achieved by introducing a Variance Expansion loss that counteracts variance collapse and leverages the adversarial interplay between reconstruction and variance expansion to achieve an adaptive balance that preserves reconstruction accuracy while improving robustness to stochastic sampling. Extensive experiments demonstrate that our approach consistently enhances generation quality across different latent diffusion architectures, confirming that robustness in latent space is a key missing ingredient for stable and faithful diffusion sampling.

cs.CV

Construction of the Global $\chi^2$ Function for the Simultaneous Fitting of Correlated Energy-Dependent Cross Sections

In this paper, the global $\chi^2$ function for the simultaneous fitting of correlated energy-dependent cross sections is constructed, where the correlations between the measured cross sections of different processes and/or at different center-of-mass energy points, as well as the contributions from the integrated luminosity measurement and the center-of-mass energy measurement, are taken into account.

physics.data-an

Improving Channel Estimation via Multimodal Diffusion Models with Flow Matching

Deep generative models offer a powerful alternative to conventional channel estimation by learning complex channel distributions. By integrating the rich environmental information available in modern sensing-aided networks, this paper proposes MultiCE-Flow, a multimodal channel estimation framework based on flow matching and diffusion transformer (DiT). We design a specialized multimodal perception module that fuses LiDAR, camera, and location data into a semantic condition, while treating sparse pilots as a structural condition. These conditions guide a DiT backbone to reconstruct high-fidelity channels. Unlike standard diffusion models, we employ flow matching to learn a linear trajectory from noise to data, enabling efficient one-step sampling. By leveraging environmental semantics, our method mitigates the ill-posed nature of estimation with sparse pilots. Extensive experiments demonstrate that MultiCE-Flow consistently outperforms traditional baselines and existing generative models. Notably, it exhibits superior robustness to out-of-distribution scenarios and varying pilot densities, making it suitable for environment-aware communication systems.

cs.LG

Hierarchical Industrial Demand Forecasting with Temporal and Uncertainty Explanations

Hierarchical time-series forecasting is essential for demand prediction across various industries. While machine learning models have obtained significant accuracy and scalability on such forecasting tasks, the interpretability of their predictions, informed by application, is still largely unexplored. To bridge this gap, we introduce a novel interpretability method for large hierarchical probabilistic time-series forecasting, adapting generic interpretability techniques while addressing challenges associated with hierarchical structures and uncertainty. Our approach offers valuable interpretative insights in response to real-world industrial supply chain scenarios, including 1) the significance of various time-series within the hierarchy and external variables at specific time points, 2) the impact of different variables on forecast uncertainty, and 3) explanations for forecast changes in response to modifications in the training dataset. To evaluate the explainability method, we generate semi-synthetic datasets based on real-world scenarios of explaining hierarchical demands for over ten thousand products at a large chemical company. The experiments showed that our explainability method successfully explained state-of-the-art industrial forecasting methods with significantly higher explainability accuracy. Furthermore, we provide multiple real-world case studies that show the efficacy of our approach in identifying important patterns and explanations that help stakeholders better understand the forecasts. Additionally, our method facilitates the identification of key drivers behind forecasted demand, enabling more informed decision-making and strategic planning. Our approach helps build trust and confidence among users, ultimately leading to better adoption and utilization of hierarchical forecasting models in practice.

cs.LG

ATD: Improved Transformer with Adaptive Token Dictionary for Image Restoration

Recently, Transformers have gained significant popularity in image restoration tasks such as image super-resolution and denoising, owing to their superior performance. However, balancing performance and computational burden remains a long-standing problem for transformer-based architectures. Due to the quadratic complexity of self-attention, existing methods often restrict attention to local windows, resulting in limited receptive field and suboptimal performance. To address this issue, we propose Adaptive Token Dictionary (ATD), a novel transformer-based architecture for image restoration that enables global dependency modeling with linear complexity relative to image size. The ATD model incorporates a learnable token dictionary, which summarizes external image priors (i.e., typical image structures) during the training process. To utilize this information, we introduce a token dictionary cross-attention (TDCA) mechanism that enhances the input features via interaction with the learned dictionary. Furthermore, we exploit the category information embedded in the TDCA attention maps to group input features into multiple categories, each representing a cluster of similar features across the image and serving as an attention group. We also integrate the learned category information into the feed-forward network to further improve feature fusion. ATD and its lightweight version ATD-light, achieve state-of-the-art performance on multiple image super-resolution benchmarks. Moreover, we develop ATD-U, a multi-scale variant of ATD, to address other image restoration tasks, including image denoising and JPEG compression artifacts removal. Extensive experiments demonstrate the superiority of out proposed models, both quantitatively and qualitatively.

cs.CV

VR Calm Plus: Coupling a Squeezable Tangible Interaction with Immersive VR for Stress Regulation

While Virtual Reality (VR) is increasingly employed for stress management, most applications rely heavily on audio-visual stimuli and overlook the therapeutic potential of squeezing engagement. To address this gap, we introduce VR Calm Plus, a multimodal system that integrates a pressure-sensitive plush toy into an interactive VR environment. This interface allows users to dynamically modulate the virtual atmosphere through physical squeezing actions, fostering a deeper sense of embodied relaxation. We evaluated the system with 40 participants using PANAS-X surveys, subjective questionnaires, physiological measures (heart rate, skin conductance, pulse rate variability), and semi-structured interviews. Results demonstrate that, compared to a visual-only baseline, squeeze-based interaction significantly enhances positive affect and perceived relaxation. Physiological data further revealed a state of "active relaxation", characterized by greater reductions in heart rate and preserved autonomic flexibility (PRV), alongside sustained emotional engagement (GSR). Our findings highlight the value of coupling tangible input with immersive environments to support emotional well-being and offer design insights for future VR-based mental health tools.

cs.HC