SearcharxivSearch

arXiv subjects

Zhuo Xu

Publications and source records attributed to Zhuo Xu.

At least 19 recordsLinked to original sources

CARNet: Channel-Adaptive Receiver Network for Robust NextG Communications

Neural receivers have been recognized as a promising paradigm for the next-generation (NextG) communications. However, due to the reliance on a static network optimized for specific channel conditions, their generalization capability across diverse scenarios remains a significant challenge. To address this issue, this paper proposes a novel channel-adaptive neural receiver network (CARNet) based on the mixture-of-experts (MoE) framework. The proposed architecture employs multiple expert networks together with an efficient routing mechanism to enable signal detection in various scenarios. The experts are constructed via stacked ResNet blocks and specialize in robust signal detection within specific channel conditions, while the routing mechanism incorporates a lightweight representation learning module, which projects the coarse channel estimate into a low-dimensional latent embedding. The learned embedding characterizes task-relevant channel conditions and provides efficient guidance for accurate expert selection. Link-level simulation experiments demonstrate that the proposed CARNet achieves superior performance across diverse channel conditions.

cs.IT

Near-Field Communications with Grating Lobes for Quasi-Distributed Arrays: From ULA to MRA

Extremely large-scale antenna array (ELAA) has emerged as a common feature of many key candidate technologies for 6G, where the near-field characteristics become dominant. The quasi-distributed array can further extend the near-field range and utilize the near-field benefits to improve the system performance. However, its typical implementation with modular arrays suffers from severe grating lobes that cause non-negligible inter-user interferences. To solve this problem, we propose the modular minimum-redundancy array (M-MRA) to suppress near-field grating lobes by redesigning the subarray configuration. Specifically, we first characterize the beam pattern of the conventional modular uniform linear array (M-ULA). Contrary to the common belief that grating lobes only exist in the angle domain, we reveal that near-field grating lobes may also occur in the distance domain. We further analyze how to suppress near-field grating lobes for the M-ULA. The results demonstrate that increasing the number of antennas per module can suppress grating lobes. In particular, the number of antennas required grows linearly with the inter-module spacing, thus the grating lobe interferences are severe under a limited number of antennas. This limitation inspires us to propose the M-MRA by redesigning the subarray structure. For M-MRA, the nonuniform antenna spacing within each subarray provides a narrower spatial envelope, allowing it to suppress near-field grating lobes in the angle and distance domains simultaneously. Simulation results verify that the proposed M-MRA can significantly improve the spectrum efficiency of multi-user near-field communications under the same number of antennas.

eess.SP

Accuracy-Configurable Floating-Point Multiplier Design for SRAM-Based Compute-in-Memory

Digital Compute-in-Memory (DCiM) reduces data movement and has become a promising solution for energy-efficient edge AI. However, most existing DCiM frameworks still primarily target integer or fixed-point arithmetic, and provide limited support for compiler-integrated and accuracy-configurable floating-point computation. Directly integrating conventional IEEE 754 floating-point units into dense SRAM-based DCiM arrays, however, incurs high area and power overhead. To address this challenge, this work presents an accuracy-configurable floating-point multiplier integrated into the OpenACM framework for SRAM-based DCiM. An exact IEEE~754-compliant multiplier is first implemented as a baseline, and a mantissa-segmentation-based approximate multiplier is then proposed to reduce hardware cost while preserving numerical fidelity. Post-layout results show up to 69% logic area reduction and 72% power savings over exact floating-point designs without delay overhead. Evaluations on image processing tasks and ResNet-18 inference further demonstrate negligible accuracy degradation. These results indicate that compiler-integrated approximate floating-point multiplication is a practical approach for enabling efficient and configurable floating-point support in SRAM-based DCiM systems. The Floating-Point Multiplier is available on https://github.com/ShenShan123/OpenACM

cs.AR

Observation antibunching with classical light in a linear interferometer

Understanding the boundary between classical and nonclassical phenomena is important for both fundamental researches in quantum optics and applications in quantum information. One of the most interesting research directions in this field is exploring nonclassical effects with classical light. In this paper, we will show that it is possible to observe antibunching with thermal light in a Hanbury Brown-Twiss interferometer by treating single-photon detectors as photon-number-resolving detectors to perform photon-number projection measurements. Both temporal and spatial antibunching is observed via the correlation of two detectors detecting one and zero photon, respectively. By comparing the measured results of thermal and laser light, it is found that the observed antibunching arises from the combined effect of photon statistics of thermal light and photon-number projection measurement.The classical and nonclassical nature of the observed antibunching is analyzed. The results are helpful to understand the connection between classical and nonclassical correlation and may find applications in multiphoton interference and quantum imaging.

quant-ph

Ghost imaging with zero photons

Ghost imaging was first demonstrated with entangled photon pairs and well-known for its peculiar properties. The signal beam that illuminates the object possesses no spatial resolution, whereas the reference beam, which never interacts with the object, is spatially resolved. Either beam alone cannot retrieve the image, which can only be obtained when the signal and reference beams are correlated. Here we will report a ghost imaging experiment with even more peculiar properties, in which the image can be reconstructed when no photon interacts with the object or even no photon in neither signal nor reference beam. All the photons interacted with the object are discarded. Only the time bins with zero photon are employed to retrieve the image, a process referred to as "ghost imaging with zero photons" hereafter. The reason why ghost image can be retrieved with zero photons is jointly determined by photon-number projection measurement and photon statistics of thermal light. The results are helpful to resolve the debate on the physics of ghost imaging and understand the relation between quantum and classical correlations.

quant-ph

Real-to-Sim Robot Policy Evaluation with Gaussian Splatting Simulation of Soft-Body Interactions

Robotic manipulation policies are advancing rapidly, but their direct evaluation in the real world remains costly, time-consuming, and difficult to reproduce, particularly for tasks involving deformable objects. Simulation provides a scalable and systematic alternative, yet existing simulators often fail to capture the coupled visual and physical complexity of soft-body interactions. We present a real-to-sim policy evaluation framework that constructs soft-body digital twins from real-world videos and renders robots, objects, and environments with photorealistic fidelity using 3D Gaussian Splatting. We validate our approach on representative deformable manipulation tasks, including plush toy packing, rope routing, and T-block pushing, demonstrating that simulated rollouts correlate strongly with real-world execution performance and reveal key behavioral patterns of learned policies. Our results suggest that combining physics-informed reconstruction with high-quality rendering enables reproducible, scalable, and accurate evaluation of robotic manipulation policies. Website: https://real2sim-eval.github.io/

cs.RO

OpenHAIV: A Framework Towards Practical Open-World Learning

Substantial progress has been made in various techniques for open-world recognition. Out-of-distribution (OOD) detection methods can effectively distinguish between known and unknown classes in the data, while incremental learning enables continuous model knowledge updates. However, in open-world scenarios, these approaches still face limitations. Relying solely on OOD detection does not facilitate knowledge updates in the model, and incremental fine-tuning typically requires supervised conditions, which significantly deviate from open-world settings. To address these challenges, this paper proposes OpenHAIV, a novel framework that integrates OOD detection, new class discovery, and incremental continual fine-tuning into a unified pipeline. This framework allows models to autonomously acquire and update knowledge in open-world environments. The proposed framework is available at https://haiv-lab.github.io/openhaiv .

cs.CV

Empowering Near-Field Communications in Low-Altitude Economy with LLM: Fundamentals, Potentials, Solutions, and Future Directions

The low-altitude economy (LAE) is gaining significant attention from academia and industry. Fortunately, LAE naturally aligns with near-field communications in extremely large-scale MIMO (XL-MIMO) systems. By leveraging near-field beamfocusing, LAE can precisely direct beam energy to unmanned aerial vehicles, while the additional distance dimension boosts overall spectrum efficiency. However, near-field communications in LAE still face several challenges, such as the increase in signal processing complexity and the necessity of distinguishing between far and near-field users. Inspired by the large language models (LLM) with powerful ability to handle complex problems, we apply LLM to solve challenges of near-field communications in LAE. The objective of this article is to provide a comprehensive analysis and discussion on LLM-empowered near-field communications in LAE. Specifically, we first introduce fundamentals of LLM and near-field communications, including the key advantages of LLM and key characteristics of near-field communications. Then, we reveal the opportunities and challenges of near-field communications in LAE. To address these challenges, we present a LLM-based scheme for near-field communications in LAE, and provide a case study which jointly distinguishes far and near-field users and designs multi-user precoding matrix. Finally, we outline and highlight several future research directions and open issues.

eess.SP

Input-to-state type Stability for Simplified Fluid-Particle Interaction System

In this paper, we study the well-posedness and the input-to-state type stability of a one-dimensional fluid-particle interaction system. A distinctive feature, not yet considered in the ISS literature, is that our system involves a free boundary. More precisely, the fluid is described by the viscous Burgers equation, and the motion of the particle obeys Newton second law. The point mass is subject to both a feedback control and an open-loop control. We first establish the well-posedness of the system for any open-loop input in the L2(0, infinity) space. Assuming the input also belongs to the L1(0,infinity) space, we prove that the particle's position remains uniformly bounded and that the system is input-to-state type stable. The proof is based on the construction of a Lyapunov functional derived from a special test function.

math.AP

EDENet: Echo Direction Encoding Network for Place Recognition Based on Ground Penetrating Radar

Ground penetrating radar (GPR) based localization has gained significant recognition in robotics due to its ability to detect stable subsurface features, offering advantages in environments where traditional sensors like cameras and LiDAR may struggle. However, existing methods are primarily focused on small-scale place recognition (PR), leaving the challenges of PR in large-scale maps unaddressed. These challenges include the inherent sparsity of underground features and the variability in underground dielectric constants, which complicate robust localization. In this work, we investigate the geometric relationship between GPR echo sequences and underground scenes, leveraging the robustness of directional features to inform our network design. We introduce learnable Gabor filters for the precise extraction of directional responses, coupled with a direction-aware attention mechanism for effective geometric encoding. To further enhance performance, we incorporate a shift-invariant unit and a multi-scale aggregation strategy to better accommodate variations in di-electric constants. Experiments conducted on public datasets demonstrate that our proposed EDENet not only surpasses existing solutions in terms of PR performance but also offers advantages in model size and computational efficiency.

cs.CV

OpenEarthSensing: Large-Scale Fine-Grained Benchmark for Open-World Remote Sensing

The advancement of remote sensing, including satellite systems, facilitates the continuous acquisition of remote sensing imagery globally, introducing novel challenges for achieving open-world tasks. Deployed models need to continuously adjust to a constant influx of new data, which frequently exhibits diverse shifts from the data encountered during the training phase. To effectively handle the new data, models are required to detect semantic shifts, adapt to covariate shifts, and continuously update their parameters without forgetting learned knowledge, as has been considered in works on a variety of open-world tasks. However, existing studies are typically conducted within a single dataset to simulate realistic conditions, with a lack of large-scale benchmarks capable of evaluating multiple open-world tasks. In this paper, we introduce \textbf{OpenEarthSensing (OES)}, a large-scale fine-grained benchmark for open-world remote sensing. OES includes 189 scene and object categories, covering the vast majority of potential semantic shifts that may occur in the real world. Additionally, to provide a more comprehensive testbed for evaluating the generalization performance, OES encompasses five data domains with significant covariate shifts, including two RGB satellite domains, one RGB aerial domain, one multispectral RGB domain, and one infrared domain. We evaluate the baselines and existing methods for diverse tasks on OES, demonstrating that it serves as a meaningful and challenging benchmark for open-world remote sensing. The proposed dataset OES is available at https://haiv-lab.github.io/OES.

cs.CV

Global Exponential Stabilization for a Simplified Fluid-Particle Interaction System

This work considers a system coupling a viscous Burgers equation (aimed to describe a simplified model of $1D$ fluid flow) with the ODE describing the motion of a point mass moving inside the fluid. The point mass is possibly under the action of a feedback control. Our main contributions are that we prove two global exponential stability results. More precisely, we first show that the velocity field corresponding to the free dynamics case is globally exponentially stable. We next show that, in the presence of the feedback control both the velocity field and the distance from the mass point to a prescribed target position decay exponentially. The proofs of these results heavily rely on the use of a special test function allowing both to prove that the mass point stays away from the boundary and to construct a perturbed Lyapunov function.

math.AP

PO-GVINS: Tightly Coupled GNSS-Visual-Inertial Integration with Pose-Only Representation

Accurate and reliable positioning is crucial for perception, decision-making, and other high-level applications in autonomous driving, unmanned aerial vehicles, and intelligent robots. Given the inherent limitations of standalone sensors, integrating heterogeneous sensors with complementary capabilities is one of the most effective approaches to achieving this goal. In this paper, we propose a filtering-based, tightly coupled global navigation satellite system (GNSS)-visual-inertial positioning framework with a pose-only formulation applied to the visual-inertial system (VINS), termed PO-GVINS. Specifically, multiple-view imaging used in current VINS requires a priori of 3D feature, then jointly estimate camera poses and 3D feature position, which inevitably introduces linearization error of the feature as well as facing dimensional explosion. However, the pose-only (PO) formulation, which is demonstrated to be equivalent to the multiple-view imaging and has been applied in visual reconstruction, represent feature depth using two camera poses and thus 3D feature position is removed from state vector avoiding aforementioned difficulties. Inspired by this, we first apply PO formulation in our VINS, i.e., PO-VINS. GNSS raw measurements are then incorporated with integer ambiguity resolved to achieve accurate and drift-free estimation. Extensive experiments demonstrate that the proposed PO-VINS significantly outperforms the multi-state constrained Kalman filter (MSCKF). By incorporating GNSS measurements, PO-GVINS achieves accurate, drift-free state estimation, making it a robust solution for positioning in challenging environments.

cs.RO

Anomalous switching pattern in the ferrimagnetic memory cell

Replacing the ferromagnet with ferrimagnet (FiM) in the magnetic tunnel junction (MTJ) allows faster magnetization switching in picoseconds. The operation of a memory cell that consists of the MTJ and a transistor requires reversable magnetization switching. When a constant voltage is applied, we find that the spin-transfer torque can only switch the FiM-MTJ from parallel to antiparallel state. This stems from the small switching window of FiM and the dynamic resistance variation during the magnetization switching. We find the resulting current variation can be suppressed by reducing the magnetoresistance ratio. Furthermore, we demonstrate that the switching window can be expanded by adjusting the amount of Gd in FiM. We predict that the polarity of both switching current (Jc,switch) and oscillation current (Jc,osc) reverses at the angular momentum compensation point but not the magnetization compensation point. This anomalous dynamic behavior is attributed to the different physical nature of magnetization switching and oscillation in FiM, which must be considered when designing FiM-based MRAM.

cond-mat.mes-hall

Orthogonal Spin-Orbit Torque-Induced Deterministic Switching in NiO

The electrical switching of antiferromagnet (AFM) is very important for the development of ultrafast magnetic random-access memory (MRAM). This task becomes more difficult in antiferromagnetic oxide NiO which has complex anisotropy. We show that by utilizing two spin-orbit torques (SOT) from orthogonal currents, one can deterministically switch the magnetic moments of NiO in two electrical distinguishable states that can be read out using the spin Hall magnetoresistance. This deterministic switching relies on the symmetry of SOT on different sublattices, where the sign reversal of magnetic moments leads to constructive torques in the beginning and balanced torques in the end. In addition, we show that the easy-plane anisotropy plays a key role in the switching, which has been ignored in some previous works. The uniform magnetic dynamics in this work provides a clear physical picture in understanding the SOT switching of NiO. Furthermore, the electrical writing and reading function in our device advances the development of AFM-MRAM.

cond-mat.mes-hall

Vision Language Models are In-Context Value Learners

Predicting temporal progress from visual trajectories is important for intelligent robots that can learn, adapt, and improve. However, learning such progress estimator, or temporal value function, across different tasks and domains requires both a large amount of diverse data and methods which can scale and generalize. To address these challenges, we present Generative Value Learning (\GVL), a universal value function estimator that leverages the world knowledge embedded in vision-language models (VLMs) to predict task progress. Naively asking a VLM to predict values for a video sequence performs poorly due to the strong temporal correlation between successive frames. Instead, GVL poses value estimation as a temporal ordering problem over shuffled video frames; this seemingly more challenging task encourages VLMs to more fully exploit their underlying semantic and temporal grounding capabilities to differentiate frames based on their perceived task progress, consequently producing significantly better value predictions. Without any robot or task specific training, GVL can in-context zero-shot and few-shot predict effective values for more than 300 distinct real-world tasks across diverse robot platforms, including challenging bimanual manipulation tasks. Furthermore, we demonstrate that GVL permits flexible multi-modal in-context learning via examples from heterogeneous tasks and embodiments, such as human videos. The generality of GVL enables various downstream applications pertinent to visuomotor policy learning, including dataset filtering, success detection, and advantage-weighted regression -- all without any model training or finetuning.

cs.RO

Imagined Potential Games: A Framework for Simulating, Learning and Evaluating Interactive Behaviors

Interacting with human agents in complex scenarios presents a significant challenge for robotic navigation, particularly in environments that necessitate both collision avoidance and collaborative interaction, such as indoor spaces. Unlike static or predictably moving obstacles, human behavior is inherently complex and unpredictable, stemming from dynamic interactions with other agents. Existing simulation tools frequently fail to adequately model such reactive and collaborative behaviors, impeding the development and evaluation of robust social navigation strategies. This paper introduces a novel framework utilizing distributed potential games to simulate human-like interactions in highly interactive scenarios. Within this framework, each agent imagines a virtual cooperative game with others based on its estimation. We demonstrate this formulation can facilitate the generation of diverse and realistic interaction patterns in a configurable manner across various scenarios. Additionally, we have developed a gym-like environment leveraging our interactive agent model to facilitate the learning and evaluation of interactive navigation algorithms.

cs.RO

Unified Error Correction Code Transformer with Low Complexity

Channel coding is vital for reliable sixth-generation (6G) data transmission, employing diverse error correction codes for various application scenarios. Traditional decoders require dedicated hardware for each code, leading to high hardware costs. Recently, artificial intelligence (AI)-driven approaches, such as the error correction code Transformer (ECCT) and its enhanced version, the foundation error correction code Transformer (FECCT), have been proposed to reduce the hardware cost by leveraging the Transformer to decode multiple codes. However, their excessively high computational complexity of $\mathcal{O}(N^2)$ due to the self-attention mechanism in the Transformer limits scalability, where $N$ represents the sequence length. To reduce computational complexity, we propose a unified Transformer-based decoder that handles multiple linear block codes within a single framework. Specifically, a standardized unit is employed to align code length and code rate across different code types, while a redesigned low-rank unified attention module, with computational complexity of $\mathcal{O}(N)$, is shared across various heads in the Transformer. Additionally, a sparse mask, derived from the parity-check matrix's sparsity, is introduced to enhance the decoder's ability to capture inherent constraints between information and parity-check bits, improving decoding accuracy and further reducing computational complexity by $86\%$. Extensive experimental results demonstrate that the proposed unified Transformer-based decoder outperforms existing methods and provides a high-performance, low-complexity solution for next-generation wireless communication systems.

cs.IT