SearcharxivSearch

arXiv subjects

Xuemin Hong

Publications and source records attributed to Xuemin Hong.

9 recordsLinked to original sources

IONext: Unlocking the Next Era of Inertial Odometry

Researchers have increasingly adopted Transformer-based models for inertial odometry. While Transformers excel at modeling long-range dependencies, their limited sensitivity to local, fine-grained motion variations and lack of inherent inductive biases often hinder localization accuracy and generalization. Recent studies have shown that incorporating large-kernel convolutions and Transformer-inspired architectural designs into CNN can effectively expand the receptive field, thereby improving global motion perception. Motivated by these insights, we propose a novel CNN-based module called the Dual-wing Adaptive Dynamic Mixer (DADM), which adaptively captures both global motion patterns and local, fine-grained motion features from dynamic inputs. This module dynamically generates selective weights based on the input, enabling efficient multi-scale feature aggregation. To further improve temporal modeling, we introduce the Spatio-Temporal Gating Unit (STGU), which selectively extracts representative and task-relevant motion features in the temporal domain. This unit addresses the limitations of temporal modeling observed in existing CNN approaches. Built upon DADM and STGU, we present a new CNN-based inertial odometry backbone, named Next Era of Inertial Odometry (IONext). Extensive experiments on six public datasets demonstrate that IONext consistently outperforms state-of-the-art (SOTA) Transformer- and CNN-based methods. For instance, on the RNIN dataset, IONext reduces the average ATE by 10% and the average RTE by 12% compared to the representative model iMOT.

cs.CV

FTIN: Frequency-Time Integration Network for Inertial Odometry

Inertial odometry (IO) leverages inertial measurement unit (IMU) signals for cost-effective localization. However, high IMU sampling rates introduce substantial redundancy that impedes IO's ability to attend to salient components, thereby creating an information bottleneck. To address this challenge, we propose a cross-domain IO framework that fuses information from the frequency and time domains. Specifically, we exploit the global context and energy-compaction properties of frequency-domain representations to capture holistic motion patterns and alleviate the bottleneck. To the best of our knowledge, this is among the first attempts to incorporate frequency-domain feature processing into IO. Experimental results on multiple public datasets demonstrate the effectiveness of the proposed frequency--time-domain fusion strategy.

cs.RO

StarIO: A Lightweight Inertial Odometry for Nonlinear Motion

Inertial odometry (IO) directly estimates the position of a carrier from inertial sensor measurements and serves as a core technology for the widespread deployment of consumer grade localization systems. While existing IO methods can accurately reconstruct simple and near linear motion trajectories, they often fail to account for drift errors caused by complex motion patterns such as turning. This limitation significantly degrades localization accuracy and restricts the applicability of IO systems in real world scenarios. To address these challenges, we propose a lightweight IO framework. Specifically, inertial data is projected into a high dimensional implicit nonlinear feature space using the Star Operation method, enabling the extraction of complex motion features that are typically overlooked. We further introduce a collaborative attention mechanism that jointly models global motion dynamics across both channel and temporal dimensions. In addition, we design Multi Scale Gated Convolution Units to capture fine grained dynamic variations throughout the motion process, thereby enhancing the model's ability to learn rich and expressive motion representations. Extensive experiments demonstrate that our proposed method consistently outperforms SOTA baselines across six widely used inertial datasets. Compared to baseline models on the RoNIN dataset, it achieves reductions in ATE ranging from 2.26% to 65.78%, thereby establishing a new benchmark in the field.

cs.RO

CKANIO: Learnable Chebyshev Polynomials for Inertial Odometry

Inertial odometry (IO) relies exclusively on signals from an inertial measurement unit (IMU) for localization and offers a promising avenue for consumer grade positioning. However, accurate modeling of the nonlinear motion patterns present in IMU signals remains the principal limitation on IO accuracy. To address this challenge, we propose CKANIO, an IO framework that integrates Chebyshev based Kolmogorov-Arnold Networks (Chebyshev KAN). Specifically, we design a novel residual architecture that leverages the nonlinear approximation capabilities of Chebyshev polynomials within the KAN framework to more effectively model the complex motion characteristics inherent in IMU signals. To the best of our knowledge, this work represents the first application of an interpretable KAN model to IO. Experimental results on five publicly available datasets demonstrate the effectiveness of CKANIO.

cs.RO

The Rate-Distortion-Perception-Classification Tradeoff: Joint Source Coding and Modulation via Inverse-Domain GANs

The joint source-channel coding (JSCC) framework leverages deep learning to learn from data the best codes for source and channel coding. When the output signal, rather than being binary, is directly mapped onto the IQ domain (complex-valued), we call the resulting framework joint source coding and modulation (JSCM). We consider a JSCM scenario and show the existence of a strict tradeoff between channel rate, distortion, perception, and classification accuracy, a tradeoff that we name RDPC. We then propose two image compression methods to navigate that tradeoff: the RDPCO algorithm which, under simple assumptions, directly solves the optimization problem characterizing the tradeoff, and an algorithm based on an inverse-domain generative adversarial network (ID-GAN), which is more general and achieves extreme compression. Simulation results corroborate the theoretical findings, showing that both algorithms exhibit the RDPC tradeoff. They also demonstrate that the proposed ID-GAN algorithm effectively balances image distortion, perception, and classification accuracy, and significantly outperforms traditional separation-based methods and recent deep JSCM architectures in terms of one or more of these metrics.

cs.LG

Capacity and Delay Tradeoff of Secondary Cellular Networks with Spectrum Aggregation

Cellular communication networks are plagued with redundant capacity, which results in low utilization and cost-effectiveness of network capital investments. The redundant capacity can be exploited to deliver secondary traffic that is ultra-elastic and delay-tolerant. In this paper, we propose an analytical framework to study the capacity-delay tradeoff of elastic/secondary traffic in large scale cellular networks with spectrum aggregation. Our framework integrates stochastic geometry and queueing theory models and gives analytical insights into the capacity-delay performance in the interference limited regime. Closed-form results are obtained to characterize the mean delay and delay distribution as functions of per user throughput capacity. The impacts of spectrum aggregation, user and base station (BS) densities, traffic session payload, and primary traffic dynamics on the capacity-delay tradeoff relationship are investigated. The fundamental capacity limit is derived and its scaling behavior is revealed. Our analysis shows the feasibility of providing secondary communication services over cellular networks and highlights some critical design issues.

cs.NI

Capacity analysis of a multi-cell multi-antenna cooperative cellular network with co-channel interference

Characterization and modeling of co-channel interference is critical for the design and performance evaluation of realistic multi-cell cellular networks. In this paper, based on alpha stable processes, an analytical co-channel interference model is proposed for multi-cell multiple-input multi-output (MIMO) cellular networks. The impact of different channel parameters on the new interference model is analyzed numerically. Furthermore, the exact normalized downlink average capacity is derived for a multi-cell MIMO cellular network with co-channel interference. Moreover, the closed-form normalized downlink average capacity is derived for cell-edge users in the multi-cell multiple-input single-output (MISO) cooperative cellular network with co-channel interference. From the new co-channel interference model and capacity, the impact of cooperative antennas and base stations on cell-edge user performance in the multi-cell multi-antenna cellular network is investigated by numerical methods. Numerical results show that cooperative transmission can improve the capacity performance of multi-cell multi-antenna cooperative cellular networks, especially in a scenario with a high density of interfering base stations. The capacity performance gain is degraded with the increased number of cooperative antennas or base stations.

cs.NI

Interference Mitigation for Cognitive Radio MIMO Systems Based on Practical Precoding

In this paper, we propose two subspace-projection-based precoding schemes, namely, full-projection (FP)- and partial-projection (PP)-based precoding, for a cognitive radio multiple-input multiple-output (CR-MIMO) network to mitigate its interference to a primary time-division-duplexing (TDD) system. The proposed precoding schemes are capable of estimating interference channels between CR and primary networks, and incorporating the interference from the primary to the CR system into CR precoding via a novel sensing approach. Then, the CR performance and resulting interference of the proposed precoding schemes are analyzed and evaluated. By fully projecting the CR transmission onto a null space of the interference channels, the FP-based precoding scheme can effectively avoid interfering the primary system with boosted CR throughput. While, the PP-based scheme is able to further improve the CR throughput by partially projecting its transmission onto the null space.

cs.IT

Aggregate Interference Modeling in Cognitive Radio Networks with Power and Contention Control

In this paper, we present an interference model for cognitive radio (CR) networks employing power control, contention control or hybrid power/contention control schemes. For the first case, a power control scheme is proposed to govern the transmission power of a CR node. For the second one, a contention control scheme at the media access control (MAC) layer, based on carrier sense multiple access with collision avoidance (CSMA/CA), is proposed to coordinate the operation of CR nodes with transmission requests. The probability density functions of the interference received at a primary receiver from a CR network are first derived numerically for these two cases. For the hybrid case, where power and contention controls are jointly adopted by a CR node to govern its transmission, the interference is analyzed and compared with that of the first two schemes by simulations. Then, the interference distributions under the first two control schemes are fitted by log-normal distributions with greatly reduced complexity. Moreover, the effect of a hidden primary receiver on the interference experienced at the receiver is investigated. It is demonstrated that both power and contention controls are effective approaches to alleviate the interference caused by CR networks. Some in-depth analysis of the impact of key parameters on the interference of CR networks is given via numerical studies as well.

cs.IT