Searcharxiv⌕ Search

arXiv subjects

Wolfgang Utschick

Publications and source records attributed to Wolfgang Utschick.

At least 19 recordsLinked to original sources

OTFS Channel Estimation Utilizing Sparse Bayesian Generative Modelling

One of the key challenges of future wireless communication systems is ensuring reliability in high-speed mobile scenarios, where accurate recovery of channel state information (CSI) is essential. Many recent studies have concluded that orthogonal time-frequency space (OTFS) modulation is a promising technology for addressing this challenge. Additionally, machine learning (ML)-based methods have the potential to improve channel estimation performance by leveraging ambient information more effectively than classical estimation techniques. This paper particularly addresses channel estimation for OTFS by employing a compressive sensing (CS)-based sparse Bayesian generative model (SBGM), namely the recently introduced compressive sensing Gaussian mixture model (CSGMM). We show that our proposed approach yields significant improvement in normalized mean squared error (NMSE) over the next-best-performing baseline. We additionally provide insights into the theoretical potential of the model to optimally approximate complex channel distributions with arbitrary precision within the Doppler-delay (DD) domain. To summarize, this work establishes the OTFS-CSGMM framework as a promising solution for high mobility wireless channel estimation.

eess.SP↗

Lightweight Beam Index Map Using Coupled Gaussian Mixture Models

This paper addresses the beam alignment problem in MIMO systems from a decentralized, mobile terminal (MT)-centric perspective. We propose a lightweight machine learning approach that leverages position information to perform beam selection without relying on exhaustive search or strong base station coordination. Specifically, we model the joint distribution of MT positions and channel observations using a coupled Gaussian mixture model (GMM), enabling the construction of a beam index map (BIM) that directly associates spatial locations with codebook entries. To account for practical hardware constraints, we introduce a refinement procedure that adapts the learned statistical model to fixed codebooks. The resulting method is computationally efficient and suitable for deployment on resource-constrained devices. Simulation results on the DeepMIMO and QuaDRiGa datasets demonstrate that the proposed approach outperforms clustering-based fingerprinting methods and achieves competitive performance compared to exhaustive search, while significantly reducing complexity and overhead.

eess.SP↗

Joint Access Point Selection and Precoder Design under Statistical CSI

This work addresses joint access point (AP) selection and precoding for sum-rate maximization under statistical channel state information (CSI) in multi-AP multi-user systems. To this end, we propose two approaches. The first method is an iterative alternating optimization algorithm that updates the precoding vectors via the stochastic WMMSE (SWMMSE) algorithm and the assignment variables via a projected gradient descent step. The second method is a graph neural network (GNN)-based framework that solves the same problem in a single forward pass during inference. Building on an attention-based Edge-GNN architecture, we extend it to a multi-AP scenario, enabling the joint learning of assignment variables and precoding vectors from statistical CSI alone. Results show that the GNN outperforms the iterative algorithm across the tested signal-to-noise ratio (SNR) range and generalizes to varying numbers of users with comparable performance. Both approaches are also compared to various baseline techniques.

eess.SP↗

Efficient Channel Prediction based on Gram-Square-Root Factorization using GMMs

Accurate channel state information (CSI) is critical for downlink (DL)-multi-user (MU)-multiple-input multiple-output (MIMO) systems, where feedback delays and mobility can degrade precoding performance. To ensure reliable beamforming and interference mitigation, CSI prediction is required. In practical systems, full CSI feedback is often infeasible due to signaling overhead, so transmitters rely on partial CSI reported by the receivers. In this work, we propose a Gaussian mixture model (GMM)-based prediction framework for MIMO-orthogonal frequency-division multiplexing (OFDM) channels under partial feedback using Gram-square-root factorization. To address the high dimensionality, we introduce an efficient parameter reduction technique that exploits structured covariance matrices, significantly lowering complexity without noticeable performance degradation. This reduction is based on the Gram-square-root factorization and remains of interest even when full CSI is available. Simulation results demonstrate that GMMs achieve the highest prediction accuracy and correctly capture the underlying channel subspaces, which is essential for effective MU-precoding. The proposed method outperforms classical baselines such as zero-order hold (ZOH), first-order hold (FOH), and linear minimum mean squared error (LMMSE) predictors, and an advanced neural network (NN)-based predictor. Notably, the parameter-reduced partial CSI GMM achieves performance comparable to that of full CSI prediction, highlighting its ability to efficiently model the channel structure under limited feedback.

eess.SP↗

Physics-Regularized Machine Learning for Proprioceptive Vehicle Localization Using Onboard Sensors

Accurate and robust localization is essential for autonomous mobility systems in real-world environments. While fusing Inertial Measurement Unit (IMU) data with satellite-based correction signals provides precise vehicle pose estimates, performance degrades substantially during outages. Recent studies indicate that Machine Learning (ML) can improve IMU-based proprioceptive localization, highlighting untapped potential for onboard sensors readily available in production vehicles. This paper introduces Physics-Regularized Machine Learning for Localization (PRML2), a hybrid framework that combines the complementary strengths of Kalman filtering and data-driven learning to estimate vehicle pose directly from onboard sensors. A key aspect of PRML2 is its physics-regularized learning, enabled by end-to-end training of an ML model through a differentiable Kalman filter. This improves consistency with vehicle motion models, thereby enhancing both localization accuracy and generalization across driving conditions. We evaluate the performance limits of ML-enhanced onboard odometry on a publicly available dataset and show that PRML2 achieves superior localization accuracy and demonstrates real-time capability. This work also introduces a novel dataset to support vehicle localization research under low-friction conditions. The proposed framework provides a robust and cost-effective solution for vehicle localization under degraded sensing conditions by integrating learning with physics-based priors.

cs.RO↗

Uncertainty-Aware Velocity Correction for Proprioceptive Vehicle Localization using Evidential Mamba

Reliable localization in GNSS-denied environments remains a fundamental challenge for intelligent vehicles, as inertial navigation systems accumulate unbounded drift without external correction. Existing approaches provide drift correction through dedicated infrastructure, expensive external sensors, or complex multi-sensor fusion, each introducing practical deployment barriers. We propose Evidential Velocity Correction using Mamba (EVC-Mamba), a learning-based architecture that transforms onboard vehicle sensor data into a virtual velocity sensor for IMU drift correction without additional hardware. A Mamba-based selective state space model captures the temporal dynamics of vehicle motion, while evidential deep learning with a Normal-Inverse-Gamma distribution provides principled uncertainty quantification. The resulting uncertainty-aware velocity estimate is incorporated as a virtual correction measurement into an Error-State Extended Kalman Filter to reduce position drift. Evaluation on real-world vehicle data demonstrates that inertial navigation using the proposed velocity correction achieves localization accuracy within 10% of a dedicated external velocity sensor across different outage durations. The proposed architecture supports real-time onboard deployment at 40 Hz on edge hardware, enabling reliable localization during prolonged GNSS outages.

cs.RO↗

Autoregressive-Gaussian Mixture Models: Efficient Generative Modeling of WSS Signals

This work addresses the challenge of making generative models suitable for resource-constrained environments like mobile wireless communication systems. We propose a generative model that integrates Autoregressive (AR) parameterization into a Gaussian Mixture Model (GMM) for modeling Wide-Sense Stationary (WSS) processes. By exploiting model-based insights allowing for structural constraints, the approach significantly reduces parameters while maintaining high modeling accuracy. Channel estimation experiments show that the model can outperform standard GMMs and variants using Toeplitz or circulant covariances, particularly with small sample sizes. For larger datasets, it matches the performance of conventional methods while improving computational efficiency and reducing the memory requirements.

eess.SP↗

Behavior-Centric Extraction of Scenarios from Highway Traffic Data and their Domain-Knowledge-Guided Clustering using CVQ-VAE

Approval of ADS depends on evaluating its behavior within representative real-world traffic scenarios. A common way to obtain such scenarios is to extract them from real-world data recordings. These can then be grouped and serve as basis on which the ADS is subsequently tested. This poses two central challenges: how scenarios are extracted and how they are grouped. Existing extraction methods rely on heterogeneous definitions, hindering scenario comparability. For the grouping of scenarios, rule-based or ML-based methods can be utilized. However, while modern ML-based approaches can handle the complexity of traffic scenarios, unlike rule-based approaches, they lack interpretability and may not align with domain-knowledge. This work contributes to a standardized scenario extraction based on the Scenario-as-Specification concept, as well as a domain-knowledge-guided scenario clustering process. Experiments on the highD dataset demonstrate that scenarios can be extracted reliably and that domain-knowledge can be effectively integrated into the clustering process. As a result, the proposed methodology supports a more standardized process for deriving scenario categories from highway data recordings and thus enables a more efficient validation process of automated vehicles.

cs.CV↗

Is Lattice Reduction Necessary for Vector Perturbation Precoding?

Vector perturbation (VP) precoding is an effective nonlinear precoding technique in the downlink (DL) with modulo channels, providing an approximation of dirty paper coding (DPC) which is capacity-achieving. Especially, when combined with Lattice reduction (LR), low-complexity algorithms achieve a very promising performance, outperforming other popular non-linear precoding techniques like Tomlinson-Harashima precoding (THP). However, these results are based on the symbol error rate (SER) or bit error rate (BER). When shifting the focus to the mutual information as the figure of merit, we show that this is different and that the underlying lattice problem has a unique structural property. For lattice problems with this special structure, we show for a whole class of algorithms that LR does not have any impact on the solution vector. At the same time, algorithms are identified which benefit from LR, even if this lattice structure arises. The provided structural analysis has strong implications on the performance evaluation of VP. In particular, we re-evaluate popular Lenstra-Lenstra-Lovász (LLL)-aided methods like the LLL-aided nearest plane (NP) algorithm and show that they do not outperform conventional THP, highlighting the effectiveness of the THP method. This is in contrast to the existing results based on SER and BER where these methods clearly outperform THP.

cs.IT↗

Context-Aware CSI Prediction for Access Point Selection Utilizing Conditional VAEs

Indoor wireless communication environments are strongly influenced by dynamic conditions, which affect channel state information (CSI) and, consequently, the precoding strategy and the selection of the access point (AP). Device-free sensing and localization functionalities can provide information about these conditions, including, for example, the user's position and the position of mobile blocking objects. To model the statistical relationship between the CSI and the provided conditions, we employ a conditional variational autoencoder (cVAE). We treat the user and object positions - referred to as context information - as conditional inputs to the cVAE. The proposed model does not rely on ground-truth CSI and is trained directly on noisy data. Once trained, the framework can infer channel statistics solely from user and blocking object positions, enabling proactive AP selection based on inferred statistical CSI without requiring continuous CSI estimation. Extensive simulations with the state-of-the-art ray-tracing tool Sionna validate the proposed method.

eess.SP↗

Uncertainty-Aware Diffusion Model for Multimodal Highway Trajectory Prediction via DDIM Sampling

Accurate and uncertainty-aware trajectory prediction remains a core challenge for autonomous driving, driven by complex multi-agent interactions, diverse scene contexts and the inherently stochastic nature of future motion. Diffusion-based generative models have recently shown strong potential for capturing multimodal futures, yet existing approaches such as cVMD suffer from slow sampling, limited exploitation of generative diversity and brittle scenario encodings. This work introduces cVMDx, an enhanced diffusion-based trajectory prediction framework that improves efficiency, robustness and multimodal predictive capability. Through DDIM sampling, cVMDx achieves up to a 100x reduction in inference time, enabling practical multi-sample generation for uncertainty estimation. A fitted Gaussian Mixture Model further provides tractable multimodal predictions from the generated trajectories. In addition, a CVQ-VAE variant is evaluated for scenario encoding. Experiments on the publicly available highD dataset show that cVMDx achieves higher accuracy and significantly improved efficiency over cVMD, enabling fully stochastic, multimodal trajectory prediction.

cs.LG↗

UrbanIng-V2X: A Large-Scale Multi-Vehicle, Multi-Infrastructure Dataset Across Multiple Intersections for Cooperative Perception

Recent cooperative perception datasets have played a crucial role in advancing smart mobility applications by enabling information exchange between intelligent agents, helping to overcome challenges such as occlusions and improving overall scene understanding. While some existing real-world datasets incorporate both vehicle-to-vehicle and vehicle-to-infrastructure interactions, they are typically limited to a single intersection or a single vehicle. A comprehensive perception dataset featuring multiple connected vehicles and infrastructure sensors across several intersections remains unavailable, limiting the benchmarking of algorithms in diverse traffic environments. Consequently, overfitting can occur, and models may demonstrate misleadingly high performance due to similar intersection layouts and traffic participant behavior. To address this gap, we introduce UrbanIng-V2X, the first large-scale, multi-modal dataset supporting cooperative perception involving vehicles and infrastructure sensors deployed across three urban intersections in Ingolstadt, Germany. UrbanIng-V2X consists of 34 temporally aligned and spatially calibrated sensor sequences, each lasting 20 seconds. All sequences contain recordings from one of three intersections, involving two vehicles and up to three infrastructure-mounted sensor poles operating in coordinated scenarios. In total, UrbanIng-V2X provides data from 12 vehicle-mounted RGB cameras, 2 vehicle LiDARs, 17 infrastructure thermal cameras, and 12 infrastructure LiDARs. All sequences are annotated at a frequency of 10 Hz with 3D bounding boxes spanning 13 object classes, resulting in approximately 712k annotated instances across the dataset. We provide comprehensive evaluations using state-of-the-art cooperative perception methods and publicly release the codebase, dataset, HD map, and a digital twin of the complete data collection environment.

cs.CV↗

On the Optimality of Rate Balancing for Max-Min Fair Multicasting

The max-min fair (MMF) multicasting problem is known to be NP-hard. In this work, we analytically derive the optimal solution to this NP-hard problem and establish the equivalence between rate balancing and the optimal MMF multicasting solution under certain conditions. Based on this theoretical insight, we propose a low-complexity algorithm for MMF multicasting that yields closed-form solutions. Simulation results validate our analysis and demonstrate that the proposed algorithm outperforms the state-of-the-art methods while being computationally more efficient.

eess.SP↗

Reconstructing Patched or Partial Holograms to allow for Whole Slide Imaging with a Self-Referencing Holographic Microscope

The last decade has seen significant advances in computer-aided diagnostics for cytological screening, mainly through the improvement and integration of scanning techniques such as whole slide imaging (WSI) and the combination with deep learning. Simultaneously, new imaging techniques such as quantitative phase imaging (QPI) are being developed to capture richer cell information with less sample preparation. So far, the two worlds of WSI and QPI have not been combined. In this work, we present a reconstruction algorithm which makes whole slide imaging of cervical smears possible by using a self-referencing three-wave digital holographic microscope. Since a WSI is constructed by combining multiple patches, the algorithm is adaptive and can be used on partial holograms and patched holograms. We present the algorithm for a single shot hologram, the adaptations to make it flexible to various inputs and show that the algorithm performs well for the tested epithelial cells. This is a preprint of our paper, which has been accepted for publication in 2026 IEEE International Symposium on Biomedical Imaging (ISBI).

eess.SP↗

On the Asymptotic MSE-Optimality of Parametric Bayesian Channel Estimation in mmWave Systems

The mean square error (MSE)-optimal estimator is known to be the conditional mean estimator (CME). This paper introduces a parametric channel estimation technique based on Bayesian estimation. This technique uses the estimated channel parameters to parameterize the well-known LMMSE channel estimator. We first derive an asymptotic CME formulation that holds for a wide range of priors on the channel parameters. Based on this, we show that parametric Bayesian channel estimation is MSE-optimal for high signal-to-noise ratio (SNR) and/or long coherence intervals, i.e., many noisy observations provided within one coherence interval. Numerical simulations validate the derived formulations.

eess.SP↗

Semi-Blind Strategies for MMSE Channel Estimation Utilizing Generative Priors

This paper investigates semi-blind channel estimation for massive multiple-input multiple-output (MIMO) systems. To this end, we first estimate a subspace based on all received symbols (pilot and payload) to provide additional information for subsequent channel estimation. This additional information enhances minimum mean square error (MMSE) channel estimation. Two variants of the linear MMSE (LMMSE) estimator are formulated, where the first one solves the estimation within the subspace, and the second one uses a subspace projection as a preprocessing step. Theoretical derivations show the latter method's superior estimation performance in terms of mean square error for uncorrelated Rayleigh fading. Further, we provide asymptotic insights on how the proposed MMSE-based channel estimation strategy outperforms the unbiased Cramer-Rao bound. Subsequently, we introduce parameterizations of these semi-blind LMMSE estimators based on two different conditional Gaussian latent models, i.e., the Gaussian mixture model and the variational autoencoder. Both models learn the propagation environment's underlying channel distribution based on training data and serve as generative priors for our semi-blind channel estimation. Extensive simulations for real-world measurement data and spatial channel models show the proposed methods' superior performance compared to state-of-the-art semi-blind channel estimators in terms of MSE.

eess.SP↗

Wireless Channel Modeling for Machine Learning -- A Critical View on Standardized Channel Models

Standardized (link-level) channel models such as the 3GPP TDL and CDL models are frequently used to evaluate machine learning (ML)-based physical-layer methods. However, in this work, we argue that a link-level perspective incorporates limiting assumptions, causing unwanted distributional shifts or necessitating impractical online training. An additional drawback is that this perspective leads to (near-)Gaussian channel characteristics. Thus, ML-based models, trained on link-level channel data, do not outperform classical approaches for a variety of physical-layer applications. Particularly, we demonstrate the optimality of simple linear methods for channel compression, estimation, and modeling, revealing the unsuitability of link-level channel models for evaluating ML models. On the upside, adopting a scenario-level perspective offers a solution to this problem and unlocks the relative gains enabled by ML.

eess.SP↗

Precoder Design in Multi-User FDD Systems with VQ-VAE and GNN

Robust precoding is efficiently feasible in frequency division duplex (FDD) systems by incorporating the learnt statistics of the propagation environment through a generative model. We build on previous work that successfully designed site-specific precoders based on a combination of Gaussian mixture models (GMMs) and graph neural networks (GNNs). In this paper, by utilizing a vector quantized-variational autoencoder (VQ-VAE), we circumvent one of the key drawbacks of GMMs, i.e., the number of GMM components scales exponentially to the feedback bits. In addition, the deep learning architecture of the VQ-VAE allows us to jointly train the GNN together with VQ-VAE along with pilot optimization forming an end-to-end (E2E) model, resulting in considerable performance gains in sum rate for multi-user wireless systems. Simulations demonstrate the superiority of the proposed frameworks over the conventional methods involving the sub-discrete Fourier transform (DFT) pilot matrix and iterative precoder algorithms enabling the deployment of systems characterized by fewer pilots or feedback bits.

cs.IT↗