Searcharxiv⌕ Search

arXiv subjects

Bumsu Park

Publications and source records attributed to Bumsu Park.

7 recordsLinked to original sources

Thinking in Tokens, Talking in Bits: A Practical Interface for Token Communication

Advanced artificial intelligence models think in tokens; contemporary communication systems carry bits. The direct way to bridge this gap is to transmit tokens, but that makes a model-specific representation part of the air interface, coupling the endpoints through a shared tokenizer, codebook, and often a neural transceiver. We take a different route: keep bits in the payload and let tokens control how those bits are generated and protected. The resulting token-bit interface transition aligns task-side tokens with source- and channel-coding units, translates token relevance into codec controls, and preserves the induced priority order across the coding chain. We instantiate it for image classification, where a vision transformer scores the task relevance of each image region from its attention maps: those scores steer block-wise JPEG rate allocation, then group the compressed bits for protection at different polar-code rates. The payload remains an explicit, reconstructable bitstream recovered by a correspondingly configured decoder. Over-the-air experiments on a software-defined radio testbed show improved accuracy--latency tradeoffs over separate source-channel coding, performance competitive with far more memory-intensive neural joint source-channel coding, and graceful degradation under channel mismatch. Token communication, then, need not transmit tokens explicitly; what it needs is an interface through which tokens determine how bits are communicated.

eess.SP↗

PrismQuant: Rate-Distortion-Optimal Vector Quantization for Gaussian-Mixture Sources

For a Gaussian source under mean-squared error (MSE), classical transform coding is rate--distortion (RD) optimal: the Karhunen--Loeve transform (KLT) diagonalizes the covariance, reverse waterfilling allocates the bits, and scalar quantization closes the loop. This elegant story breaks down for multimodal sources, where no single covariance can capture heterogeneous local geometries, and the RD function loses its closed form. We revisit this problem through Gaussian-mixture sources and develop a constructive RD theory for them. Our key finding is that the mixture structure incurs only a component label cost. Conditioned on the active mixture component, each branch is Gaussian; the challenge is allocating bits across heterogeneous branches. We prove that the genie-aided conditional RD function is governed by a single global reverse-waterfilling level shared across all components and eigenmodes. Building on this result, we introduce PrismQuant, which transmits the component label losslessly and encodes the residual using the component-matched KLT, followed by scalar quantization, achieving a rate of H(C)/n bits per source dimension of the converse, with a vanishing asymptotic gap. We further develop a practical implementation based on EM-driven Gaussian-mixture learning, component-adaptive KLTs, and entropy-constrained scalar quantization (ECSQ). Experiments on synthetic Gaussian mixtures show that PrismQuant closely approaches the theoretical RD bound, while experiments on real-world channel-state-information (CSI) data demonstrate competitive or superior performance compared with transformer-based learned codecs at more than one order of magnitude smaller model size.

cs.IT↗

CSI Feedback Under Basis Mismatch: Rate-Splitting Transform Coding for FDD Massive MIMO

In frequency division duplex massive multiple-input multiple-output systems, downlink channel state information must be fed back within a limited uplink budget. While transform coding with Karhunen-Loeve transform and reverse water-filling is rate-distortion optimal for Gaussian channels, its performance is limited by basis mismatch between the user and base station. We analyze this mismatch and propose a practical architecture separating long-term basis feedback from short-term coefficient quantization. Using a random vector quantization, we derive a closed-form end-to-end mean square error expression. This allows us to characterize the optimal rate split and identify a phase transition threshold for basis updates. Simulations on correlated Gaussian and COST2100 channels demonstrate near-optimal performance, robustness to update overhead, and significant complexity reduction compared to deep-learning-based autoencoders.

cs.IT↗

CSI Compression for Massive MIMO-OFDM: Mismatch-Aware Rate-Distortion Trade-offs

We study channel state information (CSI) compression for wideband frequency division duplex massive multiple-input multiple-output (MIMO) when the base station (BS) reconstructs CSI using an imperfect covariance model. Under matched second-order statistics, remote rate--distortion theory yields transform coding with reverse water-filling (RWF) over covariance eigenmodes. With decoder-side covariance mismatch, however, this allocation is no longer end-to-end optimal. We derive an achievable mismatched Gaussian rate--distortion characterization based on a Gaussian test channel and a mismatched minimum mean square error (MMSE) reconstruction rule. In a shared-eigenvector regime (common eigenbasis, mismatched eigenvalues), the problem decouples across modes and leads to a robust reverse water-filling (RRWF) allocation computable via bisection and per-mode root finding. Simulations using wideband massive MIMO covariance models show that RRWF consistently improves reconstruction distortion and end-to-end mean square error relative to conventional RWF under mismatch.

cs.IT↗

Fundamental Limits of CSI Compression in FDD Massive MIMO

Channel state information (CSI) feedback in frequency-division duplex (FDD) massive multiple-input multiple-output (MIMO) systems is fundamentally limited by the high dimensionality of wideband channels. In this paper, we model the stacked wideband CSI vector as a Gaussian-mixture source with a latent geometry state that represents different propagation environments. Each component corresponds to a locally stationary regime characterized by a correlated proper complex Gaussian distribution with its own covariance matrix. This representation captures the multimodal nature of practical CSI datasets while preserving the analytical tractability of Gaussian models. Motivated by this structure, we propose Gaussian-mixture transform coding (GMTC), a practical CSI feedback architecture that combines state inference with state-adaptive TC. The mixture parameters are learned offline from channel samples and stored as a shared statistical dictionary at both the user equipment (UE) and the base station. For each CSI realization, the UE identifies the most likely geometry state, encodes the corresponding label using a lossless source code, and compresses the CSI using the Karhunen-Loeve transform matched to that state. We further characterize the fundamental limits of CSI compression under this model by deriving analytical converse and achievability bounds on the rate-distortion (RD) function. A key structural result is that the optimal bit allocation across all mixture components is governed by a single global reverse-waterfilling level. Simulations on the COST2100 dataset show that GMTC significantly improves the RD tradeoff relative to neural transform coding approaches while requiring substantially smaller model memory and lower inference complexity. These results indicate that near-optimal CSI compression can be achieved through state-adaptive TC without relying on large neural encoders.

cs.IT↗

Transformer-Based Nonlinear Transform Coding for Multi-Rate CSI Compression in MIMO-OFDM Systems

We propose a novel approach for channel state information (CSI) compression in multiple-input multiple-output orthogonal frequency division multiplexing (MIMO-OFDM) systems, where the frequency-domain channel matrix is treated as a high-dimensional complex-valued image. Our method leverages transformer-based nonlinear transform coding (NTC), an advanced deep-learning-driven image compression technique that generates a highly compact binary representation of the CSI. Unlike conventional autoencoder-based CSI compression, NTC optimizes a nonlinear mapping to produce a latent vector while simultaneously estimating its probability distribution for efficient entropy coding. By exploiting the statistical independence of latent vector entries, we integrate a transformer-based deep neural network with a scalar nested-lattice uniform quantization scheme, enabling low-complexity, multi-rate CSI feedback that dynamically adapts to varying feedback channel conditions. The proposed multi-rate CSI compression scheme achieves state-of-the-art rate-distortion performance, outperforming existing techniques with the same number of neural network parameters. Simulation results further demonstrate that our approach provides a superior rate-distortion trade-off, requiring only 6% of the neural network parameters compared to existing methods, making it highly efficient for practical deployment.

eess.SP↗

Multi-Rate Variable-Length CSI Compression for FDD Massive MIMO

For frequency-division-duplexing (FDD) systems, channel state information (CSI) should be fed back from the user terminal to the base station. This feedback overhead becomes problematic as the number of antennas grows. To alleviate this issue, we propose a flexible CSI compression method using variational autoencoder (VAE) with an entropy bottleneck structure, which can support multi-rate and variable-length operation. Numerical study confirms that the proposed method outperforms the existing CSI compression techniques in terms of normalized mean squared error.

cs.IT↗