SearcharxivSearch

arXiv · 2609.15256

Thinking in Tokens, Talking in Bits: A Practical Interface for Token Communication

Abstract

Advanced artificial intelligence models think in tokens; contemporary communication systems carry bits. The direct way to bridge this gap is to transmit tokens, but that makes a model-specific representation part of the air interface, coupling the endpoints through a shared tokenizer, codebook, and often a neural transceiver. We take a different route: keep bits in the payload and let tokens control how those bits are generated and protected. The resulting token-bit interface transition aligns task-side tokens with source- and channel-coding units, translates token relevance into codec controls, and preserves the induced priority order across the coding chain. We instantiate it for image classification, where a vision transformer scores the task relevance of each image region from its attention maps: those scores steer block-wise JPEG rate allocation, then group the compressed bits for protection at different polar-code rates. The payload remains an explicit, reconstructable bitstream recovered by a correspondingly configured decoder. Over-the-air experiments on a software-defined radio testbed show improved accuracy--latency tradeoffs over separate source-channel coding, performance competitive with far more memory-intensive neural joint source-channel coding, and graceful degradation under channel mismatch. Token communication, then, need not transmit tokens explicitly; what it needs is an interface through which tokens determine how bits are communicated.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Chanho Park, Bumsu Park, Soonhee Kwon, Sangrim Lee, Namyoon Lee. 2026-09-14. Thinking in Tokens, Talking in Bits: A Practical Interface for Token Communication. https://arxiv.org/abs/2609.15256

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Towards optimal algorithms for the recovery of low-dimensional models with linear rates

We consider the problem of recovering elements of a low-dimensional model from linear measurements. From signal and image processing to inverse problems in data science, this question has been at the center of many applications. Lately, with the success of models and methods relying on deep neural networks, there has been a multiplication of different algorithms and recovery results. Comparing the performance of recovery algorithms becomes a complex task without a unifying framework. In this article, as a first step for the study of general algorithms for low-dimensional recovery, we study a class of generalized projected gradient descent algorithms that can recover a given low-dimensional model with linear rates. The obtained rates decouple the impact of the quality of the measurements with respect to the model from the geometry of the properties of the chosen generalized projection: we can directly measure performance through a restricted Lipschitz constant of the projection with respect to the low dimensional model. By optimizing this constant, we define an optimal generalized projected gradient descent. Our general approach provides an optimality result in the case of sparse recovery. Moreover, our framework allows for a common interpretation of linear rates of recovery in the context of both sparse models and models induced by some ``plug-and-play'' imaging methods that rely on deep neural networks. These rates of recovery are observed in experiments on synthetic and real data.

eess.SP

Tracking Driving Stressors through Multimodal Physiological Monitoring

Understanding and mitigating driving stress is important for improving road safety and driver well-being. Reliable estimation, however, requires distinguishing biobehavioral responses to individual stressors from gradual physiological and contextual changes. We collected physiological data and vehicle telemetry from 31 participants across 44 simulated-driving sessions containing controlled stressor events. Under cross-validation, a multimodal classifier achieved an AUROC of 0.768 when distinguishing the stressor phase from an earlier baseline, reflecting both stressor effects and temporal drift. Controlling for drift retained an AUROC of 0.661, but revealed stronger responses to sustained than brief stressors, and shifted feature attribution toward phasic cardiac and electrodermal markers. We further quantified the interaction between model-estimated physiological stress and observable changes in vehicle control through simulation telemetry. Our findings show that stressor-aware modeling can identify physiologically grounded responses that correspond to meaningful changes in driving behavior.

eess.SP

Resolution-Aliasing Trade-off in Near-Field Localisation

Extremely Large-scale MIMO (XL-MIMO) systems operating in Near-Field (NF) introduce new degrees of freedom for accurate source localisation, but make dense arrays impractical. Sparse or distributed arrays can reduce hardware complexity while maintaining high resolution, yet sub-Nyquist spatial sampling introduces aliasing artefacts in the localisation ambiguity function. This paper presents a unified framework to jointly characterise resolution and aliasing in NF localisation and study the trade-off between the two. Leveraging the concept of local chirp spatial frequency, we derive analytical expressions linking array geometry and sampling density to the spatial bandwidth of the received field. We introduce two geometric tools--Critical Antenna Elements (CAEs) and the Non-Contributive Zone (NCZ)--to intuitively identify how individual antennas contribute to resolution and/or aliasing. Our analysis reveals that resolution and aliasing are not always strictly coupled, e.g., increasing the array aperture can improve resolution without necessarily aggravating aliasing. These results provide practical guidelines for designing NF arrays that optimally balance resolution and aliasing, supporting efficient XL-MIMO deployment.

eess.SP