SearcharxivSearch

arXiv subjects

Holger Boche

Publications and source records attributed to Holger Boche.

At least 37 records · Page 2Linked to original sources

TSR: Trajectory-Search Rollouts for Multi-Turn RL of LLM Agents

Advances in large language models (LLMs) are driving a shift toward using reinforcement learning (RL) to train agents from iterative, multi-turn interactions across tasks. However, multi-turn RL remains challenging as rewards are often sparse or delayed, and environments can be stochastic. In this regime, naive trajectory sampling can hinder exploitation and induce mode collapse. We propose TSR (Trajectory-Search Rollouts), a training-time approach that repurposes test-time scaling ideas for improved per-turn rollout generation. TSR performs lightweight tree-style search to construct high-quality trajectories by selecting high-scoring actions at each turn using state-based feedback. This improves rollout quality and stabilizes learning while remaining compatible with standard policy gradient optimizers, making TSR optimizer-agnostic. We instantiate TSR with best-of-N, beam, and shallow lookahead search, and pair it with PPO and GRPO, achieving up to 15% performance gains and more stable learning on Sokoban, FrozenLake, and WebShop tasks at a modest, one-time increase in training compute. By moving search from inference time to the rollout stage of training, TSR provides a modular and general mechanism for stronger multi-turn agent learning, complementary to existing frameworks and rejection-sampling-style selection methods.

cs.AI

Rate-Reliability Tradeoff for Deterministic Identification over Gaussian Channels

We extend the recent analysis of the rate-reliability tradeoff in deterministic identification (DI) to general linear Gaussian channels, marking the first such analysis for channels with continuous output. Because DI provides a framework that can substantially enhance communication efficiency, and since the linear Gaussian model underlies a broad range of physical communication systems, our results offer both theoretical insights and practical relevance for the performance evaluation of DI in future networks. Moreover, the structural parallels observed between the Gaussian and discrete-output cases suggest that similar rate-reliability behaviour may extend to wider classes of continuous channels.

cs.IT

Practical quantum tokens: challenges and perspectives

The concept of quantum tokens dates back alongside quantum cryptography to Stephen Wiesner's seminal work in 1983[1]. Already this initial work proposes society-relevant applications such as secure quantum banknotes, which can be exchanged between a bank and a customer. This quantum currency is based on various physical states that can be easily verified but is protected from being copied by the fundamental quantum laws. Four decades later, these ideas have flourished in the field of quantum information, and the concept of quantum banknotes has not only adopted many varying names, such as quantum money, quantum coins, quantum-digital payments, and quantum tokens, but also reached its first experimental demonstrations. In this perspective article, we discuss the current state-of-the-art of quantum tokens in the field of quantum information, as well as their future perspectives. We present a number of physical realizations of quantum tokens with integrated quantum memories and their applicability scenarios in detail. Finally, we discuss how quantum tokens fit into the information security ecosystem and consider their relationship to post-quantum cryptography.

quant-ph

On (Im)possibility of Network Oblivious Transfer via Noisy Channels and Non-Signaling Correlations

This work investigates the fundamental limits of implementing network oblivious transfer via noisy multiple access channels and broadcast channels between honest-but-curious parties when the parties have access to general tripartite non-signaling correlations. By modeling the shared resource as an arbitrary tripartite non-signaling box, we obtain a unified perspective on both the channel behavior and the resulting correlations. Our main result demonstrates that perfect oblivious transfer is impossible. In the asymptotic regime, we further show that even negligible leakage cannot be achieved, as repeated use of the resource amplifies the receiver(s)'s ability to distinguish messages that were not intended for him/them. In contrast, the receiver(s)'s own privacy is not subject to a universal impossibility limitation.

cs.IT

Implementation of Oblivious Transfer over Binary-Input AWGN Channels by Polar Codes

We develop a one-out-of-two oblivious transfer protocol over the binary-input additive white Gaussian noise (BI-AWGN) channel using polar codes. The scheme uses two decoder views linked by automorphisms of the polar transform and publicly draws the encoder at random from the corresponding automorphism group. This yields perfect secrecy for Bob at any blocklength. Secrecy for Alice is obtained asymptotically via channel polarization combined with privacy amplification. Because the construction deliberately injects randomness into selected bad bit-channels, we derive a relaxed reliability criterion, which is empirically certified via Monte-Carlo simulations. We also evaluate finite-blocklength performance. Finally, we characterize the polar-transform automorphisms as bit-level permutations of bit-channel indices, and exploit this structure to derive and optimize an achievable finite-blocklength rate.

cs.IT

Lossy Source Coding with Broadcast Side Information

This paper considers the source coding problem with broadcast side information. The side information is sent to two receivers through a noisy broadcast channel. We provide an outer bound of the rate--distortion--bandwidth (RDB) quadruples and achievable RDB quadruples when the helper uses a separation-based scheme. Some special cases with full characterization are also provided. We then compare the separation-based scheme with the uncoded scheme in the quadratic Gaussian case.

cs.IT

Arithmetic Complexity of Solutions of the Dirichlet Problem

The classical Dirichlet problem on the unit disk can be solved by different numerical approaches. The two most common and popular approaches are the integration of the associated Poisson integral and, by applying Dirichlet's principle, solving a particular minimization problem. For practical use, these procedures need to be implemented on concrete computing platforms. This paper studies the realization of these procedures on Turing machines, the fundamental model for any digital computer. We show that on this computing platform both approaches to solve Dirichlet's problem yield generally a solution that is not Turing computable, even if the boundary function is computable. Then the paper provides a precise characterization of this non-computability in terms of the Zheng--Weihrauch hierarchy. For both approaches, we derive a lower and an upper bound on the degree of non-computability in the Zheng--Weihrauch hierarchy.

cs.CC

Real-Time Multi-Target Detection and Tracking with mmWave 5G NR Waveforms on RFSoC

We demonstrate a real-time implementation of multi-target detection and tracking using 5G New Radio (NR) physical downlink shared channel (PDSCH) waveform with 400 MHz bandwidth at 28 GHz carrier frequency. The hardware platform is built on a radio frequency system-on-chip (RFSoC) 4x2 board connected with a pair of Sivers EVK02001 mmWave beamformers for transmission and reception. The entire sensing transceiver processing and fast beam control are realized purely in the programmable logic (PL) part of the RFSoC, enabling low-latency and fully hardware-accelerated operation. The continuously acquired sensing data constitute 3D range-angle (RA) tensors, which are processed on a host PC using adaptive background subtraction, cell-averaging constant false alarm rate (CA-CFAR) detection with density-based spatial clustering of applications with noise (DBSCAN) clustering, and extended Kalman filtering (EKF), to detect and track targets in the environment. Our software-defined radio (SDR) testbed integrates heterogeneous computing resources, including CPUs, GPUs, and FPGAs, thereby providing design flexibility for a wide range of tasks.

eess.SP

Secure Event-triggered MolecularvCommunication - Information Theoretic Perspective and Optimal Performance

Molecular Communication (MC) is an emerging field of research focused on understanding how cells in the human body communicate and exploring potential medical applications. In theoretical analysis, the goal is to investigate cellular communication mechanisms and develop nanomachine-assisted therapies to combat diseases. Since cells transmit information by releasing molecules at varying intensities, this process is commonly modeled using Poisson channels. In our study, we consider a discrete-time Poisson channel (DTPC). MC is often event-driven, making traditional Shannon communication an unsuitable performance metric. Instead, we adopt the identification framework introduced by Ahlswede and Dueck. In this approach, the receiver is only concerned with detecting whether a specific message of interest has been transmitted. Unlike Shannon transmission codes, the size of identification (ID) codes for a discrete memoryless channel (DMC) increases doubly exponentially with blocklength when using randomized encoding. This remarkable property makes the ID paradigm significantly more efficient than classical Shannon transmission in terms of energy consumption and hardware requirements. Another critical aspect of MC, influenced by the concept of the Internet of Bio-NanoThings, is security. In-body communication must be protected against potential eavesdroppers. To address this, we first analyze the DTPC for randomized identification (RI) and then extend our study to secure randomized identification (SRI). We derive capacity formulas for both RI and SRI, providing a comprehensive understanding of their performance and security implications.

cs.IT

When Data is the Algorithm: A Systematic Study and Curation of Preference Optimization Datasets

Aligning large language models (LLMs) is a central objective of post-training, often achieved through reward modeling and reinforcement learning methods. Among these, direct preference optimization (DPO) has emerged as a widely adopted technique that fine-tunes LLMs on preferred completions over less favorable ones. While most frontier LLMs do not disclose their curated preference pairs, the broader LLM community has released several open-source DPO datasets, including TuluDPO, ORPO, UltraFeedback, HelpSteer, and Code-Preference-Pairs. However, systematic comparisons remain scarce, largely due to the high computational cost and the lack of rich quality annotations, making it difficult to understand how preferences were selected, which task types they span, and how well they reflect human judgment on a per-sample level. In this work, we present the first comprehensive, data-centric analysis of popular open-source DPO corpora. We leverage the Magpie framework to annotate each sample for task category, input quality, and preference reward, a reward-model-based signal that validates the preference order without relying on human annotations. This enables a scalable, fine-grained inspection of preference quality across datasets, revealing structural and qualitative discrepancies in reward margins. Building on these insights, we systematically curate a new DPO mixture, UltraMix, that draws selectively from all five corpora while removing noisy or redundant samples. UltraMix is 30% smaller than the best-performing individual dataset yet exceeds its performance across key benchmarks. We publicly release all annotations, metadata, and our curated mixture to facilitate future research in data-centric preference optimization.

cs.CL

Computability of the Optimizer for Rate Distortion Functions

Rate distortion theory treats the problem of encoding a source with minimum codebook size while at the same time allowing for a certain amount of errors in the reconstruction measured by a fidelity criterion and distortion level. Similar to the channel coding problem the optimal rate of the codebook with respect to the blocklength is given by a convex optimization problem involving information theoretic quantities like mutual information. The value of the rate in dependence of the distortion level as well as the optimizer used in the codebook construction are of theoretical and practical importance in communication and information theory. In this paper the behavior of the rate distortion function regarding the computability of the optimizing test channel is investigated. We find that comparable with known results about the optimizer for other information theoretic problems a similar result is found to be true also regarding the computability of the optimizer for rate distortion functions. It turns out that while the rate distortion function is usually computable the optimizer for this problem is in general non-computable even for simple distortion measures.

cs.IT

Network Oblivious Transfer via Noisy Broadcast Channels

This paper investigates information-theoretic oblivious transfer via a discrete memoryless broadcast channel with one sender and two receivers. We analyze both non-colluding and colluding honest-but-curious user models and establish general upper bounds on the achievable oblivious transfer capacity region for each case. Two explicit oblivious transfer protocols are proposed. The first ensures correctness and privacy for independent, non-colluding receivers by leveraging the structure of binary erasure broadcast channels. The second protocol, secure even under receiver collusion, introduces additional entropy-sharing and privacy amplification mechanisms to preserve secrecy despite information leakage between users. Our results show that for the non-colluding case, the upper and lower bounds on oblivious transfer capacity coincide, providing a complete characterization of the achievable region. The work provides a unified theoretical framework bridging network information theory and cryptographic security, highlighting the potential of noisy broadcast channels as powerful primitives for multi-user privacy-preserving communication.

cs.IT

A Variational Framework for the Complexity of PDE Solutions

Partial Differential Equations (PDEs) are fundamental mathematical models for describing physical phenomena, yet most PDEs of practical interest require numerical approximations. The feasibility of such methods is constrained by existing computational models. Since digital computers are the primary realizations of numerical computations, and Turing machines define their theoretical limits, computability of PDE solutions is of fundamental significance. It provides a rigorous framework to distinguish equations that are effectively solvable from those that encode undecidable or non-computable behavior. Once computability is established, complexity theory quantifies the resources required to approximate PDE solutions. In this work, we present a novel framework based on least-squares variational formulations and associated gradient flows to analyze the computability and complexity of PDE solutions from an optimization perspective. Our approach approximates PDE solution operators via discrete gradient flows, linking PDE properties, such as coercivity, ellipticity, and convexity, to solution complexity. Within this setting, we characterize representation- and discretization-dependent sufficient conditions for regimes where PDEs admit polynomial-time approximations, as well as regimes exhibiting complexity blowup, where polynomial-time input data produce solutions with super-polynomial complexity. In summary, this paper develops a variational framework for analyzing computability and computational complexity of PDE solution classes. The results show how PDE structure and solution regularity influence their complexity, by establishing sufficient conditions for computability and complexity bounds. Beyond the theoretical characterization, the framework provides guidelines for effective numerical methods and contributes to understanding the limitations of digital computation for PDE problems.

math.NA

Variational Secret Common Randomness Extraction

This paper studies the problem of extracting common randomness (CR) or secret keys from correlated random sources observed by two legitimate parties, Alice and Bob, through public discussion in the presence of an eavesdropper, Eve. We propose a practical two-stage CR extraction framework. In the first stage, the variational probabilistic quantization (VPQ) step is introduced, where Alice and Bob employ probabilistic neural network (NN) encoders to map their observations into discrete, nearly uniform random variables (RVs) with high agreement probability while minimizing information leakage to Eve. This is realized through a variational learning objective combined with adversarial training. In the second stage, a secure sketch using code-offset construction reconciles the encoder outputs into identical secret keys, whose secrecy is guaranteed by the VPQ objective. As a representative application, we study physical layer key (PLK) generation. Beyond the traditional methods, which rely on the channel reciprocity principle and require two-way channel probing, thus suffering from large protocol overhead and being unsuitable in high mobility scenarios, we propose a sensing-based PLK generation method for integrated sensing and communications (ISAC) systems, where paired range-angle (RA) maps measured at Alice and Bob serve as correlated sources. The idea is verified through both end-to-end simulations and real-world software-defined radio (SDR) measurements, including scenarios where Eve has partial knowledge about Bob's position. The results demonstrate the feasibility and convincing performance of both the proposed CR extraction framework and sensing-based PLK generation method.

cs.IT

Quantum Physical Unclonable Function based on Chaotic Hamiltonians

Quantum Physical Unclonable Functions (QPUFs) are hardware-based cryptographic primitives with strong theoretical security. This security stems from their modeling as Haar-random unitaries. However, implementing such unitaries on Intermediate-Scale Quantum devices is challenging due to exponential simulation complexity. Previous work tackled this using pseudo-random unitary designs but only under limited adversarial models with only black-box query access. In this paper, we propose a new QPUF construction based on chaotic quantum dynamics. We modeled the QPUF as a unitary time evolution under a chaotic Hamiltonian and proved that this approach offers security comparable to Haar-random unitaries. Intuitively, we show that while chaotic dynamics generate less randomness than ideal Haar unitaries, the randomness is still sufficient to make the QPUF unclonable in polynomial time. Moreover, we show that the evolution time required to achieve security scales linearly with number of qudits used in the scheme and can be kept public. We identified the Sachdev-Ye-Kitaev (SYK) model as a candidate for the QPUF Hamiltonian. Recent experiments using nuclear spins and cold atoms have shown progress toward achieving this goal. Inspired by recent experimental advances, we present a schematic architecture for realizing our proposed QPUF device based on optical Kagome Lattice with disorder. For adversaries with only query access, we also introduce an efficiently simulable pseudo-chaotic QPUF. Our results lay the preliminary groundwork for bridging the gap between theoretical security and the practical implementation of QPUFs for the first time.

quant-ph

Secure authentication via Quantum Physical Unclonable Functions: a review

Quantum Physical Unclonable Functions (QPUFs) offer a physically grounded approach to secure authentication, extending the capabilities of classical PUFs. This review covers their theoretical foundations and key implementation challenges - such as quantum memories and Haar-randomness -, and distinguishes QPUFs from Quantum Readout PUFs (QR-PUFs), more experimentally accessible yet less robust against quantum-capable adversaries. A co-citation-based selection method is employed to trace the evolution of QPUF architectures, from early QR-PUFs to more recent Hybrid PUFs (HPUFs). This method further supports a discussion on the role of information-theoretic analysis in mitigating inconsistencies in QPUF responses, underscoring the deep connection between secret-key generation and authentication. Despite notable advances, achieving practical and robust QPUF-based authentication remains an open challenge.

quant-ph

Fixing It in Post: A Comparative Study of LLM Post-Training Data Quality and Model Performance

Recent work on large language models (LLMs) has increasingly focused on post-training and alignment with datasets curated to enhance instruction following, world knowledge, and specialized skills. However, most post-training datasets used in leading open- and closed-source LLMs remain inaccessible to the public, with limited information about their construction process. This lack of transparency has motivated the recent development of open-source post-training corpora. While training on these open alternatives can yield performance comparable to that of leading models, systematic comparisons remain challenging due to the significant computational cost of conducting them rigorously at scale, and are therefore largely absent. As a result, it remains unclear how specific samples, task types, or curation strategies influence downstream performance when assessing data quality. In this work, we conduct the first comprehensive side-by-side analysis of two prominent open post-training datasets: Tulu-3-SFT-Mix and SmolTalk. Using the Magpie framework, we annotate each sample with detailed quality metrics, including turn structure (single-turn vs. multi-turn), task category, input quality, and response quality, and we derive statistics that reveal structural and qualitative similarities and differences between the two datasets. Based on these insights, we design a principled curation recipe that produces a new data mixture, TuluTalk, which contains 14% fewer samples than either source dataset while matching or exceeding their performance on key benchmarks. Our findings offer actionable insights for constructing more effective post-training datasets that improve model performance within practical resource limits. To support future research, we publicly release both the annotated source datasets and our curated TuluTalk mixture.

cs.CL

"Don't Do That!": Guiding Embodied Systems through Large Language Model-based Constraint Generation

Recent advancements in large language models (LLMs) have spurred interest in robotic navigation that incorporates complex spatial, mathematical, and conditional constraints from natural language into the planning problem. Such constraints can be informal yet highly complex, making it challenging to translate into a formal description that can be passed on to a planning algorithm. In this paper, we propose STPR, a constraint generation framework that uses LLMs to translate constraints (expressed as instructions on ``what not to do'') into executable Python functions. STPR leverages the LLM's strong coding capabilities to shift the problem description from language into structured and interpretable code, thus circumventing complex reasoning and avoiding potential hallucinations. We show that these LLM-generated functions accurately describe even complex mathematical constraints, and apply them to point cloud representations with traditional search algorithms. Experiments in a simulated Gazebo environment show that STPR ensures full compliance across several constraints and scenarios, while having short runtimes. We also verify that STPR can be used with smaller code LLMs, making it applicable to a wide range of compact models with low inference cost.

cs.AI