Searcharxiv⌕ Search

arXiv subjects

Tolga M. Duman

Publications and source records attributed to Tolga M. Duman.

At least 19 recordsLinked to original sources

Improved Lower Bounds on the Capacity of the Binary Deletion Channel via a Learning Approach to Run-Length Inputs

The best constructive lower bounds on the capacity of the binary deletion channel come from random codes with independent run lengths, yet the two strongest such bounds, due to Drinea and Mitzenmacher and to Venkataramanan, Tatikonda, and Ramchandran, have been evaluated mainly for geometric or low-parameter run-length laws. We let a learning algorithm choose the run-length law freely, which raises three challenges. First, the Drinea-Mitzenmacher functional is an infinite sum over the ways deletions merge runs; we show that it depends on the law only through its mean, its full-deletion probability, and bilinear forms in the law and its renewal weights, so that gradients of truncations are exact and every truncation can only lower the bound. Second, for non-Markov inputs the output is no longer Markov and the Venkataramanan-Tatikonda-Ramchandran analysis breaks down; we show that the residual length of the current input run turns the output into a hidden Markov chain, which extends the bound to every finite-support law. Third, its correction term counts only output runs formed from three input runs; we prove a larger correction that accounts for every output run formed by several input runs, which improves the published bound even for truncated geometric laws. Certified by interval arithmetic, the new bounds exceed all previously published deterministic lower bounds at every tabulated deletion probability, by up to $6.86\times10^{-3}$ bits per channel use and 6.8%. At large deletion probability the learned laws concentrate on run-length clusters with survivor counts about two standard deviations apart, like a pulse-amplitude constellation.

cs.IT↗

One Burst of t-Deletion and One Burst of t-Substitution Error-Correcting Codes

Synchronization errors, including insertions, deletions, and substitutions, may occur in bursts in communication systems such as DNA data storage, file synchronization, and magnetic recording. In this paper, we study an error model con?sisting of one burst of t-deletions and one burst of t-substitutions. By reformulating the original sequence into a matrix form, we propose an explicit construction of error-correcting codes capable of correcting one burst of t-deletions and one burst of t-substitutions with O(log n) redundancy.

cs.IT↗

Error-Correcting Codes for Two Bursts of t1-Deletion-t2-Insertion with Reduced Complexity

Burst errors involving simultaneous insertions, deletions, and substitutions occur in practical scenarios, including DNA data storage and document synchronization, motivating the development of channel codes that can correct such errors. In this paper, we construct error-correcting codes (ECCs) capable of handling multiple bursts of t1-deletion-t2-insertion ((t1, t2)-DI) errors, where each burst consists of t1 deletions followed by t2 insertions in a binary sequence. We make three key contributions: First, we establish the fundamental equivalence among (i) ECCs correcting two bursts of (t1, t2)-DI errors, (ii) ECCs correcting two bursts of (t2, t1)-DI errors, and (iii) ECCs correcting one burst of (t1, t2)-DI together with one burst of (t2, t1)-DI errors. Then, we derive lower and upper bounds on the code size of two-burst (t1, t2)-DI ECCs, which can naturally be extended to the case of multiple bursts. Finally, we present constructions of ECCs correcting two bursts of (t1, t2)-DI errors. Compared with codes obtained via the direct application of the syndrome compression technique, the proposed constructions achieve substantially improved computational efficiency.

cs.IT↗

Cluster-Aware Over-the-Air Federated Learning with Energy-Harvesting Devices: From Global Training to Model Personalization

Federated learning (FL) enables distributed optimization and learning across decentralized edge devices while preserving data privacy, but its performance is fundamentally constrained by heterogeneous data distributions, limited communication resources, and energy availability. In practical wireless networks, mobile devices (MDs) often exhibit diverse data and learning objectives, naturally forming clusters of users with jointly trainable models. When devices rely on energy harvesting (EH), stochastic energy arrivals further complicate participation and scheduling under communication constraints. In this work, we study over-the-air (OTA) FL with EH MDs under heterogeneous data distributions, and investigate two closely related learning objectives within a unified framework: one aiming for a more representative global model by reducing data bias, and the other learning more personalized cluster-specific models by exploiting this bias. In the global training mode, cluster information guides energy- and diversity-aware scheduling, ensuring that the scheduled active users provide a more representative aggregate update. In the personalization mode, the same cluster structure defines cluster-level learning objectives and OTA recovery targets, enabling the parameter server to train multiple cluster-specific models through simultaneous transmissions over the wireless multiple-access channel. Numerical results demonstrate that the proposed unified framework improves fairness or personalization, depending on the operating mode, while reducing communication overhead.

cs.LG↗

Bounds on Multiple $b$-Burst Deletion-Correcting Codes

Motivated by their applications in DNA-based storage systems, codes capable of correcting consecutive deletions have attracted significant attention. An important class of such codes consists of those that can correct multiple consecutive deletion errors, commonly referred to as multiple $b$-burst deletion-correcting codes. In this paper, we investigate the fundamental limits of multiple $b$-burst deletion-correcting codes. Specifically, we first characterize several structural properties of the associated deletion balls. Then, leveraging these properties, we derive several upper bounds and a combinatorial lower bound on the maximum size of such codes. As a consequence, our bounds improve upon the previously known results for general parameter regimes and are shown to be asymptotically optimal for certain cases.

cs.IT↗

Channels with Markov Synchronization Errors: Information Stability and Capacity Bounds

Particularly motivated by DNA storage channels, we consider channels with synchronization errors modeled as insertions and deletions, along with substitutions. We focus on the case where the synchronization error process has memory and investigate the information stability of these channels, hence the existence of their Shannon capacity. We assume that the synchronization errors are governed by a stationary and ergodic finite state Markov chain and prove that such a channel is information-stable, which implies the existence of a coding scheme that achieves the limit of mutual information. This result implies the existence of the Shannon capacity for a wide range of channels with synchronization errors, with different applications, including DNA storage. We also provide specific examples of deletion channels with Markov memory and numerically evaluate their capacity bounds, thereby allowing us to quantify the capacity difference between memoryless deletion channels and those with memory with the same deletion probability and reveal that having memory increases the channel capacity.

cs.IT↗

Unsourced Random Access: A Comprehensive Survey

Multiple access communication systems enable numerous users to share common communication resources, playing a crucial role in wireless networks. With the emergence of the sixth generation (6G) and beyond communication networks, supporting massive machine-type communications with sporadic activity patterns is expected to become a critical challenge. Unsourced random access (URA) has emerged as a promising paradigm to address this challenge by decoupling user identification from data transmission through the use of a common codebook. This survey offers a comprehensive overview of URA solutions, encompassing both theoretical foundations and practical applications. We present a systematic classification of URA solutions across three primary channel models: Gaussian multiple access channels (GMACs), single-antenna fading channels, and multiple-input multiple-output (MIMO) fading channels. For each category, we analyze and compare state-of-the-art solutions in terms of performance, complexity, and practical feasibility. Additionally, we discuss critical challenges such as interference management, computational complexity, and synchronization. The survey concludes with promising future research directions and potential methods to address existing limitations, providing a roadmap for researchers and practitioners in this rapidly evolving field.

cs.IT↗

Characterization of Deletion/Substitution Channel Capacity for Small Deletion and Substitution Probabilities

We consider binary input deletion/substitution channels, which model certain types of synchronization errors encountered in practice. Specifically, we focus on the regime of small deletion and substitution probabilities, and by extending an approach developed for the deletion-only channel, we obtain an asymptotic characterization of the channel capacity for independent and identically distributed (i.i.d.) deletion/substitution channels. To do so, given a target probability of successful decoding, we first develop an upper bound on the codebook size for arbitrary but fixed numbers of deletions and substitutions, and then extend the result to the case of random deletions and substitutions to obtain a bound on the channel capacity. Our final result is: The i.i.d. deletion/substitution channel capacity is approximately \(1 - H(p_d) - H(p_s)\), for \(p_d, p_s \approx0\), where \(p_d\) and \(p_s\) are the deletion and substitution probabilities, respectively.

cs.IT↗

Communication via Sensing

We present an alternative take on the recently popularized concept of `\textit{joint sensing and communications}', which focuses on using communication resources also for sensing. Here, we propose the opposite, where we utilize the receiver's sensing capabilities for communication. Our goal is to characterize the fundamental limits of communication over such a channel, which we call `\textit{communication via sensing}'. We assume that changes in the sensed attributes, such as location and speed, are limited due to practical constraints, which are captured by assuming a finite-state channel (FSC) with an input cost constraint. We first formulate an upper bound on the \(N\)-letter capacity as a cost-constrained optimization problem over the input sequence distribution, and then convert it to an equivalent problem over the state sequence distribution. Moreover, by breaking a walk on the underlying Markov chain into a weighted sum of traversed graph cycles in the long walk limit, we obtain a compact single-letter formulation of the capacity upper bound. Finally, for a specific case of a two-state FSC with noisy sensing characterized by a binary symmetric channel (BSC), we obtain a closed-form expression for the capacity upper bound. Comparison with an existing numerical lower bound shows that our proposed upper bound is very tight for all crossover probabilities.

cs.IT↗

Error Correcting Codes for Segmented Burst-Deletion Channels

We study segmented burst-deletion channels motivated by the observation that synchronization errors commonly occur in a bursty manner in real-world settings. In this channel model, transmitted sequences are implicitly divided into non-overlapping segments, each of which may experience at most one burst of deletions. In this paper, we develop error correction codes for segmented burst-deletion channels over arbitrary alphabets under the assumption that each segment may contain only one burst of t-deletions. The main idea is to encode the input subsequence corresponding to each segment using existing one-burst deletion codes, with additional constraints that enable the decoder to identify segment boundaries during the decoding process from the received sequence. The resulting codes achieve redundancy that scales as O(log b), where b is the length of each segment.

cs.IT↗

Half-Marker Codes for Deletion Channels with Applications in DNA Storage

DNA storage systems face significant challenges, including insertion, deletion, and substitution (IDS) errors. Therefore, designing effective synchronization codes, i.e., codes capable of correcting IDS errors, is essential for DNA storage systems. Marker codes are a favorable choice for this purpose. In this paper, we extend the notion of marker codes by making the following key observation. Since each DNA base is equivalent to a 2-bit storage unit, one bit can be reserved for synchronization, while the other is dedicated to data transmission. Using this observation, we propose a new class of marker codes, which we refer to as half-marker codes. We demonstrate that this extension has the potential to significantly increase the mutual information between the input symbols and the soft outputs of an IDS channel modeling a DNA storage system. Specifically, through examples, we show that when concatenated with an outer error-correcting code, half-marker codes outperform standard marker codes and significantly reduce the end-to-end bit error rate of the system.

cs.IT↗

Constrained Error-Correcting Codes for Efficient DNA Synthesis

DNA synthesis is considered as one of the most expensive components in current DNA storage systems. In this paper, focusing on a common synthesis machine, which generates multiple DNA strands in parallel following a fixed supersequence,we propose constrained codes with polynomial-time encoding and decoding algorithms. Compared to the existing works, our codes simultaneously satisfy both l-runlength limited and ε-balanced constraints. By enumerating all valid sequences, our codes achieve the maximum rate, matching the capacity. Additionally, we design constrained error-correcting codes capable of correcting one insertion or deletion in the obtained DNA sequence while still adhering to the constraints.

cs.IT↗

Update Estimation and Scheduling for Over-the-Air Federated Learning with Energy Harvesting Devices

We study over-the-air (OTA) federated learning (FL) for energy harvesting devices with heterogeneous data distribution over wireless fading multiple access channel (MAC). To address the impact of low energy arrivals and data heterogeneity on global learning, we propose user scheduling strategies. Specifically, we develop two approaches: 1) entropy-based scheduling for known data distributions and 2) least-squares-based user representation estimation for scheduling with unknown data distributions at the parameter server. Both methods aim to select diverse users, mitigating bias and enhancing convergence. Numerical and analytical results demonstrate improved learning performance by reducing redundancy and conserving energy.

cs.LG↗

A Deep Learning Based Decoder for Concatenated Coding over Deletion Channels

In this paper, we introduce a deep learning-based decoder designed for concatenated coding schemes over a deletion/substitution channel. Specifically, we focus on serially concatenated codes, where the outer code is either a convolutional or a low-density parity-check (LDPC) code, and the inner code is a marker code. We utilize Bidirectional Gated Recurrent Units (BI-GRUs) as log-likelihood ratio (LLR) estimators and outer code decoders for estimating the message bits. Our results indicate that decoders powered by BI-GRUs perform comparably in terms of error rates with the MAP detection of the marker code. We also find that a single network can work well for a wide range of channel parameters. In addition, it is possible to use a single BI-GRU based network to estimate the message bits via one-shot decoding when the outer code is a convolutional code.

cs.IT↗

A Practical Concatenated Coding Scheme for Noisy Shuffling Channels with Coset-based Indexing

Noisy shuffling channels capture the main characteristics of DNA storage systems where distinct segments of data are received out of order, after being corrupted by substitution errors. For realistic schemes with short-length segments, practical indexing and channel coding strategies are required to restore the order and combat the channel noise. In this paper, we develop a finite-length concatenated coding scheme that employs Reed-Solomon (RS) codes as outer codes and polar codes as inner codes, and utilizes an implicit indexing method based on cosets of the polar code. We propose a matched decoding method along with a metric for detecting the index that successfully restores the order, and correct channel errors at the receiver. Residual errors that are not corrected by the matched decoder are then corrected by the outer RS code. We derive analytical approximations for the frame error rate of the proposed scheme, and also evaluate its performance through simulations to demonstrate that the proposed implicit indexing method outperforms explicit indexing.

cs.IT↗

Capacity Bounds for the Poisson-Repeat Channel

We develop bounds on the capacity of Poisson-repeat channels (PRCs) for which each input bit is independently repeated according to a Poisson distribution. The upper bounds are obtained by considering an auxiliary channel where the output lengths corresponding to input blocks of a given length are provided as side information at the receiver. Numerical results show that the resulting upper bounds are significantly tighter than the best known one for a large range of the PRC parameter $λ$ (specifically, for $λ\ge 0.35$). We also describe a way of obtaining capacity lower bounds using information rates of the auxiliary channel and the entropy rate of the provided side information.

cs.IT↗

RIS-Aided Unsourced Multiple Access (RISUMA): Coding Strategy and Performance Limits

This paper considers an unsourced random access (URA) set-up equipped with a passive reconfigurable intelligent surface (RIS), where a massive number of unidentified users (only a small fraction of them being active at any given time) are connected to the base station (BS). We introduce a slotted coding scheme for which each active user chooses a slot at random for transmitting its signal, consisting of a pilot part and a randomly spread polar codeword. The proposed decoder operates in two phases. In the first phase, called the RIS configuration phase, the BS detects the transmitted pilots. The detected pilots are then utilized to estimate the corresponding users' channel state information, using which the BS suitably selects RIS phase shift employing the proposed RIS design algorithms. The proposed channel estimator offers the capability to obtain the channel coefficients of the users whose pilots interfere with each other without prior access to the list of transmitted pilots or the number of active users. In the second phase, called the data phase, transmitted messages of active users are decoded. Moreover, we establish an approximate achievability bound for the RIS-based URA scheme, providing a valuable benchmark. Computer simulations show that the proposed scheme outperforms the state-of-the-art for RIS-aided URA.

cs.IT↗

Next Generation Advanced Transceiver Technologies for 6G and Beyond

To accommodate new applications such as extended reality, fully autonomous vehicular networks and the metaverse, next generation wireless networks are going to be subject to much more stringent performance requirements than the fifth-generation (5G) in terms of data rates, reliability, latency, and connectivity. It is thus necessary to develop next generation advanced transceiver (NGAT) technologies for efficient signal transmission and reception. In this tutorial, we explore the evolution of NGAT from three different perspectives. Specifically, we first provide an overview of new-field NGAT technology, which shifts from conventional far-field channel models to new near-field channel models. Then, three new-form NGAT technologies and their design challenges are presented, including reconfigurable intelligent surfaces, flexible antennas, and holographic multi-input multi-output (MIMO) systems. Subsequently, we discuss recent advances in semantic-aware NGAT technologies, which can utilize new metrics for advanced transceiver designs. Finally, we point out other promising transceiver technologies for future research.

cs.IT↗