SearcharxivSearch

arXiv subjects

Gerhard Bauch

Publications and source records attributed to Gerhard Bauch.

At least 19 recordsLinked to original sources

Mixed Block Markov Superposition Transmission Codes

Block Markov superposition transmission (BMST) codes provide a flexible framework for constructing codes with near-capacity performance and low-complexity sliding-window decoding. However, existing BMST variants show contrasting performance limitations: recursive BMST (rBMST) codes suffer from error propagation but avoid high error floors, whereas non-recursive BMST codes exhibit the opposite behavior. Motivated by these complementary characteristics, we combine recursive and non-recursive components through parallel and serial concatenation, yielding mixed BMST (mBMST) codes. The proposed framework subsumes existing BMST variants and enables new BMST structures. Simulations show that these structures improve FER and BER performance with lower memory requirements than rBMST.

cs.IT

Generalized Framework for a Fair Comparison of Cellular and Cooperative Massive MIMO Systems

Cooperative massive multiple-input multiple-output (MIMO) promises large gains over cellular deployments, but existing comparisons of different architectures often mix antenna distribution, inter-site coordination, and processing assumptions. This paper introduces a graph-based framework for fair comparison of cellular, coordinated, and cell-free massive-MIMO systems. We differentiate between two key properties, namely antenna distribution and inter-site cooperation, which yields seven representative system types. We derive compatible uplink and downlink spectral efficiency (SE) expressions, including an uplink bound for detectors with mixed instantaneous and statistical effective channel state information (CSI), and adapt scalable user association and processing rules to all considered architectures. We evaluate these systems using extensive numerical simulations and show that for a fair comparison much larger simulation areas (at least 2.5 $\times$ 2.5 km2) than commonly used are required. We introduce the relative capacity, which measures how closely each architecture approaches centralized cell-free processing. The results show that coordinated, phase-aligned beamforming across spatially distributed antennas is the main source of cooperation gains. In dense deployments with few antennas per access point (AP), coordinated Distributed Antenna System (DAS) and hybrid cell-free architectures achieve much of the centralized cell-free performance while requiring substantially weaker midhaul assumptions.

cs.IT

Turbo Equalization with Coarse Quantization using the Information Bottleneck Method

This paper proposes a turbo equalizer for intersymbol interference channels (ISI) that uses coarsely quantized messages across all receiver components. Lookup tables (LUTs) carry out compression operations designed with the information bottleneck method aiming to maximize relevant mutual information. The turbo setup consists of an equalizer and a decoder that provide extrinsic information to each other over multiple turbo iterations. We develop simplified LUT structures to incorporate the decoder feedback in the equalizer with significantly reduced complexity. The proposed receiver is optimized for selected ISI channels. A conceptual hardware implementation is developed to compare the area efficiency and error correction performance. A thorough analysis reveals that LUT-based configurations with very coarse quantization can achieve higher area efficiency than conventional equalizers. Moreover, the proposed turbo setups can outperform the respective non-turbo setups regarding area efficiency and error correction capability.

cs.IT

Memory-Assisted Quantized LDPC Decoding

We enhance coarsely quantized LDPC decoding by reusing computed check node messages from previous iterations. Typically, variable and check nodes update and replace old messages every iteration. We show that, under coarse quantization, discarding old messages entails a significant loss of mutual information. The loss is avoided with additional memory, improving performance by up to 0.23 dB. We optimize quantization with a modified information bottleneck algorithm that considers the statistics of old messages. A simple merge operation reduces memory requirements. Depending on channel conditions and code rate, memory assistance enables up to 32 % better area efficiency for 2-bit decoding.

cs.IT

Region-Specific Coarse Quantization with Check Node Awareness in 5G LDPC Decoding

This paper presents novel techniques for improving the error correction performance and reducing the complexity of coarsely quantized 5G-LDPC decoders. The proposed decoder design supports arbitrary message-passing schedules on a base-matrix level by modeling exchanged messages with entry-specific discrete random variables. Variable nodes (VNs) and check nodes (CNs) involve compression operations designed using the information bottleneck method to maximize preserved mutual information between code bits and quantized messages. We introduce alignment regions that assign the messages to groups with aligned reliability levels to decrease the number of individual design parameters. Group compositions with degree-specific separation of messages improve performance by up to 0.4 dB. Further, we generalize our recently proposed CN-aware quantizer design to irregular LDPC codes and layered schedules. The method optimizes the VN quantizer to maximize preserved mutual information at the output of the subsequent CN update, enhancing performance by up to 0.2 dB. A schedule optimization modifies the order of layer updates, reducing the average iteration count by up to 35%. We integrate all new techniques in a rate-compatible decoder design by extending the alignment regions along a rate-dimension. Our complexity analysis shows that 2-bit decoding can double the area efficiency over 4-bit decoding at comparable performance.

cs.IT

Finite Alphabet Fast List Decoders for Polar Codes

The so-called fast polar decoding schedules are meant to improve the decoding speed of the sequential-natured successive cancellation list decoders. The decoding speedup is achieved by replacing various parts of the serial decoding process with efficient special-purpose decoder nodes. This work incorporates the fast decoding schedules for polar codes into their quantized finite alphabet decoding. In a finite alphabet successive cancellation list decoder, the log-likelihood ratio computations are replaced with lookup operations on low-resolution integer messages. The lookup tables are designed using the information bottleneck method. It is shown that the finite alphabet decoders can also leverage the special decoder nodes found in the literature. Besides their inherent decoding speed improvement, the use of these special decoder nodes drastically reduces the number of lookup tables required to perform the finite alphabet decoding. In order to perform quantized decoding using lookup operations, the proposed decoders require up to 93% less unique lookup tables as compared to the ones that use the conventional successive cancellation schedule. Moreover, the proposed decoders exhibit negligible loss in error correction performance without necessitating alterations to the lookup table design process.

cs.IT

Unrolled and Pipelined Decoders based on Look-Up Tables for Polar Codes

Unrolling a decoding algorithm allows to achieve extremely high throughput at the cost of increased area. Look-up tables (LUTs) can be used to replace functions otherwise implemented as circuits. In this work, we show the impact of replacing blocks of logic by carefully crafted LUTs in unrolled decoders for polar codes. We show that using LUTs to improve key performance metrics (e.g., area, throughput, latency) may turn out more challenging than expected. We present three variants of LUT-based decoders and describe their inner workings as well as circuits in detail. The LUT-based decoders are compared against a regular unrolled decoder, employing fixed-point representations for numbers, with a comparable error-correction performance. A short systematic polar code is used as an illustration. All resulting unrolled decoders are shown to be capable of an information throughput of little under 10 Gbps in a 28 nm FD-SOI technology clocked in the vicinity of 1.4 GHz to 1.5 GHz. The best variant of our LUT-based decoders is shown to reduce the area requirements by 23% compared to the regular unrolled decoder while retaining a comparable error-correction performance.

cs.IT

Implementation-Efficient Finite Alphabet Decoding of Polar Codes

An implementation-efficient finite alphabet decoder for polar codes relying on coarsely quantized messages and low-complexity operations is proposed. Typically, finite alphabet decoding performs concatenated compression operations on the received channel messages to aggregate compact reliability information for error correction. These compression operations or mappings can be considered as lookup tables. For polar codes, the finite alphabet decoder design boils down to constructing lookup tables for the upper and lower branches of the building blocks within the code structure. A key challenge is to realize a hardware-friendly implementation of the lookup tables. This work uses the min-sum implementation for the upper branch lookup table and, as a novelty, a computational domain implementation for the lower branch lookup table. The computational domain approach drastically reduces the number of implementation parameters. Furthermore, a restriction to uniform quantization in the lower branch allows a very hardware-friendly compression via clipping and bit-shifting. Its behavior is close to the optimal non-uniform quantization, whose implementation would require multiple high-resolution threshold comparisons. Simulation results confirm excellent performance for the developed decoder. Unlike conventional fixed-point decoders, the proposed method involves an offline design that explicitly maximizes the preserved mutual information under coarse quantization.

cs.IT

Low-Resolution Horizontal and Vertical Layered Mutual Information Maximizing LDPC Decoding

We investigate iterative low-resolution message-passing algorithms for quasi-cyclic LDPC codes with horizontal and vertical layered schedules. Coarse quantization and layered scheduling are highly relevant for hardware implementations to reduce the bit width of messages and the number of decoding iterations. As a novelty, this paper compares the two scheduling variants in combination with mutual information maximizing compression operations in variable and check nodes. We evaluate the complexity and error rate performance for various configurations. Dedicated hardware architectures for regular quasi-cyclic LDPC decoders are derived on a conceptual level. The hardware-resource estimates confirm that most of the complexity lies within the routing network operations. Our simulations reveal similar error rate performance for both layered schedules but a slightly lower average iteration count for the horizontal decoder.

cs.IT

Uniform vs. Non-Uniform Coarse Quantization in Mutual Information Maximizing LDPC Decoding

Recently, low-resolution LDPC decoders have been introduced that perform mutual information maximizing signal processing. However, the optimal quantization in variable and check nodes requires expensive non-uniform operations. Instead, we propose to use uniform quantization with a simple hardware structure, which reduces the complexity of individual node operations approximately by half and shortens the decoding delay significantly. Our analysis shows that the loss of preserved mutual information resulting from restriction to uniform quantization is very small. Furthermore, the error rate simulations with regular LDPC codes confirm that the uniformly quantized decoders cause only minor performance degradation within 0.01 dB compared to the non-uniform alternatives. Due to the complexity reduction, especially the proposed 3-bit decoder is a promising candidate to replace 4-bit conventional decoders.

cs.IT

A Variable Node Design with Check Node Aware Quantization Leveraging 2-Bit LDPC Decoding

For improving coarsely quantized decoding of LDPC codes, we propose a check node aware design of the variable node update. In contrast to previous works, we optimize the variable node to explicitly maximize the mutual information preserved in the check-to-variable instead of the variable-to-check node messages. The extended optimization leads to a significantly different solution for the compression operation at the variable node. Simulation results for regular LDPC codes confirm that the check node aware design, especially for very coarse quantization with 2- or 3-bit messages, achieves performance gains of up to 0.2 dB - without additional hardware costs. We also show that the 2-bit message resolution enables a very efficient implementation of the check node update, which requires only 2/9 of the 3-bit check node's transistor count and reduces the signal propagation delay by a factor of 4.

cs.IT

Reconstruction-Computation-Quantization (RCQ): A Paradigm for Low Bit Width LDPC Decoding

This paper uses the reconstruction-computation-quantization (RCQ) paradigm to decode low-density parity-check (LDPC) codes. RCQ facilitates dynamic non-uniform quantization to achieve good frame error rate (FER) performance with very low message precision. For message-passing according to a flooding schedule, the RCQ parameters are designed by discrete density evolution (DDE). Simulation results on an IEEE 802.11 LDPC code show that for 4-bit messages, a flooding MinSum RCQ decoder outperforms table-lookup approaches such as information bottleneck (IB) or Min-IB decoding, with significantly fewer parameters to be stored. Additionally, this paper introduces layer-specific RCQ (LS-RCQ), an extension of RCQ decoding for layered architectures. LS-RCQ uses layer-specific message representations to achieve the best possible FER performance. For LS-RCQ, this paper proposes using layered DDE featuring hierarchical dynamic quantization (HDQ) to design LS-RCQ parameters efficiently. Finally, this paper studies field-programmable gate array (FPGA) implementations of RCQ decoders. Simulation results for a (9472, 8192) quasi-cyclic (QC) LDPC code show that a layered MinSum RCQ decoder with 3-bit messages achieves more than a $10\%$ reduction in LUTs and routed nets and more than a $6\%$ decrease in register usage while maintaining comparable decoding performance, compared to a 5-bit offset MinSum decoder.

eess.SP

Bounds on the Error Probability of Raptor Codes under Maximum Likelihood Decoding

In this paper upper and lower bounds on the probability of decoding failure under maximum likelihood decoding are derived for different (nonbinary) Raptor code constructions. In particular four different constructions are considered; (i) the standard Raptor code construction, (ii) a multi-edge type construction, (iii) a construction where the Raptor code is nonbinary but the generator matrix of the LT code has only binary entries, (iv) a combination of (ii) and (iii). The latter construction resembles the one employed by RaptorQ codes, which at the time of writing this article represents the state of the art in fountain codes. The bounds are shown to be tight, and provide an important aid for the design of Raptor codes.

cs.IT

Spatiotemporal Dependable Task Execution Services in MEC-enabled Wireless Systems

Multi-access Edge Computing (MEC) enables computation and energy-constrained devices to offload and execute their tasks on powerful servers. Due to the scarce nature of the spectral and computation resources, it is important to jointly consider i) contention-based communications for task offloading and ii) parallel computing and occupation of failure-prone MEC processing resources (virtual machines). The feasibility of task offloading and successful task execution with virtually no failures during the operation time needs to be investigated collectively from a combined point of view. To this end, this letter proposes a novel spatiotemporal framework that utilizes stochastic geometry and continuous time Markov chains to jointly characterize the communication and computation performance of dependable MEC-enabled wireless systems. Based on the designed framework, we evaluate the influence of various system parameters on different dependability metrics such as (i) computation resources availability, (ii) task execution retainability, and (iii) task execution capacity. Our findings showcase that there exists an optimal number of virtual machines for parallel computing at the MEC server to maximize the task execution capacity.

cs.IT

Prioritized Multi-stream Traffic in Uplink IoT Networks: Spatially Interacting Vacation Queues

Massive Internet of Things (IoT) is foreseen to introduce plethora of applications for a fully connected world. Heterogeneous traffic is envisaged, where packets generated at each IoT device should be differentiated and served according to their priority. This paper develops a novel priority-aware spatiotemporal mathematical model to characterize massive IoT networks with uplink prioritized multistream traffic (PMT). Particularly, stochastic geometry is utilized to account for the macroscopic network wide mutual interference between the coexisting IoT devices. Discrete time Markov chains (DTMCs) are employed to track the microscopic evolution of packets within each priority stream at each device. To alleviate the curse of dimensionality, we decompose the prioritized queueing model at each device to a single-queue system with server vacation. To this end, the IoT network with PMT is modeled as spatially interacting vacation queues. Interactions between queues, in terms of the packet departure probabilities, occur due to mutual interference. Service vacations occur to lower priority packets to address higher priority packets. Based on the proposed model, dedicated and shared channel access strategies for different priority classes are presented and compared. The results show that shared access provides better performance when considering the transmission success probability, queues overflow probability and latency.

cs.IT

A Reconstruction-Computation-Quantization (RCQ) Approach to Node Operations in LDPC Decoding

In this paper, we propose a finite-precision decoding method that features the three steps of Reconstruction, Computation, and Quantization (RCQ). Unlike Mutual-Information-Maximization Quantized Belief Propagation (MIM-QBP), RCQ can approximate either belief propagation or Min-Sum decoding. One problem faced by MIM-QBP decoder is that it cannot work well when the fraction of degree-2 variable nodes is large. However, sometimes a large fraction of degree-2 variable nodes is necessary for a fast encoding structure, as seen in the IEEE 802.11 standard and the DVB-S2 standard. In contrast, the proposed RCQ decoder may be applied to any off-the-shelf LDPC code, including those with a large fraction of degree-2 variable nodes.Our simulations show that a 4-bit Min-Sum RCQ decoder delivers frame error rate (FER) performance around 0.1dB of full-precision belief propagation (BP) for the IEEE 802.11 standard LDPC code in the low SNR region.The RCQ decoder actually outperforms full-precision BP in the high SNR region because it overcomes elementary trapping sets that create an error floor under BP decoding. This paper also introduces Hierarchical Dynamic Quantization (HDQ) to design the non-uniform quantizers required by RCQ decoders. HDQ is a low-complexity design technique that is slightly sub-optimal. Simulation results comparing HDQ and an optimal quantizer on the symmetric binary-input memoryless additive white Gaussian noise channel show a loss in mutual information between these two quantizers of less than $10^{-6}$ bits, which is negligible for practical applications.

eess.SP

A Spatiotemporal Framework for Information Freshness in IoT Uplink Networks

Timely message delivery is a key enabler for Internet of Things (IoT) and cyber-physical systems to support wide range of context-dependent applications. Conventional time-related metrics, such as delay, fails to characterize the timeliness of the system update or to capture the freshness of information from application perspective. Age of information (AoI) is a time.evolving measure of information freshness that has received considerable attention during the past years. In the foreseen large scale and dense IoT networks, joint temporal (i.e., queue aware) and spatial (i.e., mutual interference aware) characterization of the AoI is required. In this work we provide a spatiotemporal framework that captures the peak AoI for large scale IoT uplink network. To this end, the paper quantifies the peak AoI for large scale cellular network with Bernoulli uplink traffic. Simulation results are conducted to validate the proposed model and show the effect of traffic load and decoding threshold. Insights are driven to characterize the network stability frontiers and the location-dependent performance within the network.

cs.IT

A Spatiotemporal Model for Peak AoI in Uplink IoT Networks: Time Vs Event-triggered Traffic

Timely message delivery is a key enabler for Internet of Things (IoT) and cyber-physical systems to support wide range of context-dependent applications. Conventional time-related metrics (e.g. delay and jitter) fails to characterize the timeliness of the system update. Age of information (AoI) is a time-evolving metric that accounts for the packet inter-arrival and waiting times to assess the freshness of information. In the foreseen large-scale IoT networks, mutual interference imposes a delicate relation between traffic generation patterns and transmission delays. To this end, we provide a spatiotemporal framework that captures the peak AoI (PAoI) for large scale IoT uplink network under time-triggered (TT) and event triggered (ET) traffic. Tools from stochastic geometry and queueing theory are utilized to account for the macroscopic and microscopic network scales. Simulations are conducted to validate the proposed mathematical framework and assess the effect of traffic load on PAoI. The results unveil a counter-intuitive superiority of the ET traffic over the TT in terms of PAoI, which is due to the involved temporal interference correlations. Insights regarding the network stability frontiers and the location-dependent performance are presented. Key design recommendations regarding the traffic load and decoding thresholds are highlighted.

cs.IT