SearcharxivSearch

arXiv subjects

Andreas Burg

Publications and source records attributed to Andreas Burg.

At least 19 recordsLinked to original sources

Centralized RAN for Future Low-Power Wide-Area Networks: A LoRa Case Study

In recent years, low-power wide-area network (LPWAN) technologies have gained significant traction as a connectivity option for Internet of Things (IoT) applications. While these networks have been successful in providing long-range, low-power, and low-cost connectivity, they currently face scalability, reliability, and efficiency challenges that require immediate attention. In this paper, we first identify important challenges for LPWANs. We then advocate for the introduction of a centralized radio access network (C-RAN) architecture tailored for LPWANs and present a proof-of-concept implementation and deployment of the proposed C-RAN for the widely popular long range (LoRa) standard. We also provide experimental results to demonstrate and quantify the increased sensitivity that can be obtained from joint processing of the baseband signals of multiple receivers, enabled by the proposed centralized architecture in quasi-static scenarios and drone-mounted transmitters.

eess.SP

Enabling Ultra-Low-Power Always-On Feedforward Leakage Suppression Logic Circuits with FDSOI

The growing deployment of real-time applications on wearable and Internet of Things (IoT) edge devices has intensified the need for energy-efficient, high-performance systems that meet stringent timing and energy constraints. Events-driven architectures leverage the sparsity of real-time to further improve system energy efficiency by employing an always-on (AO) domain to monitor inputs and activate a high-performance (HP) domain only when relevant events occur. However, for low-duty-cycle applications, the energy bottleneck shifts toward the AO domain, where leakage power dominates overall consumption. To mitigate this issue, AO circuits are typically implemented using high voltage threshold (HVT) or ultra-high voltage threshold (UHVT) transistors, thereby avoiding sub- and near-threshold operation, which is highly sensitive to process, voltage, and temperature (PVT) variations. In this context, feedforward leakage suppression logic (FLSL) has recently emerged as a promising candidate, offering reduced leakage compared to conventional. However, previous studies report a significant degradation in FLSL leakage performance in technology nodes below 90 nm, primarily due to increased gate and junction leakage currents. FDSOI technology, with its ability to effectively suppress junction leakage, provides an opportunity to overcome this limitation and restore FLSL efficiency in advanced nodes. Therefore, we demonstrate in this work that FLSL implemented in a 22 nm FDSOI technology can significantly reduce the energy consumption of small AO circuits with low-frequency inputs compared to state-of-the-art ultra-low-power CMOS designs. Silicon measurements on an FIR filter and an AES cryptographic core show reduced operating voltage and up to 9.8 x and 1.83 x reductions in leakage power compared to equivalent HVT and UHVT CMOS implementations, respectively.

cs.ET

Reducing Power Consumption of Embedded Dynamic Memories with ECCs

Gain-cell embedded dynamic random-access memory (GCRAM) offers dense and energy-efficient on-chip storage, but retention-time variations force frequent refresh operations to cover worst-case bits. Error-correction codes (ECCs) can alleviate this limitation by masking bit errors from weak cells and thereby reduce refresh cost. However, the trade-off between the additional access and logic energy introduced by ECCs and the power savings from longer refresh intervals is nontrivial, especially considering the wide range of available ECC options. To optimize overall power consumption, we propose an ECC selection method that combines a refresh-interval model with power analysis to identify the minimum-power ECC configurations under a given yield constraint. Across different memory bandwidths, activity factors, and read/write ratios, the evaluation results show that the best ECC option shifts from stronger codes in refresh-dominated operating regions to lower-overhead codes in access-dominated regions and achieves 46.8% to 94.8% reduction in total power relative to the no-ECC reference.

cs.IT

High-throughput Low-latency Hardware Implementation of BCH Decoders

Two well-known decoding algorithms for BCH codes are conventional decoding, based on the Berlekamp-Massey algorithm in combination with Chien search, and direct decoding, which uses direct solutions to find the error locator polynomial and its roots. We introduce hardware architectures for conventional and direct decoding of extended BCH codes. Both architectures support implementation for any blocklength. Our conventional decoder supports any error-correction capability, whereas direct decoding is supported up to error correcting capability t = 4. To the best of our knowledge, our work is the first to implement a direct BCH decoder with an error-correction capability 4. We synthesize for the Xilinx Ultrascale+ XCZU48DR field-programmable gate-array and 16 nm FinFET for blocklengths up to 1024 bits and t = 4. We show that the direct decoder outperforms the conventional decoder in area efficiency for t = 2, t = 3, and for t = 4 for blocklengths longer than 256. Post-synthesis results for 16 nm FinFET show codeword per clock-cycle throughput at 1 GHz, achieving 239 Gb/s for the (256, 239) eBCH code and 223 Gb/s for (256, 223) eBCH code at 2 ns and 8 ns latency, respectively.

cs.IT

Maximum Coverage Chase Decoder for Optical Interconnects

We propose a low-complexity Chase decoder for optical interconnects that formulates test pattern selection as a generalized maximum coverage problem. For concatenated RS-BCH and oFEC codes, our decoder achieves the standard Chase decoding performance with 25% and 61.3% fewer test patterns, respectively.

cs.IT

Lifted Gabidulin Construction for LDPC Representations of Finite Geometry Codes

Finite geometry (FG) codes combine the algebraic properties of classical block codes with the iterative belief propagation (BP) decoding ability of low-density parity-check~(LDPC) codes. However, exploiting both advantages in practice is hindered by the fact that the standard incidence matrix between $(\mu+1)$-flats and points is dense and contains many short cycles for any flat dimension $\mu\geq 1$. In this work, we propose to sparsify the decoding matrix based on pencil selection, formulated as a constant-dimension subspace packing problem and solved explicitly using lifted Gabidulin codes. For both affine and projective geometries, sparse parity-check matrices are constructed and verified for FG codes of lengths up to $1024$. Simulations on four FG codes show no visible error floor and around $0.5$~dB gain over corresponding 5G LDPC codes at a block error rate of $10^{-7}$.

cs.IT

Dataset and UAV Propagation Channel Modeling for LoRa in the 860 MHz ISM Band

LoRa is one of the most widely used low-power wide-area network technology for the Internet of Things. To achieve long-range communication with low power consumption at a low cost, LoRa uses a chirp spread spectrum modulation and transmits in the sub-GHz unlicensed industrial, scientific, and medical (ISM) frequency bands. Due to the rapid densification of IoT networks, it is crucial to obtain tailored channel models to evaluate the performance of LoRa networks. While channel models for cellular technologies have been investigated extensively, specific characteristics of LoRa transmissions operating at long range with a rather small (~ 250kHz) bandwidth require dedicated measurement campaigns and modeling efforts. In this work, we leverage an SDR-based testbed to gather and publish a dataset of LoRa frames transmitted in a campus environment. The dataset includes IQ samples of the received frames at multiple locations and allows for the evaluation of channel variations with high time resolution. Using the gathered data, we derive empirical propagation channel models for LoRa that include receiver correlation over distance for three scenarios: unmanned aerial vehicle (UAV) line-of-sight (LoS), UAV non-LoS, and pedestrian non-LoS. Furthermore, the dataset is annotated with synchronization information, enabling the evaluation of receiver algorithms using experimental data.

eess.SP

ComplexBeat: Breathing Rate Estimation from Complex CSI

In this paper, we explore the use of channel state information (CSI) from a WiFi system to estimate the breathing rate of a person in a room. In order to extract WiFi CSI components that are sensitive to breathing, we propose to consider the delay domain channel impulse response (CIR), while most state-of-the-art methods consider its frequency domain representation. One obstacle while processing the CSI data is that its amplitude and phase are highly distorted by measurement uncertainties. We thus also propose an amplitude calibration method and a phase offset calibration method for CSI measured in orthogonal frequency-division multiplexing (OFDM) multiple-input multiple-output (MIMO) systems. Finally, we implement a complete breathing rate estimation system in order to showcase the effectiveness of our proposed calibration and CSI extraction methods.

eess.SP

Training Channel Selection for Learning-based 1-bit Precoding in Massive MU-MIMO

Learning-based algorithms have gained great popularity in communications since they often outperform even carefully engineered solutions by learning from training samples. In this paper, we show that the selection of appropriate training examples can be important for the performance of such learning-based algorithms. In particular, we consider non-linear 1-bit precoding for massive multi-user MIMO systems using the C2PO algorithm. While previous works have already shown the advantages of learning critical coefficients of this algorithm, we demonstrate that straightforward selection of training samples that follow the channel model distribution does not necessarily lead to the best result. Instead, we provide a strategy to generate training data based on the specific properties of the algorithm, which significantly improves its error floor performance.

eess.SP

LoRa Fine Synchronization with Two-Pass Time and Frequency Offset Estimation

LoRa is currently one of the most widely used low-power wide-area network (LPWAN) technologies. The physical layer leverages a chirp spread spectrum modulation to achieve long-range communication with low power consumption. Synchronization at long distances is a challenging task as the spread signal can lie multiple orders of magnitude below the thermal noise floor. Multiple research works have proposed synchronization algorithms for LoRa under different hardware impairments. However, the impact of sampling frequency offset (SFO) has mostly either been ignored or tracked only during the data phase, but it often harms synchronization. In this work, we extend existing synchronization algorithms for LoRa to estimate and compensate SFO already in the preamble and show that this early compensation has a critical impact on the estimation of other impairments such as carrier frequency offset and sampling time offset. Therefore it is critical to recover long-range signals.

eess.SP

Impact of Reactive Jamming Attacks on LoRaWAN: a Theoretical and Experimental Study

This paper investigates the impact of reactive jamming on LoRaWAN networks, focusing on showing that LoRaWAN communications can be effectively disrupted with minimal jammer exposure time. The susceptibility of LoRa to jamming is assessed through a theoretical study of how the frame success rate is impacted by only a few jamming symbols. Different jamming approaches are studied, among which repeated-symbol jamming appears to be the most disruptive, with sufficient jamming power. A key contribution of this work is the proposal of a software-defined radio (SDR)-based jamming approach implemented on GNU Radio that generates a controlled number of random symbols, independent of the standard LoRa frame structure. This approach enables precise control over jammer exposure time and provides flexibility in studying the effect of jamming symbols on network performance. The theoretical analysis is validated through experimental results, where the implemented jammer is used to assess the impact of jamming under various configurations. Our findings demonstrate that LoRa-based networks can be disrupted with a minimal number of symbols, emphasizing the need for future research on stealthy communication techniques to counter such jamming attacks.

cs.NI

An SDR-Based Monostatic Wi-Fi System with Analog Self-Interference Cancellation for Sensing

Wireless sensing offers an alternative to wearables for contactless monitoring of human activity and vital signs. However, most existing systems use bistatic setups, which suffer from phase imperfections due to unsynchronized clocks. Monostatic systems overcome this issue, but are hindered by strong self-interference (SI) that require effective cancellation. We present a monostatic Wi-Fi sensing system that uses an auxiliary transmit RF chain to achieve SI cancellation levels of 40 dB, comparable to existing solutions with custom cancellation hardware. We demonstrate that the cancellation filter weights, fine-tuned using least-mean squares, can be directly repurposed for target sensing. Moreover, we achieve stable SI cancellation over 30 minutes in an office environment without fine-tuning, enabling traditional vital sign monitoring using channel estimates derived from baseband samples without the adaptation of the cancellation affecting the sensing channel -- a significant limitation in prior work. Experimental results confirm the detection of small, slow-moving targets, representative for breathing chest movements, at distances up to 10 meters in non-line-of-sight conditions.

eess.SP

Edge-Spreading Raptor-Like LDPC Codes for 6G Wireless Systems

Next-generation channel coding has stringent demands on throughput, energy consumption, and error rate performance while maintaining key features of 5G New Radio (NR) standard codes such as rate compatibility, which is a significant challenge. Due to excellent capacity-achieving performance, spatially-coupled low-density parity-check (SC-LDPC) codes are considered a promising candidate for next-generation channel coding. In this paper, we propose an SC-LDPC code family called edge-spreading Raptor-like (ESRL) codes. Unlike other SC-LDPC codes that adopt the structure of existing rate-compatible LDPC block codes before coupling, ESRL codes maximize the possible locations of edge placement and focus on constructing an optimal coupled matrix. Moreover, a new graph representation called the unified graph is introduced. This graph offers a global perspective on ESRL codes and identifies the optimal edge reallocation to optimize the spreading strategy. We conduct comprehensive comparisons of ESRL codes and 5G-NR LDPC codes. Simulation results demonstrate that when all decoding parameters and complexity are the same, ESRL codes have obvious advantages in error rate performance and throughput compared to 5G-NR LDPC codes in some specific scenarios (low and high number of iterations), making them a promising solution towards next-generation channel coding.

cs.IT

A Node-Based Polar List Decoder with Frame Interleaving and Ensemble Decoding Support

Node-based successive cancellation list (SCL) decoding has received considerable attention in wireless communications for its significant reduction in decoding latency, particularly with 5G New Radio (NR) polar codes. However, the existing node-based SCL decoders are constrained by sequential processing, leading to complicated and data-dependent computational units that introduce unavoidable stalls, reducing hardware efficiency. In this paper, we present a frame-interleaving hardware architecture for a generalized node-based SCL decoder. By efficiently reusing otherwise idle computational units, two independent frames can be decoded simultaneously, resulting in a significant throughput gain. Based on this new architecture, we further exploit graph ensembles to diversify the decoding space, thus enhancing the error-correcting performance with a limited list size. Two dynamic strategies are proposed to eliminate the residual stalls in the decoding schedule, which eventually results in nearly 2x throughput compared to the state-of-the-art baseline node-based SCL decoder. To impart the decoder rate flexibility, we develop a novel online instruction generator to identify the generalized nodes and produce instructions on-the-fly. The corresponding 28nm FD-SOI ASIC SCL decoder with a list size of 8 has a core area of 1.28 mm2 and operates at 692 MHz. It is compatible with all 5G NR polar codes and achieves a throughput of 3.34 Gbps and an area efficiency of 2.62 Gbps/mm2 for uplink (1024, 512) codes, which is 1.41x and 1.69x better than the state-of-the-art node-based SCL decoders.

cs.AR

A Generalized Adjusted Min-Sum Decoder for 5G LDPC Codes: Algorithm and Implementation

5G New Radio (NR) has stringent demands on both performance and complexity for the design of low-density parity-check (LDPC) decoding algorithms and corresponding VLSI implementations. Furthermore, decoders must fully support the wide range of all 5G NR blocklengths and code rates, which is a significant challenge. In this paper, we present a high-performance and low-complexity LDPC decoder, tailor-made to fulfill the 5G requirements. First, to close the gap between belief propagation (BP) decoding and its approximations in hardware, we propose an extension of adjusted min-sum decoding, called generalized adjusted min-sum (GA-MS) decoding. This decoding algorithm flexibly truncates the incoming messages at the check node level and carefully approximates the non-linear functions of BP decoding to balance the error-rate and hardware complexity. Numerical results demonstrate that the proposed fixed-point GAMS has only a minor gap of 0.1 dB compared to floating-point BP under various scenarios of 5G standard specifications. Secondly, we present a fully reconfigurable 5G NR LDPC decoder implementation based on GA-MS decoding. Given that memory occupies a substantial portion of the decoder area, we adopt multiple data compression and approximation techniques to reduce 42.2% of the memory overhead. The corresponding 28nm FD-SOI ASIC decoder has a core area of 1.823 mm2 and operates at 895 MHz. It is compatible with all 5G NR LDPC codes and achieves a peak throughput of 24.42 Gbps and a maximum area efficiency of 13.40 Gbps/mm2 at 4 decoding iterations.

cs.IT

PAC Codes: Sequential Decoding vs List Decoding

In the Shannon lecture at the 2019 International Symposium on Information Theory (ISIT), Arıkan proposed to employ a one-to-one convolutional transform as a pre-coding step before the polar transform. The resulting codes of this concatenation are called polarization-adjusted convolutional (PAC) codes. In this scheme, a pair of polar mapper and demapper as pre- and postprocessing devices are deployed around a memoryless channel, which provides polarized information to an outer decoder leading to improved error correction performance of the outer code. In this paper, the list decoding and sequential decoding (including Fano decoding and stack decoding) are first adapted for use to decode PAC codes. Then, to reduce the complexity of sequential decoding of PAC/polar codes, we propose (i) an adaptive heuristic metric, (ii) tree search constraints for backtracking to avoid exploration of unlikely sub-paths, and (iii) tree search strategies consistent with the pattern of error occurrence in polar codes. These contribute to the reduction of the average decoding time complexity from 50% to 80%, trading with 0.05 to 0.3 dB degradation in error correction performance within FER=10^-3 range, respectively, relative to not applying the corresponding search strategies. Additionally, as an important ingredient in Fano decoding of PAC/polar codes, an efficient computation method for the intermediate LLRs and partial sums is provided. This method is effective in backtracking and avoids storing the intermediate information or restarting the decoding process. Eventually, all three decoding algorithms are compared in terms of performance, complexity, and resource requirements.

cs.IT

Spreading Factor assisted LoRa Localization with Deep Reinforcement Learning

Most of the developed localization solutions rely on RSSI fingerprinting. However, in the LoRa networks, due to the spreading factor (SF) in the network setting, traditional fingerprinting may lack representativeness of the radio map, leading to inaccurate position estimates. As such, in this work, we propose a novel LoRa RSSI fingerprinting approach that takes into account the SF. The performance evaluation shows the prominence of our proposed approach since we achieved an improvement in localization accuracy by up to 6.67% compared to the state-of-the-art methods. The evaluation has been done using a fully connected deep neural network (DNN) set as the baseline. To further improve the localization accuracy, we propose a deep reinforcement learning model that captures the ever-growing complexity of LoRa networks and copes with their scalability. The obtained results show an improvement of 48.10% in the localization accuracy compared to the baseline DNN model.

eess.SP

High-Throughput Flexible Belief Propagation List Decoder for Polar Codes

Owing to its high parallelism, belief propagation (BP) decoding is highly amenable to high-throughput implementations and thus represents a promising solution for meeting the ultra-high peak data rate of future communication systems. However, for polar codes, the error-correcting performance of BP decoding is far inferior to that of the widely used CRC-aided successive cancellation list (SCL) decoding algorithm. To close the performance gap to SCL, BP list (BPL) decoding expands the exploration of candidate codewords through multiple permuted factor graphs (PFGs). From an implementation perspective, designing a unified and flexible hardware architecture for BPL decoding that supports various PFGs and code configurations presents a big challenge. In this paper, we propose the first hardware implementation of a BPL decoder for polar codes and overcome the implementation challenge by applying a hardware-friendly algorithm that generates flexible permutations on-the-fly. First, we derive the graph selection gain and provide a sequential generation (SG) algorithm to obtain a near-optimal PFG set. We further prove that any permutation can be decomposed into a combination of multiple fixed routings, and we design a low-complexity permutation network to satisfy the decoding schedule. Our BPL decoder not only has a low decoding latency by executing the decoding and permutation generation in parallel, but also supports an arbitrary list size without any area overhead. Experimental results show that, for length-1024 polar codes with a code rate of one-half, our BPL decoder with 32 PFGs has a similar error-correcting performance to SCL with a list size of 4 and achieves a throughput of 25.63 Gbps and an area efficiency of 29.46 Gbps/mm$^{2}$ at SNR=4.0dB, which is 1.82$\times$ and 4.33$\times$ faster than the state-of-the-art BP flip and SCL decoders,~respectively

cs.IT