Searcharxiv⌕ Search

arXiv subjects

Mahmoud Al-Qutayri

Publications and source records attributed to Mahmoud Al-Qutayri.

At least 19 recordsLinked to original sources

From Connectivity to Multi-Orbit Intelligence: Space-Based Data Center Architectures for 6G and Beyond

Direct handset-to-satellite (DHTS) communication is emerging as a core capability of 6G non-terrestrial networks, enabling standard devices to directly access low Earth orbit (LEO) satellites. While LEO provides the physical access layer for DHTS, large-scale device connectivity introduces challenges in mobility management, interference control, spectrum efficiency, and constellation-wide coordination. Relay-only LEO architectures are insufficient to manage massive handset access under dynamic traffic and energy constraints. This article introduces a hierarchical architecture in which direct handset-to-LEO access is supported by multi-orbit space-based data centers (SBDCs) spanning LEO, medium Earth orbit (MEO), and geostationary Earth orbit (GEO). In this framework, LEO satellites handle radio access and real-time inference, while higher orbital layers provide regional aggregation, global orchestration, and compute-aware routing. By embedding distributed in-orbit computing, energy-aware scheduling, and AI-driven hierarchical control, the constellation evolves from a passive relay network into an intelligent multi-layer system capable of supporting large-scale DHTS services. We discuss key enabling technologies, envisioned multi-orbit integrated Earth-space compute architecture, and open research challenges in integrating multi-orbit computing, highlighting pathways toward scalable and resilient 6G DHTS networks.

cs.ET↗

Hybrid Dynamic Pruning: A Pathway to Efficient Transformer Inference

In the world of deep learning, Transformer models have become very significant, leading to improvements in many areas from understanding language to recognizing images, covering a wide range of applications. Despite their success, the deployment of these models in real-time applications, particularly on edge devices, poses significant challenges due to their quadratic computational intensity and memory demands. To overcome these challenges we introduce a novel Hybrid Dynamic Pruning (HDP), an efficient algorithm-architecture co-design approach that accelerates transformers using head sparsity, block sparsity and approximation opportunities to reduce computations in attention and reduce memory access. With the observation of the huge redundancy in attention scores and attention heads, we propose a novel integer-based row-balanced block pruning to prune unimportant blocks in the attention matrix at run time, also propose integer-based head pruning to detect and prune unimportant heads at an early stage at run time. Also we propose an approximation method that reduces attention computations. To efficiently support these methods with lower latency and power efficiency, we propose a HDP co-processor architecture.

cs.LG↗

Area and Power Efficient FFT/IFFT Processor for FALCON Post-Quantum Cryptography

Quantum computing is an emerging technology on the verge of reshaping industries, while simultaneously challenging existing cryptographic algorithms. FALCON, a recent standard quantum-resistant digital signature, presents a challenging hardware implementation due to its extensive non-integer polynomial operations, necessitating FFT over the ring $\mathbb{Q}[x]/(x^n+1)$. This paper introduces an ultra-low power and compact processor tailored for FFT/IFFT operations over the ring, specifically optimized for FALCON applications on resource-constrained edge devices. The proposed processor incorporates various optimization techniques, including twiddle factor compression and conflict-free scheduling. In an ASIC implementation using a 22 nm GF process, the proposed processor demonstrates an area occupancy of 0.15 mm$^2$ and a power consumption of 12.6 mW at an operating frequency of 167 MHz. Since a hardware implementation of FFT/IFFT over the ring is currently non-existent, the execution time achieved by this processor is compared to the software implementation of FFT/IFFT of FALCON on a Raspberry Pi 4 with Cortex-A72, where the proposed processor achieves a speedup of up to 2.3$\times$. Furthermore, in comparison to dedicated state-of-the-art hardware accelerators for classic FFT, this processor occupies 42\% less area and consumes 83\% less power, on average. This suggests that the proposed hardware design offers a promising solution for implementing FALCON on resource-constrained devices.

eess.SP↗

Number Systems for Deep Neural Network Architectures: A Survey

Deep neural networks (DNNs) have become an enabling component for a myriad of artificial intelligence applications. DNNs have shown sometimes superior performance, even compared to humans, in cases such as self-driving, health applications, etc. Because of their computational complexity, deploying DNNs in resource-constrained devices still faces many challenges related to computing complexity, energy efficiency, latency, and cost. To this end, several research directions are being pursued by both academia and industry to accelerate and efficiently implement DNNs. One important direction is determining the appropriate data representation for the massive amount of data involved in DNN processing. Using conventional number systems has been found to be sub-optimal for DNNs. Alternatively, a great body of research focuses on exploring suitable number systems. This article aims to provide a comprehensive survey and discussion about alternative number systems for more efficient representations of DNN data. Various number systems (conventional/unconventional) exploited for DNNs are discussed. The impact of these number systems on the performance and hardware design of DNNs is considered. In addition, this paper highlights the challenges associated with each number system and various solutions that are proposed for addressing them. The reader will be able to understand the importance of an efficient number system for DNN, learn about the widely used number systems for DNN, understand the trade-offs between various number systems, and consider various design aspects that affect the impact of number systems on DNN performance. In addition, the recent trends and related research opportunities will be highlighted

cs.NE↗

An Effective Spatial Modulation Based Scheme for Indoor VLC Systems

We propose an enhanced spatial modulation (SM)-based scheme for indoor visible light communication systems. This scheme enhances the achievable throughput of conventional SM schemes by transmitting higher order complex modulation symbol, which is decomposed into three different parts. These parts carry the amplitude, phase, and quadrant components of the complex symbol, which are then represented by unipolar pulse amplitude modulation (PAM) symbols. Superposition coding is exploited to allocate a fraction of the total power to each part before they are all multiplexed and transmitted simultaneously, exploiting the entire available bandwidth. At the receiver, a two-step decoding process is proposed to decode the active light emitting diode index before the complex symbol is retrieved. It is shown that at higher spectral efficiency values, the proposed modulation scheme outperforms conventional SM schemes with PAM symbols in terms of average symbol error rate (ASER), and hence, enhancing the system throughput. Furthermore, since the performance of the proposed modulation scheme is sensitive to the power allocation factors, we formulated an ASER optimization problem and propose a sub-optimal solution using successive convex programming (SCP). Notably, the proposed algorithm converges after only few iterations, whilst the performance with the optimized power allocation coefficients outperforms both random and fixed power allocation.

eess.SP↗

Space-Time Block Coded Spatial Modulation for Indoor Visible Light Communications

Visible light communication (VLC) has been recognized as a promising technology for handling the continuously increasing quality of service and connectivity requirements in modern wireless communications, particularly in indoor scenarios. In this context, the present work considers the integration of two distinct modulation schemes, namely spatial modulation (SM) with space time block codes (STBCs), aiming at improving the overall VLC system reliability. Based on this and in order to further enhance the achievable transmission data rate, we integrate quasi-orthogonal STBC (QOSTBC) with SM, since relaxing the orthogonality condition of OSTBC ultimately provides a higher coding rate. Then, we generalize the developed results to any number of active light-emitting diodes (LEDs) and any M-ary pulse amplitude modulation size. Furthermore, we derive a tight and tractable upper bound for the corresponding bit error rate (BER) by considering a simple two-step decoding procedure to detect the indices of the transmitting LEDs and then decode the signal domain symbols. Notably, the obtained results demonstrate that QOSTBC with SM enhances the achievable BER compared to SM with repetition coding (RC-SM). Finally, we compare STBC-SM with both multiple active SM (MASM) and RC-SM in terms of the achievable BER and overall data rate, which further justifies the usefulness of the proposed scheme.

cs.IT↗

Towards Federated Learning-Enabled Visible Light Communication in 6G Systems

Visible light communication (VLC) technology was introduced as a key enabler for the next generation of wireless networks, mainly thanks to its simple and low-cost implementation. However, several challenges prohibit the realization of the full potentials of VLC, namely, limited modulation bandwidth, ambient light interference, optical diffuse reflection effects, devices non-linearity, and random receiver orientation. On the contrary, centralized machine learning (ML) techniques have demonstrated a significant potential in handling different challenges relating to wireless communication systems. Specifically, it was shown that ML algorithms exhibit superior capabilities in handling complicated network tasks, such as channel equalization, estimation and modeling, resources allocation, and opportunistic spectrum access control, to name a few. Nevertheless, concerns pertaining to privacy and communication overhead when sharing raw data of the involved clients with a server constitute major bottlenecks in the implementation of centralized ML techniques. This has motivated the emergence of a new distributed ML paradigm, namely federated learning (FL), which can reduce the cost associated with transferring raw data, and preserve privacy by training ML models locally and collaboratively at the clients' side. Hence, it becomes evident that integrating FL into VLC networks can provide ubiquitous and reliable implementation of VLC systems. With this motivation, this is the first in-depth review in the literature on the application of FL in VLC networks. To that end, besides the different architectures and related characteristics of FL, we provide a thorough overview on the main design aspects of FL based VLC systems. Finally, we also highlight some potential future research directions of FL that are envisioned to substantially enhance the performance and robustness of VLC systems.

cs.AI↗

Deep Neural Networks Based Weight Approximation and Computation Reuse for 2-D Image Classification

Deep Neural Networks (DNNs) are computationally and memory intensive, which makes their hardware implementation a challenging task especially for resource constrained devices such as IoT nodes. To address this challenge, this paper introduces a new method to improve DNNs performance by fusing approximate computing with data reuse techniques to be used for image recognition applications. DNNs weights are approximated based on the linear and quadratic approximation methods during the training phase, then, all of the weights are replaced with the linear/quadratic coefficients to execute the inference in a way where different weights could be computed using the same coefficients. This leads to a repetition of the weights across the processing element (PE) array, which in turn enables the reuse of the DNN sub-computations (computational reuse) and leverage the same data (data reuse) to reduce DNNs computations, memory accesses, and improve energy efficiency albeit at the cost of increased training time. Complete analysis for both MNIST and CIFAR 10 datasets is presented for image recognition , where LeNet 5 revealed a reduction in the number of parameters by a factor of 1211.3x with a drop of less than 0.9% in accuracy. When compared to the state of the art Row Stationary (RS) method, the proposed architecture saved 54% of the total number of adders and multipliers needed. Overall, the proposed approach is suitable for IoT edge devices as it reduces the memory size requirement as well as the number of needed memory accesses.

cs.CV↗

UNSAIL: Thwarting Oracle-Less Machine Learning Attacks on Logic Locking

Logic locking aims to protect the intellectual property (IP) of integrated circuit (IC) designs throughout the globalized supply chain. The SAIL attack, based on tailored machine learning (ML) models, circumvents combinational logic locking with high accuracy and is amongst the most potent attacks as it does not require a functional IC acting as an oracle. In this work, we propose UNSAIL, a logic locking technique that inserts key-gate structures with the specific aim to confuse ML models like those used in SAIL. More specifically, UNSAIL serves to prevent attacks seeking to resolve the structural transformations of synthesis-induced obfuscation, which is an essential step for logic locking. Our approach is generic; it can protect any local structure of key-gates against such ML-based attacks in an oracle-less setting. We develop a reference implementation for the SAIL attack and launch it on both traditionally locked and UNSAIL-locked designs. In SAIL, a change-prediction model is used to determine which key-gate structures to restore using a reconstruction model. Our study on benchmarks ranging from the ISCAS-85 and ITC-99 suites to the OpenRISC Reference Platform System-on-Chip (ORPSoC) confirms that UNSAIL degrades the accuracy of the change-prediction model and the reconstruction model by an average of 20.13 and 17 percentage points (pp) respectively. When the aforementioned models are combined, which is the most powerful scenario for SAIL, UNSAIL reduces the attack accuracy of SAIL by an average of 11pp. We further demonstrate that UNSAIL thwarts other oracle-less attacks, i.e., SWEEP and the redundancy attack, indicating the generic nature and strength of our approach. Detailed layout-level evaluations illustrate that UNSAIL incurs minimal area and power overheads of 0.26% and 0.61%, respectively, on the million-gate ORPSoC design.

cs.CR↗

Rate-Splitting Multiple Access: Unifying NOMA and SDMA in MISO VLC Channels

The increased proliferation of connected devices requires a paradigm shift towards the development of innovative technologies for the next generation of wireless systems. One of the key challenges, however, is the spectrum scarcity, owing to the unprecedented broadband penetration rate in recent years. Visible light communications (VLC) has recently emerged as a possible solution to enable high-speed short-range communications. However, VLC systems suffer from several limitations, including the limited modulation bandwidth of light-emitting diodes. Consequently, several multiple access techniques (MA), e.g., space-division multiple access (SDMA) and non-orthogonal multiple access (NOMA), have been considered for VLC networks. Despite their promising multiplexing gains, their performance is somewhat limited. In this article, we first provide an overview of the key MA technologies used in VLC systems. Then, we introduce rate-splitting multiple access (RSMA), which was initially proposed for RF systems and discuss its potentials in VLC systems. Through system modeling and simulations of an RSMA-based two-use scenario, we illustrate the flexibility of RSMA in generalizing NOMA and SDMA, as well as its superiority in terms of weighted sum rate (WSR) in VLC. Finally, we discuss challenges, open issues, and research directions, which will enable the practical realization of RSMA in VLC.

eess.SP↗

Non-Orthogonal Multiple Access for Hybrid VLC-RF Networks with Imperfect Channel State Information

The present contribution proposes a general framework for the energy efficiency analysis of a hybrid visible light communication (VLC) and Radio Frequency (RF) wireless system, in which both VLC and RF subsystems utilize nonorthogonal multiple access (NOMA) technology. The proposed framework is based on realistic communication scenarios as it takes into account the mobility of users, and assumes imperfect channel-state information (CSI). In this context, tractable closed-form expressions are derived for the corresponding average sum rate of NOMA-VLC and its orthogonal frequency division multiple access (OFDMA)-VLC counterparts. It is shown extensively that incurred CSI errors have a considerable impact on the average energy efficiency of both NOMA-VLC and OFDMAVLC systems and hence, they should not be neglected in practical designs and deployments. Interestingly, we further demonstrate that the average energy efficiency of the hybrid NOMA-VLCRF system outperforms NOMA-VLC system under imperfect CSI. Respective computer simulations corroborate the derived analytic results and interesting theoretical and practical insights are provided, which will be useful in the effective design and deployment of conventional VLC and hybrid VLC-RF systems.

cs.IT↗

Optical Rate-Splitting Multiple Access for Visible Light Communications

The proliferation of connected devices and emergence of internet-of-everything represent a major challenge for broadband wireless networks. This requires a paradigm shift towards the development of innovative technologies for next generation wireless systems. One of the key challenges is the scarcity of spectrum, owing to the unprecedented broadband penetration rate in recent years. A promising solution is the proposal of visible light communications (VLC), which explores the unregulated visible light spectrum to enable high-speed communications, in addition to efficient lighting. This solution offers a wider bandwidth that can accommodate ubiquitous broadband connectivity to indoor users and offload data traffic from cellular networks. Although VLC is secure and able to overcome the shortcomings of RF systems, it suffers from several limitations, e.g., limited modulation bandwidth. In this respect, solutions have been proposed recently to overcome this limitation. In particular, most common orthogonal and non-orthogonal multiple access techniques initially proposed for RF systems, e.g., space-division multiple access (SDMA) and NOMA, have been considered in the context of VLC. In spite of their promising gains, the performance of these techniques is somewhat limited. Consequently, in this article a new and generalized multiple access technique, called rate-splitting multiple access (RSMA), is introduced and investigated for the first time in VLC networks. We first provide an overview of the key multiple access technologies used in VLC systems. Then, we propose the first comprehensive approach to the integration of RSMA with VLC systems. In our proposed framework, SINR expressions are derived and used to evaluate the weighted sum rate (WSR) of a two-user scenario. Our results illustrate the flexibility of RSMA in generalizing NOMA and SDMA, and its WSR superiority in the VLC context.

cs.IT↗

Performance Analysis of SWIPT Relaying Systems in the Presence of Impulsive Noise

We develop an analytical framework to characterize the effect of impulsive noise on the performance of relay-assisted simultaneous wireless information and power transfer (SWIPT) systems. We derive novel closed-form expressions for the pairwise error probability (PEP) considering two variants based on the availability of channel state information (CSI), namely, blind relaying and CSI-assisted relaying. We further consider two energy harvesting (EH) techniques, i.e., instantaneous EH (IEH) and average EH (AEH). Capitalizing on the derived analytical results, we present a detailed numerical investigation of the diversity order for the underlying scenarios under the impulsive noise assumption. For the case when two relays and the availability of a direct link, it is demonstrated that the considered SWIPT system with blind AEH-relaying is able to achieve an asymptotic diversity order of less than 3, which is equal to the diversity order achieved by CSI-assisted IEH-relaying. This result suggests that, by employing the blind AEH relaying, the power consumption of the network can be reduced, due to eliminating the need of CSI estimation. This can be achieved without any performance loss. Our results further show that placing the relays close to the source can significantly mitigate the detrimental effects of impulsive noise. Extensive Monte Carlo simulation results are presented to validate the accuracy of the proposed analytical framework.

cs.IT↗

ScanSAT: Unlocking Static and Dynamic Scan Obfuscation

While financially advantageous, outsourcing key steps, such as testing, to potentially untrusted Outsourced Assembly and Test (OSAT) companies may pose a risk of compromising on-chip assets. Obfuscation of scan chains is a technique that hides the actual scan data from the untrusted testers; logic inserted between the scan cells, driven by a secret key, hides the transformation functions that map the scan-in stimulus (scan-out response) and the delivered scan pattern (captured response). While static scan obfuscation utilizes the same secret key, and thus, the same secret transformation functions throughout the lifetime of the chip, dynamic scan obfuscation updates the key periodically. In this paper, we propose ScanSAT: an attack that transforms a scan obfuscated circuit to its logic-locked version and applies the Boolean satisfiability (SAT) based attack, thereby extracting the secret key. We implement our attack, apply on representative scan obfuscation techniques, and show that ScanSAT can break both static and dynamic scan obfuscation schemes with 100% success rate. Moreover, ScanSAT is effective even for large key sizes and in the presence of scan compression.

cs.CR↗

Unified Analysis of SWIPT Relay Networks with Noncoherent Modulation

Simultaneous wireless information and power transfer (SWIPT) relay networks represent a paradigm shift in the development of wireless networks, enabling simultaneous radio frequency (RF) energy harvesting (EH) and information processing. Different from conventional SWIPT relaying schemes, which typically assume the availability of perfect channel state information (CSI), here we consider the application of noncoherent modulation in order to avoid the need of instantaneous CSI estimation/tracking and minimise the energy consumption. We propose a unified and comprehensive analytical framework for the analysis of time switching (TS) and power splitting (PS) receiver architectures with the amplify-and-forward (AF) protocol. In particular, we adopt a moments-based approach to derive novel expressions for the outage probability, achievable throughput, and average symbol error rate (ASER) of the considered SWIPT system. We quantify the impact of several system parameters, involving relay location, energy conversion efficiency, and TS and PS ratio assumptions, imposed on the EH relay terminal. Our results reveal that the throughput performance of the TS protocol is superior to that of the PS protocol at lower receive signal-to-noise (SNR) values, which is in contrast to the point-to-point SWIPT systems. An extensive Monte Carlo simulation study is presented to corroborate the proposed analysis.

cs.IT↗

A Novel Hybrid Fast Switching Adaptive No Delay Tanlock Loop Frequency Synthesizer

This paper presents a new fast switching hybrid frequency synthesizer with wide locking range. The hybrid synthesizer is based on the tanlock loop with no delay block (NDTL) and is capable of integer as well as fractional frequency division. The system maintains the in-lock state following the division process using an efficient adaptation mechanism. The fast switching and acquisition as well as the wide locking range and the robust jitter performance of the new hybrid NDTL synthesizer outperforms conventional time-delay tanlock loop (TDTL) synthesizer by orders of magnitude, making it attractive for synthesis even in Doppler environment. The performance of the hybrid synthesizer was evaluated under various conditions and the results demonstrate that it achieves the desired frequency division.

cs.IT↗

Performance Analysis of Energy Detection over Mixture Gamma based Fading Channels with Diversity Reception

The present paper is devoted to the evaluation of energy detection based spectrum sensing over different multipath fading and shadowing conditions. This is realized by means of a unified and versatile approach that is based on the particularly flexible mixture gamma distribution. To this end, novel analytic expressions are firstly derived for the probability of detection over MG fading channels for the conventional single-channel communication scenario. These expressions are subsequently employed in deriving closed-form expressions for the case of square-law combining and square-law selection diversity methods. The validity of the offered expressions is verified through comparisons with results from respective computer simulations. Furthermore, they are employed in analyzing the performance of energy detection over multipath fading, shadowing and composite fading conditions, which provides useful insighs on the performance and design of future cognitive radio based communication systems.

cs.IT↗

Energy Detection of Unknown Signals over Cascaded Fading Channels

Energy detection is a favorable mechanism in several applications relating to the identification of deterministic unknown signals such as in radar systems and cognitive radio communications. The present work quantifies the detrimental effects of cascaded multipath fading on energy detection and investigates the corresponding performance capability. A novel analytic solution is firstly derived for a generic integral that involves a product of the Meijer $G-$function, the Marcum $Q-$function and arbitrary power terms. This solution is subsequently employed in the derivation of an exact closed-form expression for the average probability of detection of unknown signals over $N$*Rayleigh channels. The offered results are also extended to the case of square-law selection, which is a relatively simple and effective diversity method. It is shown that the detection performance is considerably degraded by the number of cascaded channels and that these effects can be effectively mitigated by a non-substantial increase of diversity branches.

cs.IT↗