Searcharxiv⌕ Search

arXiv subjects

M. Cenk Gursoy

Publications and source records attributed to M. Cenk Gursoy.

At least 19 recordsLinked to original sources

Constrained Deep Reinforcement Learning for Cognitive Radar Resource Management

In this paper, multi-target tracking and scanning are considered in a radar system operating in the track-while-scan mode. Specifically, time allocation for radar scanning and tracking of multiple maneuvering targets under a time budget constraint is addressed, aiming to jointly optimize the performance of both tracking and scanning in a cognitive radar. We first present the details of the model for tracking and scanning and formulate the time management task as a constrained optimization problem. Subsequently, we design a \gls{cdrl} framework to find the time allocation strategy for the problem. In the proposed \gls{cdrl} framework, the parameters of the neural networks and the dual variable are learned simultaneously. The deep deterministic policy gradient (DDPG) algorithm is introduced to tackle continuous action space and its performance is compared with deep Q-learning, heuristic approaches, and an optimization-based approach. Numerical results show that the radar with the proposed \gls{cdrl} framework can autonomously allocate more time to the tracking task that requires greater attention while providing time for scanning and also constraining the total time budget below the predefined threshold.

eess.SY↗

Optimizing Finite Structures to Suppress the Photonic Density of States

We propose a topology-optimization framework for optimizing finite structures of arbitrary shape by combining density-based methods with level-set approaches. We first optimize regular polygonal structures to suppress the photonic density of states and find that the best performing polygon is consistent with a tiling of space with hexagonal unit cells. We next show that introducing cavities into hexagonal structures further suppresses the photonic density of states, particularly when the cavity is also hexagonal. Such a result would find application in the design of fiber-optic cables. We then describe an approach for optimizing arbitrary x-simple or y-simple designs that can recover finite supercells of a hexagonal unit cell. Our approach can therefore discover the symmetry of photonic-crystal primitive unit cells that significantly suppress the photonic density of states for a given set of material parameters within a single optimization.

physics.optics↗

Structure-Adaptive Topology Optimization Framework for Photonic Band Gaps with TE-Polarized Sources

Leveraging our structure-adaptive topology optimization framework based on the integration of the photonic density of states over a frequency window for the TM polarization of light [see A. Bahulikar et al., arXiv:2411.09165 (2025)], we show that the $Γ$-point and full Brillouin zone integration schemes can also recover two-dimensional photonic crystals for TE polarization. For the $Γ$-point formalism, we employ the scalar magnetic field formulation of the electromagnetic wave equation with independent sources polarized in the x and y directions. For the full Brillouin zone formalism, we employ the vector electric field formulation of the electromagnetic wave equation, again with independent sources polarized in the x and y directions. This work can simultaneously treat frequency-dependent optical response, allow for targeted optimization for a given frequency and reciprocal lattice vector pair, and inherently encourage binarized designs.

physics.optics↗

Adaptive Resource Management in Cognitive Radar via Deep Deterministic Policy Gradient

In this paper, scanning for target detection, and multi-target tracking in a cognitive radar system are considered, and adaptive radar resource management is investigated. In particular, time management for radar scanning and tracking of multiple maneuvering targets subject to budget constraints is studied with the goal to jointly maximize the tracking and scanning performances of a cognitive radar. We tackle the constrained optimization problem of allocating the dwell time to track individual targets by employing a deep deterministic policy gradient (DDPG) based reinforcement learning approach. We propose a constrained deep reinforcement learning (CDRL) algorithm that updates the DDPG neural networks and dual variables simultaneously. Numerical results show that the radar can autonomously allocate time appropriately so as to maximize the reward function without exceeding the time constraint.

eess.SP↗

Explainable AI for Radar Resource Management: Modified LIME in Deep Reinforcement Learning

Deep reinforcement learning has been extensively studied in decision-making processes and has demonstrated superior performance over conventional approaches in various fields, including radar resource management (RRM). However, a notable limitation of neural networks is their ``black box" nature and recent research work has increasingly focused on explainable AI (XAI) techniques to describe the rationale behind neural network decisions. One promising XAI method is local interpretable model-agnostic explanations (LIME). However, the sampling process in LIME ignores the correlations between features. In this paper, we propose a modified LIME approach that integrates deep learning (DL) into the sampling process, which we refer to as DL-LIME. We employ DL-LIME within deep reinforcement learning for radar resource management. Numerical results show that DL-LIME outperforms conventional LIME in terms of both fidelity and task performance, demonstrating superior performance with both metrics. DL-LIME also provides insights on which factors are more important in decision making for radar resource management.

cs.LG↗

Learning-Based Resource Management in Integrated Sensing and Communication Systems

In this paper, we tackle the task of adaptive time allocation in integrated sensing and communication systems equipped with radar and communication units. The dual-functional radar-communication system's task involves allocating dwell times for tracking multiple targets and utilizing the remaining time for data transmission towards estimated target locations. We introduce a novel constrained deep reinforcement learning (CDRL) approach, designed to optimize resource allocation between tracking and communication under time budget constraints, thereby enhancing target communication quality. Our numerical results demonstrate the efficiency of our proposed CDRL framework, confirming its ability to maximize communication quality in highly dynamic environments while adhering to time constraints.

cs.LG↗

Multi-Objective Reinforcement Learning for Cognitive Radar Resource Management

The time allocation problem in multi-function cognitive radar systems focuses on the trade-off between scanning for newly emerging targets and tracking the previously detected targets. We formulate this as a multi-objective optimization problem and employ deep reinforcement learning to find Pareto-optimal solutions and compare deep deterministic policy gradient (DDPG) and soft actor-critic (SAC) algorithms. Our results demonstrate the effectiveness of both algorithms in adapting to various scenarios, with SAC showing improved stability and sample efficiency compared to DDPG. We further employ the NSGA-II algorithm to estimate an upper bound on the Pareto front of the considered problem. This work contributes to the development of more efficient and adaptive cognitive radar systems capable of balancing multiple competing objectives in dynamic environments.

cs.LG↗

QMGeo: Differentially Private Federated Learning via Stochastic Quantization with Mixed Truncated Geometric Distribution

Federated learning (FL) is a framework which allows multiple users to jointly train a global machine learning (ML) model by transmitting only model updates under the coordination of a parameter server, while being able to keep their datasets local. One key motivation of such distributed frameworks is to provide privacy guarantees to the users. However, preserving the users' datasets locally is shown to be not sufficient for privacy. Several differential privacy (DP) mechanisms have been proposed to provide provable privacy guarantees by introducing randomness into the framework, and majority of these mechanisms rely on injecting additive noise. FL frameworks also face the challenge of communication efficiency, especially as machine learning models grow in complexity and size. Quantization is a commonly utilized method, reducing the communication cost by transmitting compressed representation of the underlying information. Although there have been several studies on DP and quantization in FL, the potential contribution of the quantization method alone in providing privacy guarantees has not been extensively analyzed yet. We in this paper present a novel stochastic quantization method, utilizing a mixed geometric distribution to introduce the randomness needed to provide DP, without any additive noise. We provide convergence analysis for our framework and empirically study its performance.

cs.LG↗

Feature-based Federated Transfer Learning: Communication Efficiency, Robustness and Privacy

In this paper, we propose feature-based federated transfer learning as a novel approach to improve communication efficiency by reducing the uplink payload by multiple orders of magnitude compared to that of existing approaches in federated learning and federated transfer learning. Specifically, in the proposed feature-based federated learning, we design the extracted features and outputs to be uploaded instead of parameter updates. For this distributed learning model, we determine the required payload and provide comparisons with the existing schemes. Subsequently, we analyze the robustness of feature-based federated transfer learning against packet loss, data insufficiency, and quantization. Finally, we address privacy considerations by defining and analyzing label privacy leakage and feature privacy leakage, and investigating mitigating approaches. For all aforementioned analyses, we evaluate the performance of the proposed learning scheme via experiments on an image classification task and a natural language processing task to demonstrate its effectiveness.

cs.LG↗

Robust and Decentralized Reinforcement Learning for UAV Path Planning in IoT Networks

Unmanned aerial vehicle (UAV)-based networks and Internet of Things (IoT) are being considered as integral components of current and next-generation wireless networks. In particular, UAVs can provide IoT devices with seamless connectivity and high coverage and this can be accomplished with effective UAV path planning. In this article, we study robust and decentralized UAV path planning for data collection in IoT networks in the presence of other noncooperative UAVs and adversarial jamming attacks. We address three different practical scenarios, including single UAV path planning, UAV swarm path planning, and single UAV path planning in the presence of an intelligent mobile UAV jammer. We advocate a reinforcement learning framework for UAV path planning in these three scenarios under practical constraints. The simulation results demonstrate that with learning-based path planning, the UAVs can complete their missions with high success rates and data collection rates. In addition, the UAVs can adapt and execute different trajectories as a defensive measure against the intelligent jammer.

eess.SY↗

Learning-Based UAV Path Planning for Data Collection with Integrated Collision Avoidance

Unmanned aerial vehicles (UAVs) are expected to be an integral part of wireless networks, and determining collision-free trajectory in multi-UAV non-cooperative scenarios while collecting data from distributed Internet of Things (IoT) nodes is a challenging task. In this paper, we consider a path planning optimization problem to maximize the collected data from multiple IoT nodes under realistic constraints. The considered multi-UAV non-cooperative scenarios involve random number of other UAVs in addition to the typical UAV, and UAVs do not communicate or share information among each other. We translate the problem into a Markov decision process (MDP) with parameterized states, permissible actions, and detailed reward functions. Dueling double deep Q-network (D3QN) is proposed to learn the decision making policy for the typical UAV, without any prior knowledge of the environment (e.g., channel propagation model and locations of the obstacles) and other UAVs (e.g., their missions, movements, and policies). The proposed algorithm can adapt to various missions in various scenarios, e.g., different numbers and positions of IoT nodes, different amount of data to be collected, and different numbers and positions of other UAVs. Numerical results demonstrate that real-time navigation can be efficiently performed with high success rate, high data collection rate, and low collision rate.

eess.SP↗

Resilient Path Planning for UAVs in Data Collection under Adversarial Attacks

In this paper, we investigate jamming-resilient UAV path planning strategies for data collection in Internet of Things (IoT) networks, in which the typical UAV can learn the optimal trajectory to elude such jamming attacks. Specifically, the typical UAV is required to collect data from multiple distributed IoT nodes under collision avoidance, mission completion deadline, and kinematic constraints in the presence of jamming attacks. We first design a fixed ground jammer with continuous jamming attack and periodical jamming attack strategies to jam the link between the typical UAV and IoT nodes. Defensive strategies involving a reinforcement learning (RL) based virtual jammer and the adoption of higher SINR thresholds are proposed to counteract against such attacks. Secondly, we design an intelligent UAV jammer, which utilizes the RL algorithm to choose actions based on its observation. Then, an intelligent UAV anti-jamming strategy is constructed to deal with such attacks, and the optimal trajectory of the typical UAV is obtained via dueling double deep Q-network (D3QN). Simulation results show that both non-intelligent and intelligent jamming attacks have significant influence on the UAV's performance, and the proposed defense strategies can recover the performance close to that in no-jammer scenarios.

cs.NI↗

Anomaly Detection via Learning-Based Sequential Controlled Sensing

In this paper, we address the problem of detecting anomalies among a given set of binary processes via learning-based controlled sensing. Each process is parameterized by a binary random variable indicating whether the process is anomalous. To identify the anomalies, the decision-making agent is allowed to observe a subset of the processes at each time instant. Also, probing each process has an associated cost. Our objective is to design a sequential selection policy that dynamically determines which processes to observe at each time with the goal to minimize the delay in making the decision and the total sensing cost. We cast this problem as a sequential hypothesis testing problem within the framework of Markov decision processes. This formulation utilizes both a Bayesian log-likelihood ratio-based reward and an entropy-based reward. The problem is then solved using two approaches: 1) a deep reinforcement learning-based approach where we design both deep Q-learning and policy gradient actor-critic algorithms; and 2) a deep active inference-based approach. Using numerical experiments, we demonstrate the efficacy of our algorithms and show that our algorithms adapt to any unknown statistical dependence pattern of the processes.

cs.LG↗

Robust Network Slicing: Multi-Agent Policies, Adversarial Attacks, and Defensive Strategies

In this paper, we present a multi-agent deep reinforcement learning (deep RL) framework for network slicing in a dynamic environment with multiple base stations and multiple users. In particular, we propose a novel deep RL framework with multiple actors and centralized critic (MACC) in which actors are implemented as pointer networks to fit the varying dimension of input. We evaluate the performance of the proposed deep RL algorithm via simulations to demonstrate its effectiveness. Subsequently, we develop a deep RL based jammer with limited prior information and limited power budget. The goal of the jammer is to minimize the transmission rates achieved with network slicing and thus degrade the network slicing agents' performance. We design a jammer with both listening and jamming phases and address jamming location optimization as well as jamming channel optimization via deep RL. We evaluate the jammer at the optimized location, generating interference attacks in the optimized set of channels by switching between the jamming phase and listening phase. We show that the proposed jammer can significantly reduce the victims' performance without direct feedback or prior knowledge on the network slicing policies. Finally, we devise a Nash-equilibrium-supervised policy ensemble mixed strategy profile for network slicing (as a defensive measure) and jamming. We evaluate the performance of the proposed policy ensemble algorithm by applying on the network slicing agents and the jammer agent in simulations to show its effectiveness.

cs.LG↗

Maximum Knowledge Orthogonality Reconstruction with Gradients in Federated Learning

Federated learning (FL) aims at keeping client data local to preserve privacy. Instead of gathering the data itself, the server only collects aggregated gradient updates from clients. Following the popularity of FL, there has been considerable amount of work, revealing the vulnerability of FL approaches by reconstructing the input data from gradient updates. Yet, most existing works assume an FL setting with unrealistically small batch size, and have poor image quality when the batch size is large. Other works modify the neural network architectures or parameters to the point of being suspicious, and thus, can be detected by clients. Moreover, most of them can only reconstruct one sample input from a large batch. To address these limitations, we propose a novel and completely analytical approach, referred to as the maximum knowledge orthogonality reconstruction (MKOR), to reconstruct clients' input data. Our proposed method reconstructs a mathematically proven high quality image from large batches. MKOR only requires the server to send secretly modified parameters to clients and can efficiently and inconspicuously reconstruct the input images from clients' gradient updates. We evaluate MKOR's performance on the MNIST, CIFAR-100, and ImageNet dataset and compare it with the state-of-the-art works. The results show that MKOR outperforms the existing approaches, and draws attention to a pressing need for further research on the privacy protection of FL so that comprehensive defense approaches can be developed.

cs.LG↗

Temporal Detection of Anomalies via Actor-Critic Based Controlled Sensing

We address the problem of monitoring a set of binary stochastic processes and generating an alert when the number of anomalies among them exceeds a threshold. For this, the decision-maker selects and probes a subset of the processes to obtain noisy estimates of their states (normal or anomalous). Based on the received observations, the decisionmaker first determines whether to declare that the number of anomalies has exceeded the threshold or to continue taking observations. When the decision is to continue, it then decides whether to collect observations at the next time instant or defer it to a later time. If it chooses to collect observations, it further determines the subset of processes to be probed. To devise this three-step sequential decision-making process, we use a Bayesian formulation wherein we learn the posterior probability on the states of the processes. Using the posterior probability, we construct a Markov decision process and solve it using deep actor-critic reinforcement learning. Via numerical experiments, we demonstrate the superior performance of our algorithm compared to the traditional model-based algorithms.

cs.LG↗

Low-Latency Hybrid NOMA-TDMA: QoS-Driven Design Framework

Enabling ultra-reliable and low-latency communication services while providing massive connectivity is one of the major goals to be accomplished in future wireless communication networks. In this paper, we investigate the performance of a hybrid multi-access scheme in the finite blocklength (FBL) regime that combines the advantages of both non-orthogonal multiple access (NOMA) and time-division multiple access (TDMA) schemes. Two latency-sensitive application scenarios are studied, distinguished by whether the queuing behaviour has an influence on the transmission performance or not. In particular, for the latency-critical case with one-shot transmission, we aim at a certain physical-layer quality-of-service (QoS) performance, namely the optimization of the reliability. And for the case in which queuing behaviour plays a role, we focus on the link-layer QoS performance and provide a design that maximizes the effective capacity. For both designs, we leverage the characterizations in the FBL regime to provide the optimal framework by jointly allocating the blocklength and transmit power of each user. In particular, for the reliability-oriented design, the original problem is decomposed and the joint convexity of sub-problems is shown via a variable substitution method. For the effective-capacity-oriented design, we exploit the method of Lagrange multipliers to formulate a solvable dual problem with strong duality to the original problem. Via simulations, we validate our analytical results of convexity/concavity and show the advantage of our proposed approaches compared to other existing schemes.

cs.IT↗

Joint Convexity of Error Probability in Blocklength and Transmit Power in the Finite Blocklength Regime

To support ultra-reliable and low-latency services for mission-critical applications, transmissions are usually carried via short blocklength codes, i.e., in the so-called finite blocklength (FBL) regime. Different from the infinite blocklength regime where transmissions are assumed to be arbitrarily reliable at the Shannon's capacity, the reliability and capacity performances of an FBL transmission are impacted by the coding blocklength. The relationship among reliability, coding rate, blocklength and channel quality has recently been characterized in the literature, considering the FBL performance model. In this paper, we follow this model, and prove the joint convexity of the FBL error probability with respect to blocklength and transmit power within a region of interest, as a key enabler for designing systems to achieve globally optimal performance levels. Moreover, we apply the joint convexity to general use cases and efficiently solve the joint optimization problem in the setting with multiple users. We also extend the applicability of the proposed approach by proving that the joint convexity still holds in fading channels, as well as in relaying networks. Via simulations, we validate our analytical results and demonstrate the advantage of leveraging the joint convexity compared to other commonly-applied approaches.

cs.IT↗