SearcharxivSearch

arXiv subjects

Hamed Kebriaei

Publications and source records attributed to Hamed Kebriaei.

At least 19 recordsLinked to original sources

A Systematic Approach to Mechanism Design with Stochastic Dynamic Stability

We consider a resource allocation problem with strategic agents that have private stochastic satisfaction functions and local constraints. To achieve a global optimal solution, we propose an incentive mechanism that induces a game among the agents. For the payment function of the mechanism, we construct a family of quadratic functions using the linear matrix inequality (LMI) approach that implements the social welfare maximizing outcome on the unique Nash equilibrium (NE) of the induced game while ensuring budget balance and individual rationality. Moreover, we propose a decentralized variable sample-size proximal best-response (VS-PBR) algorithm with Krasnoselskij iteration where only aggregate information is available to the agents. The algorithm is dynamically stable, as it is proven to converge in the mean-square sense to the NE of the game. The efficiency of the mechanism is then investigated on the Sioux Falls City transportation network, where electric vehicle (EV) users jointly select their destination and route.

eess.SY

OptiGradTrust: Byzantine-Robust Federated Learning with Multi-Feature Gradient Analysis and Reinforcement Learning-Based Trust Weighting

Federated Learning (FL) enables collaborative model training across distributed medical institutions while preserving patient privacy, but remains vulnerable to Byzantine attacks and statistical heterogeneity. We present OptiGradTrust, a comprehensive defense framework that evaluates gradient updates through a novel six-dimensional fingerprint including VAE reconstruction error, cosine similarity metrics, $L_2$ norm, sign-consistency ratio, and Monte Carlo Shapley value, which drive a hybrid RL-attention module for adaptive trust scoring. To address convergence challenges under data heterogeneity, we develop FedBN-Prox (FedBN-P), combining Federated Batch Normalization with proximal regularization for optimal accuracy-convergence trade-offs. Extensive evaluation across MNIST, CIFAR-10, and Alzheimer's MRI datasets under various Byzantine attack scenarios demonstrates significant improvements over state-of-the-art defenses, achieving up to +1.6 percentage points over FLGuard under non-IID conditions while maintaining robust performance against diverse attack patterns through our adaptive learning approach.

cs.LG

Stochastic Generalized Dynamic Games with Coupled Chance Constraints

This paper investigates stochastic generalized dynamic games with coupling chance constraints, where agents have incomplete information about uncertainties satisfying a concentration of measure property. This problem, in general, is non-convex and NP-hard. To address this, we propose a convex under-approximation by replacing chance constraints with tightened expected-value constraints, yielding a tractable game. We prove the existence of a stochastic generalized Nash equilibrium (SGNE) in this new game and show that its variational SGNE is an $\boldsymbol{\varepsilon}$-SGNE for the original game, with $\boldsymbol{\varepsilon}$ expressed via the approximation errors and Lagrange multipliers. A semi-decentralized, sampling-based algorithm with time-varying step sizes is developed, requiring no prior knowledge of the uncertainty distribution or expectation evaluations. Unlike existing methods, it avoids step-size tuning based on Lipschitz constants or adaptive rules. Under standard assumptions on the pseudo-gradient, the algorithm converges almost surely to an SGNE.

eess.SY

A Distributed Primal-Dual Method for Constrained Multi-agent Reinforcement Learning with General Parameterization

This paper proposes a novel distributed approach for solving a cooperative Constrained Multi-agent Reinforcement Learning (CMARL) problem, where agents seek to minimize a global objective function subject to shared constraints. Unlike existing methods that rely on centralized training or coordination, our approach enables fully decentralized online learning, with each agent maintaining local estimates of both primal and dual variables. Specifically, we develop a distributed primal-dual algorithm based on actor-critic methods, leveraging local information to estimate Lagrangian multipliers. We establish consensus among the Lagrangian multipliers across agents and prove the convergence of our algorithm to an equilibrium point, analyzing the sub-optimality of this equilibrium compared to the exact solution of the unparameterized problem. Furthermore, we introduce a constrained cooperative Cournot game with stochastic dynamics as a test environment to evaluate the algorithm's performance in complex, real-world scenarios.

eess.SY

Towards Opinion Shaping: A Deep Reinforcement Learning Approach in Bot-User Interactions

This paper aims to investigate the impact of interference in social network algorithms via user-bot interactions, focusing on the Stochastic Bounded Confidence Model (SBCM). This paper explores two approaches: positioning bots controlled by agents into the network and targeted advertising under various circumstances, operating with an advertising budget. This study integrates the Deep Deterministic Policy Gradient (DDPG) algorithm and its variants to experiment with different Deep Reinforcement Learning (DRL). Finally, experimental results demonstrate that this approach can result in efficient opinion shaping, indicating its potential in deploying advertising resources on social platforms.

cs.SI

Risk Sensitivity in Markov Games and Multi-Agent Reinforcement Learning: A Systematic Review

Markov games (MGs) and multi-agent reinforcement learning (MARL) are studied to model decision making in multi-agent systems. Traditionally, the objective in MG and MARL has been risk-neutral, i.e., agents are assumed to optimize a performance metric such as expected return, without taking into account subjective or cognitive preferences of themselves or of other agents. However, ignoring such preferences leads to inaccurate models of decision making in many real-world scenarios in finance, operations research, and behavioral economics. Therefore, when these preferences are present, it is necessary to incorporate a suitable measure of risk into the optimization objective of agents, which opens the door to risk-sensitive MG and MARL. In this paper, we systemically review the literature on risk sensitivity in MG and MARL that has been growing in recent years alongside other areas of reinforcement learning and game theory. We define and mathematically describe different risk measures used in MG and MARL and individually for each measure, discuss articles that incorporate it. Finally, we identify recent trends in theoretical and applied works in the field and discuss possible directions of future research.

cs.GT

Risk-Sensitive Multi-Agent Reinforcement Learning in Network Aggregative Markov Games

Classical multi-agent reinforcement learning (MARL) assumes risk neutrality and complete objectivity for agents. However, in settings where agents need to consider or model human economic or social preferences, a notion of risk must be incorporated into the RL optimization problem. This will be of greater importance in MARL where other human or non-human agents are involved, possibly with their own risk-sensitive policies. In this work, we consider risk-sensitive and non-cooperative MARL with cumulative prospect theory (CPT), a non-convex risk measure and a generalization of coherent measures of risk. CPT is capable of explaining loss aversion in humans and their tendency to overestimate/underestimate small/large probabilities. We propose a distributed sampling-based actor-critic (AC) algorithm with CPT risk for network aggregative Markov games (NAMGs), which we call Distributed Nested CPT-AC. Under a set of assumptions, we prove the convergence of the algorithm to a subjective notion of Markov perfect Nash equilibrium in NAMGs. The experimental results show that subjective CPT policies obtained by our algorithm can be different from the risk-neutral ones, and agents with a higher loss aversion are more inclined to socially isolate themselves in an NAMG.

cs.LG

A Consensus-Based Generalized Multi-Population Aggregative Game with Application to Charging Coordination of Electric Vehicles

This paper introduces a consensus-based generalized multi-population aggregative game coordination approach with application to electric vehicles charging under transmission line constraints. The algorithm enables agents to seek an equilibrium solution while considering the limited infrastructure capacities that impose coupling constraints among the users. The Nash-seeking algorithm consists of two interrelated iterations. In the upper layer, population coordinators collaborate for a distributed estimation of the coupling aggregate term in the agents' cost function and the associated Lagrange multiplier of the coupling constraint, transmitting the latest updated values to their population's agents. In the lower layer, each agent updates its best response based on the most recent information received and communicates it back to its population coordinator. For the case when the agents' best response mappings are non-expansive, we prove the algorithm's convergence to the generalized Nash equilibrium point of the game. Simulation results demonstrate the algorithm's effectiveness in achieving equilibrium in the presence of a coupling constraint.

eess.SY

Generalized Stochastic Dynamic Aggregative Game for Demand-Side Management in Microgrids with Shared Battery

In this paper, we focus on modeling and analysis of demand-side management in a microgrid where agents utilize grid energy and a shared battery charged by renewable energy sources. We model the problem as a generalized stochastic dynamic aggregative game with chance constraints that capture the effects of uncertainties in the renewable generation and agents' demands. Computing the solution of the game is a complex task due to probabilistic and coupling constraints among the agents through the state of charge of the shared battery. We investigate the Nash equilibrium of this game under uncertainty considering both the uniqueness of the solution and the effect of uncertainty on the solution. Simulation results demonstrate that the presented stochastic method is superior to deterministic methods.

eess.SY

Learning Robust Model Predictive Control for Voltage Control of Islanded Microgrid

This paper proposes a novel control design for voltage tracking of an islanded AC microgrid in the presence of {nonlinear} loads and parametric uncertainties at the primary level of control. The proposed method is based on the Tube-Based Robust Model Predictive Control (RMPC), an online optimization-based method which can handle the constraints and uncertainties as well. The challenge with this method is the conservativeness imposed by designing the tube based on the worst-case scenario of the uncertainties. This weakness is amended in this paper by employing a combination of a learning-based Gaussian Process (GP) regression and RMPC. The advantage of using GP is that both the mean and variance of the loads are predicted at each iteration based on the real data, and the resulted values of mean and the bound of confidence are utilized to design the tube in RMPC. The theoretical results are also provided to prove the recursive feasibility and stability of the proposed learning based RMPC. Finally, the simulation results are carried out on both single and multiple DG (Distributed Generation) units.

eess.SY

Continuous Reinforcement Learning-based Dynamic Difficulty Adjustment in a Visual Working Memory Game

Dynamic Difficulty Adjustment (DDA) is a viable approach to enhance a player's experience in video games. Recently, Reinforcement Learning (RL) methods have been employed for DDA in non-competitive games; nevertheless, they rely solely on discrete state-action space with a small search space. In this paper, we propose a continuous RL-based DDA methodology for a visual working memory (VWM) game to handle the complex search space for the difficulty of memorization. The proposed RL-based DDA tailors game difficulty based on the player's score and game difficulty in the last trial. We defined a continuous metric for the difficulty of memorization. Then, we consider the task difficulty and the vector of difficulty-score as the RL's action and state, respectively. We evaluated the proposed method through a within-subject experiment involving 52 subjects. The proposed approach was compared with two rule-based difficulty adjustment methods in terms of player's score and game experience measured by a questionnaire. The proposed RL-based approach resulted in a significantly better game experience in terms of competence, tension, and negative and positive affect. Players also achieved higher scores and win rates. Furthermore, the proposed RL-based DDA led to a significantly less decline in the score in a 20-trial session.

cs.HC

Incentive Mechanism in the Sponsored Content Market with Network Effect

We propose an incentive mechanism for the sponsored content provider market in which the communication of users can be represented by a graph and the private information of the users is assumed to have a continuous distribution function. The content provider stipulates incentive rewards to encourage users to reveal their private information truthfully and increase their content demand, which leads to an increase in advertising revenue. We prove that all users gain a non-negative utility and disclose their private information truthfully. Moreover, we study the effectiveness and scalability of the proposed mechanism in a case study with different network structures.

cs.GT

A Tractable Truthful Profit Maximization Mechanism Design with Autonomous Agents

Task allocation is a crucial process in modern systems, but it is often challenged by incomplete information about the utilities of participating agents. In this paper, we propose a new profit maximization mechanism for the task allocation problem, where the task publisher seeks an optimal incentive function to maximize its own profit and simultaneously ensure the truthful announcing of the agent's private information (type) and its participation in the task, while an autonomous agent aims at maximizing its own utility function by deciding on its participation level and announced type. Our mechanism stands out from the classical contract theory-based truthful mechanisms as it empowers agents to make their own decisions about their level of involvement, making it more practical for many real-world task allocation scenarios. It has been proven that by considering a linear form of incentive function consisting of two decision functions for the task publisher the mechanism's goals are met. The proposed truthful mechanism is initially modeled as a non-convex functional optimization with the double continuum of constraints, nevertheless, we demonstrate that by deriving an equivalent form of the incentive constraints, it can be reformulated as a tractable convex optimal control problem. Further, we propose a numerical algorithm to obtain the solution.

econ.TH

Generalized Wardrop Equilibrium for Charging Station Selection and Route Choice of Electric Vehicles in Joint Power Distribution and Transportation Networks

This paper presents the equilibrium analysis of a game composed of heterogeneous electric vehicles (EVs) and a power distribution system operator (DSO) as the players, and charging station operators (CSOs) and a transportation network operator (TNO) as coordinators. Each EV tries to pick a charging station as its destination and a route to get there at the same time. However, the traffic and electrical load congestion on the roads and charging stations lead to the interdependencies between the optimal decisions of EVs. CSOs and the TNO need to apply some tolling to control such congestion. On the other hand, the pricing at charging stations depends on real-time distributional locational marginal pricing, which is determined by the DSO after solving the optimal power flow over the power distribution network. This paper also takes into account the local and the coupling/infrastructure constraints of EVs, transportation and distribution networks. This problem is modeled as a generalized aggregative game, and then a decentralized learning method is proposed to obtain an equilibrium point of the game, which is known as variational generalized Wardrop equilibrium. The existence of such an equilibrium point and the convergence of the proposed algorithm to it are proven. We undertake numerical studies on the Savannah city model and the IEEE 33-bus distribution network and investigate the impact of various characteristics on demand and prices.

cs.GT

A Sufficient Condition to Guarantee Non-Simultaneous Charging and Discharging of Household Battery Energy Storage

In this letter, we model the day-ahead price-based demand response of a residential household with battery energy storage and other controllable loads, as a convex optimization problem. Further using duality theory and Karush-Kuhn-Tucker optimality conditions, we derive a sufficient criterion which guarantees non-simultaneous charging and discharging of the battery energy storage, without explicitly modelling it as a constraint

eess.SY

Local Stackelberg equilibrium seeking in generalized aggregative games

We propose a two-layer, semi-decentralized algorithm to compute a local solution to the Stackelberg equilibrium problem in aggregative games with coupling constraints. Specifically, we focus on a single-leader, multiple-follower problem, and after equivalently recasting the Stackelberg game as a mathematical program with complementarity constraints (MPCC), we iteratively convexify a regularized version of the MPCC as inner problem, whose solution generates a sequence of feasible descent directions for the original MPCC. Thus, by pursuing a descent direction at every outer iteration, we establish convergence to a local Stackelberg equilibrium. Finally, the proposed algorithm is tested on a numerical case study involving a hierarchical instance of the charging coordination of Plug-in Electric Vehicles (PEVs).

math.OC

Leader-Follower Network Aggregative Game with Stochastic Agents' Communication and Activeness

This technical note presents a leader-follower scheme for network aggregative games. The followers and leader are selfish cost minimizing agents. The cost function of each follower is affected by strategy of leader and aggregated strategies of its neighbors through a communication graph. The leader infinitely often wakes up and receives the aggregated strategy of the followers, updates its decision value and broadcasts it to all the followers. Then, the followers apply the updated strategy of the leader into their cost functions. The establishment of information exchange between each neighboring pair of followers, and the activeness of each follower to update its decision at each iteration are both considered to be drawn from two arbitrary distributions. Moreover, a distributed algorithm based on subgradient method is proposed for updating the strategies of leader and followers. The convergence of the proposed algorithm to the unique generalized Nash equilibrium point of the game is proven in both almost sure and mean square senses.

eess.SY

An Optimal Linear Dynamic Detection Method for Replay Attack in Cyber-Physical Systems

The problem of detecting replay attack to the linear stochastic system with Kalman filer state estimator and LQG controller is addressed. To this end, a dynamic attack detector method is proposed which is coupled with the dynamics of the system. While preserving stability of the main system, conditions on parameters of the attack detector dynamics are obtained such that the attack can be revealed by destabilization of a residual trajectory which is the difference between the estimated and measured output of the system. Using this method, system operator can adjust the detection rate based on the proposed scheme by changing the design parameters. Nevertheless, since the exogenous attack detector signal affects the performance of the closed loop control system, we propose an optimization problem to determine such a detector with minimum loss effect. In the simulation results, the proposed dynamical attack detector approach is compared with the well-known additive white noise watermarking method and the results confirm the superiority of the new scheme.

eess.SY