Searcharxiv⌕ Search

arXiv subjects

Saverio Bolognani

Publications and source records attributed to Saverio Bolognani.

At least 19 recordsLinked to original sources

Bi-ZOL: Bilevel Zeroth-Order Learning with Nonsmooth Responses

This paper studies lower-level-constrained bilevel optimization in a response-oracle setting, where lower-level model information is unavailable and the induced response mapping is locally Lipschitz but potentially nonsmooth. In this setting, the classical response Jacobian and reduced hypergradient may fail to exist. We propose Bilevel Zeroth-Order Learning (Bi-ZOL), a structure-guided zeroth-order method for finding stationary points of the nonsmooth reduced problem. Instead of estimating the gradient of a fully smoothed reduced hyperobjective, Bi-ZOL separates the bilevel chain-rule structure: it keeps the exact upper-level partial gradients at the queried response and uses zeroth-order sampling only to estimate the response Jacobian. This construction yields an approximate hypergradient that is more directly aligned with the Clarke chain-rule subdifferential. We show that the Bi-ZOL direction admits a partial-smoothing interpretation, quantify its pointwise structural bias, and prove finite-time convergence to a $(δ,ε)$-Bi-ZOL Frank--Wolfe stationary point. The bias is $O(δ)$ for piecewise $C^{1,1}$ responses under local regularity and vanishes for piecewise affine responses on active-cell neighborhoods. Experiments on incentive-based tracking problems show that Bi-ZOL achieves smaller stationarity gaps and lower hyperobjective values than vanilla zeroth-order smoothing under comparable response-oracle budgets.

math.OC↗

Optimal Functional Incentives for Control: The Linear-Quadratic Case with Bilinear Incentives

We study the design of functional incentive mechanisms for dynamical systems, in which a leader designs a fixed incentive function to motivate a self-interested follower to actuate the system beneficially over an extended horizon, without real-time revision of the incentive. This stands in contrast to the adaptive paradigm, in which the incentive is itself a continuously updated control variable. We formalize the problem as a discrete-time bi-level optimal control problem and derive analytical results for the linear-quadratic case with bilinear incentives and a myopic follower. Specifically, we establish a necessary and sufficient stability condition for the induced closed-loop system, derive a closed-form expression for the gradient of the expected leader cost with respect to the incentive parameter matrix, and obtain a fully closed-form cost expression in the scalar setting. Based on the latter, explicit characterizations of the optimal incentive parameter are provided in two asymptotic regimes: the infinite-horizon limit and the limit of high follower cost. For long horizons, the optimal incentive is shown to become independent of the follower's private cost parameter, with direct implications for robust mechanism design under private information.

eess.SY↗

The Limits of "Fairness'' of the Variational Generalized Nash Equilibrium

Generalized Nash equilibrium (GNE) problems are commonly used to model strategic interactions between self-interested agents who are coupled in cost and constraints. Specifically, the variational GNE, a refinement of the GNE, is often selected as the solution concept due to its non-discriminatory treatment of agents by charging a uniform ``shadow price" for shared resources. We study the fairness concept of v-GNEs from a comparability perspective and show that it makes an implicit assumption of unit comparability of agent's cost functions, one of the strongest comparability notions. Further, we introduce a new solution concept, f-GNE in which a fairness metric is chosen a priori which is compatible with the comparability at hand. We introduce an electric vehicle charging game to demonstrate the fragility of v-GNE fairness and compare it to the f-GNE under various fairness metrics.

eess.SY↗

Chess on Ice: Curling Tactical Decision-Making via Backward Induction and Deep Reinforcement Learning

Curling is often referred to as "Chess on Ice", owing to the tactical complexity of its decision-making process. Yet unlike chess, curling remains largely underexplored from a machine learning perspective, with prior work confined mainly to statistical approaches. We propose a reinforcement learning framework capable of quantitatively evaluating and comparing tactical options in curling. The game poses several modeling challenges: continuous state and action spaces, stochastic action outcomes reflecting player skill variability, and state transitions that are highly sensitive to small perturbations in the executed action. To address them, we employ the Deep Deterministic Policy Gradient actor-critic algorithm, adapted to exploit the finite-horizon structure of the game. Our experiments show that effective curling strategies can be acquired in a fully self-supervised manner, without any human-annotated data: on a reduced four-rock variant, the learned agent matches a hand-crafted expert heuristic in a regime where that heuristic is close to optimal, a parity we quantify against the intrinsic hammer advantage of the variant. Beyond the resulting policy, the learned critic provides a dense value estimate over the entire continuous action space, enabling the quantitative comparison of tactical alternatives for applications such as post-game performance analysis and decision support during athlete preparation.

cs.AI↗

A Coalitional Stable and Fair Reward Allocation for Dynamic Virtual Power Plants

This paper establishes crucial cooperation criteria for the operation of Dynamic Virtual Power Plants (DVPPs). We propose a control design and reward allocation mechanism to enable and incentivize Distributed Energy Resources (DERs) to provide dynamic ancillary services (DAS). Our results illustrate how the cooperative aggregation of heterogeneous DERs leverages technical complementarities to outperform standalone DAS provision. The proposed reward allocation fulfills critical game-theoretic criteria, including individual rationality, coalitional stability, incentive compatibility, optimality, fairness and ex-post consistency. The control design and reward allocation are validated using a case study based on the Finnish power grid.

eess.SY↗

Gray-Box Nonlinear Feedback Optimization

Feedback optimization enables autonomous optimality seeking of a dynamical system through its closed-loop interconnection with iterative optimization algorithms. Among various iteration structures, model-based approaches require the input-output sensitivity matrix of the system to construct gradients, whereas model-free approaches eliminate this need by estimating gradients from real-time objective evaluations. These approaches offer complementary benefits in sample efficiency and accuracy against model mismatch, i.e., sensitivity errors. To achieve balanced closed-loop performance, we propose a gray-box feedback optimization controller, featuring systematic incorporation of approximate sensitivities into model-free updates via a tunable convex combination. We provide unified performance characterizations covering different approaches. We elucidate how cumulative sensitivity errors (model-based) and variances due to stochastic exploration (model-free) shape the closed-loop behavior and induce a trade-off between iteration and dimensional dependence. The proposed controller retains sample efficiency and provable (local) optimality for nonconvex problems despite inaccurate sensitivities. We further develop and characterize a running gray-box controller that handles constrained time-varying problems with changing objectives and steady-state input-output maps.

math.OC↗

Welfarist Control Design -- How to fulfill the societal mandate in multi-agent control?

At the core of most socio-technical systems lies a scarce resource that is allocated among agents: highway lanes, public transit, road space, water rights, energy access, grid capacity, user attention, pollution rights, etc. With further automation of the underlying allocation processes, control engineers are increasingly tasked to make decisive assumptions regarding what society wants. In practice to date, design choices are largely driven by industry norms and conventions rather than a result of conscientiously responsible and ethical design. In this paper, we look at tools available to control engineers to design systems in a more principled manner in order to match the societal mandate. We consider three control design paradigms: online feedback optimization, control of Markov decision processes, and model predictive control. Beginning with aggregating individual agents' preferences into control design objectives, subsequently ensuring and certifying the fulfillment of those specifications, we argue that the feedback nature of control systems enables appropriate allocation of the shared resources in ways hitherto unparalleled.

eess.SY↗

From droop to optimality: The potential of volt/var control for power distribution grid enhancement

When high amounts of active power are injected into power distribution grids, the overall power flow is limited because voltages reach their upper acceptable limits. Volt/var control aims to raise this power flow limit without physically reinforcing the grid but by controlling the voltage using reactive power. We use real consumption and generation data on a low-voltage CIGRÉ grid model and an experiment on a real distribution grid feeder to analyze how different volt/var methods can enhance the grid. We show that local droop control enhances the grid but underutilizes the reactive power resources. We discuss how this inefficiency can be partly reduced by fine-tuning the droop curves through data-driven techniques but illustrate that inherent trade-off persist for any local control method. We finally demonstrate that coordinated control methods can track the optimal solution and enhance the grid to its full potential if grid-wide communication is available. Our numerical study over a whole year of real data suggests that coordinated volt/var control can enable another 10.4% of maximum active power injections compared to droop control. In a small-scale real-life experiment, coordinated control enhanced the grid by the same amount.

eess.SY↗

Dynamic Resource Allocation with Karma: An Experimental Study

We perform a behavioral experiment of karma, a class of mechanisms for repeated resource allocation with attractive fairness and efficiency properties, in theory. Individuals in these mechanisms bid non-tradable credits that flow from resource consumers to yielders, like karma. Human subjects recruited on Amazon MTurk are repeatedly and randomly paired to bid karma according to time-varying and stochastic individual preferences or urgency to acquire resources. Treatments varied in the dynamic urgency process (frequent moderate urgency versus sporadic high urgency) and the richness of the bidding scheme (binary versus full range). Results are benchmarked against random allocation, and karma achieves a (almost) Pareto improvement over random, despite the MTurk subjects deviating significantly from the theoretically optimal Nash bidding policy. Maximum improvement is attained by subjects that deviate from Nash by up to one karma bid unit on average, and positive improvement is attained with average deviations of up to 3-4 bid units. These findings hold across all treatments, among which no significant differences are found, with the exception of the sporadic high urgency process with binary bidding treatment being (weakly) favorable over others. These results offer behaviorally robust lower bounds for the expected performance of karma in human populations. They also provide guidance for future testing and implementation of karma mechanisms in the real world.

econ.GN↗

Invariant Price of Anarchy and Multiplicative Smoothness

The Price of Anarchy (PoA) is a popular measure of the costs of decentralization in terms of efficiency losses. Almost all PoA analyses operate within a framework assuming both Cardinal Full-Comparability (CFC) and smoothness, in which case any derived bounds conveniently extend beyond pure Nash to coarse correlated equilibria and no-regret learning outcomes. However, interpersonal utility comparability is an additional assumption that generally has to be justified. Without it, cardinal utilities (e.g. defined under classical von Neumann--Morgenstern framework) are unique only up to agent-specific affine transformations, rendering both the utilitarian PoA and the classical smoothness conditions representation-dependent. In this paper, we operate under a more general Cardinal Non-Comparability (CNC) framework, under which the weighted Nash welfare is a canonical admissible aggregator. We introduce multiplicative smoothness, a product-form condition matched to the multiplicative structure of Nash welfare, and obtain PoA bounds that are CNC-invariant and extend to coarse correlated equilibria. We demonstrate applicability of our framework on single-choice welfare games, deriving the bounds through simple proof relying on multiplicative retention envelope and geometric closure. The interpretation of this bound in terms of the true cost of decentralization depends crucially on interpersonal comparability of utilities.

cs.GT↗

Incentive-Based Load Curtailment with Limited Information: A Bilevel Zeroth-Order Learning Approach

Incentive-based load curtailment unlocks critical demand-side flexibility but is hindered by the limited knowledge of private user parameters and the inherent nonsmoothness of responses due to physical device constraints. We address this via a constrained bilevel optimization framework and propose the Bi-ZOL (Bilevel Zeroth-Order Learning) algorithm. Unlike conventional black-box methods, Bi-ZOL exploits the bilevel structure to decompose the hypergradient, integrating the exact analytical information of the SO's objective with a zeroth-order estimate of the unknown response sensitivity. This structural decomposition-based learning method mathematically smoothes the nonsmooth response landscape and reduces hypergradient estimation error. We provide theoretical convergence guarantees to an approximate stationary point and demonstrate through simulations that Bi-ZOL achieves near-optimal performance.

eess.SY↗

Towards Model-Free Learning in Dynamic Population Games: An Application to Karma Economies

Dynamic Population Games (DPGs) provide a tractable framework for modeling strategic interactions in large populations of self-interested agents, and have been successfully applied to the design of Karma economies, a class of fair non-monetary resource allocation mechanisms. Despite their appealing theoretical properties, existing computational tools for DPGs assume full knowledge of the game model and operate in a centralized fashion, limiting their applicability in realistic settings where agents have access only to their own private experience. This paper takes a step towards addressing this gap by studying model-free equilibrium learning in Karma DPGs. First, we analyze the setting in which a novel agent joins a Karma DPG already at its Stationary Nash Equilibrium (SNE) and learns a policy via Deep Q-Networks (DQN) without knowledge of the game model. Leveraging recent convergence results for DQN, we establish a suboptimality bound consisting of a DQN approximation error of order $O(1/\sqrt{N_s})$ and a mean field perturbation error of order $O(1/N)$, where $N_s$ is the replay buffer size and $N$ is the population size. Second, we consider the challenging problem of learning the SNE from scratch. We show empirically that combining deep RL with fictitious play and smoothed policy iteration allows agents to converge, in a model-free fashion, to a configuration close to the centrally computed SNE. Together, these contributions support the vision of Karma economies as practical tools for fair resource allocation.

cs.GT↗

A Welfarist Perspective on Fair Generation Curtailment

This paper presents a welfarist approach to fair active power curtailment in distribution grids with distributed photovoltaics. We address the lack of consistent axiomatic foundations in existing ad-hoc curtailment rules by modeling the decision as a social choice problem over feasible operating points and by deriving curtailment objectives from a set of foundational axioms that express principled stances on fairness and grid access rights. Rather than relying on the typically assumed full comparability of utilities, which can lead to undesirable outcomes in heterogeneous residential systems, we adopt a cardinal non-comparability stance on utilities. This approach requires far fewer assumptions about prosumers' private preferences while providing a rigorous basis for fair social ranking. We then present a unified framework that demonstrates that existing curtailment schemes represent specific instances of the Kalai-Smorodinsky rule applied to different normative reference points. This perspective offers grid operators an auditable, axiomatic foundation for justifying fairness in local energy systems.

eess.SY↗

Strategically Robust Aggregative Games

In many multiagent settings, such as electric vehicle charging and traffic routing, agents must make decisions in the face of uncertain behavior exhibited by others. Often, this uncertainty arises from multiple sources, such as incomplete information, limited computation, or bounded rationality, ultimately impacting the aggregate behavior. To tackle this challenge, we follow recent work on strategically robust game theory and postulate that agents seek protection directly against deviations around the emergent behavior, as opposed to explicitly modeling all sources of uncertainty. Specifically, we propose that each agent protects itself against the worst-case aggregate behavior within an optimal-transport-based ambiguity set centered at the emergent aggregate population behavior. This leads to a novel equilibrium concept, called strategically robust Wardrop equilibrium, that enables one to interpolate between standard Wardrop equilibria (no robustness) and security strategies (maximum robustness). In the setting of convex aggregative games, we establish the existence of a pure strategically robust Wardrop equilibrium and provide tractable computational tools for computing it. Through an application in electric vehicle charging, we demonstrate that strategically robust Wardrop equilibria lead to better decisions, protecting agents against the uncertain aggregate behavior of the population. Remarkably, we also observe that strategic robustness can lead to lower equilibrium costs for all agents, uncovering a "coordination-via-robustification" effect.

cs.GT↗

Towards Fair and Efficient allocation of Mobility-on-Demand resources through a Karma Economy

Mobility-on-demand systems like ride-hailing have transformed urban transportation, but they have also exacerbated socio-economic inequalities in access to these services, also due to surge pricing strategies. Although several fairness-aware frameworks have been proposed in smart mobility, they often overlook the temporal and situational variability of user urgency that shapes real-world transportation demands. This paper introduces a non-monetary, Karma-based mechanism that models endogenous urgency, allowing user time-sensitivity to evolve in response to system conditions as well as external factors. We develop a theoretical framework maintaining the efficiency and fairness guarantees of classical Karma economies, while accommodating this realistic user behavior modeling. Applied to a simplified simulated mobility-on-demand scenario, we provide a proof-of-concept illustration of the proposed framework, showing that it exhibits promising behavior in terms of system efficiency and equitable resource allocation, while acknowledging that a full treatment of realistic MoD complexity remains an important direction for future work.

eess.SY↗

Invariant Price of Anarchy: a Metric for Welfarist Traffic Control

The Price of Anarchy (PoA) is a standard metric for quantifying inefficiency in socio-technical systems, widely used to guide policies like traffic tolling. Conventional PoA analysis relies on exact numerical costs. However, in many settings, costs represent agents' preferences and may be defined only up to possibly arbitrary scaling and shifting, representing informational and modeling ambiguities. We observe that while such transformations preserve equilibrium and optimal outcomes, they change the PoA value. To resolve this issue, we rely on results from Social Choice Theory and define the Invariant PoA. By connecting admissible transformations to degrees of comparability of agents' costs, we derive the specific social welfare functions which ensure that efficiency evaluations do not depend on arbitrary rescalings or translations of individual costs. Case studies on a toy example and the Zurich network demonstrate that identical tolling strategies can lead to substantially different efficiency estimates depending on the assumed comparability. Our framework thus demonstrates that explicit axiomatic foundations are necessary in order to define efficiency metrics and to appropriately guide policy in large-scale infrastructure design robustly and effectively.

cs.GT↗

Subtransmission Grid Control via Online Feedback Optimization

The increasing electric power consumption and the shift towards renewable energy resources demand for new ways to operate transmission and subtransmission grids. Online Feedback Optimization (OFO) is a feedback real-time control method that can be employed to enable optimal operation of these grids. Such controllers can maximize grid efficiency (e.g., minimizing curtailment) while satisfying grid constraints like voltage and current limits. The OFO control method is tailored and extended to handle discrete inputs and it is explained how to design an OFO controller for the subtransmission grid. A novel benchmark is presented and published that corresponds to the real French subtransmission grid on which the proposed controller is analyzed in terms of robustness against model mismatch, constraint satisfaction, and tracking performance. It is shown that OFO controllers can help utilize the grid to its full extent, virtually reinforce it, and operate it optimally and in real-time by using the flexibility offered by renewable generators connected to distribution grids.

eess.SY↗

Welfare and Cost Aggregation for Multi-Agent Control: When to Choose Which Social Cost Function, and Why?

Many multi-agent socio-technical systems rely on aggregating heterogeneous agents' costs into a social cost function (SCF) to coordinate resource allocation in domains like energy grids, water allocation, or traffic management. The choice of SCF often entails implicit assumptions and may lead to undesirable outcomes if not rigorously justified. In this paper, we demonstrate that what determines which SCF ought to be used is the degree to which individual costs can be compared across agents and which axioms the aggregation shall fulfill. Drawing on the results from social choice theory, we provide guidance on how this process can be used in control applications. We demonstrate which assumptions about interpersonal utility comparability - ranging from ordinal level comparability to full cardinal comparability - together with a choice of desirable axioms, inform the selection of a correct SCF, be it the classical utilitarian sum, the Nash SCF, or maximin. We then demonstrate how the proposed framework can be applied for principled allocations of water and transportation resources.

math.OC↗