SearcharxivSearch

arXiv subjects

Luca Schenato

Publications and source records attributed to Luca Schenato.

At least 19 recordsLinked to original sources

Pursuing Optimal Stepsize in Adaptive Gradient-Based Quadratic Optimization

In this paper, we address the problem of achieving fast convergence in gradient descent for quadratic functions without relying on a priori knowledge of global function parameters. Inspired by adaptive stepsize algorithms for smooth convex functions, we propose a computationally lightweight strategy based on running estimates of minimal and maximal local curvatures. We prove that our proposed algorithm converges to the optimal constant stepsize which achieves the fastest convergence. Simulations show that the convergence rate achieved by our proposed algorithm is comparable or superior to recent adaptive approaches both in the quadratic case under consideration and in a preliminary test on logistic classification.

math.OC

Adaptive Stepsizes With Certified Convergence in Distributed Gradient Tracking With Quadratic Costs

In this work, we propose an adaptive stepsize rule with guaranteed convergence for Distributed Gradient Tracking applied to scalar quadratic problems with heterogeneous curvatures. Most distributed gradient-based algorithms require a suitable stepsize selection. Available theoretical bounds are often overly conservative, while practical implementations typically rely on empirically tuned heuristics. Online adaptive strategies have only recently emerged for general distributed convex optimization, but their properties and performance remain only partially understood. To gain analytical insight, we focus on the informative setting of scalar quadratic costs, which allows us to explicitly capture the interplay between network topology and curvature heterogeneity. We derive a convergence bound parameterized only by the essential spectral radius of the consensus matrix and the heterogeneity of the local cost curvatures, both computable online without any prior knowledge of the optimization problem. Optimizing this bound yields a computationally tractable surrogate for the convergence rate and the optimal constant stepsize. The resulting stepsize admits an analytical interpretation, guarantees convergence for arbitrary network topologies and curvature heterogeneity, and is provably tight for complete graphs and homogeneous curvatures. Finally, extensive numerical simulations demonstrate that the proposed distributed adaptive strategy significantly outperforms existing offline and online stepsize selection rules in the considered setting.

math.OC

Timescale Separation Through the Lens of Operator Theory

Timescale separation is a powerful tool for analyzing interconnected dynamical systems. Meanwhile, operator theory provides a general framework for studying the convergence of iterative methods formulated as fixed-point iterations, including algorithms arising in optimization, learning, and control. In this paper, we bridge these two areas by establishing timescale separation results for fixed-point iterations induced by both deterministic and stochastic operators. As customary in timescale separation, our results involve auxiliary systems that arise from the original interconnection in the limit as the timescale parameter tends to zero and separately capture the dynamics induced by the slow and fast operators. The proposed operator-theoretic framework yields explicit and readily checkable bounds on this tunable parameter, expressed in terms of standard operator constants. To illustrate the applicability of our results, we employ them to prove the convergence properties of a feedback optimization scheme in both deterministic and stochastic settings.

math.OC

The Essential Role Of Ribosomal Feedback In Bacterial Cell Growth And Metabolic Load -- A Systems Biology Approach For Unveiling Shared Resources Regulation Within Synthetic Genetic Circuits

Modeling growth in bacterial cells is a major issue in systems and synthetic biology. Despite several growth rate functions proposed in the literature, most focus on nutrient composition without explicitly accounting for the possible perturbation provided by the expression of recombinant genes, an effect known as cell load or burden. On the other hand, mathematical models that attempt to provide mechanistic details on the phenomena, leveraging ribosome partitioning and nutrient availability, are generally too detailed and complex to be easily applied to the rational design of synthetic genetic circuits. A bottom-up approach is adopted herein to identify and analyze the minimal model structure, thereby unveiling the fundamental role of negative feedback in ribosomal synthesis in predicting the effects of cell load on both gene expression and growth rate. Indeed, to ensure cellular efficiency, ribosome synthesis must be finely regulated. While an increased number of ribosomes generally enhances protein production and cellular performance, their synthesis incurs a high energetic cost. For this reason, cells have evolved mechanisms to tightly control ribosome synthesis, avoiding unnecessary accumulation. One of the key regulatory strategies, usually neglected in previous cell models, involves a negative feedback loop that modulates the production of ribosomal components. This feedback ensures that ribosomes are produced only in the amount strictly needed, balancing functionality and energy expenditure. This work evaluates the individual contribution of this feedback under heterologous expression conditions using minimal gene-circuit models, explicitly linking ribosome allocation, hidden couplings between protein synthesis levels, and growth rate.

q-bio.QM

On Convergence Analysis of Network-GIANT: An approximate Hessian-based fully distributed optimization algorithm

This paper presents a detailed convergence and performance analysis of a recently developed approximate Newton-type fully distributed optimization method for \(L\)-smooth, \(\mu\)-strongly convex local loss functions, called Network-GIANT (inspired by the Federated learning algorithm GIANT possessing mixed linear-quadratic convergence properties). Network-GIANT has been empirically seen to achieve faster linear convergence properties compared to its gradient-based counterparts, and several other existing second order distributed algorithms, while having the same communication complexity (per iteration) as its first order distributed counterparts. We first explicitly characterize a \emph{global linear convergence rate} for Network-GIANT, which can be computed as the spectral radius of a $3 \times 3$ matrix dependent on $L$, $\mu$, and the spectral norm ($\sigma$) of the consensus matrix of the underlying undirected graph. We provide an explicit bound on the step size parameter $\eta$, below which this spectral radius is guaranteed to be less than $1$. Furthermore, we derive a mixed linear-quadratic inequality based upper bound for the optimality gap norm, and provide a rigorous proof of a local asymptotic convergence rate of \(1 - \eta \big(1 - \frac{\gamma}{\mu}\big)\) given the Hessian approximation error $\gamma < \mu$, which formally explains the faster convergence rate of Network-GIANT. Numerical experiments are carried out with a reduced CovType dataset for binary logistic regression over a variety of graphs, including heterogeneous data distributions, to illustrate the above theoretical results.

math.OC

HBNET-GIANT: A communication-efficient accelerated Newton-type fully distributed optimization algorithm

This article presents a second-order fully distributed optimization algorithm, HBNET-GIANT, driven by heavy-ball momentum, for $L$-smooth and $\mu$-strongly convex objective functions. A rigorous convergence analysis is performed, and we demonstrate global linear convergence under certain sufficient conditions. Through extensive numerical experiments, we show that HBNET-GIANT with heavy-ball momentum achieves acceleration, and the corresponding rate of convergence is strictly faster than its non-accelerated version, NETWORK-GIANT. Moreover, we compare HBNET-GIANT with several state-of-the-art algorithms, both momentum-based and without momentum, and report significant performance improvement in convergence to the optimum. We believe that this work lays the groundwork for a broader class of second-order Newton-type algorithms with momentum and motivates further investigation into open problems, including an analytical proof of local acceleration in the fully distributed setting for convex optimization problems.

math.OC

Distributed clustering in partially overlapping feature spaces

We introduce and address a novel distributed clustering problem where each participant has a private dataset containing only a subset of all available features, and some features are included in multiple datasets. This scenario occurs in many real-world applications, such as in healthcare, where different institutions have complementary data on similar patients. We propose two different algorithms suitable for solving distributed clustering problems that exhibit this type of feature space heterogeneity. The first is a federated algorithm in which participants collaboratively update a set of global centroids. The second is a one-shot algorithm in which participants share a statistical parametrization of their local clusters with the central server, who generates and merges synthetic proxy datasets. In both cases, participants perform local clustering using algorithms of their choice, which provides flexibility and personalized computational costs. Pretending that local datasets result from splitting and masking an initial centralized dataset, we identify some conditions under which the proposed algorithms are expected to converge to the optimal centralized solution. Finally, we test the practical performance of the algorithms on three public datasets.

cs.DS

The role of communication delays in the optimal control of spatially invariant systems

We study optimal proportional feedback controllers for spatially invariant systems when the controller has access to delayed state measurements received from different spatial locations. We analyze how delays affect the spatial locality of the optimal feedback gain leveraging the problem decoupling in the spatial frequency domain. For the cases of expensive control and small delay, we provide exact expressions of the optimal controllers in the limit for infinite control weight and vanishing delay, respectively. In the expensive control regime, the optimal feedback control law decomposes into a delay-aware filtering of the delayed state and the optimal controller in the delay-free setting. Under small delays, the optimal controller is a perturbation of the delay-free one which depends linearly on the delay. We illustrate our analytical findings with a reaction-diffusion process over the real line and a multi-agent system coupled through circulant matrices, showing that delays reduce the effectiveness of optimal feedback control and may require each subsystem within a distributed implementation to communicate with farther-away locations.

math.OC

Optimal Control Selection over the Edge-Cloud Continuum

The emerging computing continuum paves the way for exploiting multiple computing devices, ranging from the edge to the cloud, to implement the control algorithm. Different computing units over the continuum are characterized by different computational capabilities and communication latencies, thus resulting in different control performances and advocating for an effective trade-off. To this end, in this work, we first introduce a multi-tiered controller and we propose a simple network delay compensator. Then we propose a control selection policy to optimize the control cost taking into account the delay and the disturbances. We theoretically investigate the stability of the switching system resulting from the proposed control selection policy. Accurate simulations show the improvements of the considered setup.

eess.SY

Multi-Agent Optimization and Learning: A Non-Expansive Operators Perspective

Multi-agent systems are increasingly widespread in a range of application domains, with optimization and learning underpinning many of the tasks that arise in this context. Different approaches have been proposed to enable the cooperative solution of these optimization and learning problems, including first- and second-order methods, and dual (or Lagrangian) methods, all of which rely on consensus and message-passing. In this article we discuss these algorithms through the lens of non-expansive operator theory, providing a unifying perspective. We highlight the insights that this viewpoint delivers, and discuss how it can spark future original research.

math.OC

Humans-in-the-Building: Getting Rid of Thermostats for Optimal Thermal Comfort Control in Energy Management Systems

Given the widespread attention to individual thermal comfort, coupled with significant energy-saving potential inherent in energy management systems for optimizing indoor environments, this paper aims to introduce advanced "Humans-in-the-building" control techniques to redefine the paradigm of indoor temperature design. Firstly, we innovatively redefine the role of individuals in the control loop, establishing a model for users' thermal comfort and constructing discomfort signals based on individual preferences. Unlike traditional temperature-centric approaches, "thermal comfort control" prioritizes personalized comfort. Then, considering the diversity among users, we propose a novel method to determine the optimal indoor temperature range, thus minimizing discomfort for various users and reducing building energy consumption. Finally, the efficacy of the "thermal comfort control" approach is substantiated through simulations conducted using Matlab.

eess.SY

Stochastic Approximation with Delayed Updates: Finite-Time Rates under Markovian Sampling

Motivated by applications in large-scale and multi-agent reinforcement learning, we study the non-asymptotic performance of stochastic approximation (SA) schemes with delayed updates under Markovian sampling. While the effect of delays has been extensively studied for optimization, the manner in which they interact with the underlying Markov process to shape the finite-time performance of SA remains poorly understood. In this context, our first main contribution is to show that under time-varying bounded delays, the delayed SA update rule guarantees exponentially fast convergence of the \emph{last iterate} to a ball around the SA operator's fixed point. Notably, our bound is \emph{tight} in its dependence on both the maximum delay $\tau_{max}$, and the mixing time $\tau_{mix}$. To achieve this tight bound, we develop a novel inductive proof technique that, unlike various existing delayed-optimization analyses, relies on establishing uniform boundedness of the iterates. As such, our proof may be of independent interest. Next, to mitigate the impact of the maximum delay on the convergence rate, we provide the first finite-time analysis of a delay-adaptive SA scheme under Markovian sampling. In particular, we show that the exponent of convergence of this scheme gets scaled down by $\tau_{avg}$, as opposed to $\tau_{max}$ for the vanilla delayed SA rule; here, $\tau_{avg}$ denotes the average delay across all iterations. Moreover, the adaptive scheme requires no prior knowledge of the delay sequence for step-size tuning. Our theoretical findings shed light on the finite-time effects of delays for a broad class of algorithms, including TD learning, Q-learning, and stochastic gradient descent under Markovian sampling.

cs.LG

VREM-FL: Mobility-Aware Computation-Scheduling Co-Design for Vehicular Federated Learning

Assisted and autonomous driving are rapidly gaining momentum and will soon become a reality. Artificial intelligence and machine learning are regarded as key enablers thanks to the massive amount of data that smart vehicles will collect from onboard sensors. Federated learning is one of the most promising techniques for training global machine learning models while preserving data privacy of vehicles and optimizing communications resource usage. In this article, we propose vehicular radio environment map federated learning (VREM-FL), a computation-scheduling co-design for vehicular federated learning that combines mobility of vehicles with 5G radio environment maps. VREM-FL jointly optimizes learning performance of the global model and wisely allocates communication and computation resources. This is achieved by orchestrating local computations at the vehicles in conjunction with transmission of their local models in an adaptive and predictive fashion, by exploiting radio channel maps. The proposed algorithm can be tuned to trade training time for radio resource usage. Experimental results demonstrate that VREM-FL outperforms literature benchmarks for both a linear regression model (learning time reduced by 28%) and a deep neural network for semantic image segmentation (doubling the number of model updates within the same time window).

eess.SY

FedZeN: Towards superlinear zeroth-order federated learning via incremental Hessian estimation

Federated learning is a distributed learning framework that allows a set of clients to collaboratively train a model under the orchestration of a central server, without sharing raw data samples. Although in many practical scenarios the derivatives of the objective function are not available, only few works have considered the federated zeroth-order setting, in which functions can only be accessed through a budgeted number of point evaluations. In this work we focus on convex optimization and design the first federated zeroth-order algorithm to estimate the curvature of the global objective, with the purpose of achieving superlinear convergence. We take an incremental Hessian estimator whose error norm converges linearly, and we adapt it to the federated zeroth-order setting, sampling the random search directions from the Stiefel manifold for improved performance. In particular, both the gradient and Hessian estimators are built at the central server in a communication-efficient and privacy-preserving way by leveraging synchronized pseudo-random number generators. We provide a theoretical analysis of our algorithm, named FedZeN, proving local quadratic convergence with high probability and global linear convergence up to zeroth-order precision. Numerical simulations confirm the superlinear convergence rate and show that our algorithm outperforms the federated zeroth-order methods available in the literature.

cs.LG

Visibility-Constrained Control of Multirotor via Reference Governor

For safe vision-based control applications, perception-related constraints have to be satisfied in addition to other state constraints. In this paper, we deal with the problem where a multirotor equipped with a camera needs to maintain the visibility of a point of interest while tracking a reference given by a high-level planner. We devise a method based on reference governor that, differently from existing solutions, is able to enforce control-level visibility constraints with theoretically assured feasibility. To this end, we design a new type of reference governor for linear systems with polynomial constraints which is capable of handling time-varying references. The proposed solution is implemented online for the real-time multirotor control with visibility constraints and validated with simulations and an actual hardware experiment.

cs.RO

Q-SHED: Distributed Optimization at the Edge via Hessian Eigenvectors Quantization

Edge networks call for communication efficient (low overhead) and robust distributed optimization (DO) algorithms. These are, in fact, desirable qualities for DO frameworks, such as federated edge learning techniques, in the presence of data and system heterogeneity, and in scenarios where internode communication is the main bottleneck. Although computationally demanding, Newton-type (NT) methods have been recently advocated as enablers of robust convergence rates in challenging DO problems where edge devices have sufficient computational power. Along these lines, in this work we propose Q-SHED, an original NT algorithm for DO featuring a novel bit-allocation scheme based on incremental Hessian eigenvectors quantization. The proposed technique is integrated with the recent SHED algorithm, from which it inherits appealing features like the small number of required Hessian computations, while being bandwidth-versatile at a bit-resolution level. Our empirical evaluation against competing approaches shows that Q-SHED can reduce by up to 60% the number of communication rounds required for convergence.

eess.SY

Network-GIANT: Fully distributed Newton-type optimization via harmonic Hessian consensus

This paper considers the problem of distributed multi-agent learning, where the global aim is to minimize a sum of local objective (empirical loss) functions through local optimization and information exchange between neighbouring nodes. We introduce a Newton-type fully distributed optimization algorithm, Network-GIANT, which is based on GIANT, a Federated learning algorithm that relies on a centralized parameter server. The Network-GIANT algorithm is designed via a combination of gradient-tracking and a Newton-type iterative algorithm at each node with consensus based averaging of local gradient and Newton updates. We prove that our algorithm guarantees semi-global and exponential convergence to the exact solution over the network assuming strongly convex and smooth loss functions. We provide empirical evidence of the superior convergence performance of Network-GIANT over other state-of-art distributed learning algorithms such as Network-DANE and Newton-Raphson Consensus.

math.OC

ZO-JADE: Zeroth-order Curvature-Aware Multi-Agent Convex Optimization

In this work we address the problem of convex optimization in a multi-agent setting where the objective is to minimize the mean of local cost functions whose derivatives are not available (e.g. black-box models). Moreover agents can only communicate with local neighbors according to a connected network topology. Zeroth-order (ZO) optimization has recently gained increasing attention in federated learning and multi-agent scenarios exploiting finite-difference approximations of the gradient using from $2$ (directional gradient) to $2d$ (central difference full gradient) evaluations of the cost functions, where $d$ is the dimension of the problem. The contribution of this work is to extend ZO distributed optimization by estimating the curvature of the local cost functions via finite-difference approximations. In particular, we propose a novel algorithm named ZO-JADE, that by adding just one extra point, i.e. $2d+1$ in total, allows to simultaneously estimate the gradient and the diagonal of the local Hessian, which are then combined via average tracking consensus to obtain an approximated Jacobi descent. Guarantees of semi-global exponential stability are established via separation of time-scales. Extensive numerical experiments on real-world data confirm the efficiency and superiority of our algorithm with respect to several other distributed zeroth-order methods available in the literature based on only gradient estimates.

math.OC