SearcharxivSearch

arXiv subjects

Sina Sanjari

Publications and source records attributed to Sina Sanjari.

18 recordsLinked to original sources

Optimality of Symmetric Independent Policies under Decentralized Mean-Field Information Sharing for Stochastic Teams and Equivalence with McKean-Vlasov Control of a Representative Agent

We study a class of stochastic exchangeable teams with a finite number of decision makers (DMs) as well as their mean-field limits with infinitely many DMs. In the finite population regime, we study exchangeable teams under the centralized information structure. The paper makes the following main contributions: i) For finite population exchangeable teams, we establish the existence of an optimal policy that is exchangeable (permutation invariant) and Markovian; ii) As our main result in the paper, we show that a sequence of exchangeable optimal policies for finite population settings (which satisfies a measure valued MDP formulation due to B{ä}uerle) converges to a decentralized symmetric (identical) and conditionally independent (given the mean-field) policy for the infinite population problem, which is then globally optimal under both the centralized information structure as well as the mean-field sharing information structure. (iii) This result establishes existence of a symmetric, independent, decentralized optimal randomized policy for the infinite population problem and proves the optimality of the limiting measure-valued MDP for the representative DM. Our paper thus establishes the relation between the controlled McKean-Vlasov dynamics and the optimal infinite population decentralized stochastic control problem (without an apriori restriction of symmetry in policies of individual agents), for the first time, to our knowledge (beyond several special cases). We also establish near optimality of a numerical method for solving this problem. iv) Finally, we show that symmetric, independent, decentralized optimal randomized policies are approximately optimal for the corresponding finite-population team with a large number of DMs under the centralized information structure.

math.OC

Path-independent Flow Matching for Multi-parameter Generative Dynamics

Flow Matching is a powerful framework for learning transport maps between probability distributions. Yet its standard single-parameter formulation is not designed to capture multi-parameter variations where the resulting transport should be path-independent. Path independence is crucial because it ensures that transformations depend only on the initial and target distributions, not on the specific path. In this work, we introduce Path-independent Flow Matching (PiFM), a method for learning vector fields whose induced flows yield path-independent transport between distributions. We show that PiFM generalizes Flow Matching to higher-dimensional parameter domains while enforcing structural conditions that ensure consistency of composed transformations. In addition, we show that, under suitable assumptions, PiFM approximates the Wasserstein barycenter, linking the framework to a notion of distributional interpolation. To enable practical training, we propose a tractable, simulation-free objective that regresses onto multi-parameter conditional probability paths. We showcase empirically that PiFM outperforms other approaches on both synthetic and real world data in interpolating path-independent trajectories and generating desired out of distribution samples.

cs.LG

Mean-Field Systems with Heterogeneous Subteams: Optimality of Cluster-Symmetric Independent Policies and Equivalence with Decentralized McKean-Vlasov Control of Cluster-Representative Agents

Across science and engineering, mean-field methods have been a powerful and versatile approach for the analysis of systems of many interacting elements. However, common arguments used to characterize an infinite population limit can be quite restrictive from a modeling perspective by requiring that all agents be identical (i.e. symmetric, or homogeneous). In this paper, we consider large interactive particle systems under agent heterogeneity for a class of discrete time teams composed of finitely many species of agents, grouped into symmetric subteams, called clusters. In particular, for the class of discounted, partially exchangeable cost criteria considered, we establish the optimality of centralized joint policies which are exchangeable within each cluster and depend on the agent ensemble only up to the state empirical distribution over each cluster. Following this, a generalization of De Finetti's theorem is used to demonstrate the subsequential convergence of these optimal policies to one which is decentralized (depending on only the local state and distribution over each cluster) and symmetric within each subteam as the population size approaches infinity. This solution is shown to induce a sequence of asymptotically optimal policies for the finite population problems which retain their structure and decentralization. Furthermore, our analysis justifies the optimality of a decentralized McKean-Vlasov team representation involving coupled representative agents for each of the clusters, and establishes a verification theorem/value iterations for the mean-field limit. In this way, we provide an avenue for analyzing complex, cooperative systems with finite heterogeneity and set the stage for further research on learning algorithms.

math.OC

Nonparametric Sparse Online Learning of the Koopman Operator

The Koopman operator provides a powerful framework for representing the dynamics of general nonlinear dynamical systems. However, existing data-driven approaches to learning the Koopman operator rely on batch data. In this work, we present a sparse online learning algorithm that learns the Koopman operator iteratively via stochastic approximation, with explicit control over model complexity and provable convergence guarantees. Specifically, we study the Koopman operator via its action on the reproducing kernel Hilbert space (RKHS), and address the mis-specified scenario where the dynamics may escape the chosen RKHS. In this mis-specified setting, we relate the Koopman operator to the conditional mean embeddings (CME) operator. We further establish both asymptotic and finite-time convergence guarantees for our learning algorithm in mis-specified setting, with trajectory-based sampling where the data arrive sequentially over time. Numerical experiments demonstrate the algorithm's capability to learn unknown nonlinear dynamics.

stat.ML

Decentralized Detection with Many Sensors: Optimality of Exchangeable and Identical Encoding Policies

We study a class of binary detection problems involving a single fusion center and a large or countably infinite number of sensors. Each sensor acts under a decentralized information structure, accessing only a local noisy observation related to the hypothesis. Based on this observation, sensors select policies to transmit a quantized signal through their actions to the fusion center, which makes the final decision using only these actions. This paper makes the following contributions: i) In the finitely many sensor setting, we provide a formal proof that an optimal encoding policy exists, and such an optimal policy is independent, deterministic, and of threshold type for the sensors and the maximum \emph{a posteriori} probability type for the fusion center; ii) For the finitely many sensor setting, we further show that an optimal encoding policy exhibits an exchangeability (permutation invariance) property; iii) We establish that an optimal encoding policy exists that is symmetric (identical) and independent across sensors in the infinitely many sensor setting under the error exponent cost; iv) Finally, we show that a symmetric optimal policy for the infinite population regime with the error exponent cost is approximately optimal for the large but finite sensor regime under the same cost criterion. We anticipate that the mathematical program used in the paper will find applications in several other massive communications applications.

math.OC

Quantum Markov Decision Processes: Dynamic and Semi-Definite Programs for Optimal Solutions

In this paper, building on the formulation of quantum Markov decision processes (q-MDPs) presented in our previous work [{\sc N.~Saldi, S.~Sanjari, and S.~Yüksel}, {\em Quantum Markov Decision Processes: General Theory, Approximations, and Classes of Policies}, SIAM Journal on Control and Optimization, 2024], our focus shifts to the development of semi-definite programming approaches for optimal policies and value functions of both open-loop and classical-state-preserving closed-loop policies. First, by using the duality between the dynamic programming and the semi-definite programming formulations of any q-MDP with open-loop policies, we establish that the optimal value function is linear and there exists a stationary optimal policy among open-loop policies. Then, using these results, we establish a method for computing an approximately optimal value function and formulate computation of optimal stationary open-loop policy as a bi-linear program. Next, we turn our attention to classical-state-preserving closed-loop policies. Dynamic programming and semi-definite programming formulations for classical-state-preserving closed-loop policies are established, where duality of these two formulations similarly enables us to prove that the optimal policy is linear and there exists an optimal stationary classical-state-preserving closed-loop policy. Then, similar to the open-loop case, we establish a method for computing the optimal value function and pose computation of optimal stationary classical-state-preserving closed-loop policies as a bi-linear program.

quant-ph

Nonparametric Sparse Online Learning of the Koopman Operator

The Koopman operator provides a powerful framework for representing the dynamics of general nonlinear dynamical systems. Data-driven techniques to learn the Koopman operator typically assume that the chosen function space is closed under system dynamics. In this paper, we study the Koopman operator via its action on the reproducing kernel Hilbert space (RKHS), and explore the mis-specified scenario where the dynamics may escape the chosen function space. We relate the Koopman operator to the conditional mean embeddings (CME) operator and then present an operator stochastic approximation algorithm to learn the Koopman operator iteratively with control over the complexity of the representation. We provide both asymptotic and finite-time last-iterate guarantees of the online sparse learning algorithm with trajectory-based sampling with an analysis that is substantially more involved than that for finite-dimensional stochastic approximation. Numerical examples confirm the effectiveness of the proposed algorithm.

stat.ML

Quantum Markov Decision Processes: General Theory, Approximations, and Classes of Policies

In this paper, the aim is to develop a quantum counterpart to classical Markov decision processes (MDPs). Firstly, we provide a very general formulation of quantum MDPs with state and action spaces in the quantum domain, quantum transitions, and cost functions. Once we formulate the quantum MDP (q-MDP), our focus shifts to establishing the verification theorem that proves the sufficiency of Markovian quantum control policies and provides a dynamic programming principle. Subsequently, a comparison is drawn between our q-MDP model and previously established quantum MDP models (referred to as QOMDPs) found in the literature. Furthermore, approximations of q-MDPs are obtained via finite-action models, which can be formulated as QOMDPs. Finally, classes of open-loop and classical-state-preserving closed-loop policies for q-MDPs are introduced, along with structural results for these policies. In summary, we present a novel quantum MDP model aiming to introduce a new framework, algorithms, and future research avenues. We hope that our approach will pave the way for a new research direction in discrete-time quantum control.

quant-ph

Decentralized Exchangeable Stochastic Dynamic Teams in Continuous-time, their Mean-Field Limits and Optimality of Symmetric Policies

We study a class of stochastic exchangeable teams comprising a finite number of decision makers (DMs) as well as their mean-field limits involving infinite numbers of DMs. In the finite population regime, we study exchangeable teams under the centralized information structure. For the infinite population setting, we study exchangeable teams under the decentralized mean-field information sharing. The paper makes the following main contributions: i) For finite population exchangeable teams, we establish the existence of a randomized optimal policy that is exchangeable (permutation invariant) and Markovian; ii) As our main result in the paper, we show that a sequence of exchangeable optimal policies for finite population settings converges to a conditionally symmetric (identical), independent, and decentralized randomized policy for the infinite population problem, which is globally optimal for the infinite population problem. This result establishes the existence of a symmetric, independent, decentralized optimal randomized policy for the infinite population problem. Additionally, this proves the optimality of the limiting measure-valued MDP for the representative DM; iii) Finally, we show that symmetric, independent, decentralized optimal randomized policies are approximately optimal for the corresponding finite-population team with a large number of DMs under the centralized information structure. Our paper thus establishes the relation between the controlled McKean-Vlasov dynamics and the optimal infinite population decentralized stochastic control problem (without an apriori restriction of symmetry in policies of individual agents), for the first time, to our knowledge.

math.OC

Incentive Designs for Stackelberg Games with a Large Number of Followers and their Mean-Field Limits

We study incentive designs for a class of stochastic Stackelberg games with one leader and a large number of (finite as well as infinite population of) followers. We investigate whether the leader can craft a strategy under a dynamic information structure that induces a desired behavior among the followers. For the finite population setting, under convexity of the leader's cost and other sufficient conditions, we show that there exist symmetric \emph{incentive} strategies for the leader that attain approximately optimal performance from the leader's viewpoint and lead to an approximate symmetric (pure) Nash best response among the followers. Leveraging functional analytic tools, we further show that there exists a symmetric incentive strategy, which is affine in the dynamic part of the leader's information, comprising partial information on the actions taken by the followers. Driving the follower population to infinity, we arrive at the interesting result that in this infinite-population regime the leader cannot design a smooth ``finite-energy'' incentive strategy, namely, a mean-field limit for such games is not well-defined. As a way around this, we introduce a class of stochastic Stackelberg games with a leader, a major follower, and a finite or infinite population of minor followers. For this class of problems, we establish the existence of an incentive strategy and the corresponding mean-field Stackelberg game. Examples of quadratic Gaussian games are provided to illustrate both positive and negative results. In addition, as a byproduct of our analysis, we establish the existence of a randomized incentive strategy for the class mean-field Stackelberg games, which in turn provides an approximation for an incentive strategy of the corresponding finite population Stackelberg game.

cs.GT

Isomorphism Properties of Optimality and Equilibrium Solutions under Equivalent Information Structure Transformations: Stochastic Dynamic Games and Teams

Static reduction of information structures (ISs) is a method that is commonly adopted in stochastic control, team theory, and game theory. One approach entails change of measure arguments, which has been crucial for stochastic analysis and has been an effective method for establishing existence and approximation results for optimal policies. Another approach entails utilization of invertibility properties of measurements, with further generalizations of equivalent IS reductions being possible. In this paper, we demonstrate the limitations of such approaches for a wide class of stochastic dynamic games and teams, and present a systematic classification of static reductions for which both positive and negative results on equivalence properties of equilibrium solutions can be obtained: (i) those that are policy-independent, (ii) those that are policy-dependent, and (iii) a third type that we will refer to as static measurements with control-sharing reduction (where the measurements are static although control actions are shared according to the partially nested IS). For the first type, we show that there is a bijection between Nash equilibrium (NE) policies under the original IS and their policy-independent static reductions, and establish sufficient conditions under which stationary solutions are also isomorphic between these ISs. For the second type, however, we show that there is generally no isomorphism between NE (or stationary) solutions under the original IS and their policy-dependent static reductions. Sufficient conditions (on the cost functions and policies) are obtained to establish such an isomorphism relationship between Nash equilibria of dynamic non-zero-sum games and their policy-dependent static reductions. For zero-sum games and teams, these sufficient conditions can be further relaxed.

math.OC

Nash Equilibria for Exchangeable Team against Team Games, their Mean Field Limit, and Role of Common Randomness

We study stochastic mean-field games among finite number of teams with large finite as well as infinite number of decision makers. For this class of games within static and dynamic settings, we establish the existence of a Nash equilibrium, and show that a Nash equilibrium exhibits exchangeability in the finite decision maker regime and symmetry in the infinite one. To arrive at these existence and structural theorems, we endow the set of randomized policies with a suitable topology under various decentralized information structures, which leads to the desired convexity and compactness of the set of randomized policies. Then, we establish the existence of a randomized Nash equilibrium that is exchangeable (not necessarily symmetric) among decision makers within each team for a general class of exchangeable stochastic games. As the number of decision makers within each team goes to infinity (that is for the mean-field game among teams), using a de Finetti representation theorem, we show existence of a randomized Nash equilibrium that is symmetric (i.e., identical) among decision makers within each team and also independently randomized. Finally, we establish that a Nash equilibrium for a class of mean-field games among teams (which is symmetric) constitutes an approximate Nash equilibrium for the corresponding pre-limit (exchangeable) game among teams with large but finite number of decision makers. We thus show that common randomness is not necessary for large team-against-team games, unlike the case with small sized teams.

math.OC

Isomorphism Properties of Optimality and Equilibrium Solutions under Equivalent Information Structure Transformations I: Stochastic Dynamic Teams

In stochastic optimal control, change of measure arguments have been crucial for stochastic analysis. Such an approach is often called static reduction in dynamic team theory (or decentralized stochastic control) and has been an effective method for establishing existence and approximation results for optimal policies. In this paper, we place such static reductions into three categories: (i) those that are policy-independent (as those introduced by Witsenhausen), (ii) those that are policy-dependent (as those introduced by Ho and Chu for partially nested dynamic teams), and (iii) those that we will refer to as static measurements with control-sharing reduction (where the measurements are static although control actions are shared according to the partially nested information structure). For the first type, we show that there is a bijection between person-by-person optimal (globally optimal) policies of dynamic teams and their policy-independent static reductions. For the second type, although there is a bijection between globally optimal policies of dynamic teams with partially nested information structures and their static reductions, in general there is no bijection between person-by-person optimal policies of dynamic teams and their policy-dependent static reductions. We also establish a stronger negative result concerning stationary solutions. We present sufficient conditions under which bijection relationships hold. Under static measurements with control-sharing reduction, connections between optimality concepts can be established under relaxed conditions. An implication is a convexity characterization of dynamic team problems under static measurements with control-sharing reduction. Finally, we introduce multi-stage refinements of such reductions. Part II of the paper addresses similar issues in the context of stochastic dynamic games, where further subtleties arise.

math.OC

Optimality of Independently Randomized Symmetric Policies for Exchangeable Stochastic Teams with Infinitely Many Decision Makers

We study stochastic team (known also as decentralized stochastic control or identical interest stochastic dynamic game) problems with large or countably infinite number of decision makers, and characterize existence and structural properties for (globally) optimal policies. We consider both static and dynamic non-convex team problems where the cost function and dynamics satisfy an exchangeability condition. To arrive at existence and structural results on optimal policies, we first introduce a topology on control policies, which involves various relaxations given the decentralized information structure. This is then utilized to arrive at a de Finetti type representation theorem for exchangeable policies. This leads to a representation theorem for policies which admit an infinite exchangeability condition. For a general setup of stochastic team problems with $N$ decision makers, under exchangeability of observations of decision makers and the cost function, we show that without loss of global optimality, the search for optimal policies can be restricted to those that are $N$-exchangeable. Then, by extending $N$-exchangeable policies to infinitely-exchangeable ones, establishing a convergence argument for the induced costs, and using the presented de Finetti type theorem, we establish the existence of an optimal decentralized policy for static and dynamic teams with countably infinite number of decision makers, which turns out to be symmetric (i.e., identical) and randomized. In particular, unlike prior work, convexity of the cost in policies is not assumed. Finally, we show near optimality of symmetric independently randomized policies for finite $N$-decision maker team problems and thus establish approximation results for $N$-decision maker weakly coupled stochastic teams.

math.OC

Optimal Policies for Convex Symmetric Stochastic Dynamic Teams and their Mean-field Limit

This paper studies convex stochastic dynamic team problems with finite and infinite time horizons under decentralized information structures. First, we introduce two notions called exchangeable teams and symmetric information structures. We show that in convex exchangeable team problems an optimal policy exhibits a symmetry structure. We give a characterization for such symmetrically optimal teams for a general class of convex dynamic team problems under a mild conditional independence condition. In addition, through concentration of measure arguments, we establish the convergence of optimal policies for teams with $N$ decision makers to the corresponding optimal policies for symmetric mean-field teams with infinitely many decision makers. As a by-product, we present an existence result for convex mean-field teams, where the main contribution of our paper is with respect to the information structure in the system when compared with the related results in the literature that have either assumed a classical information structure or a static information structure. We also apply these results to the important special case of Linear Quadratic Gaussian (LQG) team problems, where while for partially nested LQG team problems with finite time horizons it is known that the optimal policies are linear, for infinite horizon problems the linearity of optimal policies has not been established in full generality. We also study average cost finite and infinite horizon dynamic team problems with a symmetric partially nested information structure and obtain globally optimal solutions where we establish linearity of optimal policies.

math.OC

Optimal Solutions to Infinite-Player Stochastic Teams and Mean-Field Teams

We study stochastic static teams with countably infinite number of decision makers, with the goal of obtaining (globally) optimal policies under a decentralized information structure. We present sufficient conditions to connect the concepts of team optimality and person by person optimality for static teams with countably infinite number of decision makers. We show that under uniform integrability and uniform convergence conditions, an optimal policy for static teams with countably infinite number of decision makers can be established as the limit of sequences of optimal policies for static teams with $N$ decision makers as $N \to \infty$. Under the presence of a symmetry condition, we relax the conditions and this leads to optimality results for a large class of mean-field optimal team problems where the existing results have been limited to person-by-person-optimality and not global optimality (under strict decentralization). In particular, we establish the optimality of symmetric (i.e., identical) policies for such problems. As a further condition, this optimality result leads to an existence result for mean-field teams. We consider a number of illustrative examples where the theory is applied to setups with either infinitely many decision makers or an infinite-horizon stochastic control problem reduced to a static team.

math.OC

Finite-time Stability Analysis for Random Nonlinear Systems

This paper presents an analysis approach to finite-time attraction in probability concerns with nonlinear systems described by nonlinear random differential equations (RDE). RDE provide meticulous physical interpreted models for some applications contain stochastic disturbance. The existence and the path-wise uniqueness of the finite-time solution are investigated through nonrestrictive assumptions. Then a finite-time attraction analysis is considered through the definition of the stochastic settling time function and a Lyapunov based approach. A Lyapunov theorem provides sufficient conditions to guarantee finite-time attraction in probability of random nonlinear systems. A Lyapunov function ensures stability in probability and a finiteness of the expectation of the stochastic settling time function. Results are demonstrated employing the method for two examples to show potential of the proposed technique.

eess.SY

Sliding Mode Control Design: a Sum of Squares Approach

This paper presents an approach to systematically design sliding mode control and manifold to stabilize nonlinear uncertain systems. The objective is also accomplished to enlarge the inner bound of region of attraction for closed-loop dynamics. The method is proposed to design a control that guarantees both asymptotic and finite time stability given helped by (bilinear) sum of squares programming. The approach introduces an iterative algorithm to search over sliding mode manifold and Lyapunov function simultaneity. In the case of local stability it concludes also the subset of estimated region of attraction for reduced order sliding mode dynamics. The sliding mode manifold and the corresponding Lyapunov function are obtained if the iterative SOS optimization program has a solution. Results are demonstrated employing the method for several examples to show potential of the proposed technique.

eess.SY