SearcharxivSearch

arXiv subjects

Silun Zhang

Publications and source records attributed to Silun Zhang.

17 recordsLinked to original sources

A No-Regret Framework for Adaptive Incentive Design

Incentive design studies how a central authority can influence strategic agents through payments, subsidies, or taxes, so that individual objectives align with collective welfare. This paper introduces a No-Regret Adaptive Incentive Design (RAID) framework for nonlinear games with continuous action spaces and private agent costs. In this framework, the authority (planner) designs incentives that regulate the Nash equilibrium toward a socially optimal action profile, while simultaneously learning agents' unknown preferences from repeated strategic responses. We formulate the RAID problem and construct a least-squares estimator whose strong consistency requires only diminishing excitation. Leveraging this weak excitation requirement, we propose a switching incentive policy that alternates between probing (exploration) and estimate-based (exploitation) incentives. The resulting policy achieves an $O(t^{-0.5})$ parameter estimation rate and accumulates $O(t^{0.5}\log t)$ squared social-cost regret, almost surely. We further extend the framework to an endogenous-noise response model, where standard least-squares estimation is biased due to an error-in-variables correlation between the noise and agent responses. We utilize a repeated-sampling estimator and corresponding switching policy that retain the same almost-sure convergence and regret rates. Numerical experiments validate the effectiveness and predicted convergence rates of the method.

math.OC

Hyperedge approximation for stochastic processes on higher-order networks

Graphs are a standard framework for describing dynamical processes shaped by pairwise interactions among agents. But many systems involve interactions in groups of three or more agents. Here, we develop a method of "$\ell$-hyperedge approximation", a framework to analyze stochastic population processes on regular hypergraphs, in which each individual belongs to $k$ groups of size $\ell$. The framework accommodates both higher-order interactions that determine payoffs and higher-order processes for updating states in response to payoffs. Applied to evolutionary game dynamics, the framework generalizes the classical pairwise result on benefits and costs, $b/c>k$, that favors the spread of cooperation; and it provides critical benefit-to-cost ratios for nonlinear $\ell$-player public goods games that cannot be reduced to pairwise interactions. Applied to complex contagions, where inheritance of states occurs within hyperedges rather than along parent-offspring edges, the framework gives a closed-form result for the fixation probability, which shows how a complexity parameter governs the spread of rare types. Coupling the two processes produces a single stochastic model of payoff-biased complex contagion in structured populations. These results extend pair approximation from graphs to hypergraphs, accommodating multi-way interactions and inheritance structures with no pairwise analog.

physics.soc-ph

Incentive Design without Hypergradients: A Social-Gradient Method

Incentive design problems consider a system planner who steers self-interested agents toward a socially optimal Nash equilibrium by issuing incentives in the presence of information asymmetry, that is, uncertainty about the agents' cost functions. A common approach formulates the problem as a Mathematical Program with Equilibrium Constraints (MPEC) and optimizes incentives using hypergradients-the total derivatives of the planner's objective with respect to incentives. However, computing or approximating the hypergradients typically requires full or partial knowledge of equilibrium sensitivities to incentives, which is generally unavailable under information asymmetry. In this paper, we propose a hypergradient-free incentive law, called the social-gradient flow, for incentive design when the planner's social cost depends on the agents' joint actions. We prove that the social cost gradient is always a descent direction for the planner's objective, irrespective of the agent cost landscape. In the idealized setting where equilibrium responses are observable, the social-gradient flow converges to the unique socially optimal incentive. When equilibria are not directly observable, the social-gradient flow emerges as the slow-timescale limit of a two-timescale interaction, in which agents' strategies evolve on a faster timescale. It is established that the joint strategy-incentive dynamics converge to the social optimum for any agent learning rule that asymptotically tracks the equilibrium. Theoretical results are also validated via numerical experiments.

math.OC

Stochastic Adaptive Control for Systems with Nonlinear Parameterization: Almost Sure Stability and Tracking

This paper concerns the adaptive control problem for a class of nonlinear stochastic systems in which the state update is given by a nonlinear function of linear dynamics plus additive stochastic noise. Such systems arise in a wide range of applications, including recurrent neural networks, social dynamics, and signal processing. Despite their importance, adaptive control for these systems remains relatively unexplored in the literature. This gap is primarily due to the inherently nonconvex dependence of the system dynamics on unknown parameters, which significantly complicates both controller design and analysis. To address these challenges, we propose an online nonlinear weighted least-squares (WLS)-based parameter estimation algorithm and establish the global strong consistency of the resulting parameter estimates. In contrast to most existing results, our consistency analysis does not rely on restrictive assumptions such as persistent excitation conditions of the trajectory data, making it applicable to stochastic adaptive control settings. Building on the proposed estimator, we further develop an adaptive control algorithm with an attenuating excitation signal that can effectively combine adaptive estimation and feedback control. Finally, we are able to show that the resulting closed-loop system is globally stable and that the system trajectory can track, in a long-run average sense, the reference trajectory generated with the true system parameters. The proposed methods and theoretical results are finally validated through simulations in two nonlinear interaction network applications.

eess.SY

Adaptive Incentive Design with Regret Minimization

Incentive design constitutes a foundational paradigm for influencing the behavior of strategic agents, wherein a system planner (principal) publicly commits to an incentive mechanism designed to align individual objectives with collective social welfare. This paper introduces the Regret-Minimizing Adaptive Incentive Design (RAID) problem, which aims to synthesize incentive laws under information asymmetry and achieve asymptotically minimal regret compared to an oracle with full information. To this end, we develop the RAID algorithm, which employs a switching policy alternating between probing (exploration) and estimate-based incentivization (exploitation). The associated type estimator relies only on a weaker excitation condition required for strong consistency in least squares estimation, substantially relaxing the persistence-of-excitation assumptions previously used in adaptive incentive design. In addition, we establish the strong consistency of the proposed type estimator and prove that the incentive obtained asymptotically minimizes the planner's average regret almost surely. Numerical experiments illustrate the convergence rate of the proposed methodology.

math.OC

Self-Identifying Internal Model-Based Online Optimization

In this paper, we propose a novel online optimization algorithm built by combining ideas from control theory and system identification. The foundation of our algorithm is a control-based design that makes use of the internal model of the online problem. Since such prior knowledge of this internal model might not be available in practice, we incorporate an identification routine that learns this model on the fly. The algorithm is designed starting from quadratic online problems but can be applied to general problems. For quadratic cases, we characterize the asymptotic convergence to the optimal solution trajectory. We compare the proposed algorithm with existing approaches, and demonstrate how the identification routine ensures its adaptability to changes in the underlying internal model. Numerical results also indicate strong performance beyond the quadratic setting.

math.OC

Learning Ordinal Response Policies in Rank-Based Stochastic Prize-Collecting Games

The Team Orienteering Problem (TOP) generalizes many real-world multi-agent scheduling and routing tasks that occur in autonomous mobility, aerial logistics, and surveillance applications. While many flavors of the TOP exist for planning in multi-agent systems, they assume that all the agents cooperate toward a single objective; therefore, they do not extend to settings when they compete in reward-scarce environments. We propose Stochastic Prize-Collecting Orienteering Games (SPCOG) as an extension of the TOP to plan in the presence of self-interested agents operating on a graph, under energy constraints and stochastic transitions. A theoretical discussion on complete and star graphs establishes that there is a unique pure Nash equilibrium in SPCOGs that coincides with the optimal routing solution of an equivalent TOP under rank-based conflict resolution. We propose the concept of Ordinal Rank (OR) as a concise representation of an agents' global rank and its location within a topological, well-defined neighborhood. Empirical evaluations conducted on real-world, road-network graphs under both dynamic and stationary prize distributions show that in parameter-sharing settings, the policies that leverage local information can outperform those policies leverage global information when the former is conditioned on the OR rather than the global rank, indicating that the OR acts as a strong inductive bias in multi-agent games on graphs. The OR-conditioned policies also generalize much better to games with large number of agents compared to global-rank conditioned policies. Finally, we also propose we propose Fictitious Ordinal Response Learning (FORL) as an entropy-regulated algorithm to obtain convergent policies in independent-learning settings in prize-collecting games on graphs.

cs.RO

Collective decision-making dynamics in hypernetworks

This work describes a collective decision-making dynamical process in a multiagent system under the assumption of cooperative higher-order interactions within the community, modeled as a hypernetwork. The nonlinear interconnected system is characterized by saturated nonlinearities that describe how agents transmit their opinion state to their neighbors in the hypernetwork, and by a bifurcation parameter representing the community's social effort. We show that the presence of higher-order interactions leads to the unfolding of a pitchfork bifurcation, introducing an interval for the social effort parameter in which the system exhibits bistability. With equilibrium points representing collective decisions, this implies that, depending on the initial conditions, the community will either remain in a deadlock state (with the origin as the equilibrium point) or reach a nontrivial decision. A numerical example is given to illustrate the results.

math.OC

Online Learning for Nonlinear Dynamical Systems without the I.I.D. Condition

This paper investigates online identification and prediction for nonlinear stochastic dynamical systems. In contrast to offline learning methods, we develop online algorithms that learn unknown parameters from a single trajectory. A key challenge in this setting is handling the non-independent data generated by the closed-loop system. Existing theoretical guarantees for such systems are mostly restricted to the assumption that inputs are independently and identically distributed (i.i.d.), or that the closed-loop data satisfy a persistent excitation (PE) condition. However, these assumptions are often violated in applications such as adaptive feedback control. In this paper, we propose an online projected Newton-type algorithm for parameter estimation in nonlinear stochastic dynamical systems, and develop an online predictor for system outputs based on online parameter estimates. By using both the stochastic Lyapunov function and martingale estimation methods, we demonstrate that the average regret converges to zero without requiring traditional persistent excitation (PE) conditions. Furthermore, we establish a novel excitation condition that ensures global convergence of the online parameter estimates. The proposed excitation condition is applicable to a broader class of system trajectories, including those violating the PE condition.

eess.SY

Network Consensus with Privacy: A Secret Sharing Method

In this work, inspired by secret sharing schemes, we introduce a privacy-preserving approach for network consensus, by which all nodes in a network can reach an agreement on their states without exposing the individual state to neighbors. With the privacy degree defined for the agents, the proposed method makes the network resistant to the collusion of any given number of neighbors, and protects the consensus procedure from communication eavesdropping. Unlike existing works, the proposed privacy-preserving algorithm is resilient to node failures. When a node fails, the method offers the possibility of rebuilding the lost node via the information kept in its neighbors, even though none of the neighbors knows the exact state of the failing node. Moreover, it is shown that the proposed method can achieve consensus and average consensus almost surely, when the agents have arbitrary privacy degrees and a common privacy degree, respectively. To illustrate the theory, two numerical examples are presented.

eess.SY

Consensus with Preserved Privacy against Neighbor Collusion

This paper proposes a privacy-preserving algorithm to solve the average consensus problem based on Shamir's secret sharing scheme, in which a network of agents reach an agreement on their states without exposing their individual state until an agreement is reached. Unlike other methods, the proposed algorithm renders the network resistant to the collusion of any given number of neighbors (even with all neighbors' colluding). Another virtue of this work is that such a method can protect the network consensus procedure from eavesdropping.

cs.CR

Modeling collective behaviors: A moment-based approach

In this work we introduce an approach for modeling and analyzing collective behavior of a group of agents using moments. We represent the group of agents via their distribution and derive a method to estimate the dynamics of the moments. We use this to predict the evolution of the distribution of agents by first computing the moment trajectories and then use this to reconstruct the distribution of the agents. In the latter an inverse problem is solved in order to reconstruct a nominal distribution and to recover the macro-scale properties of the group of agents. The proposed method is applicable for several types of multi-agent systems, e.g., leader-follower systems. We derive error bounds for the moment trajectories and describe how to take these error bounds into account for computing the moment dynamics. The convergence of the moment dynamics is also analyzed for cases with monomial moments. To illustrate the theory, two numerical examples are given. In the first we consider a multi-agent system with interactions and compare the proposed methods for several types of moments. In the second example we apply the framework to a leader-follower problem for modeling pedestrian crowd dynamics.

math.OC

An Intrinsic Approach to Formation Control of Regular Polyhedra for Reduced Attitudes

This paper addresses formation control of reduced attitudes in which a continuous control protocol is proposed for achieving and stabilizing all regular polyhedra (also known as Platonic solids) under a unified framework. The protocol contains only relative reduced attitude measurements and does not depend on any particular parametrization as is usually used in the literature. A key feature of the control proposed is that it is intrinsic in the sense that it does not need to incorporate any information of the desired formation. Instead, the achieved formation pattern is totally attributed to the geometric properties of the space and the designed inter-agent connection topology. Using a novel coordinates transformation, asymptotic stability of the desired formations is proven by studying stability of a constrained nonlinear system. In addition, a methodology to investigate stability of such constrained systems is also presented.

math.OC

Intrinsic Reduced Attitude Formation with Ring Inter-Agent Graph

This paper investigates the reduced attitude formation control problem for a group of rigid-body agents using feedback based on relative attitude information. Under both undirected and directed cycle graph topologies, it is shown that reversing the sign of a classic consensus protocol yields asymptotical convergence to formations whose shape depends on the parity of the group size. Specifically, in the case of even parity the reduced attitudes converge asymptotically to a pair of antipodal points and distribute equidistantly on a great circle in the case of odd parity. Moreover, when the inter-agent graph is an undirected ring, the desired formation is shown to be achieved from almost all initial states.

math.OC

Finite-time attitude synchronization with distributed discontinuous protocols

The finite-time attitude synchronization problem is considered in this paper, where the rotation of each rigid body is expressed using the axis-angle representation. Two discontinuous and distributed controllers using the vectorized signum function are proposed, which guarantee almost global and local convergence, respectively. Filippov solutions and non-smooth analysis techniques are adopted to handle the discontinuities. Sufficient conditions are provided to guarantee finite-time convergence and boundedness of the solutions. Simulation examples are provided to verify the performances of the control protocols designed in this paper.

math.OC

Finite-time attitude synchronization with a discontinuous protocol

A finite-time attitude synchronization problem is considered in this paper where the rotation of each rigid body is expressed using the axis-angle representation. One simple discontinuous and distributed controller using the vectorized signum function is proposed. This controller only involves the sign of the state differences of adjacent neighbors. In order to avoid the singularity introduced by the axis-angular representation, an extra constraint is added to the initial condition. It is proved that for some initial conditions, the control law achieves finite-time attitude synchronization. One simulated example is provided to verify the usage of the control protocol designed in this paper.

math.OC

Intrinsic Tetrahedron Formation of Reduced Attitude

In this paper, formation control for reduced attitude is studied, in which both stationary and rotating regular tetrahedron formation can be achieved and are asymptotically stable under a large family of gain functions in the control. Moreover, by further restriction on the control gain, almost global stability of the stationary formation is obtained. In addition, the control proposed is an intrinsic protocol that only uses relative information and does not need to contain any information of the desired formation beforehand. The constructed formation pattern is totally attributed to the geometric properties of the space and the designed inter-agent connection topology. Besides, a novel coordinates transformation is proposed to represent the relative reduced attitudes in S^2, which is shown to be an efficient approach to reduced attitude formation problems.

math.OC