SearcharxivSearch

arXiv subjects

Yang Zhu

Publications and source records attributed to Yang Zhu.

6 recordsLinked to original sources

Critic-Free Policy Iteration for Continuous-Time Zero-Sum Games: A Policy-Space Riccati Approach

This paper develops a critic-free policy iteration (PI) method for continuous-time linear zero-sum games. The central idea is to characterize the saddle-point policies directly in the joint policy space, rather than treating the quadratic value matrix as an iterative variable. A policy game Riccati equation (PGRE) is introduced whose unknowns are policy gains only. Its solutions are shown to be in one-to-one correspondence with the symmetric solutions of the game algebraic Riccati equation. Based on a stabilizing anchor, PI is performed directly in the actor space. The actor-space Jacobian is nonsingular at every stabilizing policy, and the resulting policy sequence coincides with that of simultaneous PI. For unknown dynamics, a data-driven algorithm uses a single batch of data and nullspace projection of endpoint increments to eliminate the value matrix, yielding an actor-only regression. A necessary and sufficient rank condition for unique policy recovery is established and shown to be iteration-invariant, demonstrating that critic identifiability is unnecessary. A power systems frequency-regulation example verifies convergence and policy recovery, while scalability tests demonstrate substantial reductions in computational and memory requirements.

eess.SY

Data-Driven Critic-Free Policy Iteration for Continuous-Time Linear Quadratic Regulation

For continuous-time linear quadratic regulation with unknown system matrices, data-driven off-policy policy iteration typically estimates the value matrix and the improved feedback gain through a joint critic--actor regression. We show that the critic is not needed in the policy-improvement step. The key is to anchor the Riccati equation at a known stabilizing gain and express optimality as a policy-space residual. An endpoint null-space projection then removes the value-matrix term from the integral data equation. This yields a critic-free, actor-only least-squares update computed directly from input-state data. Under a verifiable projected rank condition, the resulting data equation is equivalent to the policy-space residual equation, and each update coincides with the Kleinman iteration. Thus, the stabilizing and convergence properties of Kleinman iteration are retained without a critic regression. We further show that the conventional off-policy full-rank condition decomposes into an endpoint critic rank condition and a projected actor rank condition. The proposed method removes the rank requirement needed for critic identification while retaining the one needed for policy improvement. The repeated least-squares dimension is reduced from $n(n+1)/2+mn$ to $mn$. Finally, comparative simulations validate the effectiveness of the proposed algorithm.

eess.SY

When classical predictors fail: an exact network outflowpredictor for heterogeneous multi-agent systems with communication delays

This paper develops an information outflow predictor-feedback framework for the exact compensation of constant but nonuniform communication delays in the output synchronization problem of discrete-time heterogeneous multi-agent systems. The delays considered here act exclusively on communication channels between neighboring agents, rather than on the plant state, input or output, and thus fall outside the scope of classical predictor-feedback formulations. We show that standard approaches based on time inversion and variational construction of predictor states are insufficient, since even for simple directed acyclic communication graphs the required prediction horizon exceeds the delay window. We further illustrate that exact predictor-based compensation is achievable only in the discrete-time setting, while in continuous-time configuration the predictor state becomes non-computable due to the need for a continuum inaccessible future neighboring-agent information. Motivated by these observations, we propose a distributed prediction architecture in which a prediction-based distributed observer exactly reconstructs the flow of information exchanged between agents across the network. Furthermore, the layer-by-layer information flow eliminates the effect of communication delays after a finite number of steps. Based on this structure, we design prediction-based distributed state-feedback and dynamic output-feedback controllers by combining standard feedback designs with prediction-based distributed observers serving as feedforward components, such that the output of each agent asymptotically tracks the trajectory generated by the exosystem, thereby ensuring output synchronization.

eess.SY

Three-dimentional reconstruction of complex, dynamic population canopy architecture for crops with a novel point cloud completion model: A case study in Brassica napus rapeseed

Quantitative descriptions of the complete canopy architecture are essential for accurately evaluating crop photosynthesis and yield performance to guide ideotype design. Although various sensing technologies have been developed for three-dimensional (3D) reconstruction of individual plants and canopies, they failed to obtain an accurate description of canopy architectures due to severe occlusion among complex canopy architectures. We proposed an effective method for 3D reconstruction of complex, dynamic population canopy architecture for rapeseed crops with a novel point cloud completion model. A complete point cloud generation framework was developed for automated annotation of the training dataset by distinguishing surface points from occluded points within canopies. The crop population point cloud completion network (CP-PCN) was then designed with a multi-resolution dynamic graph convolutional encoder (MRDG) and a point pyramid decoder (PPD) to predict occluded points. To further enhance feature extraction, a dynamic graph convolutional feature extractor (DGCFE) module was proposed to capture structural variations over the whole rapeseed growth period. The results demonstrated that CP-PCN achieved chamfer distance (CD) values of 3.35 cm -4.51 cm over four growth stages, outperforming the state-of-the-art transformer-based method (PoinTr). Ablation studies confirmed the effectiveness of the MRDG and DGCFE modules. Moreover, the validation experiment demonstrated that the silique efficiency index developed from CP-PCN improved the overall accuracy of rapeseed yield prediction by 11.2% compared to that of using incomplete point clouds. The CP-PCN pipeline has the potential to be extended to other crops, significantly advancing the quantitatively analysis of in-field population canopy architectures.

cs.CV

Adaptive Compatible Performance Control for Spacecraft Attitude Control under Motion Constraints with Guaranteed Accuracy

This paper focuses on the problem of spacecraft attitude control in the presence of time-varying parameter uncertainties and multiple constraints, accounting for angular velocity limitation, performance requirements, and input saturation. To tackle this problem, we propose a modified framework called Compatible Performance Control (CPC), which integrates the Prescribed Performance Control (PPC) scheme with a contradiction detection and alleviation strategy. Firstly, by introducing the Zeroing Barrier Function (ZBF) concept, we propose a detection strategy to yield judgment on the compatibility between the angular velocity constraint and the performance envelope constraint. Subsequently, we propose a projection operator-governed dynamical system with a varying upper bound to generate an appropriate bounded performance envelope-modification signal if a contradiction exists, thereby alleviating the contradiction and promoting compatibility within the system. Next, a dynamical filter technique is introduced to construct a bounded reference velocity signal to address the angular velocity limitation. Furthermore, we employ a time-varying gain technique to address the challenge posed by time-varying parameter uncertainties, further developing an adaptive strategy that exhibits robustness on disturbance rejection. By utilizing the proposed CPC scheme and time-varying gain adaptive strategy, we construct an adaptive CPC controller, which guarantees the ultimate boundedness of the system, and all constraints are satisfied simultaneously during the whole control process. Finally, numerical simulation results are presented to show the effectiveness of the proposed framework.

eess.SY

Iterative Mode-Dropping for the Sum Capacity of MIMO-MAC with Per-Antenna Power Constraint

We propose an iterative mode-dropping algorithm that optimizes input signals to achieve the sum capacity of the MIMO-MAC with per-antenna power constraint. The algorithm successively optimizes each user's input covariance matrix by applying mode-dropping to the equivalent single-user MIMO rate maximization problem. Both analysis and simulation show fast convergence. We then use the algorithm to briefly highlight the difference in MIMO-MAC capacities under sum and per-antenna power constraints.

cs.IT