SearcharxivSearch

arXiv subjects

Zhiyong Yu

Publications and source records attributed to Zhiyong Yu.

At least 19 recordsLinked to original sources

Classical solution to second-order Hamilton-Jacobi-Bellman equation and optimal feedback control for linear-convex problem

In this paper, we are concerned with the classical solvability of a class of second-order Hamilton-Jacobi-Bellman equations (HJB equations) arising from stochastic optimal control problems with linear dynamics and uniformly convex cost functionals. By introducing the Hamiltonian system and extending the gradient descent method to a Hilbert space, we prove the existence and uniqueness of the optimal control under the uniform convexity condition. The regularity of the solution to the Hamiltonian system is obtained, including the derivatives with respect to the initial state and the Malliavin derivatives. The connection between the Hamiltonian system and the value function is subsequently proven, enabling us to derive regularity properties of the value function via probabilistic techniques. Finally, by the dynamic programming principle, the value function is verified to be the unique classical solution to the HJB equation and the optimal feedback control is provided. These results generalize the classical linear-quadratic theory and provide a new insight into the regularity of the value function.

math.OC

From Parameter to Representation: A Closed-Form Approach for Controllable Model Merging

Model merging combines expert models for multitask performance but faces challenges from parameter interference. This has sparked recent interest in controllable model merging, giving users the ability to explicitly balance performance trade-offs. Existing approaches employ a compile-then-query paradigm, performing a costly offline multi-objective optimization to enable fast, preference-aware model generation. This offline stage typically involves iterative search or dedicated training, with complexity that grows exponentially with the number of tasks. To overcome these limitations, we shift the perspective from parameter-space optimization to a direct correction of the model's final representation. Our approach models this correction as an optimal linear transformation, yielding a closed-form solution that replaces the entire offline optimization process with a single-step, architecture-agnostic computation. This solution directly incorporates user preferences, allowing a Pareto-optimal model to be generated on-the-fly with complexity that scales linearly with the number of tasks. Experimental results show our method generates a superior Pareto front with more precise preference alignment and drastically reduced computational cost.

cs.LG

Collaborative Scheduling of Time-dependent UAVs,Vehicles and Workers for Crowdsensing in Disaster Response

Frequent natural disasters cause significant losses to human society, and timely, efficient collection of post-disaster environmental information is the foundation for effective rescue operations. Due to the extreme complexity of post-disaster environments, existing sensing technologies such as mobile crowdsensing suffer from weak environmental adaptability, insufficient professional sensing capabilities, and poor practicality of sensing solutions. Therefore, this paper explores a heterogeneous multi-agent online collaborative scheduling algorithm, HoCs-MPQ, to achieve efficient collection of post-disaster environmental information. HoCs-MPQ models collaboration and conflict relationships among multiple elements through weighted undirected graph construction, and iteratively solves the maximum weight independent set based on multi-priority queues, ultimately achieving collaborative sensing scheduling of time-dependent UA Vs, vehicles, and workers. Specifically, (1) HoCs-MPQ constructs weighted undirected graph nodes based on collaborative relationships among multiple elements and quantifies their weights, then models the weighted undirected graph based on conflict relationships between nodes; (2) HoCs-MPQ solves the maximum weight independent set based on iterated local search, and accelerates the solution process using multi-priority queues. Finally, we conducted detailed experiments based on extensive real-world and simulated data. The experiments show that, compared to baseline methods (e.g., HoCs-GREEDY, HoCs-K-WTA, HoCs-MADL, and HoCs-MARL), HoCs-MPQ improves task completion rates by an average of 54.13%, 23.82%, 14.12%, and 12.89% respectively, with computation time for single online autonomous scheduling decisions not exceeding 3 seconds.

cs.MA

Linear-Quadratic Zero-Sum Stochastic Differential Game with Partial Observation

This paper is concerned with a kind of linear-quadratic (LQ, for short) two-person zero-sum stochastic differential game problems with partial observation. We propose the notions of explicit and implicit feedback laws under partial observation. With the help of a class of conditional mean-field stochastic differential equations (CMF-SDEs, for short), the separation principle, filtering techniques, and the method of completion of squares, we construct a saddle point in the form of feedback laws for the two players. Finally, the theoretical results are applied to investigate a duopoly competition problem with partial observation.

math.OC

Indefinite Linear-Quadratic Optimal Control Problems of Backward Stochastic Differential Equations with Partial Information

This paper is concerned with a kind of linear-quadratic (LQ) optimal control problem of backward stochastic differential equation (BSDE) with partial information. The cost functional includes cross terms between the state and control, and the weighting matrices are allowed to be indefinite. Through variational methods and stochastic filtering techniques, we derive the necessary and sufficient conditions for the optimal control, where a Hamiltonian system plays a crucial role. Moreover, to construct the optimal control, we introduce a matrix-valued differential equation and a BSDE with filtering, and establish their solvability under the assumption that the cost functional is uniformly convex. Finally, we present explicit forms of the optimal control and value function.

math.OC

Autonomous Collaborative Scheduling of Time-dependent UAVs, Workers and Vehicles for Crowdsensing in Disaster Response

Natural disasters have caused significant losses to human society, and the timely and efficient acquisition of post-disaster environmental information is crucial for the effective implementation of rescue operations. Due to the complexity of post-disaster environments, existing sensing technologies face challenges such as weak environmental adaptability, insufficient specialized sensing capabilities, and limited practicality of sensing solutions. This paper explores the heterogeneous multi-agent online autonomous collaborative scheduling algorithm HoAs-PALN, aimed at achieving efficient collection of post-disaster environmental information. HoAs-PALN is realized through adaptive dimensionality reduction in the matching process and local Nash equilibrium game, facilitating autonomous collaboration among time-dependent UAVs, workers and vehicles to enhance sensing scheduling. (1) In terms of adaptive dimensionality reduction during the matching process, HoAs-PALN significantly reduces scheduling decision time by transforming a five-dimensional matching process into two categories of three-dimensional matching processes; (2) Regarding the local Nash equilibrium game, HoAs-PALN combines the softmax function to optimize behavior selection probabilities and introduces a local Nash equilibrium determination mechanism to ensure scheduling decision performance. Finally, we conducted detailed experiments based on extensive real-world and simulated data. Compared with the baselines (GREEDY, K-WTA, MADL and MARL), HoAs-PALN improves task completion rates by 64.12%, 46.48%, 16.55%, and 14.03% on average, respectively, while each online scheduling decision takes less than 10 seconds, demonstrating its effectiveness in dynamic post-disaster environments.

cs.MA

Microservice Deployment in Space Computing Power Networks via Robust Reinforcement Learning

With the growing demand for Earth observation, it is important to provide reliable real-time remote sensing inference services to meet the low-latency requirements. The Space Computing Power Network (Space-CPN) offers a promising solution by providing onboard computing and extensive coverage capabilities for real-time inference. This paper presents a remote sensing artificial intelligence applications deployment framework designed for Low Earth Orbit satellite constellations to achieve real-time inference performance. The framework employs the microservice architecture, decomposing monolithic inference tasks into reusable, independent modules to address high latency and resource heterogeneity. This distributed approach enables optimized microservice deployment, minimizing resource utilization while meeting quality of service and functional requirements. We introduce Robust Optimization to the deployment problem to address data uncertainty. Additionally, we model the Robust Optimization problem as a Partially Observable Markov Decision Process and propose a robust reinforcement learning algorithm to handle the semi-infinite Quality of Service constraints. Our approach yields sub-optimal solutions that minimize accuracy loss while maintaining acceptable computational costs. Simulation results demonstrate the effectiveness of our framework.

cs.NI

Infinite Horizon Mean-Field Linear Quadratic Optimal Control Problems with Jumps and the related Hamiltonian Systems

In this work, we focus on an infinite horizon mean-field linear-quadratic stochastic control problem with jumps. Firstly, the infinite horizon linear mean-field stochastic differential equations and backward stochastic differential equations with jumps are studied to support the research of the control problem. The global integrability properties of their solution processes are studied by introducing a kind of so-called dissipation conditions suitable for the systems involving the mean-field terms and jumps. For the control problem, we conclude a sufficient and necessary condition of open-loop optimal control by the variational approach. Besides, a kind of infinite horizon fully coupled linear mean-field forward-backward stochastic differential equations with jumps is studied by using the method of continuation. Such a research makes the characterization of the open-loop optimal controls more straightforward and complete.

math.OC

Collaborative Route Planning of UAVs, Workers and Cars for Crowdsensing in Disaster Response

Efficiently obtaining the up-to-date information in the disaster-stricken area is the key to successful disaster response. Unmanned aerial vehicles (UAVs), workers and cars can collaborate to accomplish sensing tasks, such as data collection, in disaster-stricken areas. In this paper, we explicitly address the route planning for a group of agents, including UAVs, workers, and cars, with the goal of maximizing the task completion rate. We propose MANF-RL-RP, a heterogeneous multi-agent route planning algorithm that incorporates several efficient designs, including global-local dual information processing and a tailored model structure for heterogeneous multi-agent systems. Global-local dual information processing encompasses the extraction and dissemination of spatial features from global information, as well as the partitioning and filtering of local information from individual agents. Regarding the construction of the model structure for heterogeneous multi-agent, we perform the following work. We design the same data structure to represent the states of different agents, prove the Markovian property of the decision-making process of agents to simplify the model structure, and also design a reasonable reward function to train the model. Finally, we conducted detailed experiments based on the rich simulation data. In comparison to the baseline algorithms, namely Greedy-SC-RP and MANF-DNN-RP, MANF-RL-RP has exhibited a significant improvement in terms of task completion rate.

cs.AI

Exact Controllability for Mean-Field Type Linear Game-Based Control Systems

Motivated by the self-pursuit of controlled objects, we consider the exact controllability of a linear mean-field type game-based control system (MF-GBCS, for short) generated by a linear-quadratic (LQ, for short) Nash game. A Gram-type criterion for the general timevarying coefficients case and a Kalman-type criterion for the special time-invariant coefficients case are obtained. At the same time, the equivalence between the exact controllability of this MF-GBCS and the exact observability of a dual system is established. Moreover, an admissible control that can steer the state from any initial vector to any terminal random variable is constructed in closed form.

math.OC

Dynamic surface tension of the pure liquid-vapor interface subjected to the cyclic loads

We demonstrate a methodology for computationally investigating the mechanical response of a pure molten lead surface system to the lateral mechanical cyclic loads and try to answer the question: how dose the dynamically driven liquid surface system follow the classical physics of the elastic-driven oscillation? The steady-state oscillation of the dynamic surface tension under cyclic load, including the excitation of high frequency vibration mode at different driving frequencies and amplitudes, was compared with the classical theory of single-body driven damped oscillator. Under the highest studied frequency (50 GHz) and amplitude (5%) of the load, the increase of the (mean value) dynamic surface tension could reach ~5%. The peak and trough values of the instantaneous dynamic surface tension could reach (up to) 40% increase and (up to) 20% decrease compared to the equilibrium surface tension, respectively. The extracted generalized natural frequencies and the generalized damping constants seem to be intimately related to the intrinsic timescales of the atomic temporal-spatial correlation functions of the liquids both in the bulk region and in the outermost surface layers. These insights uncovered could be helpful for quantitative manipulation of the liquid surface tension using ultrafast shockwaves or laser pulses.

physics.comp-ph

Atomistic characterization of the SiO2 high-density liquid/low-density liquid interface

The equilibrium silica liquid-liquid interface between the high-density liquid (HDL) phase and the low-density liquid (LDL) phase is examined using molecular-dynamics simulation. The structure, thermodynamics, and dynamics within the interfacial region are characterized in detail and compared with previous studies on the liquid-liquid phase transition (LLPT) in bulk silica, as well as traditional crystal-melt interfaces. We find that the silica HDL-LDL interface exhibits a spatial fragile-to-strong transition across the interface. Calculations of dynamics properties reveal three types of dynamical heterogeneity hybridizing within the silica HDL-LDL interface. We also observe that as the interface is traversed from HDL to LDL, the Si/O coordination number ratio jumps to an unexpectedly large value, defining a thin region of the interface where HDL and LDL exhibit significant mixing. In addition, the LLPT phase coexistence is interpreted in the framework of the traditional thermodynamics of alloys and phase equilibria.

cond-mat.mtrl-sci

Ultrafast modulation of the molten metal surface tension under femtosecond laser irradiation

We predict ultrafast modulation of the pure molten metal surface stress fields under the irradiation of the single femtosecond laser pulse through the two-temperature model molecular-dynamics simulations. High-resolution and precision calculations are used to resolve the ultrafast laser-induced anisotropic relaxations of the pressure components on the time-scale comparable to the intrinsic liquid density relaxation time. The magnitudes of the dynamic surface tensions are found being modulated sharply within picoseconds after the irradiation, due to the development of the nanometer scale non-hydrostatic regime behind the exterior atomic layer of the liquid surfaces. The reported novel regulation mechanism of the liquid surface stress field and the dynamic surface tension hints at levitating the manipulation of liquid surfaces, such as ultrafast steering the surface directional transport and patterning.

cond-mat.mtrl-sci

Mean-Field Type FBSDEs under Domination-Monotonicity Conditions and Application to LQ Problems

This paper is concerned with a class of mean-field type coupled forward-backward stochastic differential equations (MF-FBSDEs, for short), in which the coupling appears in integral terms, terminal terms, and initial terms. Inspired by various mean-field type linear-quadratic (MF-LQ,for short) optimal control problems, we proposed a type of randomized domination-monotonicity conditions, under which and the usual Lipschitz condition, we obtain a well-posedness result on MF-FBSDEs in the sense of square integrability including the unique solvability, an estimate of the solution, and the related continuous dependence property of the solution on the coefficients.The result of MF-FBSDEs in turn extends MF-LQ problems in the literature to a general situation where the initial states or the terminal states are also controlled at the same time, and gives explicit expressions of the related unique optimal controls.

math.OC

Object Tracking by Least Spatiotemporal Searches

Tracking a car or a person in a city is crucial for urban safety management. How can we complete the task with minimal number of spatiotemporal searches from massive camera records? This paper proposes a strategy named IHMs (Intermediate Searching at Heuristic Moments): each step we figure out which moment is the best to search according to a heuristic indicator, then at that moment search locations one by one in descending order of predicted appearing probabilities, until a search hits; iterate this step until we get the object's current location. Five searching strategies are compared in experiments, and IHMs is validated to be most efficient, which can save up to 1/3 total costs. This result provides an evidence that "searching at intermediate moments can save cost".

cs.AI

Indefinite Mean-Field Type Linear-Quadratic Stochastic Optimal Control Problems

This paper focuses on indefinite stochastic mean-field linear-quadratic (MF-LQ, for short) optimal control problems, which allow the weighting matrices for state and control in the cost functional to be indefinite. The solvability of stochastic Hamiltonian system and Riccati equations is presented under both positive definite case and indefinite case. The optimal controls in open-loop form and closed-loop form are obtained, respectively. Moreover, the dynamic mean-variance problem can be solved within the framework of the indefinite MF-LQ problem. Other two examples shed light on the theoretical results established.

math.OC

Linear Quadratic Stochastic Optimal Control Problems with Operator Coefficients: Open-Loop Solutions

An optimal control problem is considered for linear stochastic differential equations with quadratic cost functional. The coefficients of the state equation and the weights in the cost functional are bounded operators on the spaces of square integrable random variables. The main motivation of our study is linear quadratic optimal control problems for mean-field stochastic differential equations. Open-loop solvability of the problem is investigated, which is characterized as the solvability of a system of linear coupled forward-backward stochastic differential equations (FBSDE, for short) with operator coefficients. Under proper conditions, the well-posedness of such an FBSDE is established, which leads to the existence of an open-loop optimal control. Finally, as an application of our main results, a general mean-field linear quadratic control problem in the open-loop case is solved.

math.OC

A Game Problem for Heat Equation

In this paper, we consider a two-person game problem governed by a linear heat equation. The existence of Nash equilibrium for this problem is considered. Moreover, the bang-bang property of Nash equilibrium is discussed.

math.OC