SearcharxivSearch

arXiv subjects

Yougang Bian

Publications and source records attributed to Yougang Bian.

6 recordsLinked to original sources

Safety-Regulated Transfer Reinforcement Learning with Adaptive Teacher Guidance

We propose Safety-Regulated Adaptive Transfer Reinforcement Learning (SRATRL), a teacher--student framework that combines safety-triggered intervention, safety-adaptive value shaping, and policy-compatibility-based optimization for efficient target-domain adaptation. First, a safety-triggered closed-loop intervention strategy is developed that activates teacher guidance according to the instantaneous safety cost and adaptively adjusts the intervention threshold based on the student policy's recent safety performance, thereby providing timely safety supervision while progressively restoring student autonomy as its safety improves. Next, a safety-adaptive teacher-guided value-shaping scheme is introduced, in which a teacher-consistency signal is incorporated into the critic target, and its contribution is dynamically regulated by the safety-constraint multiplier, enabling stronger teacher guidance under elevated safety risks and gradually weakening such guidance as the safety constraint is better satisfied. In addition, a teacher-student policy-compatibility weighting approach is proposed to alleviate the adverse optimization effects caused by policy mismatch. It reweights teacher-intervened transitions according to the relative likelihood of the executed action under the teacher and student policies, thereby improving policy-update stability. Experimental results demonstrate that compared with a Proximal Policy Optimization with Lagrangian constraint baseline, the proposed method improves the average velocity by 6.90%, and reduces the crash ratio by 75.00%. These results demonstrate that the proposed method can reduce safety costs while maintaining competitive task efficiency.

cs.LG

Multimodal Classification Network Guided Trajectory Planning for Four-Wheel Independent Steering Autonomous Parking Considering Obstacle Attributes

Four-wheel Independent Steering (4WIS) vehicles have attracted increasing attention for their superior maneuverability. Human drivers typically choose to cross or drive over the low-profile obstacles (e.g., plastic bags) to efficiently navigate through narrow spaces, while existing planners neglect obstacle attributes, leading to suboptimal efficiency or planning failures. To address this issue, we propose a novel multimodal trajectory planning framework that employs a neural network for scene perception, combines 4WIS hybrid A* search to generate a warm start, and utilizes an optimal control problem (OCP) for trajectory optimization. Specifically, a multimodal perception network fusing visual information and vehicle states is employed to capture semantic and contextual scene understanding, enabling the planner to adapt the strategy according to scene complexity (hard or easy task). For hard tasks, guided points are introduced to decompose complex tasks into local subtasks, improving the search efficiency. The multiple steering modes of 4WIS vehicles, Ackermann, diagonal, and zero-turn, are also incorporated as kinematically feasible motion primitives. Moreover, a hierarchical obstacle handling strategy, which categorizes obstacles as "non-traversable", "crossable", and "drive-over", is incorporated into the node expansion process, explicitly linking obstacle attributes to planning actions to enable efficient decisions. Furthermore, to address dynamic obstacles with motion uncertainty, we introduce a probabilistic risk field model, constructing risk-aware driving corridors that serve as linear collision constraints in OCP. Experimental results demonstrate the proposed framework's effectiveness in generating safe, efficient, and smooth trajectories for 4WIS vehicles, especially in constrained environments.

cs.RO

Observer-Based Distributed Model Predictive Control for String-Stable Multi-vehicle Systems with Markovian Switching Topology

Switching communication topologies can cause instability in vehicle platoons, as vehicle information may be lost during the dynamic switching process. This highlights the need to design a controller capable of maintaining the stability of vehicle platoons under dynamically changing topologies. However, capturing the dynamic characteristics of switching topologies and obtaining complete vehicle information for controller design while ensuring stability remains a significant challenge. In this study, we propose an observer-based distributed model predictive control (DMPC) method for vehicle platoons under directed Markovian switching topologies. Considering the stochastic nature of the switching topologies, we model the directed switching communication topologies using a continuous-time Markov chain. To obtain the leader vehicle's information for controller design, we develop a fully distributed adaptive observer that can quickly adapt to the randomly switching topologies, ensuring that the observed information is not affected by the dynamic topology switches. Additionally, a sufficient condition is derived to guarantee the mean-square stability of the observer. Furthermore, we construct the DMPC terminal update law based on the observer and formulate a string stability constraint based on the observed information. Numerical simulations demonstrate that our method can reduce tracking errors while ensuring string stability.

eess.SY

Reachable Sets-based Trajectory Planning Combining Reinforcement Learning and iLQR

The driving risk field is applicable to more complex driving scenarios, providing new approaches for safety decision-making and active vehicle control in intricate environments. However, existing research often overlooks the driving risk field and fails to consider the impact of risk distribution within drivable areas on trajectory planning, which poses challenges for enhancing safety. This paper proposes a trajectory planning method for intelligent vehicles based on the risk reachable set to further improve the safety of trajectory planning. First, we construct the reachable set incorporating the driving risk field to more accurately assess and avoid potential risks in drivable areas. Then, the initial trajectory is generated based on safe reinforcement learning and projected onto the reachable set. Finally, we introduce a trajectory planning method based on a constrained iterative quadratic regulator to optimize the initial solution, ensuring that the planned trajectory achieves optimal comfort, safety, and efficiency. We conduct simulation tests of trajectory planning in high-speed lane-changing scenarios. The results indicate that the proposed method can guarantee trajectory comfort and driving efficiency, with the generated trajectory situated outside high-risk boundaries, thereby ensuring vehicle safety during operation.

eess.SY

Distributed Model Predicted Control of Multi-agent Systems with Applications to Multi-vehicle Cooperation

This paper proposes a distributed model predicted control (DMPC) approach for consensus control of multi-agent systems (MASs) with linear agent dynamics and bounded control input constraints. Within the proposed DMPC framework, each agent exchanges assumed state trajectories with neighbors and solves a local open-loop optimization problem to obtain the optimal control input. In the optimization problem, a discrete-time consensus protocol is introduced into update law design for assumed terminal states, with which asymptotic consensus of assumed terminal states and recursive feasibility are rigorously proved. Together with the optimal cost function, an infinite series of cost-to-go functions is introduced into the design of a Lyapunov function, with which closed-loop asymptotic consensus is finally proved. Two applications including cooperation of autonomous underwater vehicles (AUVs) and connected and automated vehicles (CAVs) are used to validate the effectiveness of the proposed DMPC approach.

eess.SY

Cooperative Control of Heterogeneous Connected Vehicles with Directed Acyclic Interactions

Cooperation of multiple connected vehicles has the potential to benefit the road traffic greatly. In this paper, we consider analysis and synthesis problems of the cooperative control of a platoon of heterogeneous connected vehicles with directed acyclic interactions (characterized by directed acyclic graphs). In contrast to previous works that view heterogeneity as a type of uncertainty, this paper directly takes heterogeneity into account in the problem formulation, allowing us to develop a deeper understanding of the influence of heterogeneity on the collective behavior of a platoon of connected vehicles. Our major strategies include an application of the celebrated internal model principle and an exploitation of lower-triangular structures for platoons with directed acyclic interactions. The major findings include: 1) we explicitly highlight the tracking ability of heterogeneous platoons, showing that the followers can only track the leader's spacing and velocity; 2) we analytically derive a stability region of feedback gains for platoons with directed acyclic interactions; 3) and consequently we propose a synthesis method based on the solution to an algebraic Riccati equation that shares the dimension of single vehicle dynamics. Numerical experiments are carried out to validate the effectiveness of our results.

math.OC