SearcharxivSearch

arXiv subjects

Yongjun Shen

Publications and source records attributed to Yongjun Shen.

3 recordsLinked to original sources

FactorDrive: Adaptive Multi-Step Reasoning Driven by Planning-Critical Factors for End-to-End Autonomous Driving

Vision-language models (VLMs) have advanced scene understanding and enabled explicit reasoning in end-to-end autonomous driving. However, existing methods insufficiently integrate spatial-physical evidence into planning reasoning, while reasoning adaptation remains coarse-grained and falls short of scene-specific planning demands. Furthermore, reasoning-path optimization for higher planning quality remains largely unexplored in autonomous-driving post-training. To address these limitations, we propose FactorDrive, an end-to-end autonomous driving framework for adaptive multi-step reasoning driven by planning-critical factors (PCFs). We first perform large-scale driving-domain instruction tuning to establish foundational driving knowledge. Building on this foundation, we construct PCF-CoT, a chain-of-thought (CoT) dataset that grounds planning reasoning in trajectory-relevant spatial-physical evidence and organizes reasoning around scene-specific PCFs, enabling the composition and depth of reasoning paths to adapt to different planning demands. We further introduce Quality Search-Guided Group Relative Policy Optimization (QS-GRPO), which guides Monte Carlo Tree Search (MCTS) with trajectory-level planning rewards to discover reasoning paths with higher planning quality and uses the resulting responses to optimize the policy through GRPO, thereby improving trajectory planning performance. Extensive experiments on both open-loop (nuScenes) and closed-loop-oriented (NAVSIM) benchmarks demonstrate that FactorDrive achieves state-of-the-art planning performance.

cs.RO

LLaVAShield: Safeguarding Multimodal Multi-Turn Dialogues in Vision-Language Models

As Vision-Language Models (VLMs) move into interactive, multi-turn use, safety concerns intensify for multimodal multi-turn dialogue, which is characterized by concealment of malicious intent, contextual risk accumulation, and cross-modal joint risk. These characteristics limit the effectiveness of content moderation approaches designed for single-turn or single-modality settings. To address these limitations, we first construct the Multimodal Multi-turn Dialogue Safety (MMDS) dataset, comprising 4,484 annotated dialogues and a comprehensive risk taxonomy with 8 primary and 60 subdimensions. As part of MMDS construction, we introduce Multimodal Multi-turn Red Teaming (MMRT), an automated framework for generating unsafe multimodal multi-turn dialogues. We further propose LLaVAShield, which audits the safety of both user inputs and assistant responses under specified policy dimensions in multimodal multi-turn dialogues. Extensive experiments show that LLaVAShield significantly outperforms state-of-the-art VLMs and existing content moderation tools while demonstrating strong generalization and flexible policy adaptation. Additionally, we analyze vulnerabilities of mainstream VLMs to harmful inputs and evaluate the contribution of key components, advancing understanding of safety mechanisms in multimodal multi-turn dialogues.

cs.CV

On the Melnikov method for fractional-order systems

This paper is dedicated to clarifying and introducing the correct application of Melnikov method in fractional dynamics. Attention to the complex dynamics of hyperbolic orbits and to fractional calculus can be, respectively, traced back to Poincar\'es attack on the three-body problem a century ago and to the early days of calculus three centuries ago. Nowadays, fractional calculus has been widely applied in modeling dynamic problems across various fields due to its advantages in describing problems with non-locality. Some of these models have also been confirmed to exhibit hyperbolic orbit dynamics, and recently, they have been extensively studied based on Melnikov method, an analytical approach for homoclinic and heteroclinic orbit dynamics. Despite its decade-long application in fractional dynamics, there is a universal problem in these applications that remains to be clarified, i.e., defining fractional-order systems within finite memory boundaries leads to the neglect of perturbation calculation for parts of the stable and unstable manifolds in Melnikov analysis. After clarifying and redefining the problem, a rigorous analytical case is provided for reference. Unlike existing results, the Melnikov criterion here is derived in a globally closed form, which was previously considered unobtainable due to difficulties in the analysis of fractional-order perturbations characterized by convolution integrals with power-law type singular kernels. Finally, numerical methods are employed to verify the derived Melnikov criterion. Overall, the clarification for the problem and the presented case are expected to provide insights for future research in this topic.

nlin.CD