SearcharxivSearch

arXiv subjects

Yuhang Wang

Publications and source records attributed to Yuhang Wang.

At least 19 recordsLinked to original sources

Full Inseparability and Genuine Multipartite Entanglement Coincide for Finite-Mode Gaussian States

For general mixed states, entanglement across every bipartition need not imply genuine multipartite entanglement (GME), because a biseparable decomposition may switch the separable cut from term to term. We prove that this convex ambiguity disappears for Gaussian states of finitely many bosonic modes. More generally, for any finite family of partitions, a Gaussian density operator in the trace-norm-closed convex class generated by states separable across those partitions is already separable across one fixed partition in the family. Only the target is Gaussian; a valid decomposition may be continuous and may contain arbitrary non-Gaussian states. Thus full inseparability and GME coincide, Gaussian k-separability and k-producibility reduce to fixed-partition tests, and party-wise tensor powers cannot activate GME from a biseparable Gaussian state. The proof combines a spectral selector with a holomorphic rigidity argument that converts one product vector in the square-root range of a Gaussian state into a block-local covariance certificate. The result shows that partition mixing, a generic mixed-state mechanism, adds no new exact finite-mode Gaussian states.

quant-ph

DriveMotion: A Large-Scale Multi-Source Benchmark for Driver Motion Sequence Modeling and Forecasting

Driver motion can provide cues to ongoing behavior, attention, and near-term driving intent. However, most existing driver-centric datasets focus on recognizing predefined driver behaviors from short video clips, while human motion forecasting benchmarks largely target motion outside the vehicle. We introduce DriveMotion, a multi-source benchmark for continuous driver motion forecasting. DriveMotion contains 393 hours of 133-keypoint motion sequences at 10 Hz from 360 drivers, integrating naturalistic driving data, curated public in-cabin videos, and the AIDE dataset into a unified representation with per-joint validity masks and synchronized driving context. Naturalistic driving contains long periods of limited body movement, making uniformly sampled evaluation dominated by persistence and less sensitive to brief but behaviorally meaningful motion. To address this, we use dynamics-anchored evaluation, placing forecasting windows around vehicle maneuvers identified offline from CAN signals without providing CAN to the model at inference. Arm motion in pre-maneuver windows is 3.4x greater than in route-matched stable-driving controls. On these anchored windows, learned models reduce forecasting error over persistence by up to 15%, while maneuver-enriched training improves forecast-derived Part-State F1 by 44% over the zero-motion reference. Training on the full multi-source corpus further reduces forecasting error on held-out web drivers by 38% compared with BATON-only training. DriveMotion provides identity-disjoint splits, fixed evaluation subsets, and reference implementations for reproducible evaluation of continuous driver motion forecasting. The dataset and benchmark are available at https://huggingface.co/datasets/HenryYHW/DriveMotion

cs.CV

Observe Before You Alert: Adaptive Driver Alerting with Vision-Language Models

Driver alerting from dashcam video requires sequential decision-making under partial observability: a system must decide not only whether a scene is risky, but also when the evidence is sufficient to warn. Most existing accident anticipation models output a binary risk score, leaving ambiguous scenes to be handled by thresholding. We propose VLAlert, a vision-language alerting framework that casts warning generation as a tri-action policy over SILENT, OBSERVE, and ALERT. The OBSERVE action acts as an internal evidence-gathering decision that delays uncertain warnings and changes the next observation window, creating a lightweight perception-action loop for adaptive alerting. VLAlert uses Qwen3-VL-4B as a safety-evidence generator and pools hidden states from structured belief spans to form compact representations for danger estimation and policy prediction. We evaluate VLAlert on VLAlert-Bench, a unified per-tick benchmark from four real-world dashcam alert datasets, and further test transfer to held-out naturalistic ADAS takeover clips. On VLAlert-Bench validation, VLAlert achieves the highest deployment-oriented utility among tested baselines, with DAUS 0.4878 compared with 0.4752 for Open-BADAS, and improves AUROC, AP_tick, F1_t, and balanced accuracy from 0.610, 0.176, 0.276, and 0.581 to 0.689, 0.195, 0.297, and 0.648, respectively. On 221 held-out ADAS-TO-Critic clips, VLAlert improves R@5s from 74.2% to 88.7% and F1 from 0.585 to 0.686. These results indicate that adaptive observation and safety-focused VLM representations provide measurable gains for driver-facing alert decisions.

cs.CV

SSP-DMGTimeNet: Physics-Constrained Learning for Spatiotemporal Trajectory Prediction of Vehicle Platoons

Existing car-following prediction methods mainly optimize trajectory accuracy, while rarely considering whether predicted disturbances propagate realistically along a vehicle platoon. This limitation may lead to accurate but string-unstable predictions. We propose SSP-DMGTimeNet, a physics-constrained learning framework for spatiotemporal trajectory prediction of vehicle platoons. The model combines multi-scale temporal representations with cross-vehicle interaction features to capture complex and time-varying platoon dynamics. A propagation-delay-aware causal attention mechanism explicitly models upstream-to-downstream disturbance propagation by learning response delays between adjacent vehicles and accumulating them along the platoon. In addition, time- and frequency-domain string-stability losses relieve disturbance amplification across both adjacent vehicles and arbitrary sub-platoons during training. Experiments on HighD show that SSP-DMGTimeNet achieves an unstable-window rate of 0.65\% for five-vehicle platoons and a maximum head-to-tail amplification of 0.898 on the ground-truth excitation subset, while maintaining competitive trajectory prediction performance. In zero-shot evaluation on NGSIM US-101 and I-80, the model achieves velocity MAEs of 1.316~m/s and 1.252~m/s, with unstable-window rates of 3.90\% and 4.10\%, respectively. These results demonstrate that incorporating platoon-level physical constraints can effectively balance trajectory prediction accuracy and disturbance propagation stability.

cs.AI

Out-of-Distribution Semantic Occupancy Prediction

3D semantic occupancy prediction is crucial for autonomous driving, providing a dense, semantically rich environmental representation. However, existing methods focus on in-distribution scenes, making them susceptible to Out-of-Distribution (OoD) objects and long-tail distributions, which increase the risk of undetected anomalies and misinterpretations, posing safety hazards. To address these challenges, we introduce the task of Out-of-Distribution Semantic Occupancy Prediction, targeting OoD detection in 3D voxel space. To fill dataset gaps, we propose Realistic Anomaly Augmentation that injects synthetic anomalies while preserving realistic spatial and occlusion patterns, enabling the creation of two datasets: VAA-KITTI and VAA-KITTI-360. We then propose OccOoD, a novel framework that integrates OoD detection into 3D semantic occupancy prediction, which uses Cross-Space Semantic Refinement (CSSR) to refine semantic predictions from complementary voxel and BEV representations, improving OoD detection. Experimental results demonstrate that OccOoD achieves an AuROC of 65.50% and an AuPRCr of 31.83% within a 1.2m radius, while maintaining competitive semantic occupancy prediction accuracy, significantly improving detection sensitivity for unknown obstacles, and validating strong generalization in real-world urban driving scenes. The established datasets and source code will be made publicly available at https://github.com/7uHeng/OccOoD.

cs.CV

Scale-CDA: A Scalable Aftermarket Platform to Democratize Cooperative Driving Automation in Production Cars

Scaling cooperative driving automation (CDA) to production passenger vehicles requires an affordable retrofit platform that can accommodate heterogeneous OEM Controller Area Network (CAN) signals and advanced driver-assistance system (ADAS) commands. Scale-CDA addresses this challenge by building on OpenDBC and openpilot, which provide vehicle interfaces and Level-2 automation support for more than 300 car models. The proposed open-hardware and open-software prototype integrates commodity edge computing, camera sensing, Wi-Fi, cellular communication, and a CAN adapter for less than $1,000. It exchanges telemetry and cooperative messages through MQTT over Wi-Fi~6 or LTE, avoiding costly DSRC or C-V2X sidelink radios. In moving-vehicle experiments, Wi-Fi~6 achieved a mean round-trip time of 5.25~ms and mean negotiated physical-layer rates of 98.51~Mb/s for transmission and 109.17~Mb/s for reception, with approximately 2% of observations exceeding 50~ms. These results demonstrate feasibility for non-safety-critical CDA applications. Scale-CDA also introduces a Generative AI interface that combines camera observations, CAN data, and connectivity messages through the Model Context Protocol. The model generates semantic message intents and structured MetaActions, while deterministic adapters validate and encode cooperative messages and map admissible actions to existing Level-2 functions without allowing direct actuator control. An on-road construction-zone demonstration validates driver advisory, message generation, and speed-related MetaAction pathways. Scale-CDA provides a reproducible platform for connectivity and GenAI research on supported production vehicles.

cs.ET

Analytical Framework for Evaluating Traffic Capacity Impacts of Electric Vehicles' Regenerative Braking Dynamics

Regenerative braking (RB) significantly influences electric vehicle (EV) car-following (CF) dynamics, yet traditional traffic-flow models rarely capture these effects. We introduce a comprehensive empirical dataset comprising 197.5 hours of driving data from 25 drivers across eight EV models to systematically investigate regen-induced CF behaviors. Two primary CF patterns emerge: (i) steady-state scenarios where EVs use regenerative braking and subsequently re-accelerate to equilibrium speeds with larger spacing, and (ii) dynamic scenarios involving lead oscillations, characterized by a distinctive three-phase CF process-regenerative deceleration, transitional plateau, and rapid re-acceleration. The paper's main contribution is the development of an analytical framework that models these EV-specific CF behaviors and quantifies their impacts on traffic capacity. We derive closed-form expressions for the established $η$ function from the literature, explicitly quantifying EV driving deviations from stable CF defined by Newell's CF model and assessing their implications for roadway capacity. Validation against empirical data and simulation confirm the model's accuracy ($R^2=0.96$) in replicating real-world $η$ trajectories. Sensitivity analyses demonstrate that increased RB intensity, prolonged transitions, and shorter reaction delays significantly raise values and cumulative capacity losses. These findings highlight a clear trade-off between enhanced energy recovery through RB and reduced traffic efficiency, providing critical insights for EV-aware traffic modeling, control strategies, and transportation policy.

physics.soc-ph

A 3D VTI factored eikonal solver using six-tetrahedron pyramidal stencil

Accurate traveltime computation for the eikonal equation is essential in seismic applications such as tomography and migration. The fast sweeping method (FSM) is widely used because of its unconditional stability and computational efficiency. We develop a 3D fast sweeping solver for vertical transverse isotropic (VTI) media that combines multiplicative factorization with a six-tetrahedron pyramidal stencil. The factorization removes the point-source singularity, while the stencil improves local accuracy. In the unfactored formulation, the six-tetrahedron finite-difference scheme requires solving only quadratic equations. After factorization, however, the local update for the perturbation factor becomes quartic, and the update systems on oblique stencil faces are substantially more complicated than those on non-oblique faces. To resolve these difficulties, we solve the update equation using Ferrari's method with a robust root-selection strategy and derive complete update formulas for all oblique-face configurations. For horizontally constrained faces, the characteristic constraint reduces to a linear relation by exploiting the structure of the VTI Hamiltonian, so the local system still reduces to a quartic equation. For mixed horizontal-vertical constrained faces, we design a bisection-based iterative solver. Numerical examples show that the proposed method effectively suppresses source-related errors and improves traveltime accuracy.

physics.geo-ph

Large Scale Entanglement Structure Detection in 100-Qubit Systems via Local Joint Measurements

Identifying the entanglement structure of a many-body quantum state, namely how its constituents partition into unentangled blocks, is a central task in quantum information science, yet conventional tomography scales exponentially with system size. Here we introduce a scalable framework that recognizes large-scale entanglement structures directly from local correlation fingerprints. By choosing a representative local Pauli basis that satisfies a boundary-matching condition p_1 = p_R, the entire chain is read out in a single measurement configuration, keeping the measurement effort independent of system size. In noisy simulations, this single-basis protocol classifies GHZ-, W-, and cluster-type structures among 30 candidate partitions with a mean accuracy exceeding 95% for systems of up to 100 qubits. We further validate the protocol on a superconducting quantum processor, where it reliably classifies block structures for systems of up to 13 qubits before noise- and depth-induced degradation sets in at larger sizes. By mapping these failure modes explicitly, our results delineate the boundary of hardware-level scalability and point to a concrete strategy for characterizing entanglement structure on near-term quantum devices.

quant-ph

Physics-Grounded Materials Artificial Intelligence for Reliable Materials Discovery

Artificial intelligence (AI) is transforming materials discovery, yet conventional data-driven approaches often suffer from limited interpretability, poor extrapolation, and inconsistency with physical laws. Since materials behavior is fundamentally governed by thermodynamics, kinetics, electronic structure, transport processes, and operating environments, the next generation of materials intelligence must move beyond correlation-based prediction toward physics-grounded reasoning. In this Perspective, we systematically discuss Physics-Grounded Materials AI (PhysMat AI) as a unifying perspective for integrating physical knowledge into materials intelligence through five complementary roles: physics as prior knowledge, descriptors, constraints, verifiers, and infrastructure. Using representative examples from catalysis, solid-state electrolytes in solid-state battery, and hydrogen-storage materials, we illustrate how physical principles guide data representation, model reasoning, validation workflows, and knowledge management. We further present how AI agents can leverage these physics-aware components to perform mechanism-guided discovery within physically feasible search spaces. Finally, we outline a developmental roadmap from physics-aware AI to physics-reasoning AI and ultimately physics-autonomous AI. Looking forward, materials intelligence should evolve from predictive models toward autonomous scientific systems capable of integrating physical reasoning, multiscale simulations, experimental validation, and continuous knowledge updating for reliable materials discovery.

physics.chem-ph

A Model Merging Approach for Continual MLLM Unlearning

Multimodal large language model (MLLM) unlearning methods have been proposed to remove private, sensitive, or proprietary information from well-trained models. However, most existing MLLM unlearning methods are designed for one-shot requests and fail to adequately address continual scenarios, as repeatedly applying one-shot operations leads to cumulative utility degradation, unlearning rebound, and retention drift. We introduce Merging for Continual Unlearning (MCU), an approach that dynamically merges multiple one-shot unlearning adapters into a unified adapter upon receiving each new unlearning request.Through a leave-one-out merging analysis, we reveal that these unlearning adapters exhibit strong cross-task dependencies. Such dependencies have two contrasting effects: they can facilitate cross-task unlearning transferability, but they can also introduce severe interference that degrades unlearning effectiveness and compromises retained knowledge. To address this challenge, MCU projects the adapters into a shared representation space, preserves their dominant directions, suppresses over-concentrated coordinates, and reconfigures cross-task dependencies to mitigate interference while enhancing transferability. Experiments on ICU-Bench and MLLMU-Bench demonstrate that MCU achieves superior unlearning effectiveness while preserving both retained knowledge and general multimodal utility.

cs.LG

Self-Guided Adaptive Safety Alignment: Synthesizing and Internalizing Guidelines in Reasoning Models

Explicit safety policies can improve reasoning-model safety, but their effective coverage may lag behind evolving jailbreak strategies. We study whether a reasoning model can synthesize and internalize a task-specific safety guideline from a small set of harmful and benign examples. We introduce Self-Guided Adaptive Safety Alignment (SGASA). The model generates a guideline, refines it on its own errors, and selects a version by self-evaluation, which can then be applied in context or distilled into the model for guideline-free inference. Across two adversarial prompt datasets and three Qwen3 scales, in-context guidelines improve a combined safety and non-over-refusal score by 20.0-45.5 points on the 8B and 14B models. Self-evaluation selects the externally best refinement round in five of six settings, while guideline generation and utilization show distinct scaling patterns. Using alignment supervision derived only from WildJailbreak, internalized models retain 13.3-14.1 point gains on WildJailbreak without an inference-time guideline. These results support self-generated guidelines as a useful intermediate representation for low-resource safety adaptation.

cs.CL

ICU-Bench:Benchmarking Continual Unlearning in Multimodal Large Language Models

Privacy deletion requests often arrive sequentially, creating a continual unlearning challenge for deployed multimodal large language models (MLLMs). However, existing benchmarks mainly focus on static or short-sequence settings, offering limited support for evaluating continual privacy deletion on privacy-critical documents. To bridge this gap, we introduce ICU-Bench, an Identity-centric Continual Unlearning benchmark (pronounced ``I see you'') for privacy-critical multimodal documents. ICU-Bench contains 1,000 synthetic privacy-sensitive profiles from medical reports and labor contracts, comprising 9,500 document images, 16,000 question-answer pairs, and 100 sequential forget tasks. We further introduce history-aware evaluation protocols and sequence-aware metrics to assess current forgetting, historical forgetting preservation, retained utility, and stability throughout the unlearning sequence. Experiments with representative unlearning methods on two MLLMs reveal that methods effective on current targets often fail to preserve earlier forgetting or retained capabilities over long sequences. Some methods further obtain low forget accuracy through severe model degradation. These results expose long-horizon failure modes overlooked by conventional evaluations and highlight the need for multimodal unlearning methods explicitly designed for continual privacy deletion. Our code and benchmark are publicly available at: https://github.com/AstorYH/ICU-Bench.

cs.AI

BATON: A Multimodal Benchmark for Bidirectional Automation Transition Observation in Naturalistic Driving

Existing driving automation (DA) systems on production vehicles rely on human drivers to decide when to engage DA while requiring them to remain continuously attentive and ready to intervene. This design demands substantial situational judgment and imposes significant cognitive load, leading to steep learning curves, suboptimal user experience, and safety risks from both over-reliance and delayed takeover. Predicting when drivers hand over control to DA and when they take it back is therefore critical for designing proactive, context-aware HMI, yet existing datasets rarely capture the multimodal context, including road scene, driver state, vehicle dynamics, and route environment. To fill this gap, we introduce BATON, a large-scale naturalistic dataset capturing real-world DA usage across 127 drivers, and 136.6 hours of driving. The dataset synchronizes front-view video, in-cabin video, decoded CAN bus signals, radar-based lead-vehicle interaction, and GPS-derived route context, forming a closed-loop multimodal record around each control transition. We define three benchmark tasks: driving action understanding, handover prediction, and takeover prediction, and evaluate baselines spanning sequence models, classical classifiers, and zero-shot VLMs. Results show that visual input alone is insufficient for reliable transition prediction: front-view video captures road context but not driver state, while in-cabin video reflects driver readiness but not the external scene. Incorporating CAN and route-context signals substantially improves performance over video-only settings, indicating strong complementarity across modalities. We further find takeover events develop more gradually and benefit from longer prediction horizons, whereas handover events depend more on immediate contextual cues, revealing an asymmetry with direct implications for HMI design in assisted driving systems.

cs.HC

DriveDNA: A Large-Scale Multimodal Naturalistic Driving Dataset and Benchmark for Driving Style Identification

Driving style captures stable, driver-specific patterns in how a vehicle is driven. In naturalistic data, however, this signal is hard to isolate because drivers are observed in different vehicles, on different roads, and under different conditions, so models may mistake vehicle- or situation-specific regularities for driver-specific style. We introduce DriveDNA, a large-scale naturalistic dataset and benchmark for personalized driving-style modeling, comprising 4,121 drives from 465 drivers across 115 vehicle models and totaling 975 hours of human-controlled driving at 10 Hz with forward video, collected from community drivers in everyday use. DriveDNA defines driving style as a consistent, driver-specific behavioral pattern in how a vehicle moves under similar conditions. The benchmark evaluates this signal through three core tasks: few-shot driver re-identification, personalized behavior prediction, and condition-matched comparison, and provides behavioral annotations plus 276,248 rule-generated maneuver events across six classes with large-scale human auditing. We evaluate baselines spanning classical descriptors, supervised and self-supervised time-series encoders, multimodal fusion, probabilistic prediction, and zero-shot foundation models under a fixed multi-seed protocol. Learned representations substantially outperform classical descriptors on unseen drivers (AUROC .935 vs. .707) and retain driver-specific information under matched driving conditions, while descriptor performance approaches chance. Video-only models achieve comparable re-identification accuracy but exhibit severe route leakage, showing that strong recognition may arise from contextual shortcuts rather than driving behavior. These findings show that reliable driving-style evaluation must assess both the behavioral value of learned representations and their robustness to vehicle, drive, and condition confounds.

cs.LG

SWE-Pruner Pro: The Coder LLM Already Knows What to Prune

Pruning long context for coding agents has been a vital technology for efficient context management. While existing context pruning methods such as SWE-Pruner realize this by attaching a separate code classifier, we find the agent itself encodes internal representations indicating the relevance of code context when reading tool output. Based on this finding, we propose SWE-Pruner Pro, which prunes tool outputs directly inside the agent. Concretely, a small head turns the agent's own internal representations into a keep-or-prune label for each line, with a length-aware embedding keyed to each tool output's line count. Across two open-weight backbones and four multi-turn benchmarks, SWE-Pruner Pro saves up to 39% of prompt and completion tokens while preserving task quality, with bounded inference overhead. Notably, on MiMo-V2-Flash SWE-Pruner Pro additionally raises the SWE-Bench Verified resolve rate by +3.8% and the long-context Oolong accuracy by +2.2 points.

cs.CL

Dynamic Airspace Management for UAVs in Evolving Urban Environments: Collaborative Coordination and Human Safety

The low-altitude economy is an emerging industry with significant development potential, in which the safety of unmanned aerial vehicle (UAV) operations is a critical challenge. Particularly within complex urban topographies and human-populated environments, UAV airspace management must prioritize collision avoidance and human safety. We propose Pharos, a collaborative multi-UAV airspace management system. Pharos lies between the distributed local perception paradigm and the centralized fine-grained control paradigm. Pharos coordinates the safe parallel execution of UAVs in shared airspace while innovatively accounting for the impact of human fear. Pharos is implemented using the MAPPO algorithm due to its faster convergence and higher rewards than other typical MARL algorithms (HAPPO and HATRPO). To evaluate Pharos, we developed a 3D simulation system using real urban data. Visualization results demonstrate its effective airspace coordination capability. Regarding performance verification, Pharos reduced human fear by 52.72% compared to the benchmark Ipopt. Moreover, we designed spatial entropy as a system evaluation metric to quantify space utilization, which improved performance by 70.82% and 2.03% compared to the benchmarks Ipopt and A-star, respectively. The source code is available at an anonymized repository: https://github.com/pharos-anonymized/source-code.git.

cs.MA

Search for Truth from Reasoning: A Dynamic Representation Editing Framework for Steering LLM Trajectories

Current approaches to enhance Large Language Model (LLM) reasoning, such as Chain-of-Thought and "Wait" prompts, primarily encourage models to think more, yet often fail to guide them toward Truth. While Representation Editing (RepE) offers a intrinsic control, its application to dynamic reasoning trajectories remains underexplored. In this work, we bridge this gap by investigating the geometry of truth within unfolding reasoning chains. We uncover three critical insights: (1) Truth is encoded at the sentence level and is entangled with latent reasoning patterns; (2) Effective intervention follows an Uncertainty Principle and a Decay Effect, requiring localization to early, high-entropy forks; (3) Naive steering vectors suffer from noise, risking collateral damage to correct trajectories. Based on these findings, we propose DynaSteer, a dynamic RepE framework. DynaSteer employs pattern clustering to disentangle reasoning manifolds and utilizes Fisher-LDA to project purified truth. By dynamically monitoring lookahead entropy, it selectively steers and rolls back trajectories only when necessary. Comprehensive experimental results on several MATH benchmark verify the effectiveness of DynaSteer, and experiments on out-of-domain coding tasks further confirm its generalization ability. Our code is publicly available at https://github.com/tianlwang/DynaSteer.

cs.AI