SearcharxivSearch

arXiv subjects

Zhenhao Wang

Publications and source records attributed to Zhenhao Wang.

11 recordsLinked to original sources

Drones as Annotators: Amodal 3D Auto-Labeling for Ground LiDAR with Aerial Priors

Scaling perception data in autonomous driving is hindered by manual 3D bounding box annotation, a costly and labor intensive process requiring substantial domain expertise. Existing auto-labeling methods reduce this burden, but most of them rely on onboard sensors, where a single ground-level viewpoint yields occluded and sparse observations and inaccurate object geometry. We introduce DAA (Drones as Annotators), a drone-assisted training-free framework for amodal 3D auto-labeling. DAA augments ground LiDAR with an aerial agent that provides less occluded vehicle detections with continuous tracks. Ground LiDAR, in turn, provides precise metric 3D geometry unavailable from aerial imagery alone. DAA exploits the two complementary yet heterogeneous modalities through a two-stage framework: (i) air-ground coordinate alignment unifies aerial detections and ground LiDAR scans into a shared map frame and bootstraps coarse 3D bounding boxes from the aligned aerial detections; (ii) EM-like amodal box refinement alternates between identifying the vehicle's foreground points and updating the coarse box from the identified foreground and aerial priors. We evaluate DAA on an air-ground cooperative perception dataset, where it consistently outperforms existing auto-labeling baselines. LiDAR detectors trained on DAA-generated labels remain competitive with those trained on manual annotations at moderate IoU thresholds. Evaluation against a reference vehicle with known poses and dimensions further confirms its accuracy. We also demonstrate that DAA transfers, without parameter retuning, from onboard to stationary roadside LiDAR and to multi-agent fused point clouds. Code and dataset will be released upon publication.

eess.IV

MuRA: Multi-Rank Adaptation for Efficient and Effective Test-Time Vision-Language Generalization

Vision-language models exhibit remarkable zero-shot capabilities but suffer significant performance degradation under distribution shifts. While test-time adaptation (TTA) via Low-Rank Adaptation offers a parameter-efficient solution, we identify a fundamental bottleneck in current methods: the reliance on static rank configurations. Because visual inputs inherently possess varying information densities, a fixed rank forces an inevitable optimization compromise, leading to underfitting on complex scenes and overfitting on simple ones. To bridge this gap, we propose Multi-Rank Adaptation (MuRA), a novel framework that dynamically selects and fuses adaptation modules of varying capacities based on token-level visual complexity. MuRA synergizes Multi-Rank Orthogonal Decomposition to provide a superior, knowledge-preserving initialization, and Unified Component Fusion with Continuous Router Updating to sustainably learn semantic-to-rank mappings. Furthermore, we provide rigorous theoretical justifications mathematically proving the necessity and gradient stability of this adaptive mechanism. Crucially, MuRA's dynamic design uniquely thrives at the deepest visual layer, capitalizing on the shortest gradient backpropagation path. Extensive experiments demonstrate that MuRA achieves state-of-the-art accuracy across extensive domain generalization and cross-dataset benchmarks while significantly reducing both computational and memory overhead.

cs.CV

Confinement-controlled pattern selection in a finite population-imbalanced dipolar Bose-Einstein condensate

We study the ground-state density patterns of a population-imbalanced two-component dipolar Bose-Einstein condensate confined in a circular quasi-two-dimensional box. Using a mean-field model, we map out phase diagrams as functions of the axial confinement, interaction imbalance, and population ratio. The system supports a rich sequence of stationary morphologies, including a nearly uniform pancake state, pancake-droplet and ring-droplet coexistence states, droplet arrays, and concentric rings. These patterns show a close structural correspondence to microphase-separated morphologies in diblock-copolymer systems, with the population imbalance acting as an effective volume fraction that selects the pattern topology. Analysis of the density profiles and structure factors reveals that the modulated states possess an intrinsic nonzero characteristic wave vector, which remains essentially unchanged when the box size is varied. We also find that the characteristic pattern spacing scales linearly with the axial confinement length, indicating that the transverse thickness of the condensate controls the effective in-plane length scale. In a finite circular box, this smooth scaling is interrupted by discrete steps, reflecting geometric frustration and the integer locking of the number of rings or droplets. Our results show that box-trapped dipolar mixtures provide a controllable platform for studying finite-size pattern selection and nonlocal microphase formation in quantum fluids.

cond-mat.quant-gas

PCASim: Promptable Closed-loop Adversarial Simulation for Urban Traffic Environment

Real-world autonomous driving, particularly in urban environments with numerous corner cases, requires rigorous testing to ensure product safety and robustness. However, few studies have explored integrating adversarial scenario generation with the training of safety agents in closed-loop testing, enabling efficient co-evolution and mutual enhancement of both. To address this challenge, an adversarial behavior knowledge repository is constructed by applying rule-based filtering to an open-source dataset, combined with knowledge retrieval modules tailored for simulation environments. A large language model (LLM) is employed to integrate knowledge-, data-, and adversarial-driven approaches, generating safety-critical traffic scenarios customized to user needs. Additionally, while evaluating the generated scenarios, we employ reinforcement learning models to train the behaviors of different types of vehicles, thereby enriching scenario diversity beyond existing datasets while preserving realism. Experimental results demonstrate that the proposed framework improves the accuracy of domain-specific language generation by 12\%. Moreover, the success rate of newly generated scenario transformations increases by 8\%, while obstacle-avoidance capability is enhanced by 30\%. For the complete manuscript, please refer to: https://zhenhaooo.github.io/PCASim.github.io/

cs.RO

An efficient preconditioned conjugate-gradient solver for a two-component dipolar Bose-Einstein condensate

We develop a preconditioned nonlinear conjugate-gradient solver for ground states of binary dipolar Bose-Einstein condensates within the extended Gross-Pitaevskii equation including Lee-Huang-Yang corrections. The optimization is carried out on the product-of-spheres normalization manifold and combines a manifold-preserving analytic line search, derived from a second-order energy expansion and validated along the exact normalized path, with complementary Fourier-space kinetic and real-space diagonal (Hessian-inspired) preconditioners. The method enforces monotonic energy descent and exhibits robust convergence across droplet, stripe, and supersolid regimes while retaining spectrally accurate discretizations and FFT-based evaluation of the dipolar term. In head-to-head benchmarks against imaginary-time evolution on matched grids and tolerances, the solver reduces iteration counts by one to two orders of magnitude and overall time-to-solution, and it typically attains slightly lower energies, indicating improved resilience to metastability. We reproduce representative textures and droplet-stability windows reported for dipolar mixtures. These results establish a reliable and efficient tool for large-scale parameter scans and phase-boundary mapping, and for quantitatively linking numerically obtained metastable branches to experimentally accessible states.

cond-mat.quant-gas

Large Foundation Models for Trajectory Prediction in Autonomous Driving: A Comprehensive Survey

Trajectory prediction serves as a critical functionality in autonomous driving, enabling the anticipation of future motion paths for traffic participants such as vehicles and pedestrians, which is essential for driving safety. Although conventional deep learning methods have improved accuracy, they remain hindered by inherent limitations, including lack of interpretability, heavy reliance on large-scale annotated data, and weak generalization in long-tail scenarios. The rise of Large Foundation Models (LFMs) is transforming the research paradigm of trajectory prediction. This survey offers a systematic review of recent advances in LFMs, particularly Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs) for trajectory prediction. By integrating linguistic and scene semantics, LFMs facilitate interpretable contextual reasoning, significantly enhancing prediction safety and generalization in complex environments. The article highlights three core methodologies: trajectory-language mapping, multimodal fusion, and constraint-based reasoning. It covers prediction tasks for both vehicles and pedestrians, evaluation metrics, and dataset analyses. Key challenges such as computational latency, data scarcity, and real-world robustness are discussed, along with future research directions including low-latency inference, causality-aware modeling, and motion foundation models.

cs.RO

Generative Modeling for Adversarial Lane-Change Scenarios

Decision-making in long-tail scenarios is pivotal to autonomous-driving development, and realistic and challenging simulations play a crucial role in testing safety-critical situations. However, existing open-source datasets lack systematic coverage of long-tail scenes, and lane-change maneuvers being emblematic, rendering such data exceedingly scarce. To bridge this gap, we introduce a data mining framework that exhaustively analyzes two widely used datasets, NGSIM and INTERACTION, to identify sequences marked by hazardous behavior, thereby replenishing these neglected scenarios. Using Generative Adversarial Imitation Learning (GAIL) enhanced with Proximal Policy Optimization (PPO), and enriched by vehicular-environment interaction analytics, our method iteratively refines and parameterizes newly generated trajectories. Distinguished by a rationally adversarial and sensitivity-aware perspective, the approach optimizes the creation of challenging scenes. Experiments show that, compared to unfiltered data and baseline models, our method produces behaviors that are simultaneously both adversarial and natural, judged by collision frequency, acceleration profiles, and lane-change dynamics, offering constructive insights to amplifying long-tailed lane-change instances in datasets and advancing decision-making training.

cs.RO

Probabilistic load flow calculation of AC/DC hybrid system based on cumulant method

The operating conditions of the power system have become more complex and changeable. This paper proposes a probabilistic load flow based on the cumulant method (PLF-CM) for the voltage sourced converter high voltage direct current (VSC-HVDC) hybrid system containing photovoltaic grid-connected systems. Firstly, the corresponding control mode is set for the converter, including droop control and master-slave control. The unified iterative method is used to calculate the conventional AC/DC flow. Secondly, on the basis of the probability model of load and photovoltaic output, based on the aforementioned flow results, use correlation coefficient matrix of this paper will change the relevant sample into independent sample, the cumulants of the load and photovoltaic output are obtained; then, the probability density function (PDF) and cumulative distribution function (CDF) of state variables are obtained by using Gram-Charlie series expansion method. Finally, the mean value and standard deviation of node voltage and line power are calculated on the modified IEEE 34-bus and IEEE 57-bus transmission systems. The algorithm can reflect the inherent uncertainty of new energy sources, and replace the complex convolution operation, greatly improving the calculation speed and the convergence.

eess.SY

Ultrafast dynamic evolution of multilevel systems in medium-strength laser fields

The ultrafast dynamic evolution of an atomic system under medium-strength laser fields is studied by performing transient absorption measurement. An analytical model developed from perturbation theory with a modified transition dipole moment is presented to explain the spectral features of the multilevel system. By fitting the measured absorption spectra to the model, the system's dynamic evolution is quantified by different amplitude and phase modulation factors in the pump--probe and probe--pump scenarios. This study provides a way to understand laser--matter interaction in the transition area between the strong-field and weak-field regimes.

physics.atom-ph

A multifeature fusion approach for power system transient stability assessment using PMU data

Taking full advantage of synchrophasors provided by GPS-based wide-area measurement system (WAMS), a novel VBpMKL-based transient stability assessment (TSA) method through multifeature fusion is proposed in this paper. First, a group of classification features reflecting the transient stability characteristics of power systems are extracted from synchrophasors, and according to the different stages of the disturbance process they are broken into three nonoverlapped subsets; then a VBpMKL-based TSA model is built using multifeature fusion through combining feature spaces corresponding to each feature subset; and finally application of the proposed model to the IEEE 39-bus system and a real-world power system is demonstrated. The novelty of the proposed approach is that it improves the classification accuracy and reliability of TSA using multifeature fusion with synchrophasors. The application results on the test systems verify the effectiveness of the proposal.

eess.SP

Rule extraction based on extreme learning machine and an improved ant-miner algorithm for transient stability assessment

In order to overcome the problems of poor understandability of the pattern recognition-based transient stability assessment (PRTSA) methods, a new rule extraction method based on extreme learning machine (ELM) and an improved Ant-miner (IAM) algorithm is presented in this paper. First, the basic principles of ELM and Ant-miner algorithm are respectively introduced. Then, based on the selected optimal feature subset, an example sample set is generated by the trained ELM-based PRTSA model. And finally, a set of classification rules are obtained by IAM algorithm to replace the original ELM network. The novelty of this proposal is that transient stability rules are extracted from an example sample set generated by the trained ELM-based transient stability assessment model by using IAM algorithm. The effectiveness of the proposed method is shown by the application results on the New England 39-bus power system and a practical power system - the southern power system of Hebei province.

eess.SP