SearcharxivSearch

arXiv subjects

Yuchen Zhang

Publications and source records attributed to Yuchen Zhang.

At least 19 recordsLinked to original sources

PlannerForge: LLM Agents for Scenario-Based Testing of Motion Planners in Autonomous Driving

Ensuring the safety of autonomous driving is a critical challenge. Scenario-based testing is a systematic process used to validate Autonomous Driving Systems (ADSs), but it remains a fragmented modular pipeline in which scenario generation, retrieval, modification, ADS execution, and results analysis are performed by separate tools with little interaction. Large Language Model (LLM) agents have shown promise across ADS sub-systems such as perception, planning, and control. However, no prior work covers the whole scenario-based testing pipeline for ADSs with a unified LLM-agent framework. We present PlannerForge, an LLM-agent framework that extends all scenario-based testing stages (from Scenario Generation to ADS Assessment) and adds two further LLM-enhanced stages: ADS Enhancement and ADS Benchmarking. We evaluate PlannerForge with 10 off-the-shelf LLMs across all tasks (Generation, Selection, Modification, Module Routing, Planner Testing, and Enhancement) under 5 prompt conditions. Best-per-task scores range from 0.88 to 1.00, and open-source 20-35B backends match commercial APIs on most tasks. Open-source models such as Qwen3.6:35B match commercial APIs on three of the five tasks. Chaining the modules end-to-end retains 83% / 78% of seed queries (commercial / open). It outperforms Scenario Factory 2.0 (Finkeldei et al., 2025) on natural-language generation (193 vs. 144 executable of 200) and realises 92-96% of requested city, road and vehicle attributes. It outperforms BM25 (Robertson and Zaragoza, 2009) at rank 1 selection (92.0% vs. 67.5%) and From-Words-to-Collisions (Gao et al., 2025) on physically valid edits (>=94% vs. 31%). At N=400, cost-tuning lifts planner success from 50.4% to 70.2% and cuts collisions from 19.0% to 8.4%, without domain-specific fine-tuning.

cs.AI

VERPO: Verified Evidence Regularized Policy Optimization

Verifiable outcome rewards guide language-model post-training, but sequence-level advantages do not identify which token-level decisions should be preserved or revised. Evidence-conditioned Teachers provide denser supervision by replaying sampled trajectories with privileged feedback. Yet indiscriminate imitation risks transferring formatting or reasoning-style shifts that do not support task success. We introduce VERPO, a Verified Evidence Regularized Policy Optimization framework that treats evidence as a proposal for policy correction while retaining the outcome objective. It separates evidence-free reference restoration from signed token-level evidence corrections. Fisher Evidence Contrast attenuates corrections along an estimated evidence-presence direction. A stopped token-wise ZPD controller scales acceptance according to local reward alignment and Fisher movement cost, while the reference channel remains independent of acceptance. Across five scientific-reasoning and tool-use tasks, the best variant on each backbone exceeds the strongest compared baseline in average score. The averages rise from 0.6826 to 0.6857 on Qwen3-4B, from 0.6895 to 0.7058 on Qwen3-8B, and from 0.4751 to 0.5657 on Llama-3.2-1B.

cs.LG

SUPER ODOMETRY 2.0: Resilient Odometry via Hierarchical Adaptation

Resilient and robust odometry is crucial for autonomous systems operating in complex and dynamic environments. Existing odometry systems often struggle with severe sensory degradations and extreme conditions such as smoke, sandstorms, snow, or low-light conditions, threatening both the safety and functionality of robots. To address these challenges, we present Super Odometry, a sensor fusion framework that dynamically adapts to varying levels of environmental degradation. Super Odometry employs a hierarchical structure to integrate four core modules from lower-level to higher-level adaptability including adaptive feature selection, adaptive state direction selection, adaptive engine selection, and a novel learning- based inertial odometry. The inertial odometry, trained on over 100 hours of heterogeneous robotic platforms, captures comprehensive motion dynamics. Super Odometry elevates the inertial measurement unit (IMU) to equal importance with camera and LiDAR within the sensor fusion framework, providing a reliable fallback when exteroceptive sensors fail. Super Odometry has been validated across 200 kilometers and 800 operational hours on a fleet of aerial, wheeled, and legged robots, under diverse sensor configurations, environmental degradation, and aggressive motion profiles. It marks an important step towards safe and long-term robotic autonomy in all-degraded environments.

cs.RO

CLARA: Clip-Level Multimodal Alignment with VLM-Derived Rationales for Hateful Video Detection

Hateful video detection has become increasingly important with the rapid growth of video-centric social media platforms, given the serious risks that hate speech poses to both individual well-being and social cohesion. Compared with text or static multimodal content, hateful video detection remains underexplored and significantly more challenging, as hateful meaning often arises from complex interactions among multimodal cues, including speech, audio, and visual content. Moreover, such signals are often brief, implicit, and temporally dependent, making them difficult to capture using conventional video-level representations. In this work, we propose CLARA, a clip-level multimodal framework for hateful video detection. Instead of treating a video as a single instance, CLARA models it as a sequence of fine-grained clips, enabling more precise capture of temporally localized hateful signals. We introduce a Mixture-of-Experts clip encoder for adaptive multimodal alignment, a local-global segment contrastive objective to jointly model short-term cues and long-range temporal dependencies, and VLM-derived rationales integrated via a gated Transformer to provide high-level semantic guidance. Extensive experiments on three hateful video datasets demonstrate that CLARA consistently outperforms state-of-the-art methods. Further ablation studies and parameter analyses validate the effectiveness of each component.

cs.CV

VOLA: Improving Open-World Driving by VLM-Based Semantic Attribute Prediction

Driving in the real world is open-world: a car may encounter a fallen mattress, a deer, or other objects outside its training data. Naming them is not enough. The system must know how to treat each region: can it drive over it, and how severe would a collision be? We therefore shift scene perception from category labels to dense action-relevant attributes, where each pixel is labeled by how it should affect motion rather than by object name. We instantiate this general formulation with two ordered attributes: 7-rank drivability and 5-rank vulnerability. We read Qwen3.5 image-token hidden states directly as a spatial semantic representation. A lightweight boundary-aware decoder then turns this coarse token grid into sharp full-resolution attribute maps. The whole process requires neither autoregressive text generation nor an external mask model such as SAM. We train on dense attribute labels built in CARLA and test transfer to real scenes and to novel obstacles never seen in training. We compare with vision-only segmenters trained on the same attributes and prompted VLM segmenters. Our model matches strong vision-only segmenters on familiar categories and improves transfer to real open-world anomalies, reaching 69.4% mean vulnerability-rank recall versus 57.1% for the best vision-only baseline and 53.9% for the best prompted VLM baseline. These results show that VLM image tokens provide useful semantic cues for transferring driving attributes to objects outside the training vocabulary.

cs.CV

Observational constraints on fractional holographic dark energy in the light of DESI DR2

Based on the fractional entropy from fractional quantum mechanics, fractional holographic dark energy (FHDE) has been proposed with the Hubble horizon as the IR cutoff (FHDEH). We extend this framework by adopting the future event horizon and the particle horizon as the IR cutoff, proposing the FHDEF and FHDEP models. Using the SN+OHD+DESI DR2 dataset to constrain these models, we find that all three models provide a marginally lower $\chi^{2}_{min}$ compared to $\Lambda$CDM but without significant preference according to AIC and BIC. When CMB distance priors are included, the FHDEH and FHDEP models are strongly ruled out. We further analyze the cosmological evolution for these models, and find that only the FHDEF model predicts nearly identical evolutions of $\Omega_{m}$ and $\Omega_{de}$ to those of the $\Lambda$CDM model across cosmic history, but its deceleration parameter $q$ deviate from the $\Lambda$CDM model in the future, indicating richer late time dynamics beyond the standard $\Lambda$CDM cosmology.

gr-qc

AirFlow: Context Preserving and Multi-Rate State Modeling for Air Quality Forecasting

Accurate air quality forecasting is essential for public health and urban environmental management, but remains challenging because pollutant channels differ in periodicity and distribution drift, while their concentration trajectories contain both multi-scale dependencies and rapid changes. Recent methods have improved spatial dependency learning and meteorological covariate modeling. However, pollutant channels are still passed through the same normalization rule and temporal backbone, using a shared latent representation for channel-specific distributions and changes at different rates. To address this limitation, we propose AirFlow, a pollutant-aware dual-stream framework that operates on station multivariate observations without additional graph propagation or predefined signal decomposition. Specifically, AirFlow designs two novel blocks: (1) a statistic-guided normalization routing mechanism that selects a normalization path for each pollutant according to its 24-hour autocorrelation and distribution drift; and (2) a hierarchical dual-stream state model that combines multi-scale state space propagation with learnable response coefficients, where gated bidirectional cross-attention exchanges information and adaptively fuses the resulting representations. Experiments on real-world data from multiple cities show that AirFlow achieves the best performance in 34 of 36 metrics comparisons, with reductions of up to 11.11% root mean square error over the state-of-the-art baseline. AirFlow also requires only 0.0483M parameters and 0.0215G FLOPs, achieving high forecasting accuracy with low computational overhead.

cs.AI

Science Edge Evaluation: SEE the Missing Step Toward Real Scientific Discovery

Large language models (LLMs) are increasingly involved in scientific discovery, yet it remains unclear whether they can support complex real laboratory science. Here we introduce Science Edge Evaluation (SEE), a multimodal benchmark of expert-curated questions grounded in peer-reviewed literature and experimental practice in chemistry, biology, and materials science. Evaluation of 19 multimodal large language models (MLLMs) shows that even the best-performing model reaches only 48.7% accuracy. Moreover, general-purpose models outperform science-specialized models on average. In the visual-agent evaluation, the use of tools increases the best accuracy to 52.7%. Tool use can expand the information available to models, but more information does not necessarily lead to reliable scientific reasoning. The key challenge is whether models can manage tool-derived information within the boundaries of the original experimental evidence. Together, these findings reveal that current MLLMs still cannot reliably make justified and evidence-bounded inferences from experimental results, which is an essential capability in real scientific discovery. Bridging this gap requires MLLMs to transition from explaining established scientific concepts to deriving novel and evidence-based insights from experimental data.

cs.AI

Multi-User Localization via Active Sensing with Electromagnetically Reconfigurable Antennas

This paper investigates multi-user localization in uplink wireless systems assisted by electromagnetically reconfigurable antennas (ERAs). Unlike traditional localization schemes, we formulate an active sensing problem where a base station (BS) exploits historical pilot observations accumulated over previous sensing stages to adapt the shared ERA configuration and progressively refine position estimates. To capture both theoretical flexibility and practical hardware constraints, we establish a unified wideband geometric signal model accommodating two complementary ERA paradigms: a synthesis-based model utilizing spherical-harmonic basis functions, and a finite-state model based on measured radiation codebooks. Because analytically solving the resulting joint design problem is highly intractable due to the high-dimensional observation and the shared-aperture coupling among multiple users, we develop a learning-based active sensing framework. Specifically, pilot-matched wideband observations are compressed into compact user-wise features and sequentially accumulated by a long short-term memory (LSTM) module. These temporal features are then processed by a graph neural network (GNN) to capture multi-user shared-aperture coupling. Model-specific output heads generate either continuous synthesis coefficients or finite-state ERA selections, while a localization head produces stage-wise position estimates. Numerical results under a specific channel distribution show that the proposed ERA-assisted active sensing framework achieves progressive localization refinement across sensing stages and obtains better performance than conventional non-reconfigurable arrays and representative ablation baselines.

eess.SP

Model predictive control for laser thermal processing: operator learning, closed-loop validation, and out-of-distribution analysis

Laser-based thermal processing, such as laser powder bed fusion, requires tight regulation of the peak surface temperature: heat accumulates where the moving source re-enters previously heated material, driving the temperature out of its process window and causing defects. High-fidelity thermal models capture this physics but are too slow for online optimization, which motivates fast, differentiable, and generalizable surrogates. We develop and validate a complete surrogate-based control pipeline that regulates the maximum surface temperature of a moving laser on a 304-stainless-steel substrate. We also determine conditions under which our surrogate can be trusted inside the control loop by probing its out-of-distribution limits. A key component of our surrogate is a multi-step deep operator network bespoke for moving sources: its branch subnetwork encodes the future power and trajectory (position and velocity) sequence, while its trunk encodes the current peak temperature and the temperature at the future laser locations, yielding a one-shot five-step prediction. By way of illustration, we use this surrogate as a smooth (algebraic-rectifier) nonlinear program inside a receding-horizon model predictive controller solved in CasADi/IPOPT. The surrogate forward pass is over thousand times faster than the equivalent finite-difference steps. We show that aggregate open-loop accuracy is necessary but not sufficient for control-readiness: two surrogates with near-identical offline error behave drastically differently in closed loop. A controlled two-ensemble data design reduces a 91 K path-corner underprediction failure to 1.4 K, and a calibrated one-sided constraint margin of 13 K yields zero violations of the true upper bound on all tested paths.

math.OC

A Koszul complex in quaternionic analysis and its applications

Let $n\geqslant 1, \Omega\subset\mathbb{H}^n $ be a domain. We construct a Koszul-type complex for the ideal sheaf $\mathcal{I}_X^{(k)}$ of $k$-regular functions vanishing on $X=\{(q_0, q_1, \cdots, q_{n-1})\in \Omega: q_0=0\}$ in several quaternionic variables: $$0\to \mathcal{R}^{(k+2)}\xrightarrow{\widetilde{\mathscr{L}}^{(k)}} \mathcal{R}^{(k+1)}\oplus\mathcal{R}^{(k+1)}\xrightarrow{\mathscr{L}^{(k)}} \mathcal{I}_X^{(k)}\to 0,$$ where $k\geqslant 0$, $\mathcal{R}^{(k)}$ is the sheaf of $k$-regular functions on $\Omega$, $\widetilde{\mathscr{L}}^{(k)}=(-L_1^{(k+2)},L_0^{(k+2)})^{T}$, $\mathscr{L}^{(k)}=(L_0^{(k+1)},L_1^{(k+1)})$, and $L_0^{(k)},L_1^{(k)}$ are multiplication-like operators on $k$-regular functions. This gives the quaternionic analogue of the classical Koszul complex. And we present the long exact sequence in cohomology for the case $\Omega\cap\{q_0=0\}=\emptyset$ with explicit differential connecting maps, by applying the Cauchy-Fueter complex and cohomological methods. As an application, in the special case $n=1, k=1$, the operator pair $(L_0^{(1)}, L_1^{(1)})$ is shown to be surjective if and only if $H^3(\Omega, \mathbb{R})=0$. Furthermore, a cohomological vanishing criterion is given for $H^1(\Omega,\mathcal{I}_X^{(k)})$; under this criterion, every $k$-regular function on $\{q_0=0\}\cap\Omega$ extends to a $k$-regular function on $\Omega$.

math.CV

CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning

Modern automated audio captioning systems pair a frozen audio encoder with a large language model (LLM) via a trainable projector, incurring the encoder's inference cost and bottlenecking the model through its fixed acoustic features. We present CARD, an encoder-free audio captioning model that removes the encoder at inference: a 13.2M projector feeds a frozen LLM with merged LoRA adapters, while the teacher used to train it is discarded. CARD distills a pretrained audio teacher (CLAP-HTSAT) into the model, but rather than injecting it into the LLM alone, it routes the teacher's representations across components: perceptual stages to the projector and semantic stages to the LLM. This placement improves CIDEr-D by +12.18 over an LLM-only distilled model on AudioCaps and by +5.21 on Clotho, reaching 55.4 against a 66.4 encoder-kept upper bound with no encoder at inference, showing that where a teacher's knowledge is placed matters as much as its presence.

cs.SD

Text-Driven 3D Indoor Scene Synthesis in Non-Manhattan Environments

Large Language Models (LLMs) have demonstrated remarkable capabilities in 3D indoor synthesis for Manhattan environments. However, existing methods often fail to capture plausible object layout patterns in non-Manhattan settings, primarily because they struggle to model non-orthogonal spatial relationships, leading to high geometric violations and low physical fidelity. To address this challenge, we propose SPG-Layout, a novel text-driven framework designed to generate physically plausible indoor scenes within complex non-Manhattan environments. Specifically, we first utilize statistical priors of object distributions to guide the training process, enhancing environmental understanding and fidelity. Furthermore, mirroring human design workflows, we adopt a hierarchical layout strategy that prioritizes the placement of large objects, thereby substantially minimizing layout violations. By synergizing these components, SPG-Layout achieves a balanced optimization of semantic realism and physical plausibility. To evaluate performance in these complex settings, we constructed a new benchmark comprising 500 diverse non-Manhattan environments. Extensive experiments demonstrate that SPG-Layout consistently and significantly outperforms existing methods across both Manhattan and non-Manhattan environments. The code will be publicly released.

cs.AI

Robust 3D Alignment of Generative Reconstructions via Partial Monocular Observations

Aligning generative 3D reconstructions with partial monocular observations is a critical but under-explored challenge in computer vision. This task is inherently ill-posed due to severe asymmetries between noisy, sparse monocular inputs and dense generative priors, whose scale ambiguity and geometric hallucinations, combined with the lack of initial overlap, render traditional registration pipelines ineffective. To resolve these issues, we propose a training-free and interpretable geometric alignment framework that grounds generative 3D priors via a 3D similarity transformation (Sim(3)), which can recover accurate metric scale and pose. Specifically, we introduce an explicit scale factor to resolve metric ambiguity and employ a coarse-to-fine alignment strategy, leveraging geometry-aware descriptors for robust initialization and a decoupled closed-form solver for precision refinement. In addition, we introduce a Hallucination Filtering operation to effectively suppress outliers caused by hallucinated geometry. To evaluate alignment performance under these extreme conditions, we introduce GenPMOAlign--Where2Place, a rigorous benchmark specifically designed for Generative-to-Partial Monocular Observational Alignment. Experiments demonstrate that our method achieves stable and accurate registration, substantially outperforming both classical geometric pipelines and state-of-the-art learning-based baselines. Code and the benchmark will be publicly released.

cs.CV

Estimating Cosmological Parameters from Localized Fast Radio Bursts: A Method for Removing Milky Way Dispersion-Measure Contributions

Fast radio bursts (FRBs) are emerging as powerful probes for cosmology. However, cosmological inference based on FRB dispersion measures (DMs) is limited by uncertainties in the Milky Way contribution, including those from the Galactic interstellar medium and the Galactic halo. In this Letter, we propose a method that eliminates the Milky Way contribution by using DM differences between localized FRBs within the same sky region. The method removes the need to adopt a specific Galactic electron-density model or a prior assumption for the Galactic halo DM. We validate the reliability of the method using mock FRB samples and show that it successfully recovers the fiducial cosmological parameter. Applying the method to current localized FRB data, we obtain a constraint on $\Gamma \equiv \Omega_b H_0 f_{\rm d}$ that differs from that inferred using the conventional treatment of the Milky Way contribution. This difference highlights the importance of Milky Way DM systematics in FRB cosmology and demonstrates the potential of differential DM methods for future large samples of localized FRBs.

astro-ph.CO

How Many RF Chains Does a Microwave Linear Analog Computer (MiLAC) Need to Match the Fully-Digital Cram\'er-Rao Bound?

A microwave linear analog computer (MiLAC) is a tunable microwave network that performs linear operations directly on radio-frequency signals through wave propagation. Used as an antenna-array front end, it can map many antenna signals to a small number of active RF chains. While lossless reciprocal MiLACs have been shown to provide flexible or capacity-achieving beamforming for wireless communications, their sensing performance remains largely unexplored. We analyze direction-of-arrival estimation for $K$ far-field targets using a tunable receive-side lossless reciprocal MiLAC combiner. We show that the Fisher information matrix depends on the combiner only through the orthogonal projector onto its row space and never exceeds that of a fully digital receiver. Equality holds when the row space contains the $2K$-dimensional joint steering--derivative subspace, establishing a zero-gap threshold of two RF chains per target. A dimension-counting argument lower-bounds the number of tunable components required to achieve the digital Cram\'er--Rao bound for every target configuration. The stem-connected MiLAC attains this bound asymptotically, up to an antenna-count-independent additive overhead, while scaling linearly with the antenna and target counts. Unlike a phase-shifter front end with the same number of RF chains, MiLAC can exactly attain the fully digital bound. Numerical results validate the analysis.

cs.IT

What Shapes Emergent Misalignment? Insights from Training Dynamics, Model Priors, and Data

Emergent misalignment (EM) is a phenomenon in which models generalize with narrow fine-tuning, leading to broad (yet uneven) misalignment across evaluation questions. We study EM and its variability directly through the components of fine-tuning: training dynamics, model priors, and data. (1) We first explored how in-domain training loss relates to out-of-domain alignment scores across datasets and model families. Then, we tried to induce potential alternative local minima through different learning schedules for one narrow fine-tuning, but did not find strong runs with better broad alignment scores conditioned on similar or lower training loss. (2) We found that although the mean and standard deviations of the misaligned model scores are usually statistically different from those of the pre-trained model, there are some potential signals on overall positive correlation. The evaluation prompt-only activations from both the pre-trained and the original instruct models (prior to narrow fine-tuning) could predict fine-grained alignment scores after narrow fine-tuning. (3) Finally, we compared activation deltas before and after narrow fine-tuning and found moderate-to-high subspace overlap and similarity between the resulting activation shifts for training and evaluation prompts. Subspace overlaps between training and evaluation prompt activations correlate with their shifts' similarities when measuring with the last prompt-token activations. The train-evaluation data prompt overlap is controlled against overlap computed from random vectors and evaluation prompts activations.

cs.AI

ReMP: Low-Downtime Runtime Model-Parallelism Reconfiguration for LLM Serving

Current large language model (LLM) inference systems universally deploy ultra-large-scale models using a combination of Tensor Parallelism (TP) and Pipeline Parallelism (PP). However, existing systems treat the model parallelism topology as a static configuration that cannot be flexibly adjusted at runtime. This rigid design creates a fundamental contradiction with the dynamically changing inference workloads in real-world scenarios. State-of-the-art systems lack online reconfiguration capabilities and can only switch configurations by restarting the service, resulting in several minutes of service interruption, KV cache loss, and prohibitive recomputation overhead. To address this problem, this paper presents ReMP, a runtime model parallelism reconfiguration framework that supports low downtime. ReMP achieves dynamic adjustment through three key techniques: (1) decoupling the model parallelism topology from runtime state to avoid full service reconstruction; (2) designing a two-dimensional KV cache migration mechanism to preserve reusable cache states after TP/PP changes; and (3) implementing end-to-end online reconfiguration. Experiments demonstrate that ReMP can complete most topology switches within 1-7 seconds on models ranging from 7B to 70B parameters, achieving speedups of tens to over a hundred times compared to the restart approach. Moreover, ReMP significantly outperforms fixed configurations under dynamic workloads, delivering superior performance in terms of TTFT, TPOT, and output throughput.

cs.DC