Searcharxiv⌕ Search

arXiv subjects

Haochen Wang

Publications and source records attributed to Haochen Wang.

At least 55 records · Page 3Linked to original sources

DriveVLA-W0: World Models Amplify Data Scaling Law in Autonomous Driving

Scaling Vision-Language-Action (VLA) models on large-scale data offers a promising path to achieving a more generalized driving intelligence. However, VLA models are limited by a ``supervision deficit'': the vast model capacity is supervised by sparse, low-dimensional actions, leaving much of their representational power underutilized. To remedy this, we propose \textbf{DriveVLA-W0}, a training paradigm that employs world modeling to predict future images. This task generates a dense, self-supervised signal that compels the model to learn the underlying dynamics of the driving environment. We showcase the paradigm's versatility by instantiating it for two dominant VLA archetypes: an autoregressive world model for VLAs that use discrete visual tokens, and a diffusion world model for those operating on continuous visual features. Building on the rich representations learned from world modeling, we introduce a lightweight action expert to address the inference latency for real-time deployment. Extensive experiments on the NAVSIM v1/v2 benchmark and a 680x larger in-house dataset demonstrate that DriveVLA-W0 significantly outperforms BEV and VLA baselines. Crucially, it amplifies the data scaling law, showing that performance gains accelerate as the training dataset size increases.

cs.CV↗

Electrical Stability of Cr2O3/\b{eta}-Ga2O3 and NiOx/\b{eta}-Ga2O3 Heterojunction Diodes

This work reports the electrical characteristics comparison study between Cr2O3 and NiOx based heterojunction diodes (HJD) on halide vapor phase epitaxy (HVPE) grown \b{eta}-Ga2O3 epitaxial layers. Both as-fabricated Cr2O3 and NiOx HJDs exhibited forward current density in a range of 130-150 A/cm^2 at 5 V with rectifying ratios >10^10 and a reverse leakage current density at 10^-8 A/cm^2 at -5 V. The differential specific on-resistance of Cr2O3 and NiOx HJDs was 12.01 mΩ*cm^2 and 12.05 mΩ*cm^2, respectively. Breakdown voltages of Cr2O3 HJDs ranged from 1.4-1.9 kV and 1.5-2.3 kV for NiOx HJDs. Theoretical band alignment between Cr2O3 and \b{eta}-Ga2O3 was calculated from first principles. The ambient exposed NiOx/HVPE \b{eta}-Ga2O3 HJDs forward current density degraded after 10 days while that of Cr2O3/HVPE \b{eta}-Ga2O3 HJDs remained nearly unchanged after the same amount of time. It was later confirmed that the ambient exposed sputtered NiOx sheet resistance (Rsh) degradation gave rise to the reduction of the forward current density of the NiOx based HJDs, and water (H2O) was qualitatively determined to be the agent attributed to the forward conduction degradation by measuring the Rsh of NiOx-on-sapphire reference wafer after exposing it to different environments. The Cr2O3/HVPE \b{eta}-Ga2O3 HJD also exhibited enhanced thermal stability compared to the NiOx/\b{eta}-Ga2O3 heterostructures at elevated temperatures. Interfacial nickel gallate (Ga2NiO4) phase formation expected from phase diagrams can explain the reduced thermal stability of NiOx/\b{eta}-Ga2O3 HJDs. This study indicates that Cr2O3 is a stable p-type oxide for the realization of robust multi-kV \b{eta}-Ga2O3 HJDs.

physics.app-ph↗

CrossVid: A Comprehensive Benchmark for Evaluating Cross-Video Reasoning in Multimodal Large Language Models

Cross-Video Reasoning (CVR) presents a significant challenge in video understanding, which requires simultaneous understanding of multiple videos to aggregate and compare information across groups of videos. Most existing video understanding benchmarks focus on single-video analysis, failing to assess the ability of multimodal large language models (MLLMs) to simultaneously reason over various videos. Recent benchmarks evaluate MLLMs' capabilities on multi-view videos that capture different perspectives of the same scene. However, their limited tasks hinder a thorough assessment of MLLMs in diverse real-world CVR scenarios. To this end, we introduce CrossVid, the first benchmark designed to comprehensively evaluate MLLMs' spatial-temporal reasoning ability in cross-video contexts. Firstly, CrossVid encompasses a wide spectrum of hierarchical tasks, comprising four high-level dimensions and ten specific tasks, thereby closely reflecting the complex and varied nature of real-world video understanding. Secondly, CrossVid provides 5,331 videos, along with 9,015 challenging question-answering pairs, spanning single-choice, multiple-choice, and open-ended question formats. Through extensive experiments on various open-source and closed-source MLLMs, we observe that Gemini-2.5-Pro performs best on CrossVid, achieving an average accuracy of 50.4%. Notably, our in-depth case study demonstrates that most current MLLMs struggle with CVR tasks, primarily due to their inability to integrate or compare evidence distributed across multiple videos for reasoning. These insights highlight the potential of CrossVid to guide future advancements in enhancing MLLMs' CVR capabilities.

cs.CV↗

Cr2O3/\b{eta}-Ga2O3 Heterojunction Diodes with Orientation-Dependent Breakdown Electric Field up to 12.9 MV/cm

We report the fabrication of Cr2O3/\b{eta}-Ga2O3 heterojunction diodes using reactive magnetron sputtering of Cr2O3 on highly doped \b{eta}-Ga2O3 bulk substrates along (100), (010), (001), (110), and (011) orientation dependence of high electric field handling capability in \b{eta}-Ga2O3. Additional relative permittivity values in (110) and (011) orientations of \b{eta}-Ga2O3 were computed by using first-principles calculation methods for accurate apparent charge density (ND-NA) extraction and breakdown electric field analysis from capacitance-voltage measurements. The HJDs fabricated on n+ (110) exhibited breakdown electric fields >10 MV/cm up to 12.9 MV/cm, showing the highest experimentally observed parallel-plane junction electric field among \b{eta}-Ga2O3-based junctions. Breakdown electric fields among (100), (010), (001), and (011) orientations showed distinct distribution in the range of 5.13-5.26 MV/cm, 5.10-7.05 MV/cm, 2.70-3.33 MV/cm, and 3.88-4.38 MV/cm, respectively, validating the orientational dependence of parallel-plane junction electric field at breakdown in low-symmetry monoclinic \b{eta}-Ga2O3. The parallel-plane breakdown electric fields (EBr,||) reported in this work were extracted when the device experienced catastrophic breakdown at 100 mA/cm^2 current density compliance, and should not be confused with critical electric field (Ec) as a function of drift layer doping concentration, which accounts for electric-field dependent impact ionization coefficients in Si, SiC and GaN. This study can guide the choice of crystal orientation for high performance gallium oxide-based devices that require high electric field handling capability.

cond-mat.mtrl-sci↗

MMaDA-Parallel: Multimodal Large Diffusion Language Models for Thinking-Aware Editing and Generation

While thinking-aware generation aims to improve performance on complex tasks, we identify a critical failure mode where existing sequential, autoregressive approaches can paradoxically degrade performance due to error propagation. To systematically analyze this issue, we propose ParaBench, a new benchmark designed to evaluate both text and image output modalities. Our analysis using ParaBench reveals that this performance degradation is strongly correlated with poor alignment between the generated reasoning and the final image. To resolve this, we propose a parallel multimodal diffusion framework, MMaDA-Parallel, that enables continuous, bidirectional interaction between text and images throughout the entire denoising trajectory. MMaDA-Parallel is trained with supervised finetuning and then further optimized by Parallel Reinforcement Learning (ParaRL), a novel strategy that applies semantic rewards along the trajectory to enforce cross-modal consistency. Experiments validate that our model significantly improves cross-modal alignment and semantic consistency, achieving a 6.9\% improvement in Output Alignment on ParaBench compared to the state-of-the-art model, Bagel, establishing a more robust paradigm for thinking-aware image synthesis. Our code is open-sourced at https://github.com/tyfeld/MMaDA-Parallel

cs.CV↗

MVU-Eval: Towards Multi-Video Understanding Evaluation for Multimodal LLMs

The advent of Multimodal Large Language Models (MLLMs) has expanded AI capabilities to visual modalities, yet existing evaluation benchmarks remain limited to single-video understanding, overlooking the critical need for multi-video understanding in real-world scenarios (e.g., sports analytics and autonomous driving). To address this significant gap, we introduce MVU-Eval, the first comprehensive benchmark for evaluating Multi-Video Understanding for MLLMs. Specifically, our MVU-Eval mainly assesses eight core competencies through 1,824 meticulously curated question-answer pairs spanning 4,959 videos from diverse domains, addressing both fundamental perception tasks and high-order reasoning tasks. These capabilities are rigorously aligned with real-world applications such as multi-sensor synthesis in autonomous systems and cross-angle sports analytics. Through extensive evaluation of state-of-the-art open-source and closed-source models, we reveal significant performance discrepancies and limitations in current MLLMs' ability to perform understanding across multiple videos. The benchmark will be made publicly available to foster future research.

cs.CV↗

FastMap: Revisiting Structure from Motion through First-Order Optimization

We propose FastMap, a new global structure from motion method focused on speed and simplicity. Previous methods like COLMAP and GLOMAP are able to estimate high-precision camera poses, but suffer from poor scalability when the number of matched keypoint pairs becomes large, mainly due to the time-consuming process of second-order Gauss-Newton optimization. Instead, we design our method solely based on first-order optimizers. To obtain maximal speedup, we identify and eliminate two key performance bottlenecks: computational complexity and the kernel implementation of each optimization step. Through extensive experiments, we show that FastMap is up to 10 times faster than COLMAP and GLOMAP with GPU acceleration and achieves comparable pose accuracy.

cs.CV↗

Object-centric Video Question Answering with Visual Grounding and Referring

Video Large Language Models (VideoLLMs) have recently demonstrated remarkable progress in general video understanding. However, existing models primarily focus on high-level comprehension and are limited to text-only responses, restricting the flexibility for object-centric, multiround interactions. In this paper, we make three contributions: (i) we address these limitations by introducing a VideoLLM model, capable of performing both object referring for input and grounding for output in video reasoning tasks, i.e., allowing users to interact with videos using both textual and visual prompts; (ii) we propose STOM (Spatial-Temporal Overlay Module), a novel approach that propagates arbitrary visual prompts input at any single timestamp to the remaining frames within a video; (iii) we present VideoInfer, a manually curated object-centric video instruction dataset featuring questionanswering pairs that require reasoning. We conduct comprehensive experiments on VideoInfer and other existing benchmarks across video question answering and referring object segmentation. The results on 12 benchmarks of 6 tasks show that our proposed model consistently outperforms baselines in both video question answering and segmentation, underscoring its robustness in multimodal, object-centric video and image understanding. Project page: https://qirui-chen.github.io/RGA3-release/.

cs.CV↗

Stellar Mass-Dispersion Measure Correlations Constrain Baryonic Feedback in Fast Radio Burst Host Galaxies

Low redshift fast radio bursts (FRBs) provide robust measurements of the host-galaxy contribution to the dispersion measure (DM), which can constrain the circumgalactic medium (CGM) of the hosts. We curate a sample of 20 nearby FRBs with low scattering timescales and face-on host galaxies with stellar masses ranging from $10^9 < M^* / M_\odot < 10^{11}$. We fit the distribution of the host galaxy DM to a quadratic model as a function of stellar mass with a mass-independent scatter and find that the more massive the host, the lower its host DM. We report that this relation has a negative slope of $m = -97 \pm 44$ pc/cm$^{-3}$ per dex in stellar mass. We compare this measurement to similar fits to three sub-grid models implemented in the CAMELS suite of simulations from Astrid, IllustrisTNG, and SIMBA and find that fine-tuning of the host ISM contribution as a function of stellar mass is required in order to reconcile the observational data with the predictions of the fiducial CAMELS-Astrid model. More generally, models which attribute a positive correlation between stellar mass and host dispersion measure ($m > 0$) to the CGM are in tension with our measurement. We show that this conclusion is robust to a wide range of assumptions, such as the offset distribution of FRBs from their hosts and the statistics of the cosmic contribution to the DM budget along each sightline. Our results indirectly imply a lower limit on the strength of baryonic feedback in the Local Universe $(z < 0.2)$ in isolated $\sim L^*$ halos, complementing results from weak lensing surveys and kSZ observations which target higher halo mass and redshift ranges.

astro-ph.GA↗

Hita: Holistic Tokenizer for Autoregressive Image Generation

Vanilla autoregressive image generation models generate visual tokens step-by-step, limiting their ability to capture holistic relationships among token sequences. Moreover, because most visual tokenizers map local image patches into latent tokens, global information is limited. To address this, we introduce \textit{Hita}, a novel image tokenizer for autoregressive (AR) image generation. It introduces a holistic-to-local tokenization scheme with learnable holistic queries and local patch tokens. Hita incorporates two key strategies to better align with the AR generation process: 1) {arranging} a sequential structure with holistic tokens at the beginning, followed by patch-level tokens, and using causal attention to maintain awareness of previous tokens; and 2) adopting a lightweight fusion module before feeding the de-quantized tokens into the decoder to control information flow and prioritize holistic tokens. Extensive experiments show that Hita accelerates the training speed of AR generators and outperforms those trained with vanilla tokenizers, achieving \textbf{2.59 FID} and \textbf{281.9 IS} on the ImageNet benchmark. Detailed analysis of the holistic representation highlights its ability to capture global image properties, such as textures, materials, and shapes. Additionally, Hita also demonstrates effectiveness in zero-shot style transfer and image in-painting. The code is available at \href{https://github.com/CVMI-Lab/Hita}{https://github.com/CVMI-Lab/Hita}.

cs.CV↗

FRB 20250316A: A Brilliant and Nearby One-Off Fast Radio Burst Localized to 13 parsec Precision

Precise localizations of a small number of repeating fast radio bursts (FRBs) using very long baseline interferometry (VLBI) have enabled multiwavelength follow-up observations revealing diverse local environments. However, the 2--3\% of FRB sources that are observed to repeat may not be representative of the full population. Here we use the VLBI capabilities of the full CHIME Outriggers array for the first time to localize a nearby (40 Mpc), bright (kJy), and apparently one-off FRB source, FRB 20250316A, to its environment on 13-pc scales. We use optical and radio observations to place deep constraints on associated transient emission and the properties of its local environment. We place a $5σ$ upper limit of $L_{\mathrm{9.9~\mathrm{GHz}}} < 2.1\times10^{25}~\mathrm{erg~s^{-1}~Hz^{-1}}$ on spatially coincident radio emission, a factor of 100 lower than any known compact persistent radio source associated with an FRB. Our KCWI observations allow us to characterize the gas density, metallicity, nature of gas ionization, dust extinction and star-formation rate through emission line fluxes. We leverage the exceptional brightness and proximity of this source to place deep constraints on the repetition of FRB 20250316A, and find it is inconsistent with all well-studied repeaters given the non-detection of bursts at lower spectral energies. We explore the implications of a measured offset of 190$\pm20$ pc from the center of the nearest star-formation region, in the context of progenitor channels. FRB 20250316A marks the beginning of an era of routine localizations for one-off FRBs on tens of mas-scales, enabling large-scale studies of their local environments.

astro-ph.HE↗

Measurement of the Dispersion$\unicode{x2013}$Galaxy Cross-Power Spectrum with the Second CHIME/FRB Catalog

The dispersion of extragalactic fast radio bursts (FRBs) can serve as a powerful probe of the diffuse plasma between and surrounding galaxies, which contains most of the Universe's baryons. By cross-correlating the dispersion of background FRBs with the locations of foreground galaxies, we can study the relative spatial distributions of plasma and galaxies on scales of 0.1 to 50 Mpc, which are strongly affected by feedback processes in galaxy formation. Here we present the measurement of the dispersion$\unicode{x2013}$galaxy angular cross-power spectrum between 2873 FRBs from the Second CHIME/FRB Catalog and nearly 6 million galaxies from the Dark Energy Spectroscopic Instrument (DESI) Legacy Imaging Survey. Over five photometric galaxy redshift bins spanning $0.05 < z <0.5$ and at 5.1$σ$ significance, we make the first definitive detection of spatial correlations in FRB dispersion measure due to cosmic structure. While parameter inferences should be interpreted with caution because of incomplete modelling of both the signal and systematic errors, our data indicate that the plasma$\unicode{x2013}$galaxy cross-power spectrum cuts off relative to the matter power spectrum at a scale $k_\textrm{cut}^{-1}=0.9^{+0.4}_{-0.4}\,\textrm{Mpc}$. This scale is consistent with those X-ray stacking analyses that suggest dark-matter halos with group-scale masses are largely evacuated of their baryons by feedback processes. Our study demonstrates that FRBs are promising tools to discern the physics of baryonic structure formation and will only become more powerful as FRB surveys expand.

astro-ph.CO↗

Mitigating antenna gain errors with HyFoReS in CHIME simulations

Hybrid Foreground Residual Subtraction (HyFoReS) is a new family of algorithms designed to remove systematics-induced foreground contamination for 21-cm intensity mapping data. Previously, the algorithm was shown to be effective in mitigating beam perturbations in sky maps from the Canadian Hydrogen Intensity Mapping Experiment (CHIME). In this study, we apply HyFoReS to CHIME simulations and test the algorithm's ability to mitigate antenna gain-type systematics in polarized visibilities. Simulating a two-cylinder telescope similar to the CHIME pathfinder, we find that HyFoReS reduces foreground bias caused by bandpass perturbations to a level below the thermal noise, provided that the RMS value of the perturbations is on the order of $10^{-4}$ or lower. When tested with complex antenna-dependent gain errors, HyFoReS can reduce residual foreground bias in the power spectrum by up to three orders of magnitude. While noise bias and second-order perturbations are currently the limiting factors for the algorithm, we have demonstrated that HyFoReS can suppress gain-induced foreground leakage in polarized data from 21-cm telescopes, aiding in the detection of the 21-cm auto-power spectrum for hydrogen intensity mapping experiments.

astro-ph.IM↗

Integrating Learning-Based Manipulation and Physics-Based Locomotion for Whole-Body Badminton Robot Control

Learning-based methods, such as imitation learning (IL) and reinforcement learning (RL), can produce excel control policies over challenging agile robot tasks, such as sports robot. However, no existing work has harmonized learning-based policy with model-based methods to reduce training complexity and ensure the safety and stability for agile badminton robot control. In this paper, we introduce Hamlet, a novel hybrid control system for agile badminton robots. Specifically, we propose a model-based strategy for chassis locomotion which provides a base for arm policy. We introduce a physics-informed "IL+RL" training framework for learning-based arm policy. In this train framework, a model-based strategy with privileged information is used to guide arm policy training during both IL and RL phases. In addition, we train the critic model during IL phase to alleviate the performance drop issue when transitioning from IL to RL. We present results on our self-engineered badminton robot, achieving 94.5% success rate against the serving machine and 90.7% success rate against human players. Our system can be easily generalized to other agile mobile manipulation tasks such as agile catching and table tennis. Our project website: https://dreamstarring.github.io/HAMLET/.

cs.RO↗

Emergent Kagome lattice and non-Abelian lattice gauge field of biexcitons in t-MoTe$_2$

Non-Abelian gauge fields, characterized by their non-commutative symmetry groups, shape physical laws from the Standard Model to emergent topological matter for quantum computation. Here we find that moiré exciton dimers (biexcitons) in twisted bilayer MoTe$_2$ are governed by a genuine non-Abelian lattice gauge field. These dipolar-bound exciton dimers, formed on bonds of the honeycomb moiré superlattice, exhibit three quadrupole configurations organized into a Kagome lattice geometry, on which the valley-flip biexciton hoppings through electron-hole Coulomb exchange act as link variables of the non-Abelian lattice gauge theory. The emergence of gauge structure here is a new possibility for composite particles, where the moiré electronic structure and interactions between the electron and hole constituents jointly enforce the underlying geometric constraint. The quadrupole nature of biexciton further makes possible local gate controls to isolate designated pathways from the extended lattice for exploiting consequences of non-commutative gauge structure including the genuine non-Abelian Aharonov-Bohm effect. This also provides a new approach for quantum manipulation of excitonic valley qubit. We show path interference on a simplest loop can deterministically transform the computational basis states into Bell states.

cond-mat.mes-hall↗

The Scalability of Simplicity: Empirical Analysis of Vision-Language Learning with a Single Transformer

This paper introduces SAIL, a single transformer unified multimodal large language model (MLLM) that integrates raw pixel encoding and language decoding within a singular architecture. Unlike existing modular MLLMs, which rely on a pre-trained vision transformer (ViT), SAIL eliminates the need for a separate vision encoder, presenting a more minimalist architecture design. Instead of introducing novel architectural components, SAIL adapts mix-attention mechanisms and multimodal positional encodings to better align with the distinct characteristics of visual and textual modalities. We systematically compare SAIL's properties-including scalability, cross-modal information flow patterns, and visual representation capabilities-with those of modular MLLMs. By scaling both training data and model size, SAIL achieves performance comparable to modular MLLMs. Notably, the removal of pretrained ViT components enhances SAIL's scalability and results in significantly different cross-modal information flow patterns. Moreover, SAIL demonstrates strong visual representation capabilities, achieving results on par with ViT-22B in vision tasks such as semantic segmentation. Code and models are available at https://github.com/bytedance/SAIL.

cs.CV↗

CHIME/FRB Outriggers: Design Overview

The Canadian Hydrogen Intensity Mapping Experiment (CHIME) has emerged as the world's premier facility for studying fast radio bursts (FRBs) through its fast transient search backend CHIME/FRB\@. The CHIME/FRB Outriggers project will augment this high detection rate of 2--3 FRBs per day with the ability to precisely localize them using very long baseline interferometry (VLBI). Using three strategically located stations in North America and deploying recently developed synoptic VLBI observing techniques, the Outriggers will provide $\sim 50$~milliarcsecond localization precision for the majority of detected FRBs. This paper presents an overview of the design and implementation of the Outriggers, covering their geographic distribution, structural design, and observational capabilities. We detail the scientific objectives driving the project, including the characterization of FRB populations, host galaxy demographics, and the use of FRBs as cosmological probes. We also discuss the calibration strategies available to mitigate ionospheric and instrumental effects, ensuring high-precision localization. With two stations currently in science operations, and the third in commissioning, the CHIME/FRB Outriggers project is poised to become a cornerstone of the FRB field, offering unprecedented insights into this enigmatic cosmic phenomenon.

astro-ph.HE↗

Ross3D: Reconstructive Visual Instruction Tuning with 3D-Awareness

The rapid development of Large Multimodal Models (LMMs) for 2D images and videos has spurred efforts to adapt these models for interpreting 3D scenes. However, the absence of large-scale 3D vision-language datasets has posed a significant obstacle. To address this issue, typical approaches focus on injecting 3D awareness into 2D LMMs by designing 3D input-level scene representations. This work provides a new perspective. We introduce reconstructive visual instruction tuning with 3D-awareness (Ross3D), which integrates 3D-aware visual supervision into the training procedure. Specifically, it incorporates cross-view and global-view reconstruction. The former requires reconstructing masked views by aggregating overlapping information from other views. The latter aims to aggregate information from all available views to recover Bird's-Eye-View images, contributing to a comprehensive overview of the entire scene. Empirically, Ross3D achieves state-of-the-art performance across various 3D scene understanding benchmarks. More importantly, our semi-supervised experiments demonstrate significant potential in leveraging large amounts of unlabeled 3D vision-only data.

cs.CV↗