SearcharxivSearch

arXiv subjects

Haochen Wang

Publications and source records attributed to Haochen Wang.

At least 37 records · Page 2Linked to original sources

Fast radio burst dispersion is an unbiased tracer of matter on large scales

The dispersion of fast radio bursts (FRBs) measures the column density of free electrons, tracing the diffuse ionized gas that contains more than $90\%$ of all baryons. On linear scales the FRB dispersion field is an approximately unbiased tracer of the matter distribution, an idea long assumed in the FRB large-scale structure literature and recently formalized by Zhou and Zhang [arXiv:2510.11022]. This follows from baryon-mass conservation, which forces the total baryon field to have unit linear bias, with dispersion inheriting this bias up to small corrections from the stellar and neutral-gas components. We show these corrections can be bounded at the percent level using existing galaxy and 21 cm surveys, and confirm with the FLAMINGO hydrodynamical simulations that the electron bias varies at the percent level across a wide range of feedback prescriptions. The dispersion-galaxy cross-power spectrum at linear scales directly constrains $B_8 \equiv σ_8(Ω_b/0.05)^{1/2}$, a baryonic analog of $S_8$, independently of feedback physics. Because most of the per-object variance in dispersion is cosmological signal rather than noise, $\sim\!10^5$ localized FRBs can match the statistical power of $\sim\!10^8$ weak-lensing galaxy shape measurements. FRB dispersion thus joins weak lensing and redshift-space distortions as a new unbiased tracer of matter on large scales.

astro-ph.CO

Constraining Gas Mass Fractions in Galaxy Groups and Clusters with the First CHIME/FRB Outrigger

In recent years, localized fast radio bursts (FRBs) have emerged as a powerful tool to study the structure of the baryonic matter in the universe. Their dispersion measures (DMs) scale linearly with electron density independent of gas temperature, making them particularly well suited to studying the intragroup medium (IGrM), where traditional probes such as X-ray emission and the SZ effect are weak. Evidence suggests that the gas in group mass halos ($M_{500}$ ~ $10^{13}$ -- $10^{14}$ M$_\odot$) is strongly affected by galactic feedback, causing deviations from cluster scaling relations. Three FRBs from the first CHIME/FRB Outrigger sample come from host galaxies found within or behind galaxy clusters and groups. We estimate the DM contribution of each ICM/IGrM by integrating different halo density profiles, accounting for uncertainties in halo mass and the host galaxy line of sight distance. For the more massive halos, predicted cluster DMs agree with the extragalactic DM budget. One burst, FRB 20230703A, intersects three groups yet has a low extragalactic DM. By comparing model predictions with the measured DM, we constrain the gas mass fraction $f_g(R)$ in these halos. Comparing with published $M$--$f_g$ relations, we find consistency with recent eROSITA results at $R_{500}$ and mild tension at $R_{200}$ and with earlier X-ray--based relations. As CHIME/FRB Outriggers build a large catalog of localized FRBs, many additional sightlines through groups and clusters will be obtained. These will enable systematic tests of intragroup and intracluster gas properties and sharpen constraints on the distribution of baryons in massive halos.

astro-ph.GA

VideoZeroBench: Probing the Limits of Video MLLMs with Spatio-Temporal Evidence Verification

Recent video multimodal large language models achieve impressive results across various benchmarks. However, current evaluations suffer from two critical limitations: (1) inflated scores can mask deficiencies in fine-grained visual understanding and reasoning, and (2) answer correctness is often measured without verifying whether models identify the precise spatio-temporal evidence supporting their predictions. To address this, we present VideoZeroBench, a hierarchical benchmark designed for challenging long-video question answering that rigorously verifies spatio-temporal evidence. It comprises 500 manually annotated questions across 13 domains, paired with temporal intervals and spatial bounding boxes as evidence. To disentangle answering generation, temporal grounding, and spatial grounding, we introduce a five-level evaluation protocol that progressively tightens evidence requirements. Experiments show that even Gemini-3-Pro correctly answers fewer than 17% of questions under the standard end-to-end QA setting (Level-3). When grounding constraints are imposed, performance drops sharply: No model exceeds 1% accuracy when both correct answering and accurate spatio-temporal localization are required (Level-5), with most failing to achieve any correct grounded predictions. These results expose a significant gap between surface-level answer correctness and genuine evidence-based reasoning, revealing that grounded video understanding remains a bottleneck for long-video QA. We further analyze performance across minimal evidence spans, atomic abilities, and inference paradigms, providing insights for future research in grounded video reasoning. The benchmark and code will be made publicly available.

cs.CV

Unified Multimodal Models as Auto-Encoders

Image-to-text (I2T) understanding and text-to-image (T2I) generation are two fundamental, important yet traditionally isolated multimodal tasks. Despite their intrinsic connection, existing approaches typically optimize them independently, missing the opportunity for mutual enhancement. In this paper, we argue that the both tasks can be connected under a shared Auto-Encoder perspective, where text serves as the intermediate latent representation bridging the two directions - encoding images into textual semantics (I2T) and decoding text back into images (T2I). Our key insight is that if the encoder truly "understands" the image, it should capture all essential structure, and if the decoder truly "understands" the text, it should recover that structure faithfully. Building upon this principle, we propose Unified-GRPO, a post-training method based on reinforcement learning that jointly optimizes both modules through reconstructive rewards, maximizing the semantic consistency between the input and the generated images. Under this reconstruction objective, the encoder is encouraged to extract as much accurate and comprehensive semantic information from the input image to maximize reconstruction quality, while the decoder is simultaneously optimized to generate conditioned on the encoder's prior, enabling a self-evolving improvement. Empirically, we find that using text as the intermediate representation and training under a reconstructive RL paradigm effectively benefits both I2T and T2I. The I2T module gains stronger fine-grained visual perception, such as small-object recognition, grounding, etc, while its dense embeddings and language priors, in turn, provide richer semantic signals that improve T2I fidelity and complex instruction following. These results demonstrate that the reconstructive RL establishes a mutually reinforcing cross-modal synergy within the auto-encoding framework.

cs.CV

Detection of the Cosmological 21 cm Signal in Auto-correlation at z ~ 1 with the Canadian Hydrogen Intensity Mapping Experiment

We present the first detection of the cosmological 21 cm intensity mapping signal in auto-correlation at z ~ 1 with the Canadian Hydrogen Intensity Mapping Experiment (CHIME). Using 94 nights of observation, we have measured the 21 cm auto-power spectrum over a frequency range from 608.2 MHz to 707.8 MHz (z = 1.34 to 1.01) at 0.4 h Mpc^-1 < k < 1.5 h Mpc^-1, with a detection significance of 12.5 sigma. Our analysis employs significant improvements to the CHIME data processing pipeline compared to previous work, including novel radio frequency interference (RFI) detection and masking algorithms, achromatic beamforming techniques, and foreground filtering before time averaging to minimize spectral leakage. We establish the robustness and reliability of our detection through a comprehensive suite of validation tests. We also measure the 21 cm signal in two independent sub-bands centered at z ~ 1.08 and z ~ 1.24 with detection significance of 8.7 sigma and 9.2 sigma, respectively. We briefly discuss the theoretical interpretation of these measurements in terms of a power spectrum model, deferring the details to a companion paper. This auto-power spectrum detection demonstrates CHIME's capability to probe large-scale structure through 21 cm intensity mapping without reliance on external galaxy surveys.

astro-ph.CO

Interpretation of 21 cm Auto Power Spectrum Measurement at $z\sim 1$ by the Canadian Hydrogen Intensity Mapping Experiment

Observations with the Canadian Hydrogen Intensity Mapping Experiment (CHIME) have been used to measure the 21 cm intensity mapping auto power spectrum, at $z\sim 1$, over a frequency range from 608.2 MHz to 707.8 MHz at wavenumbers $0.4~h~{\rm Mpc}^{-1} \lesssim k \lesssim 1.5~h~{\rm Mpc}^{-1}$. In this paper, we present the results of two different approaches to interpreting this measurement. In the first approach, we use a parametric power spectrum model to constrain an amplitude parameter, defined as $\mathcal{A}^2_{\rm HI} \equiv 10^6 Ω_{\rm HI}^2(b^2_{\rm HI}+\langle f μ^2\rangle)^2$, where $Ω_{\rm HI}$ is the cosmological density parameter for atomic hydrogen ($\rm HI$), $b_{\rm HI}$ is the linear bias for $\rm HI$, and $\langle f μ^2\rangle$ incorporates the dominant large-scale impact of redshift-space distortions on the angle-averaged power spectrum. Imposing an additional prior on either $Ω_{\rm HI}$ or $b_{\rm HI}$, based on values in the literature, allows us to break the pairwise degeneracy between those two parameters. In the second approach, we compare CHIME's measurement with predictions for the power spectrum of $\rm HI$ from the IllustrisTNG simulations, finding that the measurement disagrees with the TNG100 run at $3.1σ$ and the TNG300 run at $4.0σ$. This disagreement is most likely attributable to the strength of nonlinear redshift-space clustering of $\rm HI$ in the simulations, rather than the total abundance of $\rm HI$, and invites further investigation of the physical processes in the simulations that determine the behavior of $\rm HI$ at nonlinear scales. These results exemplify the ability of 21 cm intensity mapping to provide astrophysical information using measurements at nonlinear scales.

astro-ph.CO

A spatial filter for mitigating radio interference and its application to CHIME/FRB Outriggers

The sensitivity of radio telescopes is becoming increasingly limited by the presence of radio frequency interference (RFI), which will worsen as the radio spectrum becomes more crowded. One context where this poses a challenge is the field of fast radio burst (FRB) science, where there is increasing scientific interest in capturing as large of a population of bursts as possible and accurately measuring their celestial coordinates using interferometry. With several modern radio facilities actively collecting data for large FRB surveys that will be transformative to the field, properly mitigating unwanted interference is essential for the science goals of these surveys to be met. In this work, we present variations of a spatial filter based on the Karhunen-Loeve (KL) Transform to enhance the sensitivity of radio interferometers and demonstrate its applicability to FRB detection and localization. We derive a particular variation of the filter for the case of point-like radio pulses, which we show reduces to the maximum-signal-to-noise beamformer. We apply this filter to CHIME/FRB baseband data and demonstrate its capability to enhance the sensitivity and overall localization rate of CHIME/FRB Outriggers. We compare the cross-correlation signal-to-noise obtained using the spatial filter with that obtained using a spectral-kurtosis RFI flagger for a sample of 100 FRBs recorded by CHIME and its Outriggers, and show that this filter will double the total number of FRBs successfully localized with the CHIME/FRB Outrigger telescopes. While demonstrated here in the context of CHIME/FRB Outriggers, the spatial filter presented in this work--which we have made publicly available--is broadly applicable to other interferometric radio facilities engaged in FRB science and transient detection, including next-generation telescopes such as CHORD, DSA-2000, BURSTT, and CHARTS.

astro-ph.IM

Open-o3-Video: Grounded Video Reasoning with Explicit Spatio-Temporal Evidence

Most video reasoning models only generate textual reasoning traces without indicating when and where key evidence appears. Recent models such as OpenAI-o3 have sparked wide interest in evidence-centered reasoning for images, yet extending this ability to videos is more challenging due to the need for joint temporal tracking and spatial localization across dynamic scenes. We introduce Open-o3-Video, a non-agent framework that integrates explicit spatio-temporal evidence into video reasoning by highlighting key timestamps, objects, and bounding boxes, making the reasoning process traceable and verifiable. To enable this capability, we first construct high-quality datasets STGR that provide unified spatio-temporal supervision, which is absent in existing resources. We further adopt a cold-start reinforcement learning strategy with specially designed rewards that jointly encourage answer accuracy, temporal alignment, and spatial precision. On the V-STAR benchmark, Open-o3-Video achieves state-of-the-art performance, improving mAM by 14.4% and mLGM by 24.2% over the Qwen2.5-VL baseline, and shows consistent gains across a range of video understanding benchmarks. Beyond accuracy, the grounded reasoning traces produced by Open-o3-Video support confidence-aware test-time scaling, improving answer reliability.

cs.CV

Traceable Evidence Enhanced Visual Grounded Reasoning: Evaluation and Methodology

Models like OpenAI-o3 pioneer visual grounded reasoning by dynamically referencing visual regions, just like human "thinking with images". However, no benchmark exists to evaluate these capabilities holistically. To bridge this gap, we propose TreeBench (Traceable Evidence Evaluation Benchmark), a diagnostic benchmark built on three principles: (1) focused visual perception of subtle targets in complex scenes, (2) traceable evidence via bounding box evaluation, and (3) second-order reasoning to test object interactions and spatial hierarchies beyond simple object localization. Prioritizing images with dense objects, we initially sample 1K high-quality images from SA-1B, and incorporate eight LMM experts to manually annotate questions, candidate options, and answers for each image. After three stages of quality control, TreeBench consists of 405 challenging visual question-answering pairs, even the most advanced models struggle with this benchmark, where none of them reach 60% accuracy, e.g., OpenAI-o3 scores only 54.87. Furthermore, we introduce TreeVGR (Traceable Evidence Enhanced Visual Grounded Reasoning), a training paradigm to supervise localization and reasoning jointly with reinforcement learning, enabling accurate localizations and explainable reasoning pathways. Initialized from Qwen2.5-VL-7B, it improves V* Bench (+16.8), MME-RealWorld (+12.6), and TreeBench (+13.4), proving traceability is key to advancing vision-grounded reasoning. The code is available at https://github.com/Haochen-Wang409/TreeVGR.

cs.CV

Grasp Any Region: Towards Precise, Contextual Pixel Understanding for Multimodal LLMs

While Multimodal Large Language Models (MLLMs) excel at holistic understanding, they struggle in capturing the dense world with complex scenes, requiring fine-grained analysis of intricate details and object inter-relationships. Region-level MLLMs have been a promising step. However, previous attempts are generally optimized to understand given regions in isolation, neglecting crucial global contexts. To address this, we introduce Grasp Any Region (GAR) for comprehen- sive region-level visual understanding. Empowered by an effective RoI-aligned feature replay technique, GAR supports (1) precise perception by leveraging necessary global contexts, and (2) modeling interactions between multiple prompts. Together, it then naturally achieves (3) advanced compositional reasoning to answer specific free-form questions about any region, shifting the paradigm from passive description to active dialogue. Moreover, we construct GAR-Bench, which not only provides a more accurate evaluation of single-region comprehension, but also, more importantly, measures interactions and complex reasoning across multiple regions. Extensive experiments have demonstrated that GAR-1B not only maintains the state-of-the-art captioning capabilities, e.g., outperforming DAM-3B +4.5 on DLC-Bench, but also excels at modeling relationships between multiple prompts with advanced comprehension capabilities, even surpassing InternVL3-78B on GAR-Bench-VQA. More importantly, our zero-shot GAR-8B even outperforms in-domain VideoRefer-7B on VideoRefer-BenchQ, indicating its strong capabilities can be easily transferred to videos.

cs.CV

Strain effects on $n$-type doping in AlN

Controllable doping in AlN and its alloys is essential for deep-ultraviolet light sources. Ionization energies for donors in AlN ($\mathrm{Si_{Al}}$, $\mathrm{S_N}$, $\mathrm{Se_N}$) are high. We report first-principles calculations demonstrating that strain engineering can result in a reduction in ionization energies. The donor levels for $\mathrm{S_N}$ and $\mathrm{Se_N}$ shift closer to the conduction-band minimum (CBM) under in-plane tensile strains, driven by a downward shift of the CBM. The most widely used donor, $\mathrm{Si_{Al}}$, forms a $DX$ center in AlN. We find that a 2.5% in-plane tensile strain (which would be induced by pseudomorphic growth on GaN in experiment) shifts the ($+/-$) transition level from 271 meV to 98 meV below the CBM, which would enhance the electron concentration by three orders of magnitude. These results demonstrate that strain engineering offers an effective route to enhance doping levels in AlN.

cond-mat.mtrl-sci

PairUni: Pairwise Training for Unified Multimodal Language Models

Unified Vision-Language Models (UVLMs) perform both understanding and generation within a single architecture. Since these models rely on heterogeneous data and supervision, balancing both generation and understanding in reinforcement learning (RL) is challenging. To address this challenge, we propose PairUni, a unified framework that reorganizes data into understanding-generation (UG) pairs and aligns optimization accordingly. Specifically, we construct a unified paired dataset by synthesizing aligned instances via cross-modal semantic completion and retrieving semantically related samples. These paired structures expose cross-task semantic correspondences and support consistent policy learning. To leverage this structure, we present PairGRPO, a pair-aware variant based on Group Relative Policy Optimization. It assigns a similarity score to each pair to modulate the advantage, strengthening learning from well-aligned examples and reducing task interference. Extensive experiments across diverse UVLM architectures (Autoregressive and Discrete Diffusion) and scales (1B to 14B) demonstrate that PairUni yields consistent improvements over strong baselines. Notably, our method also demonstrates strong generalization by improving performance on image editing tasks without using any editing-specific data. Codes are available at https://github.com/Haochen-Wang409/PairUni.

cs.CL

U-Net Based Image Enhancement for Short-time Muon Scattering Tomography

Muon Scattering Tomography (MST) is a promising non-invasive inspection technique, yet the practical application of short-time MST is hindered by poor image quality due to limited muon flux. To address this limitation, we propose a U-Net-based framework trained on Point of Closest Approach (PoCA) images reconstructed with simulation MST data to enhance image quality. When applied to experimental MST data, the framework significantly improves image quality, increasing the Structural Similarity Index Measure (SSIM) from 0.7232 to 0.9699 and decreasing the Learned Perceptual Image Patch Similarity (LPIPS) from 0.3604 to 0.0270. These results demonstrate that our method can effectively enhance low-statistics MST images, thereby paving the way for the practical deployment of short-time MST.

eess.IV

Discovery of a 21 cm absorption system at z=2.327 with CHIME

We report the detection of a new 21 cm absorption system associated with the radio source NVSS J164725+375218 at a redshift of z=2.327, identified through a pilot survey conducted by the Canadian Hydrogen Intensity Mapping Experiment (CHIME). This is the fifth detection of an associated system at z > 2. By analyzing a subset of available data, we conduct a spectrally blind survey for 21 cm absorption systems within the redshift range of 0.78 to 2.55 along 202 lines of sight toward known sources in the declination range of 35 to 60 degrees. We detect three 21 cm absorbers: two previously known intervening systems and one newly discovered associated system. By fitting the absorption profiles with models containing one to three Gaussian components and selecting the best model using the Bayesian information criterion, we estimate the optical depth, velocity-integrated optical depth, and the ratio between the HI column density and the spin temperature of the absorption systems. These results demonstrate CHIME's ability to discover new absorbers, even in a small subset of its full dataset.

astro-ph.GA

Exploring selection biases in FRB dispersion-galaxy cross-correlations with magnetohydrodynamical simulations

The dispersion measure (DM) of fast radio bursts (FRBs) in conjunction with their redshifts can be used as powerful probes of the distribution of extragalactic plasma. With a large enough sample, the free-electron--galaxy power spectrum $P_{eg}$ can be measured by cross-correlating FRB DMs with galaxy positions. However, a precise measurement of $P_{eg}$ requires a careful investigation of selection effects: the probability of both observing the FRB DM and obtaining a host galaxy redshift depends on their properties. We ray trace through the magnetohydrodynamic simulation IllustrisTNG to investigate the impact of expected observational selection effects on FRB dispersion--galaxy angular cross-correlations with a sample of 3000 FRBs at $0.3\leq z\leq 0.4$ . Our results show that cross-correlations with such an FRB sample are robust to properties of the FRB host galaxy: this includes DM contributions from the FRB host and optical follow-up selection effects. We also find that such cross-correlations are robust to DM-dependent and scattering selection effects specific to the CHIME/FRB survey. However, a DM-dependent selection effect that cuts off the 10\% most dispersed FRBs at a fixed redshift shell can bias the amplitude of the cross-correlation signal by over 50\% at angular scales of $\sim 0.1^\circ$, corresponding to $\sim$ Mpc physical scales. Our findings highlight the importance of both measuring and accounting for selection effects present in existing FRB surveys, as well as mitigating DM-dependent selection effects in the design of upcoming FRB surveys aiming to probe large-scale structure with FRBs.

astro-ph.CO

The Second CHIME/FRB Catalog of Fast Radio Bursts

We present a catalog of 4539 fast radio bursts (FRBs) observed with the Canadian Hydrogen Intensity Mapping Experiment (CHIME) telescope between 25 July 2018 and 15 September 2023. These bursts originate from 3641 unique sources, including 981 bursts from 83 known repeating sources. For each FRB, the catalog provides a $O(10')$ estimate of sky location along with corresponding measurements of cumulative exposure time and survey sensitivity over the observing period. It includes a total-intensity dynamic spectrum between 400 and 800 MHz at 0.983 ms resolution. From this spectrum, we constrain a model of the burst morphology and measure key parameters such as arrival time, intrinsic temporal width, dispersion measure, scattering time, and flux density. This second catalog includes all FRBs from the first catalog, with every event reprocessed using a uniform and improved analysis framework. We show that previously published inferences remain valid under the updated measurements. We assess consistency of the detection rate across observational parameters, present initial distributions of burst properties, and outline ongoing and future studies that will use this catalog to investigate the nature of FRBs and their utility as astrophysical and cosmological probes.

astro-ph.HE

MMFormalizer: Multimodal Autoformalization in the Wild

Autoformalization, which translates natural language mathematics into formal statements to enable machine reasoning, faces fundamental challenges in the wild due to the multimodal nature of the physical world, where physics requires inferring hidden constraints (e.g., mass or energy) from visual elements. To address this, we propose MMFormalizer, which extends autoformalization beyond text by integrating adaptive grounding with entities from real-world mathematical and physical domains. MMFormalizer recursively constructs formal propositions from perceptually grounded primitives through recursive grounding and axiom composition, with adaptive recursive termination ensuring that every abstraction is supported by visual evidence and anchored in dimensional or axiomatic grounding. We evaluate MMFormalizer on a new benchmark, PhyX-AF, comprising 115 curated samples from MathVerse, PhyX, Synthetic Geometry, and Analytic Geometry, covering diverse multimodal autoformalization tasks. Results show that frontier models such as GPT-5 and Gemini-3-Pro achieve the highest compile and semantic accuracy, with GPT-5 excelling in physical reasoning, while geometry remains the most challenging domain. Overall, MMFormalizer provides a scalable framework for unified multimodal autoformalization, bridging perception and formal reasoning. To the best of our knowledge, this is the first multimodal autoformalization method capable of handling classical mechanics (derived from the Hamiltonian), as well as relativity, quantum mechanics, and thermodynamics. More details are available on our project page: MMFormalizer.github.io

cs.CL

The Squeezed Bispectrum from CHIME HI Emission and Planck CMB Lensing: Current Sensitivity and Forecasts

Line intensity mapping using atomic hydrogen (HI) has the potential to efficiently map large volumes of the universe if the signal can be successfully separated from overwhelmingly bright radio foreground emission. This motivates cross-correlations, to ascertain the cosmological nature of measured HI fluctuations, and to study their connections with galaxies and the underlying matter density field. However, these same foregrounds render the cross-correlation with projected fields such as the lensing of the cosmic microwave background (CMB) difficult. Indeed, the correlated Fourier modes vary slowly along the line of sight, and are thus most contaminated by the smooth-spectrum radio continuum foregrounds. In this paper, we implement a method that avoids this issue by attempting to measure the non-linear gravitational coupling of the small-scale 21cm power from the Canadian Hydrogen Intensity Mapping Experiment (CHIME) with large-scale Planck CMB lensing. This measurement is a position-dependent power spectrum, i.e. a squeezed integrated bispectrum. Using 94 nights of CHIME data between $1.0 < z < 1.3$ and aggressive foreground filtering, we find that the expected signal is five times smaller than the current noise. We forecast that incorporating the additional nights of CHIME data already collected would enable a signal-to-noise ratio of 3, without any further improvements in filtering for foreground cleaning.

astro-ph.CO