Searcharxiv⌕ Search

arXiv subjects

Haochen Wang

Publications and source records attributed to Haochen Wang.

At least 73 records · Page 4Linked to original sources

Ross3D: Reconstructive Visual Instruction Tuning with 3D-Awareness

The rapid development of Large Multimodal Models (LMMs) for 2D images and videos has spurred efforts to adapt these models for interpreting 3D scenes. However, the absence of large-scale 3D vision-language datasets has posed a significant obstacle. To address this issue, typical approaches focus on injecting 3D awareness into 2D LMMs by designing 3D input-level scene representations. This work provides a new perspective. We introduce reconstructive visual instruction tuning with 3D-awareness (Ross3D), which integrates 3D-aware visual supervision into the training procedure. Specifically, it incorporates cross-view and global-view reconstruction. The former requires reconstructing masked views by aggregating overlapping information from other views. The latter aims to aggregate information from all available views to recover Bird's-Eye-View images, contributing to a comprehensive overview of the entire scene. Empirically, Ross3D achieves state-of-the-art performance across various 3D scene understanding benchmarks. More importantly, our semi-supervised experiments demonstrate significant potential in leveraging large amounts of unlabeled 3D vision-only data.

cs.CV↗

Transfer learning empowers material Z classification with muon tomography

Cosmic-ray muon sources exhibit distinct scattering angle distributions when interacting with materials of different atomic numbers (Z values), facilitating the identification of various Z-class materials, particularly those radioactive high-Z nuclear elements. Most of the traditional identification methods are based on complex muon event reconstruction and trajectory fitting processes. Supervised machine learning methods offer some improvement but rely heavily on prior knowledge of target materials, significantly limiting their practical applicability in detecting concealed materials. For the first time, transfer learning is introduced into the field of muon tomography in this work. We propose two lightweight neural network models for fine-tuning and adversarial transfer learning, utilizing muon tomography data of bare materials to predict the Z-class of coated materials. By employing the inverse cumulative distribution function method, more accurate scattering angle distributions could be obtained from limited data, leading to an improvement by nearly 4\% in prediction accuracy compared with the traditional random sampling based training. When applied to coated materials with limited labeled or even unlabeled muon tomography data, the proposed method achieves an overall prediction accuracy exceeding 96\%, with high-Z materials reaching nearly 99\%. Simulation results indicate that transfer learning improves prediction accuracy by approximately 10\% compared to direct prediction without transfer. This study demonstrates the effectiveness of transfer learning in overcoming the physical challenges associated with limited labeled/unlabeled data, highlights the promising potential of transfer learning in the field of muon tomography.

physics.ins-det↗

A Catalog of Local Universe Fast Radio Bursts from CHIME/FRB and the KKO Outrigger

We present the first catalog of fast radio burst (FRB) host galaxies from CHIME/FRB Outriggers, selected uniformly in the radio and the optical by localizing 81 new bursts to 2'' x ~60'' accuracy using CHIME and the KKO Outrigger, located 66 km from CHIME. Of the 81 localized bursts, we use the Probabilistic Association of Transients to their Hosts (PATH) algorithm to securely identify 21 new FRB host galaxies, and compile spectroscopic redshifts for 19 systems, 15 of which are newly obtained via spectroscopic observations. The most nearby source is FRB 20231229A, at a distance of 90 Mpc. One burst in our sample is from a previously reported repeating source in a galaxy merger (FRB 20190303A). Three new FRB host galaxies (FRBs 20230203A, 20230703A, and 20231206A) are found towards X-ray and optically selected galaxy clusters, potentially doubling the sample of known galaxy cluster FRBs. A search for radio counterparts reveals that FRB 20231128A is associated with a luminous persistent radio source (PRS) candidate with high significance ($P_{cc} \sim 10^{-2}$). If its compactness is confirmed, it would be the nearest known compact PRS at $z = 0.1079$. Our catalog significantly increases the statistics of the Macquart relation at low redshifts ($z < 0.2$). In the near future, the completed CHIME/FRB Outriggers array will produce hundreds of FRBs localized with very long baseline interferometry (VLBI). This will significantly expand the known sample and pave the way for future telescopes relying on VLBI for FRB localization.

astro-ph.HE↗

DocVideoQA: Towards Comprehensive Understanding of Document-Centric Videos through Question Answering

Remote work and online courses have become important methods of knowledge dissemination, leading to a large number of document-based instructional videos. Unlike traditional video datasets, these videos mainly feature rich-text images and audio that are densely packed with information closely tied to the visual content, requiring advanced multimodal understanding capabilities. However, this domain remains underexplored due to dataset availability and its inherent complexity. In this paper, we introduce the DocVideoQA task and dataset for the first time, comprising 1454 videos across 23 categories with a total duration of about 828 hours. The dataset is annotated with 154k question-answer pairs generated manually and via GPT, assessing models' comprehension, temporal awareness, and modality integration capabilities. Initially, we establish a baseline using open-source MLLMs. Recognizing the challenges in modality comprehension for document-centric videos, we present DV-LLaMA, a robust video MLLM baseline. Our method enhances unimodal feature extraction with diverse instruction-tuning data and employs contrastive learning to strengthen modality integration. Through fine-tuning, the LLM is equipped with audio-visual capabilities, leading to significant improvements in document-centric video understanding. Extensive testing on the DocVideoQA dataset shows that DV-LLaMA significantly outperforms existing models. We'll release the code and dataset to facilitate future research.

cs.CV↗

Reconstructive Visual Instruction Tuning

This paper introduces reconstructive visual instruction tuning (ROSS), a family of Large Multimodal Models (LMMs) that exploit vision-centric supervision signals. In contrast to conventional visual instruction tuning approaches that exclusively supervise text outputs, ROSS prompts LMMs to supervise visual outputs via reconstructing input images. By doing so, it capitalizes on the inherent richness and detail present within input images themselves, which are often lost in pure text supervision. However, producing meaningful feedback from natural images is challenging due to the heavy spatial redundancy of visual signals. To address this issue, ROSS employs a denoising objective to reconstruct latent representations of input images, avoiding directly regressing exact raw RGB values. This intrinsic activation design inherently encourages LMMs to maintain image detail, thereby enhancing their fine-grained comprehension capabilities and reducing hallucinations. Empirically, ROSS consistently brings significant improvements across different visual encoders and language models. In comparison with extrinsic assistance state-of-the-art alternatives that aggregate multiple visual experts, ROSS delivers competitive performance with a single SigLIP visual encoder, demonstrating the efficacy of our vision-centric supervision tailored for visual outputs.

cs.CV↗

The classification of real and bogus transients using active learning and semi-supervised learning

Deep-learning-based methods have been favored in astrophysics owing to their adaptability and remarkable performance and have been applied to the task of the classification of real and bogus transients. Different from most existing approaches which necessitate massive yet expensive annotated data, We aim to leverage training samples with only 1000 labels available to discover real sources that vary in brightness over time in the early stage of the WFST 6-year survey. Methods. We present a novel deep-learning method that combines active learning and semi-supervised learning to construct a competitive real/bogus classifier. Our method incorporates an active learning stage, where we actively select the most informative or uncertain samples for annotation. This stage aims to achieve higher model performance by leveraging fewer labeled samples, thus reducing annotation costs and improving the overall learning process efficiency. Furthermore, our approach involves a semi-supervised learning stage that exploits the unlabeled data to enhance the model's performance and achieve superior results compared to using only the limited labeled data.

astro-ph.IM↗

OpenSatMap: A Fine-grained High-resolution Satellite Dataset for Large-scale Map Construction

In this paper, we propose OpenSatMap, a fine-grained, high-resolution satellite dataset for large-scale map construction. Map construction is one of the foundations of the transportation industry, such as navigation and autonomous driving. Extracting road structures from satellite images is an efficient way to construct large-scale maps. However, existing satellite datasets provide only coarse semantic-level labels with a relatively low resolution (up to level 19), impeding the advancement of this field. In contrast, the proposed OpenSatMap (1) has fine-grained instance-level annotations; (2) consists of high-resolution images (level 20); (3) is currently the largest one of its kind; (4) collects data with high diversity. Moreover, OpenSatMap covers and aligns with the popular nuScenes dataset and Argoverse 2 dataset to potentially advance autonomous driving technologies. By publishing and maintaining the dataset, we provide a high-quality benchmark for satellite-based map construction and downstream tasks like autonomous driving.

cs.CV↗

Pulling Target to Source: A New Perspective on Domain Adaptive Semantic Segmentation

Domain adaptive semantic segmentation aims to transfer knowledge from a labeled source domain to an unlabeled target domain. However, existing methods primarily focus on directly learning qualified target features, making it challenging to guarantee their discrimination in the absence of target labels. This work provides a new perspective. We observe that the features learned with source data manage to keep categorically discriminative during training, thereby enabling us to implicitly learn adequate target representations by simply \textbf{pulling target features close to source features for each category}. To this end, we propose T2S-DA, which we interpret as a form of pulling Target to Source for Domain Adaptation, encouraging the model in learning similar cross-domain features. Also, considering the pixel categories are heavily imbalanced for segmentation datasets, we come up with a dynamic re-weighting strategy to help the model concentrate on those underperforming classes. Extensive experiments confirm that T2S-DA learns a more discriminative and generalizable representation, significantly surpassing the state-of-the-art. We further show that our method is quite qualified for the domain generalization task, verifying its domain-invariant property.

cs.CV↗

Towards higher electro-optic response in AlScN

Novel materials with large electro-optic (EO) coefficients are essential for developing ultra-compact broadband modulators and enabling effective quantum transduction. Compared to lithium niobate, the most widely used nonlinear optical material, wurtzite AlScN offers advantages in nano-photonic devices due to its compatibility with integrated circuits. We perform detailed first-principles calculations to investigate the electro-optic effect in $\mathrm{Al}_{1-x}\mathrm{Sc}_{x}\mathrm{N}$ alloys and superlattices. At elevated Sc concentrations in alloys, the EO coefficients increase; importantly, we find that cation ordering along the $c$ axis leads to enhanced EO response. Strain engineering can be used to further manipulate the EO coefficients of AlScN films. With applied in-plane strains, the piezoelectric contributions to the EO coefficients increase dramatically, even exceeding 251 pm/V. We also explore the possibility of EO enhancement through superlattice engineering, finding that nonpolar $a$-plane $\mathrm{(AlN)}_m/\mathrm{(ScN)}_n$ superlattices increase EO coefficients beyond 40 pm/V. Our findings provide design principles to enhance the electro-optic effect through alloy engineering and heterostructure architecture.

cond-mat.mtrl-sci↗

Unveiling the Pockels Coefficient of Ferroelectric Nitride ScAlN

Nitride ferroelectrics have recently emerged as promising alternatives to oxide ferroelectrics due to their compatibility with mainstream semiconductor processing. ScAlN, in particular, has exhibited remarkable piezoelectric coupling strength ($K^2$) comparable to that of lithium niobate (LN), making it a valuable choice for RF filters in wireless communications. Recently, ScAlN has sparked interest in its use for nanophotonic devices, chiefly due to its large bandgap facilitating operation in blue wavelengths coupled with promises of enhanced nonlinear optical properties such as a large second-order susceptibility ($χ^{(2)}$). It is still an open question whether ScAlN can outperform oxide ferroelectrics concerning the Pockels effect -- an electro-optic coupling extensively utilized in optical communications devices. In this paper, we present a comprehensive theoretical analysis and experimental demonstration of ScAlN's Pockels effect. Our findings reveal that the electro-optic coupling of ScAlN, despite being weak at low Sc concentration, may be significantly enhanced and exceed LiNbO$_3$ at high levels of Sc doping, which points the direction of continued research efforts to unlock the full potential of ScAlN.

cond-mat.mtrl-sci↗

CJEval: A Benchmark for Assessing Large Language Models Using Chinese Junior High School Exam Data

Online education platforms have significantly transformed the dissemination of educational resources by providing a dynamic and digital infrastructure. With the further enhancement of this transformation, the advent of Large Language Models (LLMs) has elevated the intelligence levels of these platforms. However, current academic benchmarks provide limited guidance for real-world industry scenarios. This limitation arises because educational applications require more than mere test question responses. To bridge this gap, we introduce CJEval, a benchmark based on Chinese Junior High School Exam Evaluations. CJEval consists of 26,136 samples across four application-level educational tasks covering ten subjects. These samples include not only questions and answers but also detailed annotations such as question types, difficulty levels, knowledge concepts, and answer explanations. By utilizing this benchmark, we assessed LLMs' potential applications and conducted a comprehensive analysis of their performance by fine-tuning on various educational tasks. Extensive experiments and discussions have highlighted the opportunities and challenges of applying LLMs in the field of education.

cs.AI↗

A VLBI Calibrator Grid at 600MHz for Fast Radio Transient Localizations with CHIME/FRB Outriggers

The Canadian Hydrogen Intensity Mapping Experiment Fast Radio Burst (CHIME/FRB) Project has a new VLBI Outrigger at the Green Bank Observatory (GBO), which forms a 3300km baseline with CHIME operating at 400-800MHz. Using 100ms long full-array baseband "snapshots" collected commensally during FRB and pulsar triggers, we perform a shallow, wide-area VLBI survey covering a significant fraction of the Northern sky targeted at the positions of compact sources from the Radio Fundamental Catalog. In addition, our survey contains calibrators detected from two 1s long trial baseband snapshots for a deeper survey with CHIME and GBO. In this paper, we present the largest catalog of compact calibrators suitable for 30-milliarcsecond-scale VLBI observations at sub-GHz frequencies to date. Our catalog consists of 200 total calibrators in the Northern Hemisphere that are compact on 30-milliarcsecond scales with fluxes above 100mJy. This calibrator grid will enable the precise localization of hundreds of FRBs a year with CHIME/FRB-Outriggers.

astro-ph.IM↗

Exploring Accessibility Trends and Challenges in Mobile App Development: A Study of Stack Overflow Questions

The proliferation of mobile applications (apps) has made it crucial to ensure their accessibility for users with disabilities. However, there is a lack of research on the real-world challenges developers face in implementing mobile accessibility features. This study presents a large-scale empirical analysis of accessibility discussions on Stack Overflow to identify the trends and challenges Android and iOS developers face. We examine the growth patterns, characteristics, and common topics mobile developers discuss. Our results show several challenges, including integrating assistive technologies like screen readers, ensuring accessible UI design, supporting text-to-speech across languages, handling complex gestures, and conducting accessibility testing. We envision our findings driving improvements in developer practices, research directions, tool support, and educational resources.

cs.SE↗

DocTabQA: Answering Questions from Long Documents Using Tables

We study a new problem setting of question answering (QA), referred to as DocTabQA. Within this setting, given a long document, the goal is to respond to questions by organizing the answers into structured tables derived directly from the document's content. Unlike traditional QA approaches which predominantly rely on unstructured text to formulate responses, DocTabQA aims to leverage structured tables as answers to convey information clearly and systematically, thereby enhancing user comprehension and highlighting relationships between data points. To the best of our knowledge, this problem has not been previously explored. In this paper, we introduce the QTabA dataset, encompassing 300 financial documents, accompanied by manually annotated 1.5k question-table pairs. Initially, we leverage Large Language Models (LLMs) such as GPT-4 to establish a baseline. However, it is widely acknowledged that LLMs encounter difficulties when tasked with generating intricate, structured outputs from long input sequences. To overcome these challenges, we present a two-stage framework, called DocTabTalk, which initially retrieves relevant sentences from extensive documents and subsequently generates hierarchical tables based on these identified sentences. DocTabTalk incorporates two key technological innovations: AlignLLaMA and TabTalk, which are specifically tailored to assist GPT-4 in tackling DocTabQA, enabling it to generate well-structured, hierarchical tables with improved organization and clarity. Comprehensive experimental evaluations conducted on both QTabA and RotoWire datasets demonstrate that our DocTabTalk significantly enhances the performances of the GPT-4 in our proposed DocTabQA task and the table generation task. The code and dataset are available at https://github.com/SmileWHC/DocTabQA for further research.

cs.CL↗

Using Unreliable Pseudo-Labels for Label-Efficient Semantic Segmentation

The crux of label-efficient semantic segmentation is to produce high-quality pseudo-labels to leverage a large amount of unlabeled or weakly labeled data. A common practice is to select the highly confident predictions as the pseudo-ground-truths for each pixel, but it leads to a problem that most pixels may be left unused due to their unreliability. However, we argue that every pixel matters to the model training, even those unreliable and ambiguous pixels. Intuitively, an unreliable prediction may get confused among the top classes, however, it should be confident about the pixel not belonging to the remaining classes. Hence, such a pixel can be convincingly treated as a negative key to those most unlikely categories. Therefore, we develop an effective pipeline to make sufficient use of unlabeled data. Concretely, we separate reliable and unreliable pixels via the entropy of predictions, push each unreliable pixel to a category-wise queue that consists of negative keys, and manage to train the model with all candidate pixels. Considering the training evolution, we adaptively adjust the threshold for the reliable-unreliable partition. Experimental results on various benchmarks and training settings demonstrate the superiority of our approach over the state-of-the-art alternatives.

cs.CV↗

Demonstration of hybrid foreground removal on CHIME data

The main challenge of 21 cm cosmology experiments is astrophysical foregrounds which are difficult to separate from the signal due to telescope systematics. An earlier study has shown that foreground residuals induced by antenna gain errors can be estimated and subtracted using the hybrid foreground residual subtraction (HyFoReS) technique which relies on cross-correlating linearly filtered data. In this paper, we apply a similar technique to the CHIME stacking analysis to subtract beam-induced foreground contamination. Using a linear high-pass delay filter for foreground suppression, the CHIME collaboration reported a $11.1σ$ detection in the 21 cm signal stacked on eBOSS quasar locations, despite foreground residual contamination mostly due to the instrument chromatic transfer function. We cross-correlate the foreground-dominated data at low delay with the contaminated signal at high delay to estimate residual foregrounds and subtract them from the signal. We find foreground residual subtraction can improve the signal-to-noise ratio of the stacked 21 cm signal by $ 10 - 20\%$ after the delay foreground filter, although some of the improvement can also be achieved with an alternative flagging technique. We have shown that it is possible to use HyFoReS to reduce beam-induced foreground contamination, benefiting the analysis of the HI auto power spectrum with CHIME and enabling the recovery of large scale modes.

astro-ph.CO↗

Faraday tomography with CHIME: the `tadpole' feature G137+7

A direct consequence of Faraday rotation is that the polarized radio sky does not resemble the total intensity sky at long wavelengths. We analyze G137+7, which is undetectable in total intensity but appears as a depolarization feature. We use the first polarization maps from the Canadian Hydrogen Intensity Mapping Experiment. Our $400-729$ MHz bandwidth and angular resolution, $17'$ to $30'$, allow us to use Faraday synthesis to analyze the polarization structure. In polarized intensity and polarization angle maps, we find a "tail" extending $10^\circ$ from the "head" and designate the combined object the "tadpole". Similar polarization angles, distinct from the background, indicate that the head and tail are physically associated. The head appears as a depolarized ring in single channels, but wideband observations show that it is a Faraday rotation feature. Our investigations of H I and H$α$ find no connections to the tadpole. The tail suggests motion of either the gas or an ionizing star through the ISM; the B2(e) star HD 20336 is a candidate. While the head features a coherent, $\sim -8$ rad m$^2$ Faraday depth, Faraday synthesis also identifies multiple components in both the head and tail. We verify the locations of the components in the spectra using QU fitting. Our results show that $\sim$octave-bandwidth Faraday rotation observations at $\sim 600$ MHz are sensitive to low-density ionized or partially-ionized gas which is undetectable in other tracers.

astro-ph.GA↗

Holographic Beam Measurements of the Canadian Hydrogen Intensity Mapping Experiment (CHIME)

We present the first results of the holographic beam mapping program for the Canadian Hydrogen Intensity Mapping Experiment (CHIME). We describe the implementation of the holographic technique as adapted for CHIME, and introduce the processing pipeline which prepares the raw holographic timestreams for analysis of beam features. We use data from six bright sources across the full 400-800\,MHz observing band of CHIME to provide measurements of the co-polar and cross-polar beam response of CHIME in both amplitude and phase for the 1024 dual-polarized feeds instrumented on CHIME. In addition, we present comparisons with independent probes of the CHIME beam which indicate the presence of polarized beam leakage in CHIME. Holographic measurements of the CHIME beam have already been applied in science with CHIME, e.g. in estimating detection significance of far sidelobe FRBs, and in validating the beam models used for CHIME's first detections of \tcm emission (in cross-correlation with measurements of large-scale structure from galaxy surveys and the Lyman-$α$ forest). Measurements presented in this paper, and future holographic results, will provide a unique data set to characterize the CHIME beam and improve the experiment's prospects for a detection of BAO.

astro-ph.IM↗