SearcharxivSearch

arXiv subjects

Huiqun Wang

Publications and source records attributed to Huiqun Wang.

10 recordsLinked to original sources

Mitigating Positional Leakage in 3D Masked Autoencoders for Robust Representation Learning

Masked autoencoding has emerged as a prominent paradigm for self-supervised learning on 3D point clouds, achieving competitive performance across downstream tasks. Unlike its 2D counterpart, 3D masked autoencoding directly reconstructs spatial coordinates, making it inherently susceptible to positional leakage. In this work, we identify that the decoder in existing 3D MAE frameworks tends to over-rely on positional information, which weakens semantic representation learning and leads to suboptimal feature quality. To address this issue, we propose MPL-MAE, a masked point learning framework that mitigates positional over-reliance while enhancing the utilization of encoder features. Specifically, we introduce a recalibrated positional embedding module that suppresses metric-dominant coordinate signals while preserving geometric topology, together with a gated positional interface module that dynamically regulates positional injection during reconstruction. These designs promote a more balanced interaction between spatial priors and semantic features, yielding robust and informative representations. Extensive experiments across downstream tasks demonstrate that MPL-MAE consistently achieves competitive performance, validating its effectiveness. Code is available at https://github.com/yanx57/MPL-MAE.

cs.CV

EgoMind: Activating Spatial Cognition through Linguistic Reasoning in MLLMs

Multimodal large language models (MLLMs) are increasingly being applied to spatial cognition tasks, where they are expected to understand and interact with complex environments. Most existing works improve spatial reasoning by introducing 3D priors or geometric supervision, which enhances performance but incurs substantial data preparation and alignment costs. In contrast, purely 2D approaches often struggle with multi-frame spatial reasoning due to their limited ability to capture cross-frame spatial relationships. To address these limitations, we propose EgoMind, a Chain-of-Thought framework that enables geometry-free spatial reasoning through Role-Play Caption, which jointly constructs a coherent linguistic scene graph across frames, and Progressive Spatial Analysis, which progressively reasons toward task-specific questions. With only 5K auto-generated SFT samples and 20K RL samples, EgoMind achieves competitive results on VSI-Bench, SPAR-Bench, SITE-Bench, and SPBench, demonstrating its effectiveness in strengthening the spatial reasoning capabilities of MLLMs and highlighting the potential of linguistic reasoning for spatial cognition. Code and data are released at https://github.com/Hyggge/EgoMind.

cs.CV

CoVFT: Context-aware Visual Fine-tuning for Multimodal Large Language Models

Multimodal large language models (MLLMs) achieve remarkable progress in cross-modal perception and reasoning, yet a fundamental question remains unresolved: should the vision encoder be fine-tuned or frozen? Despite the success of models such as LLaVA and Qwen-VL, inconsistent design choices and heterogeneous training setups hinder a unified understanding of visual fine-tuning (VFT) in MLLMs. Through a configuration-aligned benchmark, we find that existing VFT methods fail to consistently outperform the frozen baseline across multimodal tasks. Our analysis suggests that this instability arises from visual preference conflicts, where the context-agnostic nature of vision encoders induces divergent parameter updates under diverse multimodal context. To address this issue, we propose the Context-aware Visual Fine-tuning (CoVFT) framework, which explicitly incorporates multimodal context into visual adaptation. By integrating a Context Vector Extraction (CVE) and a Contextual Mixture-of-Experts (CoMoE) module, CoVFT decomposes conflicting optimization signals and enables stable, context-sensitive visual updates. Extensive experiments on 12 multimodal benchmarks demonstrate that CoVFT achieves state-of-the-art performance with superior stability. Notably, fine-tuning a 7B MLLM with CoVFT surpasses the average performance of its 13B counterpart, revealing substantial untapped potential in visual encoder optimization within MLLMs.

cs.CV

Implicit Modeling for Transferability Estimation of Vision Foundation Models

Transferability estimation identifies the best pre-trained models for downstream tasks without incurring the high computational cost of full fine-tuning. This capability facilitates deployment and advances the pre-training and fine-tuning paradigm. However, existing methods often struggle to accurately assess transferability for emerging pre-trained models with diverse architectures, training strategies, and task alignments. In this work, we propose Implicit Transferability Modeling (ITM), a novel framework that implicitly models each model's intrinsic transferability, coupled with a Divide-and-Conquer Variational Approximation (DVA) strategy to efficiently approximate embedding space evolution. This design enables generalization across a broader range of models and downstream tasks. Extensive experiments on a comprehensive benchmark--spanning extensive training regimes and a wider variety of model types--demonstrate that ITM consistently outperforms existing methods in terms of stability, effectiveness, and efficiency.

cs.CV

High-resolution geostationary satellite observations of free tropospheric NO2 over North America: implications for lightning emissions

Free tropospheric (FT) nitrogen dioxide (NO2) plays a critical role in atmospheric oxidant chemistry as a source of tropospheric ozone and of the hydroxyl radical (OH). It also contributes significantly to satellite-observed tropospheric NO2 columns, and must be subtracted when using these columns to quantify surface emissions of nitrogen oxide radicals (NOx = NO + NO2). But large uncertainties remain in the sources and chemistry of FT NO2 because observations are sparse. Here, we construct a new cloud-sliced FT NO2 (700-300 hPa) product from the TEMPO geostationary satellite instrument over North America. This product provides higher data density and quality than previous products from low Earth orbit (LEO) instruments, with the first observation of the FT NO2 diurnal cycle across seasons. Combined with coincident observations from the Geostationary Lightning Mapper (GLM), the TEMPO data demonstrate the dominance of lightning as a source of FT NO2 in non-winter seasons. Comparison of TEMPO FT NO2 data with the GEOS-CF atmospheric chemistry model shows overall consistent magnitudes, seasonality, and diurnal variation, with a midday minimum in non-winter seasons from photochemical loss. However, there are major discrepancies that we attribute to GEOS-CF's use of a standard cloud-top-height (CTH)-based scheme for the lightning NOx source. We find this scheme greatly underestimates offshore lighting flash density and misrepresents the diurnal cycle of lightning over land. Our FT NO2 product provides a unique resource for improving the lightning NOx parameterization in atmospheric models and the ability to use NO2 observations from space to quantify surface NOx emissions.

physics.ao-ph

Multi-modal Relation Distillation for Unified 3D Representation Learning

Recent advancements in multi-modal pre-training for 3D point clouds have demonstrated promising results by aligning heterogeneous features across 3D shapes and their corresponding 2D images and language descriptions. However, current straightforward solutions often overlook intricate structural relations among samples, potentially limiting the full capabilities of multi-modal learning. To address this issue, we introduce Multi-modal Relation Distillation (MRD), a tri-modal pre-training framework, which is designed to effectively distill reputable large Vision-Language Models (VLM) into 3D backbones. MRD aims to capture both intra-relations within each modality as well as cross-relations between different modalities and produce more discriminative 3D shape representations. Notably, MRD achieves significant improvements in downstream zero-shot classification tasks and cross-modality retrieval tasks, delivering new state-of-the-art performance.

cs.CV

Ozone Anomalies in Dry Intrusions Associated with Atmospheric Rivers

As a result of their important role in weather and the global hydrological cycle, understanding atmospheric rivers' (ARs) connection to synoptic-scale climate patterns and atmospheric dynamics has become increasingly important. In addition to case studies of two extreme AR events, we produce a December climatology of the three-dimensional structure of water vapor and O3 (ozone) distributions associated with ARs in the northeastern Pacific from 2004-2014 using MERRA-2 reanalysis products. Results show that positive O3 anomalies reside in dry intrusions of stratospheric air due to stratosphere-to-troposphere transport (STT) behind the intense water vapor transport of the AR. In composites, we find increased excesses of O3 concentration, as well as in the total O3 flux within the dry intrusions, with increased AR strength. We find that STT O3 flux associated with ARs over the NE Pacific accounts for up to 13 percent of total Northern Hemisphere STT O3 flux in December, and extrapolation indicates that AR-associated dry intrusions may account for as much as 32 percent of total NH STT O3 flux. This study quantifies STT of O3 in connection with ARs for the first time and improves estimates of tropospheric ozone concentration due to STT in the identification of this correlation. In light of predictions that ARs will become more intense and/or frequent with climate change, quantifying AR-related STT O3 flux is especially valuable for future radiative forcing calculations.

physics.ao-ph

iDARTS: Improving DARTS by Node Normalization and Decorrelation Discretization

Differentiable ARchiTecture Search (DARTS) uses a continuous relaxation of network representation and dramatically accelerates Neural Architecture Search (NAS) by almost thousands of times in GPU-day. However, the searching process of DARTS is unstable, which suffers severe degradation when training epochs become large, thus limiting its application. In this paper, we claim that this degradation issue is caused by the imbalanced norms between different nodes and the highly correlated outputs from various operations. We then propose an improved version of DARTS, namely iDARTS, to deal with the two problems. In the training phase, it introduces node normalization to maintain the norm balance. In the discretization phase, the continuous architecture is approximated based on the similarity between the outputs of the node and the decorrelated operations rather than the values of the architecture parameters. Extensive evaluation is conducted on CIFAR-10 and ImageNet, and the error rates of 2.25\% and 24.7\% are reported within 0.2 and 1.9 GPU-day for architecture search respectively, which shows its effectiveness. Additional analysis also reveals that iDARTS has the advantage in robustness and generalization over other DARTS-based counterparts.

cs.CV

PR-GCN: A Deep Graph Convolutional Network with Point Refinement for 6D Pose Estimation

RGB-D based 6D pose estimation has recently achieved remarkable progress, but still suffers from two major limitations: (1) ineffective representation of depth data and (2) insufficient integration of different modalities. This paper proposes a novel deep learning approach, namely Graph Convolutional Network with Point Refinement (PR-GCN), to simultaneously address the issues above in a unified way. It first introduces the Point Refinement Network (PRN) to polish 3D point clouds, recovering missing parts with noise removed. Subsequently, the Multi-Modal Fusion Graph Convolutional Network (MMF-GCN) is presented to strengthen RGB-D combination, which captures geometry-aware inter-modality correlation through local information propagation in the graph convolutional network. Extensive experiments are conducted on three widely used benchmarks, and state-of-the-art performance is reached. Besides, it is also shown that the proposed PRN and MMF-GCN modules are well generalized to other frameworks.

cs.CV

Eddy evolution during large dust storms

The evolution of eddy kinetic energy during the development of large regional dust storms on Mars is investigated using the Mars Analysis Correction Data Assimilation (MACDA) reanalysis product and the dust storm data derived from Mars Global Surveyor Mars Daily Global Maps. Transient eddies in MACDA are decomposed into different components according to their eddy periods: $P\leq1$ sol, $1<P\leq8$ sols, $8<P\leq60$ sols. This paper primarily focuses on the Mars year 24 pre-solstice "A" storm that starts with many episodes of frontal/flushing dust storms from the northern hemisphere and attains its maximum global mean opacity after dust expansion in the southern hemisphere. During the development of this storm, the dominant eddies in terms of eddy kinetic energy progress from the $1<P\leq8$ sol eddies in the northern mid/high latitudes to the $P\leq1$ sol eddies (dominated by thermal tides) in the southern mid latitudes, and the $8<P\leq60$ sol eddies show a prominent peak with the increased global-mean dust opacity. The peaks of the $1<P\leq8$ sol eddies are found to best correlate with the average area of textured frontal/flushing dust storms within 40$^\circ$N $-$ 60$^\circ$N. The region where the $1<P\leq8$ sol eddies increase the most corresponds to the main flushing channel. The eddy kinetic energy of the $P\leq1$ eddies, dominated by $P$ = 1 and its harmonics, increases with the global mean dust opacity both before and after the winter solstice in Mars year 24. The $8<P\leq60$ sol eddies briefly spike during large, regional dust storms but remain weak if dust storm sequences do not lead to a major dust storm. Zonal wavenumber analysis of eddy kinetic energy shows that the peaks of the $1<P\leq8$ eddies often result from combinations of zonal wavenumbers 1 to 3, while the $P\leq1$ eddies and $8<P\leq60$ sol eddies are each dominated by zonal wavenumber 1.

astro-ph.EP