SearcharxivSearch

arXiv subjects

Yigong Zhang

Publications and source records attributed to Yigong Zhang.

11 recordsLinked to original sources

Fine-Grained Visual Preprocessing and Dual-Stream Temporal Modeling for Multimodal Sentiment Analysis on Social Media

Multimodal sentiment analysis often remains text-dominant due to raw-video noise and insufficient temporal modeling. Using CH-SIMS v2.0S, this study proposes three improvements: the NAPS pipeline---a seven-stage system integrating face tracking,identity embedding, and normalized lip-motion analysis to reduce visual noise;DS-TANet, combining an EfficientNetB2 static stream, RAFT optical-flow motion stream, motion-guided attention, and Bi-GRU temporal modeling; and DS-TAFNet, fusing visual and MacBERT-Base textual representations via concatenation fusion. With NAPS, the static visual baseline achieves 80.98\% Macro F1, comparable to the text baseline of 80.55\%; DS-TANet improves visual Macro F1 to 82.58\%;and DS-TAFNet achieves 87.49\% accuracy and 87.48\% Macro F1. These results demonstrate that improving visual input quality and temporal representation is more effective than increasing fusion complexity under limited-data conditions.

cs.CV

An adaptive parameter optimization method for astronomical image alignment using Bayesian optimization. I. A hierarchical search strategy for FWHM and SNR

The alignment and stacking of astronomical images are fundamental steps for detecting faint objects and performing high?precision astrometry. In traditional alignment workflows, the extraction of source lists is critically dependent on key parameters such as the Full Width at Half Maximum (FWHM) and the Signal-to-Noise Ratio (SNR) threshold. These parameters are often selected manually through an inefficient trial-error process that lacks objectivity and does not guarantee optimal results. We present an adaptive method for optimizing astronomical image alignment parameters based on Bayesian Optimization (BO). We frame the parameter search as an optimization problem, with an objective function designed to maximize the number of successfully matched source pairs. By employing a hierarchical search strategy, we perform an efficient global search for FWHM and SNR to automatically determine the optimal combination for a given observational dataset. Experimental results demonstrate that our method effectively handles image data with varying seeing conditions and back?ground noise levels. It rapidly converges to a robust set of alignment parameters, achieving sub-pixel accuracy and significantly improving the automation level and success rate of the alignment process. This work may provide a useful basis for developing large-scale, automated astronomical data processing pipelines

astro-ph.IM

MV3DIS: Multi-View Mask Matching via 3D Guides for Zero-Shot 3D Instance Segmentation

Conventional 3D instance segmentation methods rely on labor-intensive 3D annotations for supervised training, which limits their scalability and generalization to novel objects. Recent approaches leverage multi-view 2D masks from the Segment Anything Model (SAM) to guide the merging of 3D geometric primitives, thereby enabling zero-shot 3D instance segmentation. However, these methods typically process each frame independently and rely solely on 2D metrics, such as SAM prediction scores, to produce segmentation maps. This design overlooks multi-view correlations and inherent 3D priors, leading to inconsistent 2D masks across views and ultimately fragmented 3D segmentation. In this paper, we propose MV3DIS, a coarse-to-fine framework for zero-shot 3D instance segmentation that explicitly incorporates 3D priors. Specifically, we introduce a 3D-guided mask matching strategy that uses coarse 3D segments as a common reference to match 2D masks across views and consolidates multi-view mask consistency via 3D coverage distributions. Guided by these view-consistent 2D masks, the coarse 3D segments are further refined into precise 3D instances. Additionally, we introduce a depth consistency weighting scheme that quantifies projection reliability to suppress ambiguities from inter-object occlusions, thereby improving the robustness of 3D-to-2D correspondence. Extensive experiments on the ScanNetV2, ScanNet200, ScanNet++, Replica, and Matterport3D datasets demonstrate the effectiveness of MV3DIS, which achieves superior performance over previous methods

cs.CV

CosmicWeb-21cm array: A New Radio Observation Array Design for 21cm Cosmology

This paper presents the CosmicWeb-21cm array, a novel radio interferometer designed to overcome the key challenges in 21 cm cosmology. Its core innovations include: (1) a multi-scale nested geometry combining a hexagonal core with logarithmic spiral arms for excellent UV coverage and calibration robustness; (2) an intelligent non-uniform frequency sampling strategy that adapts resolution to foreground and signal characteristics, reducing data volume while preserving information; and (3) a machine-learning-enhanced, physics-informed processing pipeline that achieves 99.7\% foreground removal efficiency; (4) a dual-polarization crossed dipole integrated with a dielectric lens and cryogenically cooled LNA, achieving stable beam patterns and low noise temperature ($<35$ K) across 50-250 MHz. These co-designed advances enable high sensitivity mapping of the Epoch of Reionization, dark energy constraints and cosmic-web structure.

astro-ph.IM

MonoSE(3)-Diffusion: A Monocular SE(3) Diffusion Framework for Robust Camera-to-Robot Pose Estimation

We propose MonoSE(3)-Diffusion, a monocular SE(3) diffusion framework that formulates markerless, image-based robot pose estimation as a conditional denoising diffusion process. The framework consists of two processes: a visibility-constrained diffusion process for diverse pose augmentation and a timestep-aware reverse process for progressive pose refinement. The diffusion process progressively perturbs ground-truth poses to noisy transformations for training a pose denoising network. Importantly, we integrate visibility constraints into the process, ensuring the transformations remain within the camera field of view. Compared to the fixed-scale perturbations used in current methods, the diffusion process generates in-view and diverse training poses, thereby improving the network generalization capability. Furthermore, the reverse process iteratively predicts the poses by the denoising network and refines pose estimates by sampling from the diffusion posterior of current timestep, following a scheduled coarse-to-fine procedure. Moreover, the timestep indicates the transformation scales, which guide the denoising network to achieve more accurate pose predictions. The reverse process demonstrates higher robustness than direct prediction, benefiting from its timestep-aware refinement scheme. Our approach demonstrates improvements across two benchmarks (DREAM and RoboKeyGen), achieving a notable AUC of 66.75 on the most challenging dataset, representing a 32.3% gain over the state-of-the-art.

cs.CV

MaskHOI: Robust 3D Hand-Object Interaction Estimation via Masked Pre-training

In 3D hand-object interaction (HOI) tasks, estimating precise joint poses of hands and objects from monocular RGB input remains highly challenging due to the inherent geometric ambiguity of RGB images and the severe mutual occlusions that occur during interaction.To address these challenges, we propose MaskHOI, a novel Masked Autoencoder (MAE)-driven pretraining framework for enhanced HOI pose estimation. Our core idea is to leverage the masking-then-reconstruction strategy of MAE to encourage the feature encoder to infer missing spatial and structural information, thereby facilitating geometric-aware and occlusion-robust representation learning. Specifically, based on our observation that human hands exhibit far greater geometric complexity than rigid objects, conventional uniform masking fails to effectively guide the reconstruction of fine-grained hand structures. To overcome this limitation, we introduce a Region-specific Mask Ratio Allocation, primarily comprising the region-specific masking assignment and the skeleton-driven hand masking guidance. The former adaptively assigns lower masking ratios to hand regions than to rigid objects, balancing their feature learning difficulty, while the latter prioritizes masking critical hand parts (e.g., fingertips or entire fingers) to realistically simulate occlusion patterns in real-world interactions. Furthermore, to enhance the geometric awareness of the pretrained encoder, we introduce a novel Masked Signed Distance Field (SDF)-driven multimodal learning mechanism. Through the self-masking 3D SDF prediction, the learned encoder is able to perceive the global geometric structure of hands and objects beyond the 2D image plane, overcoming the inherent limitations of monocular input and alleviating self-occlusion issues. Extensive experiments demonstrate that our method significantly outperforms existing state-of-the-art approaches.

cs.CV

See through the Dark: Learning Illumination-affined Representations for Nighttime Occupancy Prediction

Occupancy prediction aims to estimate the 3D spatial distribution of occupied regions along with their corresponding semantic labels. Existing vision-based methods perform well on daytime benchmarks but struggle in nighttime scenarios due to limited visibility and challenging lighting conditions. To address these challenges, we propose LIAR, a novel framework that learns illumination-affined representations. LIAR first introduces Selective Low-light Image Enhancement (SLLIE), which leverages the illumination priors from daytime scenes to adaptively determine whether a nighttime image is genuinely dark or sufficiently well-lit, enabling more targeted global enhancement. Building on the illumination maps generated by SLLIE, LIAR further incorporates two illumination-aware components: 2D Illumination-guided Sampling (2D-IGS) and 3D Illumination-driven Projection (3D-IDP), to respectively tackle local underexposure and overexposure. Specifically, 2D-IGS modulates feature sampling positions according to illumination maps, assigning larger offsets to darker regions and smaller ones to brighter regions, thereby alleviating feature degradation in underexposed areas. Subsequently,3D-IDP enhances semantic understanding in overexposed regions by constructing illumination intensity fields and supplying refined residual queries to the BEV context refinement process. Extensive experiments on both real and synthetic datasets demonstrate the superior performance of LIAR under challenging nighttime scenarios. The source code and pretrained models are available [here](https://github.com/yanzq95/LIAR).

cs.CV

GigaSLAM: Large-Scale Monocular SLAM with Hierarchical Gaussian Splats

Tracking and mapping in large-scale, unbounded outdoor environments using only monocular RGB input presents substantial challenges for existing SLAM systems. Traditional Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS) SLAM methods are typically limited to small, bounded indoor settings. To overcome these challenges, we introduce GigaSLAM, the first RGB NeRF / 3DGS-based SLAM framework for kilometer-scale outdoor environments, as demonstrated on the KITTI, KITTI 360, 4 Seasons and A2D2 datasets. Our approach employs a hierarchical sparse voxel map representation, where Gaussians are decoded by neural networks at multiple levels of detail. This design enables efficient, scalable mapping and high-fidelity viewpoint rendering across expansive, unbounded scenes. For front-end tracking, GigaSLAM utilizes a metric depth model combined with epipolar geometry and PnP algorithms to accurately estimate poses, while incorporating a Bag-of-Words-based loop closure mechanism to maintain robust alignment over long trajectories. Consequently, GigaSLAM delivers high-precision tracking and visually faithful rendering on urban outdoor benchmarks, establishing a robust SLAM solution for large-scale, long-term scenarios, and significantly extending the applicability of Gaussian Splatting SLAM systems to unbounded outdoor environments. GitHub: https://github.com/DengKaiCQ/GigaSLAM.

cs.RO

Predicting Astrometric Microlensing Events from Gaia DR3

Currently astrometric microlensing is the only tool that can directly measure the mass of a single star, it can also help us to detect compact objects like isolated neutron stars and black holes. The number of microlensing events that are being predicted and reported is increasing. In the paper, the potential lens stars are selected from three types of stars, high-proper-motion stars, nearby stars and high-mass stars. For each potential lens star, we select a larger search scope to find possible matching sources to avoid missing events as much as possible. Using Gaia DR3 data, we predict 4500 astrometric microlensing events with signal>0.1mas that occur between J2010.0 and J2070.0, where 1664 events are different from those found previously. There are 293 lens stars that can cause two or more events, where 5 lens stars can cause more than 50 events. We find that 116 events have the distance of background stars from the proper motion path of lens stars more than 8 arcsec in the reference epoch, where the maximum distance is 16.6 arcsec, so the cone search method of expanding the search range of sources for each potential lens star can reduce the possibility of missing events.

astro-ph.SR

Astrometric observations of a near-Earth object using the image fusion technique

The precise astrometric observation of small near-Earth objects (NEOs) is an important observational research topic in the astrometric discipline, which greatly promotes multidisciplinary research, such as the origin and evolution of the solar system, the detection and early warning of small NEOs, and deep-space navigation. The characteristics of small NEOs, such as faintness and fast moving speed, restrict the accuracy and precision of their astrometric observations. In the paper, we present a method to improve the accurate and precise astrometric positions of NEOs based on image fusion technique. The noise analysis and astrometric test from the observed images of the open cluster M23 are given. Using the image fusion technique, we obtain the sets of superimposed images and original images containing reference stars and moving targets respectively. The final fused image set includes background stars with high signal-to-noise ratios and ideal NEO images simultaneously and avoids the saturation of background stars. Using the fused images, we can reduce the influence of telescope tracking and NEO ephemeris errors on astrometric observations, and our results indicate that the accuracy and precision of NEO Eros astrometry are improved obviously after we choose suitable image fuse mode.

astro-ph.IM

CCD astrometric observations of 2017 VR12,Camillo and Midas

We have observed three near-Earth objects(NEOs), 2017VR12, Camillo, and Midas during the year 2018. The observations were made by the 1-m telescope of Yunnan Observatory over 2 nights. Their precise astrometric positions are derived from 989 CCD observations. The theoretical positions of asteroids are retrieved from the Jet Propulsion Laboratory (JPL) Horizons System and Institut de M\'{e}canique C\'{e}leste et de Calcul des \'{E}ph\'{e}m\'{e}rides (IMCCE). The positions of three asteroids are measured with respect to the stars in Gaia DR2 star catalogue. For 2017 VR12, the mean (O-C) of right ascension and declination are -0.090$^{''}$ and -0.623$^{''}$ based on the ephemeris of published JPL, but the mean (O-C) are 3.122$^{''}$ and -0.636$^{''}$ based on the ephemeris of published IMCCE. The great difference in declination could be explained by several factors. (1)The degenerated CCD images caused by the high apparent motion speed of the object leads to the reduction of positioning accuracy. (2)The poor timing system may bring the system error, especially in the high speed direction. (3)The asteroid may be perturbed by the earth when it approaches the earth too closely. The astrometric results show that the centroid centring method can reduce the dispersion of the non-Gaussian images compared with the PSF model method. For Camillo and Midas, the astrometric results are consistent based on two ephemerides. High-precision timing system, some astronomical effects and geometric distortion of CCD images should be carefully considered in the future works.

astro-ph.EP