SearcharxivSearch

arXiv subjects

Yan Gong

Publications and source records attributed to Yan Gong.

At least 19 recordsLinked to original sources

CM2: Multimodal Cultural Reasoning via an Integrated Multi-Agent Framework

Multimodal Large Language Models (MLLMs) have shown remarkable success in STEM domains, where progress is often driven by vertical, step-by-step deduction under relatively stable symbol systems. Their horizontal, interdisciplinary cultural reasoning, however, remains underexplored.We propose CM2, a multi-agent framework grounded in the cognitive pathway of human cultural interpretation. CM2 integrates multimodal perception, retrieval-augmented generation, networked reasoning, gated fusion, and reward-driven feedback.Experiments on CM2D across multiple MLLM backbones show consistent gains over CoT and typical reasoning paradigms; ablations validate each module's contribution, and conflict analyses confirm genuine cross-modal arbitration.

cs.AI

Not All History Helps: Velocity-Aware Selective Memory for Long-Horizon End-to-End Autonomous Driving

Reliable long-horizon planning remains a key challenge in end-to-end autonomous driving. By accounting for future motion evolution and potential consequences, it provides forward-looking guidance for safe and consistent driving in evolving traffic environments. Existing methods use historical planning states as temporal context. Self-generated history may become stale or conflict with the current motion stage, introducing unreliable priors. We propose StableDrive to address cross-cycle historical reliability and within-horizon motion-stage evolution. Selective Momentum Memory (SMM), implemented with a Mamba selective state-space operator, controls the influence of the preceding self-predicted planning state on the current cycle. Motion-Stage Training Scaffold (MSTS) uses motion-stage, long-horizon trajectory, and longitudinal-motion supervision to guide stage-aware future motion learning and is removed before inference. A fixed parameter midpoint between two architecture-aligned endpoints yields a single deployable SMM planner without model ensembling or extra inference-time computation. On nuScenes under the MomAD evaluation protocol, StableDrive achieves SOTA performance across all reported planning metrics from 1 to 6 s, reducing average collision rate by 23.3%, TPC by 30.9%, and L2 by 11.8% over the best previously reported value for each metric. On the curated Longitudinal-Transition nuScenes (LT-nuScenes), StableDrive reduces 6-s collision rate by 23.81%, TPC by 10.90%, and L2 by 6.37%. On NAVSIM v1 and v2, StableDrive achieves the highest PDMS/EPDMS in all three reported settings, including a 5.7-point EPDMS gain on v2 navhard over the previous best.

cs.RO

RoofGS: Roofline-Guided End-to-End Acceleration of 3D Gaussian Splatting

3D Gaussian Splatting (3DGS) enables real-time novel-view synthesis but remains limited on GPUs at high resolutions. Through a stage-wise Roofline characterization, we identify two distinct hardware bottlenecks: global memory traffic dominates the front end, whereas instruction throughput limits rasterization. Guided by this analysis, we develop RoofGS, a rendering framework that applies bottleneck-specific optimizations rather than generic kernel acceleration. For the memory-bound front end, we design a resolution-adaptive quantized depth sorting key that compresses each key to 32 bits. For the compute-bound rasterizer, we introduce a range-aware bit-level fast exponential approximation tailored to the bounded exponent range after opacity culling, with a derived per-pixel error bound. These two core techniques are complemented by additional optimizations (kernel fusion, compact attribute storage, culling, dual-pixel evaluation) that additionally reduce memory traffic and improve instruction-level parallelism. Experiments show that RoofGS achieves a 10.1$\times$ end-to-end speedup over 3DGS at 4K on an RTX 4090, increasing throughput from 61 to 616 FPS, with only a 0.028 dB PSNR loss.

cs.CV

CO Structures with Narrow Lines in Nearby Quiescent Regions

Using CO data from Phase I of the Milky Way Imaging Scroll Painting (MWISP) survey, we present a systematic study of molecular structures with narrow lines. We identify 57 CO structures, most of which exhibit low densities and subsonic/transonic turbulence. Among them, structures with large projected areas and diffuse, sheet-like geometries are identified as veil clouds. The low LSR velocities and the concentration of these CO structures toward both the Galactic center (e.g., Ophiuchus, Aquila) and anticenter (e.g., Cepheus, Taurus) regions suggest a local origin for the sample, as supported by distance measurements of about 200--300pc for a subset with relatively large angular extents. These nearby structures likely arise from large-scale compression driven by past supernova activity within the Local Bubble. The observed low-velocity-dispersion emission may trace quiescent regions where turbulence has decayed due to a lack of sustained energy injection. For diffuse veil clouds with an assumed magnetic field of ~10uG, ion-neutral friction may provide an additional mechanism for turbulent dissipation on sub-parsec scales corresponding to their thickness of 0.1--0.3pc. Tracing the atomic-to-molecular transition, veil clouds provide a unique window into the diffuse, quiescent precursor state of dense gas. They likely represent a widespread but previously overlooked component of the Galactic molecular gas reservoir, with significant implications for cloud formation and evolution, the total mass budget and spatial distribution of molecular gas, and the initial conditions of star formation as a related consequence.

astro-ph.GA

Text-Aided Multi-Modal Panoptic Symbol Spotting for CAD Floor Plan Drawings

Computer-Aided Design (CAD) floor plan drawings contain both graphical primitives and textual annotations, which provide complementary geometric and semantic cues for intelligent design understanding. Among CAD analysis tasks, panoptic symbol spotting has become increasingly important with the growing demand for industrial digitalization and deep learning-based automation. However, most existing methods remain primarily primitive-centric and underexploit textual annotations, despite their critical semantic value. Even the few text-aware approaches often treat annotations only superficially, without properly modeling complex syntax and hierarchical semantics of CAD annotations, which leads to semantic loss and suboptimal spotting performance. To address these limitations, we propose TextCAD, a multimodal framework that jointly models graphical primitives and textual annotations for panoptic symbol spotting. Specifically, we design a Type-Attribute Correlation Encoder (TACE) to explicitly encode the compositional semantics within annotations by jointly modeling their types and attributes. We further introduce a Semantic Hierarchy Alignment framework with Multi-level Semantic Filtering (MSF) and primitive downsampling, which adaptively aligns annotation semantics with graphical primitives at different semantic levels and enables accurate cross-modal semantic injection and fusion. Experiments on real-world building-design datasets show that TextCAD effectively improves symbol spotting performance and achieves state-of-the-art results.

cs.CV

Quantifying Environmental Effects on Galaxy Properties using Non-spherical Voids Identified from SDSS DR7

Cosmic voids provide a distinct low-density region for studying the environmental effects of galaxy properties. Using the SDSS DR7 catalog, we identify non-spherical voids via Voronoi tessellation and the watershed algorithm, and classify void galaxies based on their local volume. We compare and find that void galaxies classified by this method are systematically less massive, fainter, bluer, and have higher specific star formation rate (sSFR) than non-void galaxies and all galaxy samples. We then divide void and non-void galaxies into stellar mass bins to focus on the environmental dependence of $g-r$ color and sSFR. By further classifying galaxies into blue/red and star-forming/quiescent populations, we calculate the ratio of blue to red and star-forming to quiescent for void and non-void galaxies separately. Comparing the ratio of the void value to the non-void value for both metrics presents an overall decreasing trend with stellar mass $M_*$ over the $9.4-10.4$ range in $\log[M_*/\mathrm{M}_\odot]$, indicating a stronger environmental effect in lower-mass systems. These results show that our classification of void galaxies in non-spherical voids based on local volume offers a robust approach for quantifying the influence of underdense environments on galaxy evolution.

astro-ph.GA

Probing Baryonic Feedback Effect with CSST Weak Lensing and Future FRB Measurements

We explore the joint probe on the baryonic feedback effect using the weak lensing measurement from the upcoming Chinese Space-station Survey Telescope (CSST) photometric survey and the dispersion measure (DM) statistics of the fast radio bursts (FRBs) from next-generation radio telescopes, i.e., the Square Kilometre Array (SKA) and the Deep Synoptic Array (DSA-2000). By employing the baryonic halo model, we compute the matter, electron, and matter--electron power spectra, and generate mock data considering realistic noise and systematic effects based on the designs of the telescopes. These mock data are then analyzed using the Markov Chain Monte Carlo (MCMC) method to investigate the parameter constraints. We find that CSST weak lensing alone can constrain the baryonic feedback parameter $\log_{10} T_{\text{AGN}}$ to an accuracy of $3.1\%$, with the sum of neutrino mass bound $\sum m_{\nu} < 0.53\,\mathrm{eV}$. When performing the $3\times2 \,\mathrm{pt}(e,\mathrm{m})$ analysis, the inclusion of FRB DM measurements can significantly improve the precision of $\log_{10} T_{\text{AGN}}$ to $0.4\%$, and will lead to a better constraint on $\sum m_{\nu}$ with an upper limit $< 0.47\,\mathrm{eV}$ by effectively breaking the degeneracy. Our results demonstrate that the joint observation of future FRB DM and weak lensing surveys is a powerful tool for probing the baryonic feedback effect, which is helpful in obtaining robust constraints on the neutrino mass and other important cosmological parameters.

astro-ph.CO

Exploring Primordial Non-Gaussianity Measurements in the CSST Spectroscopic Survey

Primordial non-Gaussianity (PNG) is a fundamental probe of the physics of the early Universe and inflation. Here we present a comprehensive study of the constraints on the local-type PNG parameter, $f_{\rm NL}$, for the spectroscopic galaxy survey of the upcoming Chinese Space-station Survey Telescope (CSST). Utilizing the high-resolution Jiutian N-body simulation suite, we construct realistic mock catalogs for emission line galaxies (ELGs) at three representative redshifts $z=0.3$, 0.6, and 0.9. The expected CSST observational characteristics are also considered, including redshift uncertainties and selection functions based on signal-to-noise ratios of emission lines. We develop a robust analysis framework for the redshift-space galaxy power spectrum and bispectrum that accounts for redshift-space distortions, scale-dependent bias, and nonlinear effects. Through a joint Markov Chain Monte Carlo (MCMC) analysis, we find that the power spectrum alone provides competitive constraints, while the inclusion of the bispectrum, specifically targeting the squeezed-limit configurations, improves the $f_{\rm NL}$ constraint precision by approximately 5%-6%. Our joint analysis yields a constraint result of $f_{\rm NL}=-20\pm52$ for the mock data in the 1~($h^{-1}$Gpc)$^3$ comoving volume at the three redshifts, and the constraint accuracy is expected to be improved by several times or even one order of magnitude for the CSST full spectroscopic survey. This work demonstrates the potential of the Stage~IV surveys like CSST to probe inflationary physics, and highlights the importance of higher-order statistics in extracting information from large-scale structure surveys.

astro-ph.CO

Local-GS: Accelerating 3D Gaussian Splatting via Tile-Local Warp Coherence

3D Gaussian Splatting (3DGS) has significantly advanced real-time novel view synthesis by representing scenes as dense collections of anisotropic 3D Gaussian primitives. However, the irregular spatial distribution of Gaussians often leads to poor GPU utilization, as warp divergence and redundant computation degrade rendering performance. To address this, we present Local-GS, a warp-coherent rendering paradigm that, organizes Gaussian primitives with respect to SIMT (Single Instruction, Multiple Threads) execution boundaries rather than scene geometry. Specifically, we propose three warp-coherent stages: a hoisting stage that precomputes shared parameters at tile level, a culling stage that discards warps with no contribution, and a blending stage that replaces per-pixel branching with a uniform instruction stream. Across extensive benchmarks on multiple datasets, Local-GS improves efficiency without compromising quality. As a plug-and-play optimization, it provides additional performance gains to all tested baselines, culminating in a $7.76\times$ speedup on Deep Blending scenes.

cs.CV

Star Formation Drives Production of Low Energy Cosmic Rays

For over a century, the origin of low-energy cosmic rays (LECRs), the dominant heaters and ionizers of dense interstellar gas, remains elusive owing to solar modulation and uncertain transport processes. In this study, we introduce a new astrophysical approach based on HI Narrow Self-Absorption (HINSA) to obtain spatially resolved measurements of LECR ionization rates using high-fidelity HI observations toward the Orion region from the FAST telescope. The LECR ionization rate is found to scale with local star formation rate (SFR) as $log_{10}\zeta = (1.4\pm 0.70)log_{10}\mathrm{SFR} + (-10.5\pm 2.9)$. Moreover, it increases with visual extinction, and is found to exceed, toward active star-forming regions, the value predicted for diffuse regions based on \textit{Voyager} measurements and an external propagation model. These findings demonstrate that LECRs are generated in situ by star-forming activities rather than penetrating from the broader Galactic cosmic-ray population. This is further supported by \textit{Fermi}-LAT gamma-ray observations toward the Orion region. Together, these results resolve a key uncertainty in cosmic-ray origin and establish a new avenue for quantifying the energetic feedback that regulates the interstellar medium.

astro-ph.GA

Lyman-$\alpha$ forest constraints on pure and mixed fuzzy dark matter

Fuzzy dark matter (FDM), often realized as an ultralight scalar field, can suppress the growth of small-scale structures and could be strictly tested with Lyman-$\alpha$ forest measurements. In this work, we constrain both pure and mixed FDM models (PFDM and MFDM) using measurements of the one-dimensional (1D) Lyman-$\alpha$ forest flux power spectrum at $z=5.0$, 4.6, and 4.2. We perform cosmological hydrodynamical simulations with modified initial conditions and construct a two-stage neural network emulator for accurate analysis. The first stage predicts the cold dark matter (CDM) 1D flux power spectrum, while the second stage predicts the MFDM effect relative to the CDM baseline. This construction improves the sensitivity to weak FDM effects, enforces the correct CDM limit, and enables robust interpolation across a broad range of FDM masses and fractions. After marginalizing over the intergalactic medium parameters, we obtain the FDM mass $m_{\mathrm{FDM}}>1.9\times10^{-21}~\mathrm{eV}$ at 95\% credible level for the PFDM model. For the MFDM model, we find the FDM fraction of dark matter $f_{\mathrm{FDM}}<0.07$, $0.12$, and $0.65$ at 95\% credible level for $\log_{10}(m_{\mathrm{FDM}}/\mathrm{eV})=-23.0$, $-22.0$, and $-21.0$, respectively. When $\log_{10}(m_{\mathrm{FDM}}/\mathrm{eV})\gtrsim -20$, the current data do not provide an effective upper limit on $f_{\mathrm{FDM}}$.

astro-ph.CO

CSST large-scale structure analysis pipeline: IV. Cosmic Voids Identified from Galaxy Group Samples as Probes of the Large-scale Structure

Because groups are directly associated with halos, they allow for considerably simpler theoretical modeling than approaches based on individual galaxies. We therefore propose to use voids identified in galaxy group catalogs, referred to as group-voids, to investigate the cosmic large-scale structure (LSS). Using the reference mock galaxy redshift survey (MGRS) designed for the Chinese Space-station Survey Telescope (CSST), we build two galaxy group catalogs representing ideal and realistic scenarios, derived from galaxy samples with 100\% and roughly 30\% spectroscopic redshift completeness, respectively. We then identify voids in these two mock group catalogs, as well as in the underlying halo catalog, and measure two void statistics, the void size function (VSF) and the void density profile, within five redshift intervals spanning $z=0$ to $1.0$. We compare the statistics obtained from two kinds of voids: those defined by galaxy groups (group-voids) and those defined by dark matter halos (halo-voids). In the void-finding process, we adopt the brightest central galaxy (BCG) as the group center to improve the accuracy of the inferred void centers. Our analysis shows that void statistics derived from group-voids with spectroscopic redshift completeness of at least 40\% can faithfully reproduce the corresponding statistics from halo-voids. Even when the redshift completeness of galaxies falls to as low as 30\%, we can still reliably describe group-voids via halo-voids by incorporating a redshift error term. This indicates that group-voids are a promising tool for probing LSS and offer a valuable complement to standard void studies, which is especially advantageous for emulator-based methods.

astro-ph.CO

Extracting redshifts from 2D slitless spectroscopic images using deep learning for the CSST galaxy survey

Wide-field slitless spectroscopic galaxy surveys, such as the one performed by the upcoming Chinese Space Station Survey Telescope (CSST), are crucial for precision cosmology but present formidable data analysis challenges. Because spectra are dispersed directly onto the detector, they are convolved with the 2-dimensional (2D) spatial morphology, which complicates wavelength calibration and consequently degrades the fidelity of subsequent 1-dimensional (1D) spectral extraction. To overcome these limitations, we present a deep learning framework that extracts redshifts directly from 2D slitless spectral images, bypassing 1D extraction entirely. We construct a realistic mock dataset for the CSST $GV$ and $GI$ band using high-resolution images from HSC-SSP PDR3 and spectral energy distributions (SEDs) from DESI DR1. A Bayesian convolutional neural network implemented by Monte Carlo dropout is employed to map the 2D spectral images to redshift estimations while simultaneously quantifying uncertainties. We find that our model can achieve a precision $\sigma_{\rm NMAD}=0.0104$ and mean uncertainty $\langle E / (1 + z_{{\rm true}}) \rangle=0.0155$ for sources with ${\rm SNR}_{GI}\geq1$. For sources with ${\rm SNR}_{GI}$ higher than 3.0, 5.0 and 10.0, $\sigma_{\rm NMAD}$ can achieve 0.0047, 0.0037 and 0.0024 respectively, matching the redshift precision requirements for studies such as BAO using the CSST slitless spectroscopic surveys. Furthermore, by utilizing spatial augmentations, the network demonstrates resilience to wavelength calibration errors. This work provides a novel and robust pathway for data analysis of next-generation slitless spectroscopic galaxy surveys.

astro-ph.IM

Semantic Prior Guided One-View 6D Pose Estimation for Novel Objects

In many practical 6D object pose estimation scenarios, we often have access to only a single real-world RGB-D reference view per object, typically without CAD models. Existing methods largely rely on explicit 3D models or multi-view data, which limits their scalability. To address this challenging single-reference model-free setting, we propose \textbf{OneViewAll}, a semantic-prior-guided framework that performs pose estimation via a novel Project-and-Compare paradigm. Instead of relying on computationally expensive CAD-based rendering, our method directly aligns reference and query observations within a projection-equivariant space. OneViewAll progressively integrates hierarchical semantic priors across three levels: (1) \textit{category- and scene-level} priors for efficient hypothesis initialization; (2) \textit{object-level symmetry} priors for geometry completion via mirror fusion; and (3) \textit{patch-level} priors for discriminative refinement. Extensive experiments demonstrate that OneViewAll achieves \textbf{92.5\%} ADD-0.1 accuracy on the LINEMOD dataset using only one real reference view -- significantly outperforming the CVPR 2025 baseline One2Any (52.6\%). It also yields consistent improvements on YCB-V, Real275, and Toyota-Light while maintaining low inference latency. Our results underscore the efficacy of symmetry-aware projection in handling symmetric, texture-less, and occluded objects.

cs.CV

Cross-Comparison of Galaxies Detected in the CSST Spectroscopic Survey and the SKA HI Survey

We present a forward-modeling framework to forecast the galaxies detected in the Chinese Space Station Survey Telescope (CSST) spectroscopic survey and the Square Kilometre Array (SKA) HI survey. Starting from the L-Galaxies 2020 semi-analytic model run on the Millennium-II N-body simulation (MS-II), the cold gas in galaxies is partitioned into atomic and molecular components self-consistently within the model. We further model the emission-lines (H $\alpha$, H $\beta$, O III) relevant for the slitless spectrograph of the CSST in a post-processing step. We construct mock lightcones using the Mock Map Facility (MoMaF) approach, simulating the neutral hydrogen (HI) data cubes representing a 2000 hour SKA-Mid spectral line observation from redshifts 0.25--0.5, and employ the Source Finding Application 2(SOFIA-2) source-finding package to generate an HI galaxy catalog. In parallel, we apply the CSST selection function and noise model to obtain a realistic catalog of emission-line galaxies; the emission-line signal is proportional to the star formation rate. These products allow us to cross compare the galaxy samples and assess the synergy between CSST and SKA. We study the correlations of the HI and the emission-line signal with the halo mass, HI mass, and the stellar mass, and the baryonic Tully-Fisher relation (BTFR). We also perform stacking analysis of the HI signal from the CSST-selected sample, which probes the HI content in galaxies with low HI mass. Finally, we derive the optical-HI cross-correlation power spectrum of the galaxies, and measure the bias of these galaxies. These results can provide useful insight on the cold gas and stellar content of the galaxies.

astro-ph.GA

MAPRPose: Mask-Aware Proposal and Amodal Refinement for Multi-Object 6D Pose Estimation

6D object pose estimation in cluttered scenes remains challenging due to severe occlusion and sensor noise. We propose MAPRPose, a two-stage framework that leverages mask-aware correspondences for pose proposal and amodal-driven Region-of-Interest (ROI) prediction for robust refinement. In the Mask-Aware Pose Proposal (MAPP) stage, we lift 2D correspondences into 3D space to establish reliable keypoint matches and generate geometrically consistent pose hypotheses based on correspondence-level scoring, from which the top-$K$ candidates are selected. In the refinement stage, we introduce a tensorized render-and-compare pipeline integrated with an Amodal Mask Prediction and ROI Re-Alignment (AMPR) module. By reconstructing complete object geometry and dynamically adjusting the ROI, AMPR mitigates localization errors and spatial misalignment under heavy occlusion. Furthermore, our GPU-accelerated RGB-XYZ reprojection enables simultaneous refinement of all $N \times B$ pose hypotheses in a single forward pass.

cs.CV

What Heats the Dense Gas in the Galactic Center?

Previous studies using p-H$_2$CO $J=3$--$2$ transitions at 218 GHz suggested widespread high-temperature gas exceeding 60 K and even 100 K in the CMZ, with heating mechanisms possibly related to cosmic rays or turbulent dissipation. However, at temperatures above 100 K, p-H$_2$CO $J=3$--$2$ line emission may lead to significant overestimates of kinetic temperature. This study combines o-H$_2$CO $J=5$--$4$ data from JCMT with p-H$_2$CO $J=3$--$2$ data from APEX to analyze three molecular clouds (The Brick, Sgr A1, and Sgr A2) with high temperatures. We used the non-LTE radiative transfer code RADEX to model spectral lines and constrain physical parameters with multiple line ratios, obtaining more reliable kinetic temperatures. Our results show that the previously reported extreme temperatures ($>100$ K) based on p-H$_2$CO $J=3$--$2$ line ratios are revised downward, with the average kinetic temperatures now constrained to 84--95 K using o-H$_2$CO $J=5$--$4$ line ratios, indicating systematic overestimation in the earlier studies. Further analysis reveals that the relationship between temperature and gas line width aligns more closely with predictions from models incorporating both high cosmic ray ionization rate and turbulent heating, suggesting that these molecular clouds are likely heated by a combination of cosmic-ray and turbulent dissipation mechanisms.

astro-ph.GA

A Clinical Point Cloud Paradigm for In-Hospital Mortality Prediction from Multi-Level Incomplete Multimodal EHRs

Deep learning-based modeling of multimodal Electronic Health Records (EHRs) has become an important approach for clinical diagnosis and risk prediction. However, due to diverse clinical workflows and privacy constraints, raw EHRs are inherently multi-level incomplete, including irregular sampling, missing modalities, and sparse labels. These issues cause temporal misalignment, modality imbalance, and limited supervision. Most existing multimodal methods assume relatively complete data, and even methods designed for incompleteness usually address only one or two of these issues in isolation. As a result, they often rely on rigid temporal/modal alignment or discard incomplete data, which may distort raw clinical semantics. To address this problem, we propose HealthPoint (HP), a unified clinical point cloud paradigm for multi-level incomplete EHRs. HP represents heterogeneous clinical events as points in a continuous 4D space defined by content, time, modality, and case. To model interactions between arbitrary point pairs, we introduce a Low-Rank Relational Attention mechanism that efficiently captures high-order dependencies across these four dimensions. We further develop a hierarchical interaction and sampling strategy to balance fine-grained modeling and computational efficiency. Built on this framework, HP enables flexible event-level interaction and fine-grained self-supervision, supporting robust modality recovery and effective use of unlabeled data. Experiments on large-scale EHR datasets for risk prediction show that HP consistently achieves state-of-the-art performance and strong robustness under varying degrees of incompleteness.

cs.LG