SearcharxivSearch

arXiv subjects

Haijie Li

Publications and source records attributed to Haijie Li.

12 recordsLinked to original sources

KiloDA: Reconstructing kilometer-scale near-surface wind states from sparse station observations

Accurate kilometer-scale near-surface winds are important for understanding atmospheric processes over complex terrain, yet remain difficult to reconstruct from sparse and unevenly distributed observations. Here we introduce KiloDA, a diffusion framework for hourly kilometer-scale wind reconstruction from surface stations. KiloDA learns the statistical distribution and spatial structure of wind fields from historical 3-km Weather Research and Forecasting (WRF) model forecasts. At each reconstruction time, no contemporaneous WRF field is used. Instead, station observations provide the only constraints on the current atmospheric state and guide posterior sampling from the learned prior. In idealized WRF experiments, KiloDA recovers localized wind structures when only 0.24% of grid cells are observed and shows an overall advantage over conventional interpolation across terrain conditions and wind speed regimes. This capability largely transfers to real observations. In a fully withheld region, KiloDA reduces the median wind speed root mean square error (RMSE) by 19% relative to ERA5 reanalysis, using only observations outside the region, with the largest improvements over high-elevation and high-relief terrain. A random station holdout further confirms that this advantage extends across different complex-terrain locations and holdout configurations. These results show that historical model archives can provide useful structural knowledge for reconstructing kilometer-scale wind fields from sparse observations without requiring an accurate model estimate of the current atmospheric state.

physics.ao-ph

CoordRefer: Coordinate-Aware 3D Visual Grounding from Multiview Images

Multiview image-based 3D visual grounding predicts a coordinate frame to define a coordinate system and then regresses a 3D bounding box for localization. However, existing methods jointly optimize coordinate frame selection and box regression, leading to coordinate-relative box ambiguity and degraded grounding performance. This ambiguity arises because the same box admits different numerical representations across coordinate frames, creating multiple optimization targets and yielding invalid compromise predictions. To tackle this challenge, we propose CoordRefer, a coordinate-aware framework that decouples coordinate frame selection from coordinate-conditioned grounding. CoordRefer first selects a reference frame to define the coordinate system and then conditions 3D box prediction on the coordinate system. We perform coordinate-aware supervised fine-tuning to establish coordinate frame selection and coordinate-conditioned box regression, followed by Group Relative Policy Optimization with 3D IoU-based rewards to align both stages with downstream grounding quality. On ScanRefer with Qwen3-VL-2B, CoordRefer achieves gains of 11% in Acc@0.25 and 7% in Acc@0.5 over the coordinate-agnostic baseline, while its geometrically refined variant surpasses methods using explicit 3D inputs.

cs.CV

Diffusion posterior sampling enables zero shot kilometre scale wind forecasting over complex terrain

Reliable prediction of near surface wind over complex terrain is limited by the mismatch between the spatial resolution of operational forecasts and terrain controlled local wind variability. Here, we develop KiloGen, a diffusion posterior sampling framework for kilometre scale wind forecast enhancement. KiloGen learns a high resolution vector wind prior from Weather Research and Forecasting (WRF) model simulations and constrains posterior sampling with 25 km forecasts from the European Centre for Medium Range Weather Forecasts (ECMWF) at inference time. This formulation avoids paired ECMWF and WRF training samples and an explicitly learned mapping from coarse to fine resolution. Applied over Shanxi, China, a region with complex mountainous terrain, KiloGen reconstructs terrain organized wind structures and restores high wavenumber variability while retaining the large scale evolution of the operational forecast. Station verification shows that KiloGen achieves the lowest overall wind speed root mean square error (RMSE) among the evaluated products, with larger benefits at elevated and topographically complex sites. The improvement is strongest under strong wind conditions, reducing RMSE by approximately 10% for observed winds above 20 m s^-1. Across 13 distinct strong wind events, KiloGen improves upon the 0.25 degree ECMWF forecast in all cases and outperforms the 0.1 degree ECMWF forecast in most cases. These results show that diffusion posterior sampling provides an effective approach for terrain aware kilometre scale wind forecast enhancement.

physics.ao-ph

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models

Multimodal Large Language Models (MLLMs) demonstrate exceptional semantic reasoning but struggle with 3D spatial perception when restricted to pure RGB inputs. Despite leveraging implicit geometric priors from 3D reconstruction models, image-based methods still exhibit a notable performance gap compared to methods using explicit 3D data. We argue that this gap does not arise from insufficient geometric priors, but from a misalignment in the training paradigm: text-dominated fine-tuning fails to activate geometric representations within MLLMs. Existing approaches typically resort to naive feature concatenation and optimize directly for downstream tasks without geometry-specific supervision, leading to suboptimal structural utilization. To address this limitation, we propose GAP-MLLM, a Geometry-Aligned Pre-training paradigm that explicitly activates structural perception before downstream adaptation. Specifically, we introduce a visual-prompted joint task that compels the MLLMs to predict sparse pointmaps alongside semantic labels, thereby enforcing geometric awareness. Furthermore, we design a multi-level progressive fusion module with a token-level gating mechanism, enabling adaptive integration of geometric priors without suppressing semantic reasoning. Extensive experiments demonstrate that GAP-MLLM significantly enhances geometric feature fusion and consistently enhances performance across 3D visual grounding, 3D dense captioning, and 3D video object detection tasks.

cs.CV

Spark3R: Asymmetric Token Reduction Makes Fast Feed-Forward 3D Reconstruction

Feed-forward 3D reconstruction models based on Vision Transformers can directly estimate scene geometry and camera poses from a small set of input images, but scaling them to video inputs with hundreds or thousands of frames remains challenging due to the quadratic cost of global attention layers. Recent token-merging methods accelerate these models by compressing the token sequence within the global attention layers, but they apply a uniform reduction to query tokens and key-value tokens, ignoring their functionally distinct roles in 3D reconstruction. In this work, we identify a key property of feed-forward 3D reconstruction models: query tokens encode view-specific geometric requests and are sensitive to compression, while key-value tokens represent shared scene context and tolerate aggressive compression. Guided by this insight, we propose Spark3R, a training-free acceleration framework that decouples the compression of query tokens and key-value tokens by assigning distinct reduction factors, with intra-group token merging applied to query tokens and lightweight token pruning to key-value tokens. Additionally, Spark3R adaptively adjusts the key-value reduction factor across layers, further improving the quality-efficiency trade-off. As a plug-and-play framework requiring no retraining, Spark3R integrates directly into multiple pretrained feed-forward 3D reconstruction models, including VGGT, $π^3$, Depth-Anything-3, and VGGT-$Ω$, and achieves up to $28\times$ speedup on 1,000-frame inputs while maintaining competitive reconstruction quality.

cs.CV

Probabilistic reconstruction of global sea surface temperature using generative diffusion models

Accurate reconstruction of global Sea surface temperature (SST), which dominates the air-sea coupling and global climate variability, underpins climate monitoring and prediction. Existing SST reconstruction products primarily provide one deterministic field derived from heterogeneous satellite data and in situ observations, limiting their ability to represent observation uncertainty and to support probabilistic forecasting. Here, we introduce Satellite and in situ Adaptive Guided Estimation (SAGE), a diffusion-based uncertainty-aware generative framework for probabilistic SST reconstruction. SAGE learns a physically consistent prior from historical SST data and performs observation-conditioned posterior sampling without requiring satellite or in situ data during training, enabling flexible state inference from heterogeneous observations. Through a progressive data-fusion strategy, observations from two FengYun-3D polar-orbiting satellites constrain basin-scale structures, while sparse in situ measurements serve to refine local anomalies and extremes. The resulting ensemble SST fields well capture observational uncertainty and scale-dependent variability. Validation against independent in situ observations shows that SAGE substantially reduces reconstruction errors compared with widely used operational products. When used to initialize forecasting systems, SAGE-generated SST fields substantially reduce 10-day SST forecast errors relative to current operational analyses. At the climate scale, SAGE-driven forecasts of the 2023-2024 El Nino event show added value in capturing its onset and intensity evolution compared to conventional approaches. Our results demonstrate that SAGE represents a step toward a new paradigm for ocean state estimation and climate prediction.

physics.ao-ph

InstanceGaussian: Appearance-Semantic Joint Gaussian Representation for 3D Instance-Level Perception

3D scene understanding has become an essential area of research with applications in autonomous driving, robotics, and augmented reality. Recently, 3D Gaussian Splatting (3DGS) has emerged as a powerful approach, combining explicit modeling with neural adaptability to provide efficient and detailed scene representations. However, three major challenges remain in leveraging 3DGS for scene understanding: 1) an imbalance between appearance and semantics, where dense Gaussian usage for fine-grained texture modeling does not align with the minimal requirements for semantic attributes; 2) inconsistencies between appearance and semantics, as purely appearance-based Gaussians often misrepresent object boundaries; and 3) reliance on top-down instance segmentation methods, which struggle with uneven category distributions, leading to over- or under-segmentation. In this work, we propose InstanceGaussian, a method that jointly learns appearance and semantic features while adaptively aggregating instances. Our contributions include: i) a novel Semantic-Scaffold-GS representation balancing appearance and semantics to improve feature representations and boundary delineation; ii) a progressive appearance-semantic joint training strategy to enhance stability and segmentation accuracy; and iii) a bottom-up, category-agnostic instance aggregation approach that addresses segmentation challenges through farthest point sampling and connected component analysis. Our approach achieves state-of-the-art performance in category-agnostic, open-vocabulary 3D point-level segmentation, highlighting the effectiveness of the proposed representation and training strategies. Project page: https://lhj-git.github.io/InstanceGaussian/

cs.CV

Mirror-3DGS: Incorporating Mirror Reflections into 3D Gaussian Splatting

3D Gaussian Splatting (3DGS) has significantly advanced 3D scene reconstruction and novel view synthesis. However, like Neural Radiance Fields (NeRF), 3DGS struggles with accurately modeling physical reflections, particularly in mirrors, leading to incorrect reconstructions and inconsistent reflective properties. To address this challenge, we introduce Mirror-3DGS, a novel framework designed to accurately handle mirror geometries and reflections, thereby generating realistic mirror reflections. By incorporating mirror attributes into 3DGS and leveraging plane mirror imaging principles, Mirror-3DGS simulates a mirrored viewpoint from behind the mirror, enhancing the realism of scene renderings. Extensive evaluations on both synthetic and real-world scenes demonstrate that our method can render novel views with improved fidelity in real-time, surpassing the state-of-the-art Mirror-NeRF, especially in mirror regions.

cs.CV

OpenGaussian: Towards Point-Level 3D Gaussian-based Open Vocabulary Understanding

This paper introduces OpenGaussian, a method based on 3D Gaussian Splatting (3DGS) capable of 3D point-level open vocabulary understanding. Our primary motivation stems from observing that existing 3DGS-based open vocabulary methods mainly focus on 2D pixel-level parsing. These methods struggle with 3D point-level tasks due to weak feature expressiveness and inaccurate 2D-3D feature associations. To ensure robust feature presentation and 3D point-level understanding, we first employ SAM masks without cross-frame associations to train instance features with 3D consistency. These features exhibit both intra-object consistency and inter-object distinction. Then, we propose a two-stage codebook to discretize these features from coarse to fine levels. At the coarse level, we consider the positional information of 3D points to achieve location-based clustering, which is then refined at the fine level. Finally, we introduce an instance-level 3D-2D feature association method that links 3D points to 2D masks, which are further associated with 2D CLIP features. Extensive experiments, including open vocabulary-based 3D object selection, 3D point cloud understanding, click-based 3D object selection, and ablation studies, demonstrate the effectiveness of our proposed method. The source code is available at our project page: https://3d-aigc.github.io/OpenGaussian

cs.CV

Hybrid Fourier Score Distillation for Efficient One Image to 3D Object Generation

Single image-to-3D generation is pivotal for crafting controllable 3D assets. Given its under-constrained nature, we attempt to leverage 3D geometric priors from a novel view diffusion model and 2D appearance priors from an image generation model to guide the optimization process. We note that there is a disparity between the generation priors of these two diffusion models, leading to their different appearance outputs. Specifically, image generation models tend to deliver more detailed visuals, whereas novel view models produce consistent yet over-smooth results across different views. Directly combining them leads to suboptimal effects due to their appearance conflicts. Hence, we propose a 2D-3D hybrid Fourier Score Distillation objective function, hy-FSD. It optimizes 3D Gaussians using 3D priors in spatial domain to ensure geometric consistency, while exploiting 2D priors in the frequency domain through Fourier transform for better visual quality. hy-FSD can be integrated into existing 3D generation methods and produce significant performance gains. With this technique, we further develop an image-to-3D generation pipeline to create high-quality 3D objects within one minute, named Fourier123. Extensive experiments demonstrate that Fourier123 excels in efficient generation with rapid convergence speed and visually-friendly generation results.

cs.CV

Photon-induced thermal effects in superconducting coplanar waveguide resonators

We experimentally investigated the optical responses of a superconducting niobium resonator. It was found that, with increasing radiation power, the resonance frequency increases monotonically below around 500 mK, decreases monotonically above around 1 K and exhibits a nonmonotonic behavior at around 700 mK. These observations show that one can operate the irradiated resonator in three temperature regimes, depending on whether two-level system (TLS) effects or kinetic inductance effects dominate. Furthermore, we found that the optical responses at ultra-low temperatures can be qualitatively regarded as a photon-induced thermalization effect of TLSs, which could be utilized to achieve thermal sensitive photon detections.

cond-mat.mes-hall

Experimental demonstrations of high-Q superconducting coplanar waveguide resonators

We designed and successfully fabricated an absorption-type of superconducting coplanar waveguide (CPW) resonators. The resonators are made from a Niobium film (about 160 nm thick) on a high-resistance Si substrate, and each resonator is fabricated as a meandered quarter-wavelength transmission line (one end shorts to the ground and another end is capacitively coupled to a through feedline). With a vector network analyzer we measured the transmissions of the applied microwave through the resonators at ultra-low temperature (e.g., at 20 mK), and found that their loaded quality factors are significantly high, i.e., up to 10^6. With the temperature increases slowly from the base temperature (i.e., 20 mK), we observed the resonance frequencies of the resonators are blue shifted and the quality factors are lowered slightly. In principle, this type of CPW-device can integrate a series of resonators with a common feedline, making it a promising candidate of either the data bus for coupling the distant solid-state qubits or the sensitive detector of single photons.

cond-mat.mes-hall