SearcharxivSearch

arXiv subjects

Zhichao Ye

Publications and source records attributed to Zhichao Ye.

At least 19 recordsLinked to original sources

QuerySplat: Decoupling Geometry and Appearance Representations in 3DGS Prediction

While feed-forward 3D Gaussian Splatting (3DGS) enables efficient 3D reconstruction, achieving high-fidelity rendering remains challenging. Existing pixel-aligned approaches suffer from spatial inflexibility and massive structural redundancy, whereas query-based methods lack 3D priors and entangle geometry with appearance, yielding blurry, pose-dependent results. To overcome these deficiencies, we propose \textbf{QuerySplat}, a feed-forward 3DGS framework driven by geometric priors and explicit appearance decoupling. Specifically, we design a dual-branch query-based decoder: the geometry branch leverages a pretrained Vision Geometric Model for spatial understanding, which intrinsically endows QuerySplat with pose-free modeling capabilities, while the appearance branch recovers high-frequency details through a dedicated pathway separated from geometric attribute regression. Extensive experiments demonstrate that QuerySplat mitigates the blurry rendering issues of early query-based models and consistently outperforms pixel-aligned approaches in rendering fidelity. On the challenging DL3DV benchmark, it achieves state-of-the-art novel view synthesis performance, with average PSNR gains of 2.30 dB and 1.04 dB over the best pose-free and pose-required baselines, respectively. Project Page: https://inspatio.github.io/querysplat.

cs.CV

PostCam: Camera-Controllable Novel-View Video Generation with Query-Shared Cross-Attention

We propose PostCam, a streamlined framework for novel-view video generation that achieves superior detail preservation and precise camera trajectory editing in dynamic scenes. Current methods often struggle with a trade-off between pose-based control, which lacks visual detail, and rendering-based guidance, which is overly sensitive to geometric accuracy. Despite recent hybrid attempts, achieving precise motion and visual consistency remains challenging due to the lack of effective cross-modal alignment. We argue that robust control stems from the deep alignment of multimodal signals rather than increased input complexity. Our core contribution is the Query-Shared Cross-Attention mechanism, which projects 6-DoF poses and rendered features into a unified latent space. This allows the model to spontaneously achieve intrinsic consistency between motion cues and pixel-level guidance during denoising. Experiments demonstrate that PostCam maintains high-fidelity visual details while outperforming state-of-the-art methods by 20% in trajectory precision, exhibiting superior robustness in complex dynamic scenes. Our project webpage is publicly available at: https://cccqaq.github.io/PostCam.github.io/

cs.CV

InSpatio-WorldFM: An Open-Source Real-Time Generative Frame Model

We present InSpatio-WorldFM, an open-source real-time frame model for spatial intelligence. Unlike video-based world models that rely on sequential frame generation and incur substantial latency due to window-level processing, InSpatio-WorldFM adopts a frame-based paradigm that generates each frame independently, enabling low-latency real-time spatial inference. By enforcing multi-view spatial consistency through explicit 3D anchors and implicit spatial memory, the model preserves global scene geometry while maintaining fine-grained visual details across viewpoint changes. We further introduce a progressive three-stage training pipeline that transforms a pretrained image diffusion model into a controllable frame model and finally into a real-time generator through few-step distillation. Experimental results show that InSpatio-WorldFM achieves strong multi-view consistency while supporting interactive exploration on consumer-grade GPUs, providing an efficient alternative to traditional video-based world models for real-time world simulation.

cs.CV

INSPATIO-WORLD: A Real-Time 4D World Simulator via Spatiotemporal Autoregressive Modeling

Building world models with spatial consistency and real-time interactivity remains a fundamental challenge in computer vision. Current video generation paradigms often struggle with a lack of spatial persistence and insufficient visual realism, making it difficult to support seamless navigation in complex environments. To address these challenges, we propose INSPATIO-WORLD, a novel real-time framework capable of recovering and generating high-fidelity, dynamic interactive scenes from a single reference video. At the core of our approach is a Spatiotemporal Autoregressive (STAR) architecture, which enables consistent and controllable scene evolution through two tightly coupled components: Implicit Spatiotemporal Cache aggregates reference and historical observations into a latent world representation, ensuring global consistency during long-horizon navigation; Explicit Spatial Constraint Module enforces geometric structure and translates user interactions into precise and physically plausible camera trajectories. Furthermore, we introduce Joint Distribution Matching Distillation (JDMD). By using real-world data distributions as a regularizing guide, JDMD effectively overcomes the fidelity degradation typically caused by over-reliance on synthetic data. Extensive experiments demonstrate that INSPATIO-WORLD significantly outperforms existing state-of-the-art (SOTA) models in spatial consistency and interaction precision, ranking first among real-time interactive methods on the WorldScore-Dynamic benchmark, and establishing a practical pipeline for navigating 4D environments reconstructed from monocular videos.

cs.CV

SparseCam4D: Spatio-Temporally Consistent 4D Reconstruction from Sparse Cameras

High-quality 4D reconstruction enables photorealistic and immersive rendering of the dynamic real world. However, unlike static scenes that can be fully captured with a single camera, high-quality dynamic scenes typically require dense arrays of tens or even hundreds of synchronized cameras. Dependence on such costly lab setups severely limits practical scalability. To this end, we propose a sparse-camera dynamic reconstruction framework that exploits abundant yet inconsistent generative observations. Our key innovation is the Spatio-Temporal Distortion Field, which provides a unified mechanism for modeling inconsistencies in generative observations across both spatial and temporal dimensions. Building on this, we develop a complete pipeline that enables 4D reconstruction from sparse and uncalibrated camera inputs. We evaluate our method on multi-camera dynamic scene benchmarks, achieving spatio-temporally consistent high-fidelity renderings and significantly outperforming existing approaches. Project page available at https://inspatio.github.io/sparse-cam4d/

cs.CV

EGG-Fusion: Efficient 3D Reconstruction with Geometry-aware Gaussian Surfel on the Fly

Real-time 3D reconstruction is a fundamental task in computer graphics. Recently, differentiable-rendering-based SLAM system has demonstrated significant potential, enabling photorealistic scene rendering through learnable scene representations such as Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS). Current differentiable rendering methods face dual challenges in real-time computation and sensor noise sensitivity, leading to degraded geometric fidelity in scene reconstruction and limited practicality. To address these challenges, we propose a novel real-time system EGG-Fusion, featuring robust sparse-to-dense camera tracking and a geometry-aware Gaussian surfel mapping module, introducing an information filter-based fusion method that explicitly accounts for sensor noise to achieve high-precision surface reconstruction. The proposed differentiable Gaussian surfel mapping effectively models multi-view consistent surfaces while enabling efficient parameter optimization. Extensive experimental results demonstrate that the proposed system achieves a surface reconstruction error of 0.6\textit{cm} on standardized benchmark datasets including Replica and ScanNet++, representing over 20\% improvement in accuracy compared to state-of-the-art (SOTA) GS-based methods. Notably, the system maintains real-time processing capabilities at 24 FPS, establishing it as one of the most accurate differentiable-rendering-based real-time reconstruction systems. Project Page: https://zju3dv.github.io/eggfusion/

cs.CV

StarGen: A Spatiotemporal Autoregression Framework with Video Diffusion Model for Scalable and Controllable Scene Generation

Recent advances in large reconstruction and generative models have significantly improved scene reconstruction and novel view generation. However, due to compute limitations, each inference with these large models is confined to a small area, making long-range consistent scene generation challenging. To address this, we propose StarGen, a novel framework that employs a pre-trained video diffusion model in an autoregressive manner for long-range scene generation. The generation of each video clip is conditioned on the 3D warping of spatially adjacent images and the temporally overlapping image from previously generated clips, improving spatiotemporal consistency in long-range scene generation with precise pose control. The spatiotemporal condition is compatible with various input conditions, facilitating diverse tasks, including sparse view interpolation, perpetual view generation, and layout-conditioned city generation. Quantitative and qualitative evaluations demonstrate StarGen's superior scalability, fidelity, and pose accuracy compared to state-of-the-art methods. Project page: https://zju3dv.github.io/StarGen.

cs.CV

Ultraviolet astronomical spectrograph calibration with laser frequency combs from nanophotonic lithium niobate waveguides

Astronomical precision spectroscopy underpins searches for life beyond Earth, direct observation of the expanding Universe and constraining the potential variability of physical constants across cosmological scales. Laser frequency combs can provide the critically required accurate and precise calibration to the astronomical spectrographs. For cosmological studies, extending the calibration with such astrocombs to the ultraviolet spectral range is highly desirable, however, strong material dispersion and large spectral separation from the established infrared laser oscillators have made this exceedingly challenging. Here, we demonstrate for the first time astronomical spectrograph calibrations with an astrocomb in the ultraviolet spectral range below 400 nm. This is accomplished via chip-integrated highly nonlinear photonics in periodically-poled, nano-fabricated lithium niobate waveguides in conjunction with a robust infrared electro-optic comb generator, as well as a chip-integrated microresonator comb. These results demonstrate a viable route towards astronomical precision spectroscopy in the ultraviolet and may contribute to unlocking the full potential of next generation ground- and future space-based astronomical instruments.

physics.optics

Self-injection-locked optical parametric oscillator based on microcombs

Narrow-linewidth yet tunable laser oscillators are one of the most important tools for precision metrology, optical atomic clocks, sensing and quantum computing. Commonly used tunable coherent oscillators are based on stimulated emission or stimulated Brillouin scattering; as a result, the operating wavelength band is limited by the gain media. Based on nonlinear optical gain, optical parametric oscillators (OPOs) enable coherent signal generation within the whole transparency window of the medium used. However, the demonstration of OPO-based Hertz-level linewidth and tunable oscillators has remained elusive. Here, we present a tunable coherent oscillator based on a multimode coherent OPO in a high-Q microresonator, i.e., a microcomb. Single-mode coherent oscillation is realized through self-injection locking (SIL) of one selected comb line. We achieve coarse tuning up to 20 nm, and an intrinsic linewidth down to sub-Hertz level, which is three orders of magnitude lower than the pump. Furthermore, we demonstrate that this scheme results into repetition rate stabilization of the microcomb. These results open exciting possibilities for generating tunable coherent radiation where stimulated emission materials are difficult to obtain, and the stabilization of microcomb sources beyond the limits imposed by the thermorefractive noise in the cavity.

physics.optics

A wideband, high-resolution vector spectrum analyzer for integrated photonics

The analysis of optical spectra - emission or absorption -- has been arguably the most powerful approach for discovering and understanding matters. The invention and development of many kinds of spectrometers have equipped us with versatile yet ultra-sensitive diagnostic tools for trace gas detection, isotope analysis, and resolving hyperfine structures of atoms and molecules. With proliferating data and information, urgent and demanding requirements have been placed today on spectrum analysis with ever-increasing spectral bandwidth and frequency resolution. These requirements are especially stringent for broadband laser sources that carry massive information, and for dispersive devices used in information processing systems. In addition, spectrum analyzers are expected to probe the device's phase response where extra information is encoded. Here we demonstrate a novel vector spectrum analyzer (VSA) that is capable of characterizing passive devices and active laser sources in one setup. Such a dual-mode VSA can measure loss, phase response and dispersion properties of passive devices. It also can coherently map a broadband laser spectrum into the RF domain. The VSA features a bandwidth of 55.1 THz (1260 to 1640 nm), frequency resolution of 471 kHz, and dynamic range of 56 dB. Meanwhile, our fiber-based VSA is compact and robust. It requires neither high-speed modulators and photodetectors, nor any active feedback control. Finally, we successfully employ our VSA for applications including characterization of integrated dispersive waveguides, mapping frequency comb spectra, and coherent light detection and ranging (LiDAR). Our VSA presents an innovative approach for device analysis and laser spectroscopy, and can play a critical role in future photonic systems and applications for sensing, communication, imaging, and quantum information processing.

physics.optics

Vernier Microcombs for Integrated Optical Atomic Clocks

CMOS-compatible Kerr microcombs have drawn substantial interest as mass-manufacturable, compact alternatives to bulk frequency combs. This could enable deployment of many comb-reliant applications previously confined to laboratories. Particularly enticing is the prospect of microcombs performing optical frequency division in compact optical atomic clocks. Unfortunately, it is difficult to meet the self-referencing requirement of microcombs in these systems due to the $\sim$THz repetition rates typically required for octave-spanning comb generation. Additionally, it is challenging to spectrally engineer a microcomb system to align a comb mode with an atomic clock transition with sufficient signal-to-noise ratio. Here, we adopt a Vernier dual-microcomb scheme for optical frequency division of a stabilized ultranarrow-linewidth continuous-wave laser at 871 nm to a $\sim$235 MHz output frequency. In addition to enabling measurement of the comb repetition rates, this scheme brings the freedom to pick comb lines from either or both of the combs. We exploit this flexibility to shift an ultra-high-frequency ($\sim$100 GHz) carrier-envelope offset beat down to frequencies where detection is possible and to place a comb line close to the 871 nm laser - tuned so that if frequency-doubled it would fall close to the clock transition in $^{171}$Yb$^+$. Moreover, we introduce a novel scheme which suppresses frequency noise arising from interferometric phase fluctuations in our dual-comb system and reduces the frequency instability down to our measurement limit. Our dual-comb system can potentially combine with an integrated ion trap toward future chip-scale optical atomic clocks.

physics.optics

EC-SfM: Efficient Covisibility-based Structure-from-Motion for Both Sequential and Unordered Images

Structure-from-Motion is a technology used to obtain scene structure through image collection, which is a fundamental problem in computer vision. For unordered Internet images, SfM is very slow due to the lack of prior knowledge about image overlap. For sequential images, knowing the large overlap between adjacent frames, SfM can adopt a variety of acceleration strategies, which are only applicable to sequential data. To further improve the reconstruction efficiency and break the gap of strategies between these two kinds of data, this paper presents an efficient covisibility-based incremental SfM. Different from previous methods, we exploit covisibility and registration dependency to describe the image connection which is suitable to any kind of data. Based on this general image connection, we propose a unified framework to efficiently reconstruct sequential images, unordered images, and the mixture of these two. Experiments on the unordered images and mixed data verify the effectiveness of the proposed method, which is three times faster than the state of the art on feature matching, and an order of magnitude faster on reconstruction without sacrificing the accuracy. The source code is publicly available at https://github.com/openxrlab/xrsfm

cs.CV

Foundry manufacturing of tight-confinement, dispersion-engineered, ultralow-loss silicon nitride photonic integrated circuit

The foundry development of integrated photonics has revolutionized today's optical interconnect and datacenters. Over the last decade, we have witnessed the rising of silicon nitride (Si$_3$N$_4$) integrated photonics, which is currently transferring from laboratory research to foundry manufacturing. The development and transition are triggered by the ultimate need of low optical loss offered by Si$_3$N$_4$, which is beyond the reach of silicon and III-V semiconductors. Combined with modest Kerr nonlinearity, tight optical confinement and dispersion engineering, Si$_3$N$_4$ has today become the leading platform for linear and Kerr nonlinear photonics, and has enabled chip-scale lasers featuring ultralow noise on par with table-top fiber lasers. However, so far all the reported fabrication processes of tight-confinement, dispersion-engineered Si$_3$N$_4$ photonic integrated circuit (PIC) with optical loss down to few dB/m have only been developed on 4-inch or smaller wafers. Yet, to transfer these processes to established CMOS foundries that typically operate 6-inch or even larger wafers, challenges remain. In this work, we demonstrate the first foundry-standard fabrication process of Si$_3$N$_4$ PIC with only 2.6 dB/m loss, thickness above 800 nm, and near 100% fabrication yield on 6-inch wafers. Such thick and ultralow-loss Si$_3$N$_4$ PIC enables low-threshold generation of soliton frequency combs. Merging with advanced heterogeneous integration, active ultralow-loss Si$_3$N$_4$ integrated photonics could pave an avenue to addressing future demands in our increasingly information-driven society.

physics.optics

Compact lithium niobate photonic integrated circuits

Lithium niobate (LN) is a promising material for future complex photonic-electronic circuits, with wide applications in fields like communications, sensing, quantum optics, and computation. LN took a great stride toward compact photonic integrated circuits (PICs) with the development of partially-etched LN on insulator (LNOI) waveguides. However, integration density is still limited for future high-compact PICs due to the partial edge nature of their waveguides. Here, we demonstrate a fully-etched LN PIC platform which, for the first time, simultaneously achieves ultra-low propagation loss and compact circuit size. The tightly-confined fully-etched LN waveguides with smooth sidewalls allow us to bring the bending radius down to 20 $μ$m (corresponds to 1 THz FSR). We have achieved compact high-$Q$ microring resonators with $Q/V$ of 7.1 $\times$ 10$^{4}$ $μ$m$^{-3}$, almost one order of magnitude larger than previous demonstrations. The statistical mean propagation losses of our LN waveguides is 8.5 dB/m (corresponds to mean $Q$-factor of 4.9 $\times$ 10$^{6}$) even with a small bending radius of 40 $μ$m. Our compact and ultra-low-loss LN platform shows great potential in future miniaturized multifunctional integration systems. As complementary evidence to show the utility of our platform, we demonstrate soliton microcombs with an ultra-high repetition rate of 500 GHz in LN.

physics.optics

Hyperparametric oscillation via bound states in the continuum

Optical hyperparametric oscillation based on the third-order nonlinearity is one of the most significant mechanisms to generate coherent electromagnetic radiation and produce quantum states of light. Advances in dispersion-engineered high-$Q$ microresonators allow for generating signal waves far from the pump and decrease the oscillation power threshold to submilliwatt levels. However, the pump-to-signal conversion efficiency and absolute signal power are low, fundamentally limited by parasitic mode competition and attainable cavity intrinsic $Q$ to coupling $Q$ ratio, i.e., $Q_{\rm i}/Q_{\rm c}$. Here, we use Friedrich-Wintgen bound states in the continuum (BICs) to overcome the physical challenges in an integrated microresonator-waveguide system. As a result, on-chip coherent hyperparametric oscillation is generated in BICs with unprecedented conversion efficiency and absolute signal power. This work not only opens a path to generate high-power and efficient continuous-wave electromagnetic radiation in Kerr nonlinear media but also enhances the understanding of microresonator-waveguide system - an elementary unit of modern photonics.

physics.optics

Generative Category-Level Shape and Pose Estimation with Semantic Primitives

Empowering autonomous agents with 3D understanding for daily objects is a grand challenge in robotics applications. When exploring in an unknown environment, existing methods for object pose estimation are still not satisfactory due to the diversity of object shapes. In this paper, we propose a novel framework for category-level object shape and pose estimation from a single RGB-D image. To handle the intra-category variation, we adopt a semantic primitive representation that encodes diverse shapes into a unified latent space, which is the key to establish reliable correspondences between observed point clouds and estimated shapes. Then, by using a SIM(3)-invariant shape descriptor, we gracefully decouple the shape and pose of an object, thus supporting latent shape optimization of target objects in arbitrary poses. Extensive experiments show that the proposed method achieves SOTA pose estimation performance and better generalization in the real-world dataset. Code and video are available at https://zju3dv.github.io/gCasp.

cs.CV

Differential phase reconstruction of microcombs

Measuring microcombs in amplitude and phase provides unique insight into the nonlinear cavity dynamics but spectral phase measurements are experimentally challenging. Here, we report a linear heterodyne technique assisted by electro-optic downconversion that enables differential phase measurement of such spectra with unprecedented sensitivity (-50 dBm) and bandwidth coverage (> 110 nm in the telecommunications range). We validate the technique with a series of measurements, including single cavity and photonic molecule microcombs.

physics.optics

Optical linewidth of soliton microcombs

Soliton microcombs provide a versatile platform for realizing fundamental studies and technological applications. To be utilized as frequency rulers for precision metrology, soliton microcombs must display broadband phase coherence, a parameter characterized by the optical phase or frequency noise of the comb lines and their corresponding optical linewidths. Here, we analyse the optical phase-noise dynamics in soliton microcombs generated in silicon nitride high-Q microresonators and show that, because of the Raman self-frequency shift or dispersive-wave recoil, the Lorentzian linewidth of some of the comb lines can, surprisingly, be narrower than that of the pump laser. This work elucidates information about the physical limits in phase coherence of soliton microcombs and illustrates a new strategy for the generation of spectrally coherent light on chip.

physics.optics