SearcharxivSearch

arXiv subjects

Zhenlong Wu

Publications and source records attributed to Zhenlong Wu.

5 recordsLinked to original sources

GenSplatCodec: Feed-Forward Gaussian Splatting Compression via One-Step Diffusion

Feed-forward 3D Gaussian Splatting (3DGS) enables scalable scene reconstruction without per-scene optimization, yet produces dense Gaussians that are costly to store and transmit. Existing feed-forward Gaussian compression methods formulate decoding as deterministic representation recovery, which becomes inadequate at low bitrates when high-frequency textures and view-dependent appearance are discarded. Although generative models offer a promising alternative, using them as standalone post-processing decouples generation from the transmitted scene structure, thereby compromising cross-view consistency. To address these limitations, we propose GenSplatCodec, a unified feed-forward Gaussian codec that reformulates low-bitrate Gaussian compression as geometry-guided generative decoding. We present a detail-aware feed-forward Gaussian coding scheme within a dual-stream formulation, where the resulting compact Gaussian structural stream is complemented by a lightweight reference appearance stream. We further introduce a geometry-guided one-step generative decoding approach that jointly exploits decoded structural and appearance cues through hierarchical geometry control to reconstruct high-fidelity and view-consistent novel views. Finally, we develop a three-stage optimization strategy that stabilizes the learning of the unified codec and adapts the generative decoder to codec-derived structural and appearance cues. Extensive experiments across multiple datasets demonstrate that GenSplatCodec consistently achieves superior rate-distortion (RD) performance over existing methods.

cs.CV

GeoRect4D: Geometry-Compatible Generative Rectification for Dynamic Sparse-View 3D Reconstruction

Reconstructing dynamic 3D scenes from sparse multi-view videos is highly ill-posed, often leading to geometric collapse, trajectory drift, and floating artifacts. Recent attempts introduce generative priors to hallucinate missing content, yet naive integration frequently causes structural drift and temporal inconsistency due to the mismatch between stochastic 2D generation and deterministic 3D geometry. In this paper, we propose GeoRect4D, a novel unified framework for sparse-view dynamic reconstruction that couples explicit 3D consistency with generative refinement via a closed-loop optimization process. Specifically, GeoRect4D introduces a degradation-aware feedback mechanism that incorporates a robust anchor-based dynamic 3DGS substrate with a single-step diffusion rectifier to hallucinate high-fidelity details. This rectifier utilizes a structural locking mechanism and spatiotemporal coordinated attention, effectively preserving physical plausibility while restoring missing content. Furthermore, we present a progressive optimization strategy that employs stochastic geometric purification to eliminate floaters and generative distillation to infuse texture details into the explicit representation. Extensive experiments demonstrate that GeoRect4D achieves state-of-the-art performance in reconstruction fidelity, perceptual quality, and spatiotemporal consistency across multiple datasets. Project Page: https://mediax-sjtu.github.io/GeoRect4D

cs.CV

4DGCPro: Efficient Hierarchical 4D Gaussian Compression for Progressive Volumetric Video Streaming

Achieving seamless viewing of high-fidelity volumetric video, comparable to 2D video experiences, remains an open challenge. Existing volumetric video compression methods either lack the flexibility to adjust quality and bitrate within a single model for efficient streaming across diverse networks and devices, or struggle with real-time decoding and rendering on lightweight mobile platforms. To address these challenges, we introduce 4DGCPro, a novel hierarchical 4D Gaussian compression framework that facilitates real-time mobile decoding and high-quality rendering via progressive volumetric video streaming in a single bitstream. Specifically, we propose a perceptually-weighted and compression-friendly hierarchical 4D Gaussian representation with motion-aware adaptive grouping to reduce temporal redundancy, preserve coherence, and enable scalable multi-level detail streaming. Furthermore, we present an end-to-end entropy-optimized training scheme, which incorporates layer-wise rate-distortion (RD) supervision and attribute-specific entropy modeling for efficient bitstream generation. Extensive experiments show that 4DGCPro enables flexible quality and multiple bitrate within a single model, achieving real-time decoding and rendering on mobile devices while outperforming existing methods in RD performance across multiple datasets. Project Page: https://mediax-sjtu.github.io/4DGCPro

cs.CV

Impulse Response Invariant Discretization of Complex Fractional Order Integrator

When the order of an integrator is a complex number, the integrator is called a complex fractional order integrator (CFOI). The impulse response invariant discretization (IRID) method is proposed to approximately discretize the CFOI. The definition of the CFOI is introduced firstly, and the real and imaginary parts of the CFOI in frequency-domain responses are derived. The code of IRID for the CFOI based on the MATLAB language is explained. The comparisons of the impulse responses and frequency-domain responses between the CFOI and the approximate discrete/continuous transfer functions are presented to illustrate the effectiveness and correctness of the proposed discretization method. This paper offers a reliable method to implement the CFOI.

eess.SY

Fractional order [PI] Controller and Smith-like Predictor Design for A Class of High Order Systems

To handle the control difficulties caused by high-order dynamics, a control structure based on fractional order [proportional integral] (PI) controller and fractional order Smith-like predictor for a class of high order systems in the type of K/(Ts+1)n is proposed in this paper. The analysis of the tracking and disturbance rejection is illustrated based on the terminal value theorem and shows that the proposed control structure can ensure that the closed-loop system converges to the set point without static error and the closed-loop system recovers to its original state when the input disturbance occurs. Then, simulations about the influence on the control performance and control signal with different are carried out based on multi-objective genetic algorithm (MO-GA). The results show that the control performance can be improved and the energy of the control signal can be reduced simultaneously when the order is chosen no more than one. This can verify that the fractional order Smith-like predictor with has an advantage over that of the integral order Smith-like predictor.

eess.SY