SearcharxivSearch

arXiv subjects

Yuanwei Zhu

Publications and source records attributed to Yuanwei Zhu.

3 recordsLinked to original sources

ArtGen: Conditional Generative Modeling of Articulated Objects in Arbitrary Part-Level States

Generating articulated assets is crucial for robotics, digital twins, and embodied intelligence. Existing generative models often rely on single-view inputs representing closed states, resulting in ambiguous or unrealistic kinematic structures due to the entanglement between geometric shape and joint dynamics. To address these challenges, we introduce ArtGen, a conditional diffusion-based framework capable of generating articulated 3D objects with accurate geometry and coherent kinematics from single-view images or text descriptions at arbitrary part-level states. Specifically, ArtGen employs cross-state Monte Carlo sampling to explicitly enforce global kinematic consistency, reducing structural-motion entanglement. Additionally, we integrate a Chain-of-Thought reasoning module to infer robust structural priors, such as part semantics, joint types, and connectivity, guiding a sparse-expert Diffusion Transformer to specialize in diverse kinematic interactions. Furthermore, a compositional 3D-VAE latent prior enhanced with local-global attention effectively captures fine-grained geometry and global part-level relationships. Extensive experiments on the PartNet-Mobility benchmark demonstrate that ArtGen significantly outperforms state-of-the-art methods.

cs.CV

Towards Holographic Video Communications: A Promising AI-driven Solution

Real-time holographic video communications enable immersive experiences for next-generation video services in the future metaverse era. However, high-fidelity holographic videos require high bandwidth and significant computation resources, which exceed the transferring and computing capacity of 5G networks. This article reviews state-of-the-art holographic point cloud video (PCV) transmission techniques and highlights the critical challenges of delivering such immersive services. We further implement a preliminary prototype of an AI-driven holographic video communication system and present critical experimental results to evaluate its performance. Finally, we identify future research directions and discuss potential solutions for providing real-time and high-quality holographic experiences.

cs.MM

ISCom: Interest-aware Semantic Communication Scheme for Point Cloud Video Streaming

The provisioning of immersive point cloud video (PCV) streaming on pervasive mobile devices is a cornerstone for enabling immersive communication and interactions in the future 6G metaverse era. However, most streaming techniques are dedicated to efficient PCV compression and codec extending from traditional 3-DoF video services. Some emerging AI-enabled approaches are still in their infancy phase and are constrained by intensive computational and adaptive flow techniques. In this paper, we present ISCom, an Interest-aware Semantic Communication Scheme for PCV, consisting of a region-of-interest (ROI) selection module, a lightweight PCV streaming module, and an intelligent scheduler. First, we propose a two-stage efficient ROI selection method for providing interest-aware PCV streaming, which significantly reduces the data volume. Second, we design a lightweight PCV encoder-decoder network for resource-constrained devices, adapting to the heterogeneous computing capabilities of terminals. Third, we train a deep reinforcement learning (DRL)-based scheduler to adapt an optimal encoder-decoder network for various devices, considering the dynamic network environments and computing capabilities of different devices. Extensive experiments show that ISCom outperforms baselines on mobile devices at least 10 FPS and up to 22 FPS.

cs.MM