SearcharxivSearch

arXiv subjects

Cheng Lin

Publications and source records attributed to Cheng Lin.

At least 19 recordsLinked to original sources

C2Dex: Contact-Consistent Reconstruction and Retargeting for Dexterous Manipulation from Monocular Video

High-quality demonstrations for dexterous robot manipulation are costly and difficult to collect, whereas monocular human videos provide a scalable source of diverse manipulation behaviors. However, transferring such demonstrations to dexterous robots remains challenging: monocular hand-object interaction (HOI) reconstruction often produces temporally unstable contacts and physically implausible interactions, while conventional retargeting methods struggle to preserve task-relevant contacts and local interaction geometry across different hand embodiments. We present C2Dex, a video-to-dexterous-manipulation framework built around a shared interaction representation: stable object-side contacts recovered by aggregating noisy frame-wise observations in the canonical object space. These stable contacts serve a dual role: as trajectory-level constraints that guide reconstruction toward temporally coherent and physically plausible human HOI trajectories, and as explicit transfer targets for the dexterous hand, where Laplacian interaction optimization preserves the local hand-object geometry across embodiments and residual reinforcement learning refines the trajectory in simulation. Experiments on DexYCB and TACO show that C2Dex achieves end-to-end trajectory success rates of 57.78% and 26.67%, respectively, substantially outperforming the strongest baselines (17.78% and 10.00%) under identical evaluation criteria. Real-robot replay experiments further demonstrate physical feasibility across diverse contact-rich manipulation tasks. Project page: https://k-jie.github.io/C2Dex/

cs.RO

An affirmative solution to the generalized Busemann--Petty problem with subspace dimensions $2$ and $3$

The generalized Busemann--Petty problem asks whether origin-symmetric convex bodies in $\mathbb{R}^n$ having larger volume of all $m$-dimensional sections necessarily have larger volume. When $m\geq 4$, this is known to be false, but the cases $m=2, 3$ for $n\geq 5$ have remained open since the 1990s. In this paper, we resolve these cases. Together with the known results, the generalized Busemann--Petty problem is completely solved: the answer is affirmative for $m=1, 2, 3$, and negative for $m\geq 4$.

math.FA

Topological Rainbow Trapping for Spatial-frequency Demultiplexing of Underwater Acoustic Signals

Efficient separation and localization of multifrequency acoustic waves are essential for underwater target recognition and acoustic energy harvesting. The underwater implementation of topological rainbow trapping remains challenging because of complex fluid-solid interactions and the difficulty of integrating long-range transport with frequency-selective localization in an open system. Here, we theoretically develop and experimentally demonstrate two underwater spatial-frequency demultiplexing mechanisms based on the acoustic analogues of the QVHE and QSHE. Both mechanisms employ SSAWs, whose fields are confined near a structured surface and decay evanescently into the surrounding water, enabling experiments without an enclosed waveguide. In the QVHE mechanism, a spatial gradient along a valley-Hall edge channel shifts the local edge-state dispersion, causing different frequency components to become localized at distinct positions and thereby realizing spectral and spatial demultiplexing. In the QSHE mechanism, one-dimensional topological edge states are coupled to frequency-selective zero-dimensional higher-order corner states. Multifrequency signals first propagate robustly along a common boundary and are then transferred to prescribed remote corners according to frequency, producing a transport-then-confinement process. This mechanism combines defect-tolerant edge transport, frequency-selective corner localization, and remote rainbow trapping. Numerical simulations and experiments verify the frequency-dependent localization and the persistence of the designed transport pathways in the presence of structural defects. The proposed open SSAW platform performs robust frequency demultiplexing at the physical layer, reducing reliance on digital signal processing and offering potential for underwater target recognition and frequency-selective acoustic energy harvesting.

physics.app-ph

Elastic Trapped States at Dislocation Defects in Scaled Coupling and Hofstadter Models

Elastic topological dislocations provide a pathway for trapping elastic wave energy at internal defects, rather than being confined solely to external boundaries or corners, which are typically associated with topological insulators (TIs). However, two practical constraints persist. First, highly confined dislocation states based on conventional Su-Schrieffer-Heeger (SSH) dimerization usually require a large coupling contrast and a correspondingly enlarged bandgap, which may be challenging to realize. Second, some Hamiltonians with richer topological physics often contain complex hopping terms, synthetic gauge fields or nonlocal couplings, which substantially increase the geometric complexity of experimental samples. Here, dislocation-induced trapped states are demonstrated in both a scaled coupling (SC) model and a Hofstadter model (HM) within an elastic platform. In the SC model, the trapped mode is treated as a higher localized state in the continuum rather than an in-gap mode in the SSH model. Consequently, the SC-induced dislocation can trap an enhanced mode without the requirement of an enlarged bandgap. For the HM, Householder tridiagonalization is used to map the original tight-binding Hamiltonian with complex hopping terms onto a tridiagonal matrix with only positive-real-valued nearest-neighbour (NN) hopping terms. Truncation at a weak-hopping position preserves the topological phenomena and allows a dislocation defect to be constructed from the shortened aperiodic chain. The results establish a practical route for designing highly localized modes without relying solely on bandgap enlargement or complex couplings, which advance the topological physics of elastic wave systems and promise enhanced possibilities for elastic functional devices.

physics.app-ph

Chiral Landau levels induced by two in-plane pseudomagnetic fields in underwater acoustic metamaterials

The chiral zeroth Landau levels (LLs) constitute topologically protected bulk states that enable robust control of acoustic wave propagation. Given the central role of underwater acoustics in marine engineering, realizing such Landau-level physics in underwater acoustic systems is highly desirable. Nevertheless, existing studies have primarily been limited to airborne acoustic systems, and the implementation of chiral zeroth LLs in underwater acoustics remains a challenge due to the unavoidable fluid-solid interactions. In this study, we realize two kinds of chiral LLs in an open underwater spoof surface acoustic wave (SSAW) platform by introducing two perpendicular in-plane artificial pseudomagnetic fields (PMFs), oriented along the x and y directions, respectively, and reveal that scalar acoustic fields in water and vectorial elastic vibrations in solids can be jointly manipulated within a unified framework. Specifically, by strategically opening bandgaps at the Dirac points, position-dependent effective mass terms are introduced into the Dirac Hamiltonians, thereby synthesizing two in-plane PMFs. This results in the emergence of chiral LLs, which is confirmed both numerically and experimentally. The unidirectional propagation of the chiral LLs and their robustness against defects are also demonstrated. In addition, we achieve flexible manipulation of underwater ultrasonic energy carried by SSAWs, including beam splitting and arbitrary wave steering. Dual-band chiral LLs are also observed in small-scale underwater topological metamaterials. Our work provides a new route toward SSAW-based underwater ultrasonic control, opening opportunities for multiband underwater acoustic signal processing and detection, as well as underwater acoustic energy harvesting.

physics.app-ph

Experimental Realization of Type-II Quadrupole Topological Insulator

The discovery of quadrupole topological insulators (QTIs) has spurred extensive research into higher-order topological phases. Recently proposed type-II QTIs exhibit unconventional topological behaviors with 1/2 edge polarization \operatorname{p}_x and zero edge polarization \operatorname{p}_y, due to the inequivalence between Wannier-band and edge-spectrum gap closures, yet their experimental realization remains challenging owing to the long-range and complex off-site hopping terms in their tight-binding model (TBM). Here, we circumvent this difficulty via an optimized Householder tridiagonalization (OHT) mapping that reduces the complex two-dimensional lattices to one-dimensional chains with only negative-real-valued nearest-neighbor hopping terms, greatly facilitating experimental sample fabrication. Using this strategy, we experimentally verify the type-II QTI phase, type-I QTI phase and trivial phase in elastic wave platforms via simple aperiodic plate-beam chain structures, where the plates reflect the on-site potential terms and beams correspond to the off-site hopping terms in the TBM. Our approach provides a versatile route for experimentally exploring more complex and richer topological phenomena based on TBM.

physics.app-ph

CelloCut: Constructive Watertight Remeshing via Tetrahedral Cell Cuts

Watertight remeshing aims to recover a surface that induces a globally consistent interior--exterior partition of 3D space. However, for meshes with complex topology, single-layer structures, or large missing regions, inferring such a partition from local surface geometry is inherently ambiguous. As a result, existing methods often produce surface-accurate yet volumetrically inconsistent reconstructions, e.g., closely spaced double shells. The key insight of this work is that watertight remeshing should be treated as a volumetric partitioning problem rather than a surface-level repair task. To this end, we propose CelloCut, a constructive framework that formulates watertight conversion as a binary labeling problem over a Delaunay tetrahedral partition of space. We solve this via graph-cut energy minimization with one-sided constraints that preserve proxy-supported interior evidence and weighted interface penalties that discourage unsupported newly introduced boundaries. By computing a globally consistent volumetric partition, CelloCut guarantees a strictly watertight output by construction and strongly suppresses pseudo-watertight artifacts such as double shells, even under severe topological defects. Experimental results on two newly introduced challenging benchmarks, CelloScan and CelloFill, as well as standard ModelNet10 dataset, demonstrate that CelloCut significantly outperforms state-of-the-art methods, particularly in handling complex topologies and single-layer structures, producing compact and volumetrically consistent solid reconstructions. The project page is available at https://rangeryx-66.github.io/CelloCut/.

cs.GR

DecoRec: Decomposed 3D Scene Reconstruction from Single-View Images via Object-Level Diffusion

In this paper, we introduce \textit{DecoRec}, a novel system designed to elevate single-view 2D images to a decomposed 3D scene mesh. Current methods for single-view scene reconstruction typically rely on object retrieval or the regression of coarse 3D voxels or surfaces, leading to inaccuracies in capturing the appearance and geometry of the input image. The lack of high-quality large-scale scene-level datasets further complicates direct 3D scene generation from single-view images. To achieve high-quality 3D scene generation from a single-view image, DecoRec takes advantage of recent diffusion-based single-view object reconstruction methods to reconstruct individual objects separately. Subsequently, a refinement pipeline is proposed to effectively merge these reconstructed objects, enhancing appearance and geometry through a differentiable rendering technique and diffusion-guided refinement. Our results demonstrate that DecoRec facilitates high-quality single-view scene reconstruction in both geometry and novel synthesis, offering significant benefits for downstream applications like room interior design.

cs.CV

QuadLink: Autoregressive Quad-Dominant Mesh Generation via Point-Relation Learning

The generation of production-ready quad-dominant meshes is a cornerstone of modern 3D content creation. Generating anisotropic quad-dominant meshes from point clouds is challenging, as existing methods are typically limited to producing either pure triangular meshes or pure quadrilateral meshes with isotropic densities. In this paper, we present QuadLink, a unified framework consisting of three stages for quad-dominant mesh generation by linking points into structured faces. QuadLink formulates polygonal mesh generation as a hybrid centroid-conditioned vertex linking model: it first predicts a unified set of anchors (vertices and face centroids), then learns centroid-conditioned links that associate vertices with face centroids, and finally assembles polygonal faces with a quad-first strategy guided by robust geometric verification strategies. This link-based formulation enables efficient generation of sparse and anisotropic quad-dominant meshes with coherent edge flow and meanwhile supporting hybrid polygonal topology. To construct training data for this model, we further introduce a Tri-to-Quad Operator that converts artistic triangle meshes into quad-dominant training data via global merge selection. Extensive experiments show that QuadLink produces production-ready quad-dominant meshes from point clouds and achieves improved geometric fidelity and topological quality compared to prior baselines. Our method natively supports hybrid polygonal topology, generalizing to arbitrary n-gon meshes without architectural changes.

cs.GR

GaussiAnimate: Rig Animatable Categories with Level of Dynamics

We propose Skelebones, a Scaffold-Skin Rigging System built on three steps: (1) Bones compress temporally consistent Gaussian or mesh sequences into free-form bones with smooth skinning weights, approximating non-rigid deformations via linear blend skinning (LBS); (2) Skeleton extracts the Mean Curvature Skeleton (MCS) from the canonical shape and temporally refines its topology and kinematics into a compact skeletal structure; and (3) Binding connects the skeleton and bones through non-parametric Partwise Motion Matching (PartMM), which synthesizes novel bone motions by matching, retrieving, and blending existing ones. Together, these steps compress the dynamics of 4D shapes into compact skelebones that are simultaneously controllable and expressive. The resulting representation is category-agnostic, meaning template-free; motion-adaptive, with dynamic topology; and topology-correct, with a skeleton consistent with the surface geometry. PartMM requires no learning. We validate our method on both synthetic and real-world datasets, achieving substantial reanimation improvements on unseen poses: a 17.3 dB PSNR gain over LBS on DNA-Rendering and a 45.6 dB gain over Bag-of-Bones on ActorsHQ, while preserving high rendering fidelity for characters with complex non-rigid dynamics. PartMM generalizes robustly to both Gaussian and mesh representations, excelling in low-data regimes of approximately 1,000 frames, with a 48.4 RMSE improvement over LBS and improvements of more than 20 over GRU- and MLP-based methods. Code will be publicly released at https://cookmaker.cn/gaussianimate/.

cs.CV

RecGen3D: Reconstruction-Guided 3D Generation in a Shared Canonical Space

Sparse-view 3D modeling represents a fundamental tension between reconstruction fidelity and generative plausibility. While feed-forward reconstruction excels in efficiency and input alignment, it often lacks the global priors needed for structural completeness. Conversely, diffusion-based generation provides rich geometric details but struggles with multi-view consistency. We present RecGen3D, a framework that combines these two paradigms into a cooperative system. To overcome inherent conflicts in coordinate spaces, 3D representations, and training objectives, we align both models within a shared canonical space. We employ decoupled cooperative learning, which maintains stable training while enabling seamless collaboration during inference. Specifically, the reconstruction module is adapted to provide canonical geometric anchors, while the diffusion generator leverages latent-augmented conditioning to refine and complete the geometric structure. Experimental results demonstrate that RecGen3D achieves superior fidelity and robustness, outperforming existing methods in creating complete and consistent 3D models from sparse observations.

cs.CV

HGGT: Robust and Flexible 3D Hand Mesh Reconstruction from Uncalibrated Images

Recovering high-fidelity 3D hand geometry from images is a critical task in computer vision, holding significant value for domains such as robotics, animation and VR/AR. Crucially, scalable applications demand both accuracy and deployment flexibility, requiring the ability to leverage massive amounts of unstructured image data from the internet or enable deployment on consumer-grade RGB cameras without complex calibration. However, current methods face a dilemma. While single-view approaches are easy to deploy, they suffer from depth ambiguity and occlusion. Conversely, multi-view systems resolve these uncertainties but typically demand fixed, calibrated setups, limiting their real-world utility. To bridge this gap, we draw inspiration from 3D foundation models that learn explicit geometry directly from visual data. By reformulating hand reconstruction from arbitrary views as a visual-geometry grounded task, we propose a feed-forward architecture that, for the first time in literature, jointly infers 3D hand meshes and camera poses from uncalibrated views. Extensive evaluations show that our approach outperforms state-of-the-art benchmarks and demonstrates strong generalization to uncalibrated, in-the-wild scenarios. Here is the link of our project page: https://lym29.github.io/HGGT/.

cs.CV

FlexAM: Flexible Appearance-Motion Decomposition for Versatile Video Generation Control

Effective and generalizable control in video generation remains a significant challenge. While many methods rely on ambiguous or task-specific signals, we argue that a fundamental disentanglement of "appearance" and "motion" provides a more robust and scalable pathway. We propose FlexAM, a unified framework built upon a novel 3D control signal. This signal represents video dynamics as a point cloud, introducing three key enhancements: multi-frequency positional encoding to distinguish fine-grained motion, depth-aware encoding, and a flexible control signal for balancing precision and generalization. This representation allows FlexAM to effectively disentangle appearance and motion, enabling a wide range of tasks including I2V/V2V editing, camera control, and spatial object editing. Extensive experiments demonstrate that FlexAM achieves superior performance across all evaluated tasks.

cs.CV

RefAny3D: 3D Asset-Referenced Diffusion Models for Image Generation

In this paper, we propose a 3D asset-referenced diffusion model for image generation, exploring how to integrate 3D assets into image diffusion models. Existing reference-based image generation methods leverage large-scale pretrained diffusion models and demonstrate strong capability in generating diverse images conditioned on a single reference image. However, these methods are limited to single-image references and cannot leverage 3D assets, constraining their practical versatility. To address this gap, we present a cross-domain diffusion model with dual-branch perception that leverages multi-view RGB images and point maps of 3D assets to jointly model their colors and canonical-space coordinates, achieving precise consistency between generated images and the 3D references. Our spatially aligned dual-branch generation architecture and domain-decoupled generation mechanism ensure the simultaneous generation of two spatially aligned but content-disentangled outputs, RGB images and point maps, linking 2D image attributes with 3D asset attributes. Experiments show that our approach effectively uses 3D assets as references to produce images consistent with the given assets, opening new possibilities for combining diffusion models with 3D content creation.

cs.CV

TrackingWorld: World-centric Monocular 3D Tracking of Almost All Pixels

Monocular 3D tracking aims to capture the long-term motion of pixels in 3D space from a single monocular video and has witnessed rapid progress in recent years. However, we argue that the existing monocular 3D tracking methods still fall short in separating the camera motion from foreground dynamic motion and cannot densely track newly emerging dynamic subjects in the videos. To address these two limitations, we propose TrackingWorld, a novel pipeline for dense 3D tracking of almost all pixels within a world-centric 3D coordinate system. First, we introduce a tracking upsampler that efficiently lifts the arbitrary sparse 2D tracks into dense 2D tracks. Then, to generalize the current tracking methods to newly emerging objects, we apply the upsampler to all frames and reduce the redundancy of 2D tracks by eliminating the tracks in overlapped regions. Finally, we present an efficient optimization-based framework to back-project dense 2D tracks into world-centric 3D trajectories by estimating the camera poses and the 3D coordinates of these 2D tracks. Extensive evaluations on both synthetic and real-world datasets demonstrate that our system achieves accurate and dense 3D tracking in a world-centric coordinate frame.

cs.CV

LiteVGGT: Boosting Vanilla VGGT via Geometry-aware Cached Token Merging

3D vision foundation models like Visual Geometry Grounded Transformer (VGGT) have advanced greatly in geometric perception. However, it is time-consuming and memory-intensive for long sequences, limiting application to large-scale scenes beyond hundreds of images. To address this, we propose LiteVGGT, achieving up to 10x speedup and substantial memory reduction, enabling efficient processing of 1000-image scenes. We derive two key insights for 3D reconstruction: (1) tokens from local image regions have inherent geometric correlations, leading to high similarity and computational redundancy; (2) token similarity across adjacent network layers remains stable, allowing for reusable merge decisions. Guided by these, we design a simple yet efficient strategy, dubbed geometry-aware cached token merging. We analyze each token's geometric importance, optimizing anchor token selection to better preserve key information for reconstruction. We also cache and reuse merge indices across layers, substantially reducing latency with minimal accuracy impact. This strategy retains VGGT's core performance, enabling efficient fine-tuning and FP8 quantization for further gains. Extensive experiments validate LiteVGGT's effectiveness, scalability, and robustness. Project page: https://garlicba.github.io/LiteVGGT/

cs.CV

Wonder3D++: Cross-domain Diffusion for High-fidelity 3D Generation from a Single Image

In this work, we introduce \textbf{Wonder3D++}, a novel method for efficiently generating high-fidelity textured meshes from single-view images. Recent methods based on Score Distillation Sampling (SDS) have shown the potential to recover 3D geometry from 2D diffusion priors, but they typically suffer from time-consuming per-shape optimization and inconsistent geometry. In contrast, certain works directly produce 3D information via fast network inferences, but their results are often of low quality and lack geometric details. To holistically improve the quality, consistency, and efficiency of single-view reconstruction tasks, we propose a cross-domain diffusion model that generates multi-view normal maps and the corresponding color images. To ensure the consistency of generation, we employ a multi-view cross-domain attention mechanism that facilitates information exchange across views and modalities. Lastly, we introduce a cascaded 3D mesh extraction algorithm that drives high-quality surfaces from the multi-view 2D representations in only about $3$ minute in a coarse-to-fine manner. Our extensive evaluations demonstrate that our method achieves high-quality reconstruction results, robust generalization, and good efficiency compared to prior works. Code available at https://github.com/xxlong0/Wonder3D/tree/Wonder3D_Plus.

cs.CV

Fixed and periodic points of the intersection body operators of lower orders

For the intersection body operator of lower order $I_iK$ of a star body $K$ in $\mathbb{R}^n$, $i\in\{1, 2,\ldots, n-2\}$, we prove that $I_i^2K = cK$ iff $K$ is an origin-symmetric ball, and hence $I_iK = cK$ iff $K$ is an origin-symmetric ball. Combining the recent breakthrough (case $i = n-1$) of Milman, Shabelman and Yehudayoff (Invent. Math., 241 (2025), 509-558), slight modifications of two long-standing questions 8.6 and 8.7 posed by R. Gardner (Page 302, Geometric Tomography, Cambridge University Press, 1995) are completely solved. As applications, we show that for the spherical Radon transform $\mathcal{R}$, a non-negative $\rho\in L^{\infty}(\mathcal{S}^{n-1})$ satisfies $\mathcal{R}(\rho^i) = c\rho$ for some $c>0$ iff $\rho$ is constant. Also, the sharp Busemann intersection type inequalities are established.

math.MG