SearcharxivSearch

arXiv subjects

Marc Comino-Trinidad

Publications and source records attributed to Marc Comino-Trinidad.

4 recordsLinked to original sources

GNOCHI: Generative Neural mOdel for Close Human-Human Interactions

Creating realistic 3D human-human interactions in virtual environments is challenging due to the high degrees of freedom in the human body and the need for physically accurate poses that do not collide with each other. Traditional methods for human-human interaction are based on motion tracking or 3D body reconstruction, but lack generative capabilities. Recent generative methods enable the synthesis of individual or interacting motions via text or image input, but generally fall short in modeling close interactions. This paper introduces a novel generative model for close 3D human-human interactions using a conditional variational autoencoder (cVAE), which generates poses for one human conditioned on the pose of another, allowing for controlled and diverse interaction synthesis. To train our model, we address two underlying long-standing challenges in the field of human-human interaction: data scarcity, for which we propose an automated supervised data augmentation strategy that generates synthetic yet realistic interaction poses; and collision awareness in generative approaches, for which we propose a self-supervised loss based on a collision resolution technique using volumetric proxies to ensure physically correct interactions. We extensively evaluate the capabilities of our model, and demonstrate a wide variety of plausible and physically correct interactions, not possible to generate with current state-of-the-art methods.

cs.CV

Coarse-to-Real: Generative Rendering for Populated Dynamic Scenes

Traditional rendering pipelines rely on complex assets, accurate materials and lighting, and substantial computational resources to produce realistic imagery, yet they still face challenges in scalability and realism for populated dynamic scenes. We present C2R (Coarse-to-Real), a generative rendering framework that synthesizes real-style urban crowd videos from coarse 3D simulations. Our approach uses coarse 3D renderings to explicitly control scene layout, camera motion, and human trajectories, while a learned neural renderer generates realistic appearance, lighting, and fine-scale dynamics guided by text prompts. To overcome the lack of paired training data between coarse simulations and real videos, we adopt a two-stage synthetic-real domain-hedging strategy that first learns a strong generative prior from large-scale real footage, and then introduces controllability by using a small amount of paired synthetic coarse-to-fine data to anchor shared implicit spatio-temporal features across domains. The resulting system supports coarse-to-fine control, generalizes across diverse CG and game inputs, and produces temporally consistent, controllable, and realistic urban scene videos from minimal 3D input. We will release the model and project webpage at https://gonzalognogales.github.io/coarse2real/.

cs.CV

Revisiting Poisson-disk Subsampling for Massive Point Cloud Decimation

Scanning devices often produce point clouds exhibiting highly uneven distributions of point samples across the surfaces being captured. Different point cloud subsampling techniques have been proposed to generate more evenly distributed samples. Poisson-disk sampling approaches assign each sample a cost value so that subsampling reduces to sorting the samples by cost and then removing the desired ratio of samples with the highest cost. Unfortunately, these approaches compute the sample cost using pairwise distances of the points within a constant search radius, which is very costly for massive point clouds with uneven densities. In this paper, we revisit Poisson-disk sampling for point clouds. Instead of optimizing for equal densities, we propose to maximize the distance to the closest point, which is equivalent to estimating the local point density as a value inversely proportional to this distance. This algorithm can be efficiently implemented using k nearest-neighbors searches. Besides a kd-tree, our algorithm also uses a voxelization to speed up the searches required to compute per-sample costs. We propose a new strategy to minimize cost updates that is amenable for out-of-core operation. We demonstrate the benefits of our approach in terms of performance, scalability, and output quality. We also discuss extensions based on adding orientation-based and color-based terms to the cost function.

cs.GR

SMPLitex: A Generative Model and Dataset for 3D Human Texture Estimation from Single Image

We propose SMPLitex, a method for estimating and manipulating the complete 3D appearance of humans captured from a single image. SMPLitex builds upon the recently proposed generative models for 2D images, and extends their use to the 3D domain through pixel-to-surface correspondences computed on the input image. To this end, we first train a generative model for complete 3D human appearance, and then fit it into the input image by conditioning the generative model to the visible parts of the subject. Furthermore, we propose a new dataset of high-quality human textures built by sampling SMPLitex conditioned on subject descriptions and images. We quantitatively and qualitatively evaluate our method in 3 publicly available datasets, demonstrating that SMPLitex significantly outperforms existing methods for human texture estimation while allowing for a wider variety of tasks such as editing, synthesis, and manipulation

cs.CV