SearcharxivSearch

arXiv subjects

Yiying Tong

Publications and source records attributed to Yiying Tong.

At least 19 recordsLinked to original sources

Unmasking Face Embeddings: Reading, Rendering and Naming with Foundation Models

Modern face recognition (FR) owes much of its success to deep neural networks that learn to extract compact identity embeddings from face images. These models are typically trained for identity discrimination, producing embeddings that are highly effective for biometric matching but largely opaque to semantic interpretation. In contrast, foundation models, pretrained on broad visual or vision--language tasks, provide rich interfaces for describing, retrieving, generating, and organizing visual content. This contrast raises a natural question: what capabilities become available when face embeddings from domain-specific FR models are made interoperable with foundation models? Building on recent work on embedding compatibility across models, we use simple pre-computed linear transformations, estimated from paired embeddings alone, to connect existing FR models with off-the-shelf foundation models. Once aligned with a foundation model, a face embedding can be 'unmasked' in multiple ways, without training or modifying either model: it can be read in natural language, enabling free-form text queries over a gallery of FR embeddings; rendered into a face image that recovers a person's appearance, using an unmodified diffusion decoder; and converted to a name, enabling identification even in the absence of an enrolled face gallery. In effect, one linear transformation turns an identity embedding into a rich embedding for web-scale foundation models. This interoperability exposes face embeddings as semantically and visually rich biometric representations, with direct implications for interpretability, retrieval, reconstruction, and template security.

cs.CV

Interactive Stroke-based Neural SDF Sculpting

Recent advances in implicit neural representations have made them a popular choice for modeling 3D geometry. However, directly editing these representations presents challenges due to the complex relationship between model weights and surface geometry, as well as the slow optimization required to update neural fields. Among various editing tools, sculpting stands out as a valuable operation for the graphics and modeling community. While traditional mesh-based tools like ZBrush enable intuitive edits, a comparable high-performance toolkit for sculpting neural SDFs is currently lacking. We introduce a framework that enables interactive surface sculpting directly on neural implicit representations with optimized performance. Unlike previous methods, which are limited to spot edits, our approach allows users to perform stroke-based modifications on the fly, ensuring intuitive shape manipulation without switching representations. By employing tubular neighborhoods to sample strokes and customizable brush profiles, we achieve smooth deformations along user-defined curves, providing intuitive control over the sculpting process. Our method demonstrates that versatile edits can be achieved while preserving the smooth nature of implicit representations, all without compromising interactive performance.

cs.GR

Compatibility of Face Embeddings Across Deep Neural Networks

Automated face recognition has made rapid strides over the past decade due to the unprecedented rise of deep neural network (DNN) models that can be trained for domain-specific tasks. At the same time, large foundation models that are pretrained on broad vision or vision-language tasks have shown impressive generalization across diverse domains, including biometrics. This raises an important question: Do different DNN models---both domain-specific and foundation models---encode facial identity in similar ways, despite being trained on different datasets, loss functions, and architectures? In this regard, we directly analyze the geometric structure of embedding spaces imputed by different DNN models. Treating embeddings of face images as point clouds, we study whether simple affine transformations can align face representations of one model with another. Our findings reveal substantial cross-model compatibility: low-capacity linear mappings substantially improve cross-model face recognition over unaligned baselines for both identification and verification, including across foundation models never trained for face recognition. Alignment patterns generalize across datasets and vary systematically across model families, indicating representational convergence in facial identity encoding. These findings reframe independently trained templates as transferable rather than revocable, with implications for interoperability, ensemble design, and biometric template security.

cs.CV

Weighted Hodge Laplacians on Manifolds with Boundary

The spectrum of the Hodge Laplacian on differential manifolds encodes rich topological and geometric information and thus provides a powerful tool for analyzing data on manifolds. However, the classical unweighted formulation is restricted in its ability to study data with varying local features. To address this limitation, we propose a weighted Hodge Laplacian framework for manifolds with boundary, both in theory and in computation, by incorporating a weight function on the manifold. Under appropriate boundary conditions, we formulate the corresponding weighted de Rham-Hodge theory, in which the kernel of the weighted Hodge Laplacian coincides with the weighted harmonic space, and remains isomorphic to the de Rham cohomology of the underlying manifold. The harmonic spectrum of the weighted Hodge Laplacian captures the global topological information, while its non-harmonic spectrum encodes the local geometric property induced by the weight. The proposed framework therefore enables the study of topological and geometric features of data on manifolds across varying weights, and in addition, allows local structure to be highlighted by choosing weights that emphasize regions of interest. We demonstrate the effectiveness of the proposed method through proof-of-principle experiments in protein flexibility analysis, and the results show its promise.

math.DG

A Discrete Exterior Calculus of Bundle-valued Forms

The discretization of Cartan's exterior calculus of differential forms has been fruitful in a variety of theoretical and practical endeavors: from computational electromagnetics to the development of Finite-Element Exterior Calculus, the development of structure-preserving numerical tools satisfying exact discrete equivalents to Stokes' theorem or the de Rham complex for the exterior derivative have found numerous applications in computational physics. However, there has been a dearth of effort in establishing a more general discrete calculus, this time for differential forms with values in vector bundles over a combinatorial manifold equipped with a connection. In this work, we propose a discretization of the exterior covariant derivative of bundle-valued differential forms. We demonstrate that our discrete operator mimics its continuous counterpart, satisfies the Bianchi identities on simplicial cells, and contrary to previous attempts at its discretization, ensures numerical convergence to its exact evaluation with mesh refinement under mild assumptions.

math.DG

Refined Geometry-guided Head Avatar Reconstruction from Monocular RGB Video

High-fidelity reconstruction of head avatars from monocular videos is highly desirable for virtual human applications, but it remains a challenge in the fields of computer graphics and computer vision. In this paper, we propose a two-phase head avatar reconstruction network that incorporates a refined 3D mesh representation. Our approach, in contrast to existing methods that rely on coarse template-based 3D representations derived from 3DMM, aims to learn a refined mesh representation suitable for a NeRF that captures complex facial nuances. In the first phase, we train 3DMM-stored NeRF with an initial mesh to utilize geometric priors and integrate observations across frames using a consistent set of latent codes. In the second phase, we leverage a novel mesh refinement procedure based on an SDF constructed from the density field of the initial NeRF. To mitigate the typical noise in the NeRF density field without compromising the features of the 3DMM, we employ Laplace smoothing on the displacement field. Subsequently, we apply a second-phase training with these refined meshes, directing the learning process of the network towards capturing intricate facial details. Our experiments demonstrate that our method further enhances the NeRF rendering based on the initial mesh and achieves performance superior to state-of-the-art methods in reconstructing high-fidelity head avatars with such input.

cs.GR

Manifold Topological Deep Learning for Biomedical Data

Recently, topological deep learning (TDL), which integrates algebraic topology with deep neural networks, has achieved tremendous success in processing point-cloud data, emerging as a promising paradigm in data science. However, TDL has not been developed for data on differentiable manifolds, including images, due to the challenges posed by differential topology. We address this challenge by introducing manifold topological deep learning (MTDL) for the first time. To highlight the power of Hodge theory rooted in differential topology, we consider a simple convolutional neural network (CNN) in MTDL. In this novel framework, original images are represented as smooth manifolds with vector fields that are decomposed into three orthogonal components based on Hodge theory. These components are then concatenated to form an input image for the CNN architecture. The performance of MTDL is evaluated using the MedMNIST v2 benchmark database, which comprises 717,287 biomedical images from eleven 2D and six 3D datasets. MTDL significantly outperforms other competing methods, extending TDL to a wide range of data on smooth manifolds.

eess.IV

Persistent de Rham-Hodge Laplacians in Eulerian representation for manifold topological learning

Recently, topological data analysis has become a trending topic in data science and engineering. However, the key technique of topological data analysis, i.e., persistent homology, is defined on point cloud data, which does not work directly for data on manifolds. Although earlier evolutionary de Rham-Hodge theory deals with data on manifolds, it is inconvenient for machine learning applications because of the numerical inconsistency caused by remeshing the involving manifolds in the Lagrangian representation. In this work, we introduce persistent de Rham-Hodge Laplacian, or persistent Hodge Laplacian (PHL) as an abbreviation, for manifold topological learning. Our PHLs are constructed in the Eulerian representation via structure-persevering Cartesian grids, avoiding the numerical inconsistency over the multiscale manifolds. To facilitate the manifold topological learning, we propose a persistent Hodge Laplacian learning algorithm for data on manifolds or volumetric data. As a proof-of-principle application of the proposed manifold topological learning model, we consider the prediction of protein-ligand binding affinities with two benchmark datasets. Our numerical experiments highlight the power and promise of the proposed method.

math.DG

Topology-preserving Hodge Decomposition in the Eulerian Representation

The Hodge decomposition is a fundamental result in differential geometry and algebraic topology, particularly in the study of differential forms on a Riemannian manifold. Despite extensive research in the past few decades, topology-preserving Hodge decomposition of scalar and vector fields on manifolds with boundaries in the Eulerian representation remains a challenge due to the implicit incorporation of appropriate topology-preserving boundary conditions. In this work, we introduce a comprehensive 5-component topology-preserving Hodge decomposition that unifies normal and tangential components in the Cartesian representation. Implicit representations of planar and volumetric regions defined by level-set functions have been developed. Numerical experiments on various objects, including single-cell RNA velocity, validate the effectiveness of our approach, confirming the expected rigorous $L^2$-orthogonality and the accurate cohomology.

math.DG

INFAMOUS-NeRF: ImproviNg FAce MOdeling Using Semantically-Aligned Hypernetworks with Neural Radiance Fields

We propose INFAMOUS-NeRF, an implicit morphable face model that introduces hypernetworks to NeRF to improve the representation power in the presence of many training subjects. At the same time, INFAMOUS-NeRF resolves the classic hypernetwork tradeoff of representation power and editability by learning semantically-aligned latent spaces despite the subject-specific models, all without requiring a large pretrained model. INFAMOUS-NeRF further introduces a novel constraint to improve NeRF rendering along the face boundary. Our constraint can leverage photometric surface rendering and multi-view supervision to guide surface color prediction and improve rendering near the surface. Finally, we introduce a novel, loss-guided adaptive sampling method for more effective NeRF training by reducing the sampling redundancy. We show quantitatively and qualitatively that our method achieves higher representation power than prior face modeling methods in both controlled and in-the-wild settings. Code and models will be released upon publication.

cs.CV

Combinatorial and Hodge Laplacians: Similarity and Difference

As key subjects in spectral geometry and combinatorial graph theory respectively, the (continuous) Hodge Laplacian and the combinatorial Laplacian share similarities in revealing the topological dimension and geometric shape of data and in their realization of diffusion and minimization of harmonic measures. It is believed that they also both associate with vector calculus, through the gradient, curl, and divergence, as argued in the popular usage of "Hodge Laplacians on graphs" in the literature. Nevertheless, these Laplacians are intrinsically different in their domains of definitions and applicability to specific data formats, hindering any in-depth comparison of the two approaches. To facilitate the comparison and bridge the gap between the combinatorial Laplacian and Hodge Laplacian for the discretization of continuous manifolds with boundary, we further introduce Boundary-Induced Graph (BIG) Laplacians using tools from Discrete Exterior Calculus (DEC). BIG Laplacians are defined on discrete domains with appropriate boundary conditions to characterize the topology and shape of data. The similarities and differences of the combinatorial Laplacian, BIG Laplacian, and Hodge Laplacian are then examined. Through an Eulerian representation of 3D domains as level-set functions on regular grids, we show experimentally the conditions for the convergence of BIG Laplacian eigenvalues to those of the Hodge Laplacian for elementary shapes.

math.DG

Face Relighting with Geometrically Consistent Shadows

Most face relighting methods are able to handle diffuse shadows, but struggle to handle hard shadows, such as those cast by the nose. Methods that propose techniques for handling hard shadows often do not produce geometrically consistent shadows since they do not directly leverage the estimated face geometry while synthesizing them. We propose a novel differentiable algorithm for synthesizing hard shadows based on ray tracing, which we incorporate into training our face relighting model. Our proposed algorithm directly utilizes the estimated face geometry to synthesize geometrically consistent hard shadows. We demonstrate through quantitative and qualitative experiments on Multi-PIE and FFHQ that our method produces more geometrically consistent shadows than previous face relighting methods while also achieving state-of-the-art face relighting performance under directional lighting. In addition, we demonstrate that our differentiable hard shadow modeling improves the quality of the estimated face geometry over diffuse shading models.

cs.CV

Towards High Fidelity Face Relighting with Realistic Shadows

Existing face relighting methods often struggle with two problems: maintaining the local facial details of the subject and accurately removing and synthesizing shadows in the relit image, especially hard shadows. We propose a novel deep face relighting method that addresses both problems. Our method learns to predict the ratio (quotient) image between a source image and the target image with the desired lighting, allowing us to relight the image while maintaining the local facial details. During training, our model also learns to accurately modify shadows by using estimated shadow masks to emphasize on the high-contrast shadow borders. Furthermore, we introduce a method to use the shadow mask to estimate the ambient light intensity in an image, and are thus able to leverage multiple datasets during training with different global lighting intensities. With quantitative and qualitative evaluations on the Multi-PIE and FFHQ datasets, we demonstrate that our proposed method faithfully maintains the local facial details of the subject and can accurately handle hard shadows while achieving state-of-the-art face relighting performance.

cs.CV

HERMES: Persistent spectral graph software

Persistent homology (PH) is one of the most popular tools in topological data analysis (TDA), while graph theory has had a significant impact on data science. Our earlier work introduced the persistent spectral graph (PSG) theory as a unified multiscale paradigm to encompass TDA and geometric analysis. In PSG theory, families of persistent Laplacians (PLs) corresponding to various topological dimensions are constructed via a filtration to sample a given dataset at multiple scales. The harmonic spectra from the null spaces of PLs offer the same topological invariants, namely persistent Betti numbers, at various dimensions as those provided by PH, while the non-harmonic spectra of PLs give rise to additional geometric analysis of the shape of the data. In this work, we develop an open-source software package, called highly efficient robust multidimensional evolutionary spectra (HERMES), to enable broad applications of PSGs in science, engineering, and technology. To ensure the reliability and robustness of HERMES, we have validated the software with simple geometric shapes and complex datasets from three-dimensional (3D) protein structures. We found that the smallest non-zero eigenvalues are very sensitive to data abnormality.

math.AT

Evolutionary de Rham-Hodge method

The de Rham-Hodge theory is a landmark of the 20$^\text{th}$ Century's mathematics and has had a great impact on mathematics, physics, computer science, and engineering. This work introduces an evolutionary de Rham-Hodge method to provide a unified paradigm for the multiscale geometric and topological analysis of evolving manifolds constructed from a filtration, which induces a family of evolutionary de Rham complexes. While the present method can be easily applied to close manifolds, the emphasis is given to more challenging compact manifolds with 2-manifold boundaries, which require appropriate analysis and treatment of boundary conditions on differential forms to maintain proper topological properties. Three sets of unique evolutionary Hodge Laplacian operators are proposed to generate three sets of topology-preserving singular spectra, for which the multiplicities of zero eigenvalues correspond to exactly the persistent Betti numbers of dimensions 0, 1, and 2. Additionally, three sets of non-zero eigenvalues further reveal both topological persistence and geometric progression during the manifold evolution. Extensive numerical experiments are carried out via the discrete exterior calculus to demonstrate the utility and usefulness of the proposed method for data representation and shape analysis.

math.DG

The de Rham-Hodge analysis and modeling of biomolecules

Recent years have witnessed a trend that advanced mathematical tools, such as algebraic topology, differential geometry, graph theory, and partial differential equations, have been developed for describing biological macromolecules. These tools have considerably strengthened our ability to understand the molecular mechanism of macromolecular function, dynamics and transport from their structures. However, currently, there is no unified mathematical theory to analyze, describe and characterize biological macromolecular geometry, topology, flexibility and natural mode at a variety of scales. We introduce the de Rham-Hodge theory, a landmark of 20th Century's mathematics, as a unified paradigm for analyzing biological macromolecular geometry and algebraic topology, for predicting macromolecular flexibility, and for modeling macromolecular natural modes at a variety of scales. In this paradigm, macromolecular geometric characteristic and topological invariants are revealed by de Rham-Hodge spectral analysis. By using the Helmholtz-Hodge decomposition, every macromolecular vector field is split into orthogonal divergence-free, curl-free, and harmonic components with a distinct physical interpretation. We organize the eigenvalues and eigenvectors of the 0-form Laplace-de Rham operator into one of the most accurate protein B-factor predictors. By combining the 1-form Laplace-de Rham operator and the Helfrich-type curvature energy, we predict the natural modes of both X-ray protein structures and cryo-EM maps. We construct accurate and efficient three-dimensional discrete exterior calculus algorithms for the aforementioned modeling and analysis of biological macromolecules. Using extensive experiments, we validate that the proposed paradigm is one of the most versatile and powerful tools for biological macromolecular studies.

q-bio.BM

Parallel Transport Unfolding: A Connection-based Manifold Learning Approach

Manifold learning offers nonlinear dimensionality reduction of high-dimensional datasets. In this paper, we bring geometry processing to bear on manifold learning by introducing a new approach based on metric connection for generating a quasi-isometric, low-dimensional mapping from a sparse and irregular sampling of an arbitrary manifold embedded in a high-dimensional space. Geodesic distances of discrete paths over the input pointset are evaluated through "parallel transport unfolding" (PTU) to offer robustness to poor sampling and arbitrary topology. Our new geometric procedure exhibits the same strong resilience to noise as one of the staples of manifold learning, the Isomap algorithm, as it also exploits all pairwise geodesic distances to compute a low-dimensional embedding. While Isomap is limited to geodesically-convex sampled domains, parallel transport unfolding does not suffer from this crippling limitation, resulting in an improved robustness to irregularity and voids in the sampling. Moreover, it involves only simple linear algebra, significantly improves the accuracy of all pairwise geodesic distance approximations, and has the same computational complexity as Isomap. Finally, we show that our connection-based distance estimation can be used for faster variants of Isomap such as L-Isomap.

cs.LG

Divide-and-Conquer Strategy for Large-Scale Eulerian Solvent Excluded Surface

Motivation: Surface generation and visualization are some of the most important tasks in biomolecular modeling and computation. Eulerian solvent excluded surface (ESES) software provides analytical solvent excluded surface (SES) in the Cartesian grid, which is necessary for simulating many biomolecular electrostatic and ion channel models. However, large biomolecules and/or fine grid resolutions give rise to excessively large memory requirements in ESES construction. We introduce an out-of-core and parallel algorithm to improve the ESES software. Results: The present approach drastically improves the spatial and temporal efficiency of ESES. The memory footprint and time complexity are analyzed and empirically verified through extensive tests with a large collection of biomolecule examples. Our results show that our algorithm can successfully reduce memory footprint through a straightforward divide-and-conquer strategy to perform the calculation of arbitrarily large proteins on a typical commodity personal computer. On multi-core computers or clusters, our algorithm can reduce the execution time by parallelizing most of the calculation as disjoint subproblems. Various comparisons with the state-of-the-art Cartesian grid based SES calculation were done to validate the present method and show the improved efficiency. This approach makes ESES a robust software for the construction of analytical solvent excluded surfaces. Availability and implementation: http://weilab.math.msu.edu/ESES.

q-bio.QM