Searcharxiv⌕ Search

arXiv · 2610.09288

FaceKit: a Toolkit for Interpretable Facial Phenotyping, Synthetic Image Generation and Privacy Analysis in Rare Diseases

Abstract

Many rare genetic diseases are associated with recognizable craniofacial features. However, traditional approaches for describing facial morphology rely largely on qualitative clinical observation and free-text descriptions, which are often subjective, non-standardized, and difficult to reproduce across observers and institutions. Although the Human Phenotype Ontology (HPO) provides controlled terms for describing facial features, these terms are typically categorical rather than quantitative and may vary depending on examiner experience and interpretation. Here, we present FaceKit, a computational framework for quantitative facial phenotyping from frontal facial photographs. FaceKit extracts standardized measurements of facial landmarks and derived 120 morphological features, then reports feature-level z-scores representing deviation from population reference distributions. The reference distributions are built from the FairFace dataset spanning diverse ancestral groups. We evaluated FaceKit on a curated subset of the GestaltMatcher Database covering 50 rare-disease cohorts. In addition to quantitative facial analysis, FaceKit includes synthetic facial image generation to support rare disease model development and data augmentation. We also performed privacy evaluation to assess whether synthetic images reveal identifiable information from real patient photographs and could compromise patient privacy. Across disease case studies, FaceKit-derived quantitative measurements captured known facial features associated with rare genetic disorders and provided objective support for clinical phenotyping. Together, these results establish FaceKit as a useful tool for quantitative phenotyping, and has the potential to improve rare disease diagnosis, support genotype-phenotype studies, and enable more reproducible clinical characterization across diverse patient populations.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Hongzhuo Chen, Zhanliang Wang, Florent Pollet, Mian Umair Ahsan, Joshua Bie, Tzung-Chien Hsieh, Peter Krawitz, Cong Liu, Wendy K Chung, Chunhua Weng, Gamze Gürsoy, Kai Wang. 2026-10-07. FaceKit: a Toolkit for Interpretable Facial Phenotyping, Synthetic Image Generation and Privacy Analysis in Rare Diseases. https://arxiv.org/abs/2610.09288

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Cross-species representation learning aligns mouse and human neural dynamics and tracks clinical drug efficacy

Preclinical models poorly predict human drug efficacy, particularly in neurological disorders. Neural activity offers a uniquely rich source of translational information because it captures high-dimensional variation in nervous-system function that can be measured in both animals and humans. However, its high dimensionality makes it difficult to distinguish conserved disease-related features from variation arising from species, recording modality and experimental context. Here, we test whether shared neural dynamics can be identified directly from electrophysiology data by learning representations organized by biological state rather than species. We develop a dual-rule contrastive learning framework that aligns corresponding mouse and human states while preserving separation between distinct phenotypes. This framework recovered conserved sensory-response structure across species and, in epilepsy, resolved distinct relationships between three mouse models and heterogeneous human patient populations. When treated animals were projected into a frozen cross-species representation, drug-induced movement towards the human-aligned healthy state retrospectively tracked known clinical efficacy across ten model-drug combinations including a disease-specific detrimental effect. The framework also identified shared disease-associated neural dynamics between Fmr1-knockout mice and human 16p11.2 copy-number variant carriers despite differences in genetic aetiology and recording modality. Together, these findings show the potential of cross-species neural representation learning to map heterogeneous human disease onto experimentally tractable preclinical states and assess whether interventions restore human-relevant circuit function.

q-bio.QM↗

Elucidating the Space of Enzymatic Reaction: A Unified Benchmark and Pretrained Model

Existing reaction models primarily learn molecular transformations, whereas enzy- matic reactions depend jointly on molecular structure and catalytic function. We formulate this problem as learning an enzymatic reaction space linking reactants, products, and Enzyme Commission (EC) annotations. To characterize this space, we introduce VenusRX-Bench, a unified benchmark for forward reaction prediction, single-step retrosynthesis, and EC-number prediction. VenusRX-Bench integrates reactions from multiple biochemical databases with standardized curation, leakage- controlled splits, and consistent evaluation. Benchmarking representative chemical and enzymatic models reveals a clear chemical-to-enzymatic domain gap, driven by limited domain data, catalytic-context dependency, and the difficulty of modeling large biomolecular structures. To bridge this gap, we develop VenusRX, a unified T5-style sequence-to-sequence model for enzymatic reactions. VenusRX jointly learns forward prediction, ret- rosynthesis, and reaction reconstruction, with two-stage training on millions of template-expanded reactions followed by real biochemical reactions. In addition, optional EC conditioning incorporates catalytic context, while Molecule Library- Constrained Decoding improves the generation of complex biomolecules. Across benchmark tasks and challenging generalization splits, VenusRX achieves the best or competitive performance on most evaluated settings over representative chem- ical and enzymatic baselines. Moreover, EC information consistently improves reaction prediction, while learned reaction representations support accurate EC prediction, revealing a bidirectional relationship between reaction structure and catalytic function. Together, VenusRX-Bench and VenusRX provide a unified framework for elucidating and modeling enzymatic reaction space

q-bio.QM↗

Frame-invariant topological representations of trabecular bone microarchitecture for strength prediction

Directional topological representations of trabecular bone should retain interpretable structural information without depending on an arbitrary transverse coordinate frame. We develop a frame-invariant directional filtration and compare its strength prediction with signed distance persistent homology and conventional morphometry. Twenty-four human trabecular cores were analyzed using persistence images, directional Betti tensors, and nested ridge regression over 13 validation groups. The directional construction combines cone occupancy, directional covariance, and principal axis degeneracy. Equal bone removal experiments and a pair closely matched in morphometry were used to examine whether topological differences tracked changes in simulated elastic stiffness. The original combined persistence image model had a root mean squared error (RMSE) of 1.952 MPa, compared with 2.001 MPa for morphometry; the paired difference was $-0.049$ MPa with a 95% bootstrap interval of $[-0.491,0.393]$ MPa. Exploratory signed distance persistent homology in dimension zero gave an RMSE of 1.639 MPa. The post hoc frame-invariant directional dimension zero model gave 1.610 MPa, compared with 1.877 MPa for dimension one and 1.630 MPa for a harmonized signed distance dimension zero model. The paired RMSE difference between the invariant and signed distance models was $-0.021$ MPa with a 95% interval of $[-0.234,0.210]$ MPa. Localized removal reduced stiffness more than diffuse removal in 19 of 24 cores despite producing smaller $H_0$ persistence image changes. Frame invariance removes the dependence of directional topology on an arbitrary transverse coordinate frame. Connected component representations warrant external evaluation for strength prediction, while the mechanical experiments limit their interpretation as scalar stiffness surrogates.

q-bio.QM↗