SearcharxivSearch

arXiv subjects

Andreas Tiffeau-Mayer

Publications and source records attributed to Andreas Tiffeau-Mayer.

8 recordsLinked to original sources

External Clustering Validation by the Homogeneity-Parsimony Trade-off

Scalar metrics are often used to evaluate clusterings against known classes, but they can obscure a fundamental trade-off: clusterings should be informative about class labels while avoiding unnecessary fragmentation. Here we describe normalized scores of cluster homogeneity and parsimony that quantify this trade-off. These scores build on the information bottleneck principle, modified to not reward lossy compression. We show by example and mathematical proof that our definitions of these scores have the intuitive property of varying monotonically under cluster refinement in contrast to related proposals. Extending the information-theoretic framework beyond Shannon entropies, we furthermore derive set-matching and pair-based counterparts of the homogeneity and parsimony scores. These unify commonly used evaluation criteria and show that, in the pair-based setting, the homogeneity-parsimony trade-off recovers the receiver operating characteristic of binary classifiers. We demonstrate the framework's utility for feature selection and algorithm comparison, illustrating how considering scores jointly can clarify clustering operating points and identify Pareto-optimal solutions.

cs.LG

Fluctuating environments are sufficient to drive substantial variability in species abundance across locations

Species growing in environments that change in time and space will vary in their abundance across locations, even in the absence of persistent location preferences. Here we quantify this non-equilibrium effect by studying a minimal model of a spatially compartmentalised community with time-averaged-neutral competition but location-dependent environmental fluctuations. We analytically derive distributions of two-point inequality, defined as the log-ratio of a species' abundance across a pair of locations. We characterise how the balance of relaxation via migration and fluctuation strength determine the bulk and extreme value statistics of these distributions in the two-patch and infinite-patch cases. We demonstrate the existence of a noise-induced transition to bimodal inequality, which depends on the correlation timescale of the environmental fluctuations. Finally, we discuss the evolutionary benefit of finite migration rates in environments with temporal correlations.

q-bio.PE

Unifying Theories in High-Dimensional Biology: Approaches, Challenges and Opportunities

Across biological subdisciplines, the last decade has seen an explosion of high-dimensional datasets, including datasets for cells, species, immune systems, neurons and behaviour. At the ICTS workshop 'Unifying Theories in High-Dimensional Biophysics' we discussed whether this high dimensionality poses a challenge or opportunity for describing, understanding and predicting biological systems theoretically. We discussed methods, models and frameworks that can help with addressing empirical observations based on these high-dimensional datasets. We summarize the challenges and opportunities that emerged in discussions according to individual participants below.

physics.bio-ph

Data-driven Discovery of Biophysical T Cell Receptor Co-specificity Rules

The biophysical interactions between the T cell receptor (TCR) and its ligands determine the specificity of the cellular immune response. However, the immense diversity of receptors and ligands has made it challenging to discover generalizable rules across the distinct binding affinity landscapes created by different ligands. Here, we present an optimization framework for discovering biophysical rules that predict whether TCRs share specificity to a ligand. Applying this framework to TCRs associated with a collection of SARS-CoV-2 peptides we systematically characterize how co-specificity depends on the type and position of amino-acid differences between receptors. We also demonstrate that the inferred rules generalize to ligands highly dissimilar to any seen during training. Our analysis reveals that matching of steric properties between substituted amino acids is more important for receptor co-specificity than the hydrophobic properties that prominently determine evolutionary substitutability. Our analysis also quantifies the substantial importance of positions not in direct contact with the peptide for specificity. These findings highlight the potential for data-driven approaches to uncover the molecular mechanisms underpinning the specificity of adaptive immune responses.

q-bio.BM

Contrastive learning of T cell receptor representations

Computational prediction of the interaction of T cell receptors (TCRs) and their ligands is a grand challenge in immunology. Despite advances in high-throughput assays, specificity-labelled TCR data remains sparse. In other domains, the pre-training of language models on unlabelled data has been successfully used to address data bottlenecks. However, it is unclear how to best pre-train protein language models for TCR specificity prediction. Here we introduce a TCR language model called SCEPTR (Simple Contrastive Embedding of the Primary sequence of T cell Receptors), capable of data-efficient transfer learning. Through our model, we introduce a novel pre-training strategy combining autocontrastive learning and masked-language modelling, which enables SCEPTR to achieve its state-of-the-art performance. In contrast, existing protein language models and a variant of SCEPTR pre-trained without autocontrastive learning are outperformed by sequence alignment-based methods. We anticipate that contrastive learning will be a useful paradigm to decode the rules of TCR specificity.

q-bio.BM

Limits on Inferring T-cell Specificity from Partial Information

A key challenge in molecular biology is to decipher the mapping of protein sequence to function. To perform this mapping requires the identification of sequence features most informative about function. Here, we quantify the amount of information (in bits) that T-cell receptor (TCR) sequence features provide about antigen specificity. We identify informative features by their degree of conservation among antigen-specific receptors relative to null expectations. We find that TCR specificity synergistically depends on the hypervariable regions of both receptor chains, with a degree of synergy that strongly depends on the ligand. Using a coincidence-based approach to measuring information enables us to directly bound the accuracy with which TCR specificity can be predicted from partial matches to reference sequences. We anticipate that our statistical framework will be of use for developing machine learning models for TCR specificity prediction and for optimizing TCRs for cell therapies. The proposed coincidence-based information measures might find further applications in bounding the performance of pairwise classifiers in other fields.

q-bio.BM

Lightning-fast adaptive immune receptor similarity search by symmetric deletion lookup

An individual's adaptive immune receptor (AIR) repertoire records immune history due to the exquisite antigen specificity of AIRs. Reading this record requires computational approaches for inferring receptor function from sequence, as the diversity of possible receptor-antigen pairs vastly outstrips experimental knowledge. Identification of AIRs with similar sequence and thus putatively similar function is a common performance bottleneck in these approaches. Here, we benchmark the time complexity of five different algorithmic approaches to radius-based search for Levenshtein neighbors. We show that a symmetric deletion lookup approach, originally proposed for spell-checking, is particularly scalable. We then introduce XTNeighbor, a variant of this algorithm that can be massively parallelized on GPUs. For one million input sequences, XTNeighbor identifies all sequence neighbors that differ by up to two edits in seconds on commodity hardware, orders of magnitude faster than existing approaches. We also demonstrate how symmetric deletion lookup can speed up search with more complex sequence-similarity metrics such as TCRdist. Our contribution is poised to greatly speed up existing analysis pipelines and enable processing of large-scale immunosequencing data without downsampling.

q-bio.QM

Unbiased estimation of sampling variance for Simpson's diversity index

Quantification of measurement uncertainty is crucial for robust scientific inference, yet accurate estimates of this uncertainty remain elusive for ecological measures of diversity. Here, we address this longstanding challenge by deriving a closed-form unbiased estimator for the sampling variance of Simpson's diversity index. In numerical tests the estimator consistently outperforms existing approaches, particularly for applications in which species richness exceeds sample size. We apply the estimator to quantify biodiversity loss in marine ecosystems and to demonstrate ligand-dependent contributions of T cell receptor chains to specificity, illustrating its versatility across fields. The novel estimator provides researchers with a reliable method for comparing diversity between samples, essential for quantifying biodiversity trends and making informed conservation decisions.

q-bio.PE