SearcharxivSearch

arXiv subjects

Assya Trofimov

Publications and source records attributed to Assya Trofimov.

3 recordsLinked to original sources

T-cell repertoire response in individuals with post-acute sequelae of COVID-19

T-cells are central to SARS-CoV-2 clearance and immunological memory, yet their contribution to the persistence of post-acute sequelae of COVID-19 (PASC) remains poorly understood. The immunological features that distinguish individuals who develop PASC from those who recover fully are unresolved, in part due to the phenotypic heterogeneity of the condition and the likely multiplicity of its underlying mechanisms. Here, we profiled longitudinal bulk TCR$\beta$ repertoires from 120 individuals in the INCOV cohort--71 with PASC and 49 without--sampled at two to three time points spanning the acute and post-acute phases of infection. Using robust statistical modeling of repertoire composition and clonal dynamics, we found that global statistics such as V, J gene usage and CDR3 length do not differ between groups, but that locally enriched sequence motifs and differentially dynamic clones reveal distinct T-cell signatures associated with PASC status. Clones contracting following the peak of the acute response were significantly enriched for SARS-CoV-2 specificity in both groups. Interestingly, Influenza A-specific TCRs were disproportionately enriched among contracting clones in PASC{$^+$} repertoires, implicating viral co-infection as a potential contributor to early disease severity and, possibly, PASC pathogenesis. Rare public TCR clones were markedly enriched for SARS-CoV-2 specificity, with PASC{$^+$} individuals harboring a modestly but significantly higher proportion than PASC{$^-$} individuals. Together, we identified over 1,000 candidate TCR$\beta$ receptors potentially discriminating PASC{$^+$} from PASC{$^-$} immune responses, opening a path toward the identification of disease-relevant T-cell specificities and the development of T-cell-based immunological biomarkers for long COVID.

q-bio.GN

Accounting for Variance in Machine Learning Benchmarks

Strong empirical evidence that one machine-learning algorithm A outperforms another one B ideally calls for multiple trials optimizing the learning pipeline over sources of variation such as data sampling, data augmentation, parameter initialization, and hyperparameters choices. This is prohibitively expensive, and corners are cut to reach conclusions. We model the whole benchmarking process, revealing that variance due to data sampling, parameter initialization and hyperparameter choice impact markedly the results. We analyze the predominant comparison methods used today in the light of this variance. We show a counter-intuitive result that adding more sources of variation to an imperfect estimator approaches better the ideal estimator at a 51 times reduction in compute cost. Building on these results, we study the error rate of detecting improvements, on five different deep-learning tasks/architectures. This study leads us to propose recommendations for performance comparisons.

cs.LG

Towards the Latent Transcriptome

In this work we propose a method to compute continuous embeddings for kmers from raw RNA-seq data, without the need for alignment to a reference genome. The approach uses an RNN to transform kmers of the RNA-seq reads into a 2 dimensional representation that is used to predict abundance of each kmer. We report that our model captures information of both DNA sequence similarity as well as DNA sequence abundance in the embedding latent space, that we call the Latent Transcriptome. We confirm the quality of these vectors by comparing them to known gene sub-structures and report that the latent space recovers exon information from raw RNA-Seq data from acute myeloid leukemia patients. Furthermore we show that this latent space allows the detection of genomic abnormalities such as translocations as well as patient-specific mutations, making this representation space both useful for visualization as well as analysis.

q-bio.GN