Searcharxiv⌕ Search

arXiv subjects

Shanika Amarasoma

Publications and source records attributed to Shanika Amarasoma.

3 recordsLinked to original sources

Truecell reproduces R Seurat's single-cell analysis outputs natively in the Python ecosystem

Single-cell RNA-sequencing analyses predominantly use either Seurat (R) or Scanpy (Python), with the choice often driven by programming language preference. However, by default, these tools produce divergent variable features, neighbour graphs, clusters, and marker genes. Laboratories requiring Python, which supports most deep-learning, foundation-model, and agent tools, must re-implement Seurat analyses in a framework that does not replicate Seurat's results. Truecell was developed as a Python implementation of the Seurat interface, preserving function names, arguments, default settings, and the object model to facilitate seamless transfer of analyses. In eighteen paired end-to-end evaluations against R Seurat, deterministic outputs matched to floating-point precision, and fold-change order was identical across all nine differential expression tests. In a three-arm benchmark on three datasets, with fixed user parameters and 20 seeds per tool, Truecell more closely reproduced Seurat's clustering than a Seurat-configured Scanpy in all 12 combinations of dataset and resolution settings. Truecell's marker genes and enriched pathways were also closer to Seurat's than Scanpy's were, and pseudobulk DESeq2 reproduced Seurat's gene lists with a Jaccard index ranging from 0.95 to 1.00. This agreement reflects fidelity to Seurat rather than biological correctness.

q-bio.GN↗

Big Data Architecture for Large Organizations

The exponential growth of big data has transformed how large organisations leverage information to drive innovation, optimise processes, and maintain competitive advantages. However, managing and extracting insights from vast, heterogeneous data sources requires a scalable, secure, and well-integrated big data architecture. This paper proposes a comprehensive big data framework that aligns with organisational objectives while ensuring flexibility, scalability, and governance. The architecture encompasses multiple layers, including data ingestion, transformation, storage, analytics, machine learning, and security, incorporating emerging technologies such as Generative AI (GenAI) and low-code machine learning. Cloud-based implementations across Google Cloud, AWS, and Microsoft Azure are analysed, highlighting their tools and capabilities. Additionally, this study explores advancements in big data architecture, including AI-driven automation, data mesh, and Data Ocean paradigms. By establishing a structured, adaptable framework, this research provides a foundational blueprint for large organisations to harness big data as a strategic asset effectively.

cs.DC↗

An Integrated Genomics Workflow Tool: Simulating Reads, Evaluating Read Alignments, and Optimizing Variant Calling Algorithms

Next-generation sequencing (NGS) is a pivotal technique in genome sequencing due to its high throughput, rapid results, cost-effectiveness, and enhanced accuracy. Its significance extends across various domains, playing a crucial role in identifying genetic variations and exploring genomic complexity. NGS finds applications in diverse fields such as clinical genomics, comparative genomics, functional genomics, and metagenomics, contributing substantially to advancements in research, medicine, and scientific disciplines. Within the sphere of genomics data science, the execution of read simulation, mapping, and variant calling holds paramount importance for obtaining precise and dependable results. Given the plethora of tools available for these purposes, each employing distinct methodologies and options, a nuanced understanding of their intricacies becomes imperative for optimization. This research, situated at the intersection of data science and genomics, involves a meticulous assessment of various tools, elucidating their individual strengths and weaknesses through rigorous experimentation and analysis. This comprehensive evaluation has enabled the researchers to pinpoint the most accurate tools, reinforcing the alignment between the established workflow and the demonstrated efficacy of specific tools in the context of genomics data analysis. To meet these requirements, "VarFind", an open-source and freely accessible pipeline tool designed to automate the entire process has been introduced (VarFind GitHub repository: https://github.com/shanikawm/varfinder)

q-bio.GN↗