SearcharxivSearch

arXiv subjects

Frantisek Duris

Publications and source records attributed to Frantisek Duris.

5 recordsLinked to original sources

SnakeLines: integrated set of computational pipelines for sequencing reads

Background: With the rapid growth of massively parallel sequencing technologies, still more laboratories are utilizing sequenced DNA fragments for genomic analyses. Interpretation of sequencing data is, however, strongly dependent on bioinformatics processing, which is often too demanding for clinicians and researchers without a computational background. Another problem represents the reproducibility of computational analyses across separated computational centers with inconsistent versions of installed libraries and bioinformatics tools. Results: We propose an easily extensible set of computational pipelines, called SnakeLines, for processing sequencing reads; including mapping, assembly, variant calling, viral identification, transcriptomics, metagenomics, and methylation analysis. Individual steps of an analysis, along with methods and their parameters can be readily modified in a single configuration file. Provided pipelines are embedded in virtual environments that ensure isolation of required resources from the host operating system, rapid deployment, and reproducibility of analysis across different Unix-based platforms. Conclusion: SnakeLines is a powerful framework for the automation of bioinformatics analyses, with emphasis on a simple set-up, modifications, extensibility, and reproducibility. Keywords: Computational pipeline, framework, massively parallel sequencing, reproducibility, virtual environment

q-bio.GN

Innovative method for reducing uninformative calls in non-invasive prenatal testing

Non-invasive prenatal testing or NIPT is currently among the top researched topic in obstetric care. While the performance of the current state-of-the-art NIPT solutions achieve high sensitivity and specificity, they still struggle with a considerable number of samples that cannot be concluded with certainty. Such uninformative results are often subject to repeated blood sampling and re-analysis, usually after two weeks, and this period may cause a stress to the future mothers as well as increase the overall cost of the test. We propose a supplementary method to traditional z-scores to reduce the number of such uninformative calls. The method is based on a novel analysis of the length profile of circulating cell free DNA which compares the change in such profiles when random-based and length-based elimination of some fragments is performed. The proposed method is not as accurate as the standard z-score; however, our results suggest that combination of these two independent methods correctly resolves a substantial portion of healthy samples with an uninformative result. Additionally, we discuss how the proposed method can be used to identify maternal aberrations, thus reducing the risk of false positive and false negative calls. Keywords: Next-generation sequencing, Cell-free DNA, Uninformative result, Method, Trisomy, Prenatal testing

q-bio.GN

On Kummer's test of convergence and its relation to basic comparison tests

Testing convergence of infinite series is an important part of mathematics. A very basic test of convergence is to upper-bound a given series with a known series, term by term. In $19^{th}$ century, Kummer proposed a test of convergence for any positive series based on finding a suitable positive sequence $\{p_n\}$ and a suitable real constant $c$. It can be easily shown that by choosing appropriate sequence $\{p_n\}$, the Kummer's test yields other tests like Raabe's, Gauss' or Bertrand's as its special cases. In 1995, Samelson noted that there is another interesting relation between Kummer's test and basic comparison tests, particularly, that one can transform the sequence $\{p_n\}$ into a convergent bounding series, and he sketched a simple proof of this statement. In this paper, we fill the missing formal proof, although using a different approach, and we show how to construct a bounding series from the sequence $\{p_n\}$ and vice versa.

math.HO

Arguments for the Effectiveness of Human Problem Solving

The question of how humans solve problem has been addressed extensively. However, the direct study of the effectiveness of this process seems to be overlooked. In this paper, we address the issue of the effectiveness of human problem solving: we analyze where this effectiveness comes from and what cognitive mechanisms or heuristics are involved. Our results are based on the optimal probabilistic problem solving strategy that appeared in Solomonoff paper on general problem solving system. We provide arguments that a certain set of cognitive mechanisms or heuristics drive human problem solving in the similar manner as the optimal Solomonoff strategy. The results presented in this paper can serve both cognitive psychology in better understanding of human problem solving processes as well as artificial intelligence in designing more human-like agents.

cs.AI

Error bounds on the probabilistically optimal problem solving strategy

We consider a simple optimal probabilistic problem solving strategy that searches through potential solution candidates in a specific order. We are interested in what impact has interchanging the order of two solution candidates with respect to this optimal strategy on the problem solving effectivity (i.e., the solution candidates examined as well as time spent before solving the problem). Such interchange can happen in the applications with only partial information available. We derive bounds on these errors in general as well as in three special systems in which we impose some restrictions on the solution candidates.

math.OC