SearcharxivSearch

arXiv subjects

Tomas Fitzgerald

Publications and source records attributed to Tomas Fitzgerald.

2 recordsLinked to original sources

FlexLMM: a Nextflow linear mixed model framework for GWAS

Summary: Linear mixed models are a commonly used statistical approach in genome-wide association studies when population structure is present. However, naive permutations to empirically estimate the null distribution of a statistic of interest are not appropriate in the presence of population structure, because the samples are not exchangeable with each other. For this reason we developed FlexLMM, a Nextflow pipeline that runs linear mixed models while allowing for flexibility in the definition of the exact statistical model to be used. FlexLMM can also be used to set a significance threshold via permutations, thanks to a two-step process where the population structure is first regressed out, and only then are the permutations performed. We envision this pipeline will be particularly useful for researchers working on multi-parental crosses among inbred lines of model organisms or farm animals and plants. Availability and implementation: The source code and documentation for the FlexLMM is available at https://github.com/birneylab/flexlmm.

q-bio.GN

A Simple Reduction for Full-Permuted Pattern Matching Problems on Multi-Track Strings

In this paper we study a variant of string pattern matching which deals with tuples of strings known as \textit{multi-track strings}. Multi-track strings are a generalisation of strings (or \textit{single-track strings}) that have primarily found uses in problems related to searching multiple genomes and music information retrieval. A multi-track string $\mathcal{T} = (t_1, t_2, t_3, \ldots , t_N)$ of length $n$ and track count $N$ is a multi-set of $N$ strings of length $n$ with characters drawn from a common alphabet of size $\sigma_U$. Given two multi-track strings $\mathcal{T} = (t_1, t_2, t_3, \ldots , t_N)$ and $ \mathcal{P} = (p_1, p_2, p_3, \ldots , p_N)$ of length $n$ and track count $N$, there is a \textit{full-permuted-match} between $\mathcal{P}$ and $\mathcal{T}$ if $t_{r_i} = p_i$ for all $i \in \{1,2,3,\ldots N \}$ and some permutation $(r_1, r_2, r_3\ldots,r_N)$ of $(1, 2, 3,\ldots,N)$, we denote this $\mathcal{P}\asymp\mathcal{T}$. Efficient algorithms for some full-permuted-match problems on multi-track strings have recently been presented. In this paper we show a reduction from a multi-track string of length $n$ and track count $N$ with alphabet size $\sigma_U$, to a single-track string of length $2n-1$ with alphabet size $\sigma_U^N$. Through this reduction we allow any string algorithm to be used on multi-track string problems using $\asymp$ as the match relation. For polynomial time algorithms on single-track strings of length $n$ there is a multiplicative penalty of not more than $\mathcal{O}(N)$-time for the same algorithm on mt-strings of length $n$ and track count $N$.

cs.DS