SearcharxivSearch

arXiv subjects

Hannes Matuschek

Publications and source records attributed to Hannes Matuschek.

4 recordsLinked to original sources

On the Ambiguity of Interaction and Nonlinear Main Effects in a Regime of Dependent Covariates

The analysis of large experimental datasets frequently reveals significant interactions that are difficult to interpret within the theoretical framework guiding the research. Some of these interactions actually arise from the presence of unspecified nonlinear main effects and statistically dependent covariates in the statistical model. Importantly, such nonlinear main effects may be compatible (or, at least, not incompatible) with the current theoretical framework. In the present literature this issue has only been studied in terms of correlated (linearly dependent) covariates. Here we generalize to nonlinear main effects (i.e., main effects of arbitrary shape) and dependent covariates. We propose a novel nonparametric method to test for ambiguous interactions where present parametric methods fail. We illustrate the method with a set of simulations and with reanalyses (a) of effects of parental education on their children's educational expectations and (b) of effects of word properties on fixation locations during reading of natural sentences, specifically of effects of length and morphological complexity of the word to be fixated next. The resolution of such ambiguities facilitates theoretical progress.

stat.AP

Balancing Type I Error and Power in Linear Mixed Models

Linear mixed-effects models have increasingly replaced mixed-model analyses of variance for statistical inference in factorial psycholinguistic experiments. Although LMMs have many advantages over ANOVA, like ANOVAs, setting them up for data analysis also requires some care. One simple option, when numerically possible, is to fit the full variance-covariance structure of random effects (the maximal model; Barr et al. 2013), presumably to keep Type I error down to the nominal alpha in the presence of random effects. Although it is true that fitting a model with only random intercepts may lead to higher Type I error, fitting a maximal model also has a cost: it can lead to a significant loss of power. We demonstrate this with simulations and suggest that for typical psychological and psycholinguistic data, higher power is achieved without inflating Type I error rate if a model selection criterion is used to select a random effect structure that is supported by the data.

stat.AP

Fraud detection with statistics: A comment on "Evidential Value in ANOVA-Regression Results in Scientific Integrity Studies" (Klaassen, 2015)

Klaassen in (Klaassen 2015) proposed a method for the detection of data manipulation given the means and standard deviations for the cells of a oneway ANOVA design. This comment critically reviews this method. In addition, inspired by this analysis, an alternative approach to test sample correlations over several experiments is derived. The results are in close agreement with the initial analysis reported by an anonymous whistlelblower. Importantly, the statistic requires several similar experiments; a test for correlations between 3 sample means based on a single experiment must be considered as unreliable.

stat.ME

Computation of biochemical pathway fluctuations beyond the linear noise approximation using iNA

The linear noise approximation is commonly used to obtain intrinsic noise statistics for biochemical networks. These estimates are accurate for networks with large numbers of molecules. However it is well known that many biochemical networks are characterized by at least one species with a small number of molecules. We here describe version 0.3 of the software intrinsic Noise Analyzer (iNA) which allows for accurate computation of noise statistics over wide ranges of molecule numbers. This is achieved by calculating the next order corrections to the linear noise approximation's estimates of variance and covariance of concentration fluctuations. The efficiency of the methods is significantly improved by automated just-in-time compilation using the LLVM framework leading to a fluctuation analysis which typically outperforms that obtained by means of exact stochastic simulations. iNA is hence particularly well suited for the needs of the computational biology community.

q-bio.QM