SearcharxivSearch

arXiv subjects

Georg Zimmermann

Publications and source records attributed to Georg Zimmermann.

9 recordsLinked to original sources

A Recommender System Based on Binary Matrix Representations for Cognitive Disorders

Diagnosing cognitive (mental health) disorders is a delicate and complex task. Identifying the next most informative symptoms to assess, in order to distinguish between possible disorders, presents an additional challenge. This process requires comprehensive knowledge of diagnostic criteria and symptom overlap across disorders, making it difficult to navigate based on symptoms alone. This research aims to develop a recommender system for cognitive disorder diagnosis using binary matrix representations. The core algorithm utilizes a binary matrix of disorders and their symptom combinations. It filters through the rows and columns based on the patient's current symptoms to identify potential disorders and recommend the most informative next symptoms to examine. A prototype of the recommender system was implemented in Python. Using synthetic test and some real-life data, the system successfully identified plausible disorders from an initial symptom set and recommended further symptoms to refine the diagnosis. It also provided additional context on the symptom-disorder relationships. Although this is a prototype, the recommender system shows potential as a clinical support tool. A fully-developed application of this recommender system may assist mental health professionals in identifying relevant disorders more efficiently and guiding symptom-specific follow-up investigations to improve diagnostic accuracy.

cs.IR

Profile Generators: A Link between the Narrative and the Binary Matrix Representation

Mental health disorders, particularly cognitive disorders defined by deficits in cognitive abilities, are described in detail in the DSM-5, which includes definitions and examples of signs and symptoms. A simplified, machine-actionable representation was developed to assess the similarity and separability of these disorders, but it is not suited for the most complex cases. Generating or applying a full binary matrix for similarity calculations is infeasible due to the vast number of symptom combinations. This research develops an alternative representation that links the narrative form of the DSM-5 with the binary matrix representation and enables automated generation of valid symptom combinations. Using a strict pre-defined format of lists, sets, and numbers with slight variations, complex diagnostic pathways involving numerous symptom combinations can be represented. This format, called the symptom profile generator (or simply generator), provides a readable, adaptable, and comprehensive alternative to a binary matrix while enabling easy generation of symptom combinations (profiles). Cognitive disorders, which typically involve multiple diagnostic criteria with several symptoms, can thus be expressed as lists of generators. Representing several psychotic disorders in generator form and generating all symptom combinations showed that matrix representations of complex disorders become too large to manage. The MPCS (maximum pairwise cosine similarity) algorithm cannot handle matrices of this size, prompting the development of a profile reduction method using targeted generator manipulation to find specific MPCS values between disorders. The generators allow easier creation of binary representations for large matrices and make it possible to calculate specific MPCS cases between complex disorders through conditional generators.

cs.LG

Multivariate and Multiple Contrast Testing in General Covariate-adjusted Factorial Designs

Evaluating intervention effects on multiple outcomes is a central research goal in a wide range of quantitative sciences. It is thereby common to compare interventions among each other and with a control across several, potentially highly correlated, outcome variables. In this context, researchers are interested in identifying effects at both, the global level (across all outcome variables) and the local level (for specific variables). At the same time, potential confounding must be accounted for. This leads to the need for powerful multiple contrast testing procedures (MCTPs) capable of handling multivariate outcomes and covariates. Given this background, we propose an extension of MCTPs within a semiparametric MANCOVA framework that allows applicability beyond multivariate normality, homoscedasticity, or non-singular covariance structures. We illustrate our approach by analysing multivariate psychological intervention data, evaluating joint physiological and psychological constructs such as heart rate variability.

stat.ME

Resampling NANCOVA: Nonparametric Analysis of Covariance in Small Samples

Analysis of covariance is a crucial method for improving precision of statistical tests for factor effects in randomized experiments. However, existing solutions suffer from one or more of the following limitations: (i) they are not suitable for ordinal data (as endpoints or explanatory variables); (ii) they require semiparametric model assumptions; (iii) they are inapplicable to small data scenarios due to often poor type-I error control; or (iv) they provide only approximate testing procedures and (asymptotically) exact test are missing. In this paper, we investigate a resampling approach to the NANCOVA framework, which is a fully nonparametric model based on relative effects that allows for an arbitrary number of covariates and groups, where both outcome variable (endpoint) and covariates can be metric or ordinal. Thereby, we evaluate novel NANCOVA tests and a nonparametric competitor test without covariate adjustment in extensive simulations. Unlike approximate tests in the NANCOVA framework, our resampling version showed good performance in small sample scenarios and maintained the nominal type-I error well. Resampling NANCOVA also provided consistently high power: up to 26% higher than the test without covariate adjustment in a small sample scenario with 4 groups and two covariates. Moreover, we prove that resampling NANCOVA provides an asymptotically exact testing procedure, which makes it the first one in the NANCOVA framework. In summary, resampling NANCOVA can be considered a viable tool for analysis of covariance that overcomes issues (i) - (iv).

stat.ME

Choice of the hypothesis matrix for using the Wald-type-statistic

A widely used formulation for null hypotheses in the analysis of multivariate $d$-dimensional data is $\mathcal{H}_0: \boldsymbol{H} \boldsymbol{\theta} =\boldsymbol{y}$ with $\boldsymbol{H}$ $\in\mathbb{R}^{m\times d}$, $\boldsymbol{\theta}$ $\in \mathbb{R}^d$ and $\boldsymbol{y}\in\mathbb{R}^m$, where $m\leq d$. Here the unknown parameter vector $\boldsymbol{\theta}$ can, for example, be the expectation vector $\boldsymbol{\mu}$, a vector $\boldsymbol{\beta} $ containing regression coefficients or a quantile vector $\boldsymbol{q}$. Also, the vector of nonparametric relative effects $\boldsymbol{p}$ or an upper triangular vectorized covariance matrix $\textbf{v}$ are useful choices. However, even without multiplying the hypothesis with a scalar $\gamma\neq 0$, there is a multitude of possibilities to formulate the same null hypothesis with different hypothesis matrices $\boldsymbol{H}$ and corresponding vectors $\boldsymbol{y}$. Although it is a well-known fact that in case of $\boldsymbol{y}=\boldsymbol{0}$ there exists a unique projection matrix $\boldsymbol{P}$ with $\boldsymbol{H}\boldsymbol{\theta}=\boldsymbol{0}\Leftrightarrow \boldsymbol{P}\boldsymbol{\theta}=\boldsymbol{0}$, for $\boldsymbol{y}\neq \boldsymbol{0}$ such a projection matrix does not necessarily exist. Moreover, since such hypotheses are often investigated using a quadratic form as the test statistic, the corresponding projection matrices often contain zero rows; so, they are not even effective from a computational aspect. In this manuscript, we show that for the Wald-type-statistic (WTS), which is one of the most frequently used quadratic forms, the choice of the concrete hypothesis matrix does not affect the test decision. Moreover, some simulations are conducted to investigate the possible influence of the hypothesis matrix on the computation time.

math.ST

Small-sample performance and underlying assumptions of a bootstrap-based inference method for a general analysis of covariance model with possibly heteroskedastic and nonnormal errors

It is well known that the standard F test is severely affected by heteroskedasticity in unbalanced analysis of covariance (ANCOVA) models. Currently available potential remedies for such a scenario are based on heteroskedasticity-consistent covariance matrix estimation (HCCME). However, the HCCME approach tends to be liberal in small samples. Therefore, in the present manuscript, we propose a combination of HCCME and a wild bootstrap technique, with the aim of improving the small-sample performance. We precisely state a set of assumptions for the general ANCOVA model and discuss their practical interpretation in detail, since this issue may have been somewhat neglected in applied research so far. We prove that these assumptions are sufficient to ensure the asymptotic validity of the combined HCCME-wild bootstrap ANCOVA. The results of our simulation study indicate that our proposed test remedies the problems of the ANCOVA F test and its heteroskedasticity-consistent alternatives in small to moderate sample size scenarios. Our test only requires very mild conditions, thus being applicable in a broad range of real-life settings, as illustrated by the detailed discussion of a dataset from preclinical research on spinal cord injury. Our proposed method is ready-to-use and allows for valid hypothesis testing in frequently encountered settings (e.g., comparing group means while adjusting for baseline measurements in a randomized controlled clinical trial).

stat.ME

Multivariate analysis of covariance when standard assumptions are violated

In applied research, it is often sensible to account for one or several covariates when testing for differences between multivariate means of several groups. However, the "classical" parametric multivariate analysis of covariance (MANCOVA) tests (e.g., Wilks' Lambda) are based on quite restrictive assumptions (homoscedasticity and normality of the errors), which might be difficult to justify in small sample size settings. Furthermore, existing potential remedies (e.g., heteroskedasticity-robust approaches) become inappropriate in cases where the covariance matrices are singular. Nevertheless, such scenarios are frequently encountered in the life sciences and other fields, when for example, in the context of standardized assessments, a summary performance measure as well as its corresponding subscales are analyzed. In the present manuscript, we consider a general MANCOVA model, allowing for potentially heteroskedastic and even singular covariance matrices as well as non-normal errors. We combine heteroskedasticity-consistent covariance matrix estimation methods with our proposed modified MANCOVA ANOVA-type statistic (MANCATS) and apply two different bootstrap approaches. We provide the proofs of the asymptotic validity of the respective testing procedures as well as the results from an extensive simulation study, which indicate that especially the parametric bootstrap version of the MANCATS outperforms its competitors in most scenarios, both in terms of type I error rates and power. These considerations are further illustrated and substantiated by examining real-life data from standardized achievement tests.

stat.ME

Sample size calculation and blinded recalculation for analysis of covariance models with multiple random covariates

When testing for superiority in a parallel-group setting with a continuous outcome, adjusting for covariates (e.g., baseline measurements) is usually recommended, in order to reduce bias and increase power. For this purpose, the analysis of covariance (ANCOVA) is frequently used, and recently, several exact and approximate sample size calculation procedures have been proposed. However, in case of multiple covariates, the planning might pose some practical challenges and surprising pitfalls, which have not been recognized so far. Moreover, since a considerable number of parameters have to be specified in advance, the risk of making erroneous initial assumptions, leading to substantially over- or underpowered studies, is increased. Therefore, we propose a method, which allows for re-estimating the sample size at a prespecified time point during the course of the trial. Extensive simulations for a broad range of settings, including unbalanced designs, confirm that the proposed method provides reliable results in many practically relevant situations. An advantage of the reassessment procedure is that it does not require unblinding of the data. In order to facilitate the application of the proposed method, we provide some R code and discuss a real-life data example.

stat.ME

Scalar Subdivision Schemes and Box Splines

We study scalar $d$-variate subdivision schemes, with dilation matrix 2I, satisfying the sum rules of order $k$. Using the results of M\"oller and Sauer, stated for general expanding dilation matrices, we characterize the structure of the mask symbols of such schemes by showing that they must be linear combinations of shifted box spline generators of some quotient polynomial ideal. The directions of the corresponding box splines are columns of certain unimodular matrices. The quotient ideal is determined by the given order of the sum rules or, equivalently, by the order of the zero conditions. The results presented in this paper open a way to a systematic study of subdivision schemes, since box spline subdivisions turn out to be the building blocks of any reasonable multivariate subdivision scheme. As in the univariate case, the characterization we give is the proper way of matching the smoothness of the box spline building blocks with the order of polynomial reproduction of the corresponding subdivision scheme. However, due to the interaction of the building blocks, convergence and smoothness properties may change, if several convergent schemes are combined.

math.NA