SearcharxivSearch

arXiv subjects

Katja Ickstadt

Publications and source records attributed to Katja Ickstadt.

13 recordsLinked to original sources

Early Prediction of Student Performance Using Bayesian Updating with Informative Priors Across Cohorts

Early identification of at risk students in higher education depends on predictive models that maintain accuracy across successive cohorts -- a requirement that single-cohort modeling approaches fail to meet. This study evaluates Bayesian updating with informative priors from a previous cohort to improve cross-cohort prediction robustness using digital trace data. We fit weekly Bayesian linear, logistic, and ordinal regression models with either uninformative default priors or informative priors derived from posterior distributions of a preceding cohort. Models were applied to six weekly self-regulated learning (SRL)-aligned engagement indicators from two consecutive cohorts of students in a blended first-year mathematics course (N1 = 307; N2 = 323). Outcomes were exam points, final grades, and a binary at risk indicator. The models were evaluated weekly based on accuracy, sensitivity, and RMSE. In the source cohort, performance was already substantial by week 6. In the target cohort, informative priors improved early classification: Logistic models with priors reduced misclassification by 22% and false negatives by 38% in week 3 relative to the uninformative default. Ordinal models with priors similarly showed the strongest improvements in early weeks, reducing misclassification by 42% in week 2 and reaching an accuracy of .77 by week 4. Linear models showed little benefit from prior information. These findings demonstrate that Bayesian updating is a viable method for improving early classification performance across cohorts, with gains concentrated in the early weeks of the semester when current-cohort data are scarce.

stat.AP

Scalable Learning of Multivariate Distributions via Coresets

Efficient and scalable non-parametric or semi-parametric regression analysis and density estimation are of crucial importance to the fields of statistics and machine learning. However, available methods are limited in their ability to handle large-scale data. We address this issue by developing a novel coreset construction for multivariate conditional transformation models (MCTMs) to enhance their scalability and training efficiency. To the best of our knowledge, these are the first coresets for semi-parametric distributional models. Our approach yields substantial data reduction via importance sampling. It ensures with high probability that the log-likelihood remains within multiplicative error bounds of $(1\pm\varepsilon)$ and thereby maintains statistical model accuracy. Compared to conventional full-parametric models, where coresets have been incorporated before, our semi-parametric approach exhibits enhanced adaptability, particularly in scenarios where complex distributions and non-linear relationships are present, but not fully understood. To address numerical problems associated with normalizing logarithmic terms, we follow a geometric approximation based on the convex hull of input data. This ensures feasible, stable, and accurate inference in scenarios involving large amounts of data. Numerical experiments demonstrate substantially improved computational efficiency when handling large and complex datasets, thus laying the foundation for a broad range of applications within the statistics and machine learning communities.

cs.LG

Guidance for Addressing Individual Time Effects in Cohort Stepped Wedge Cluster Randomized Trials: A Simulation Study

Background: Stepped wedge cluster randomized trials (SW-CRTs) involve sequential measurements within clusters over time. Initially, all clusters start in the control condition before crossing over to the intervention on a staggered schedule. In cohort designs, secular trends, cluster-level changes, and individual-level changes (e.g., aging) must be considered. Methods: We performed a Monte Carlo simulation to analyze the influence of different time effects on the estimation of the intervention effect in cohort SW-CRTs. We compared four linear mixed models with different adjustment strategies, all including random intercepts for clustering and repeated measurements. We recorded the estimated fixed intervention effects and their corresponding model-based standard errors, derived from models both without and with cluster-robust variance estimators (CRVEs). Results: Models incorporating fixed categorical time effects, a fixed intervention effect, and two random intercepts provided unbiased estimates of the intervention effect in both closed and open cohort SW-CRTs. Fixed categorical time effects captured temporal cohort changes, while random individual effects accounted for baseline differences. However, these differences can cause large, non-normally distributed random individual effects. CRVEs provide reliable standard errors for the intervention effect, controlling the Type I error rate. Conclusions: Our simulation study is the first to assess individual-level changes over time in cohort SW-CRTs. Linear mixed models incorporating fixed categorical time effects and random cluster and individual effects yield unbiased intervention effect estimates. However, cluster-robust variance estimation is necessary when time-varying independent variables exhibit nonlinear effects. We recommend always using CRVEs.

stat.ME

MCBench: A Benchmark Suite for Monte Carlo Sampling Algorithms

In this paper, we present MCBench, a benchmark suite designed to assess the quality of Monte Carlo (MC) samples. The benchmark suite enables quantitative comparisons of samples by applying different metrics, including basic statistical metrics as well as more complex measures, in particular the sliced Wasserstein distance and the maximum mean discrepancy. We apply these metrics to point clouds of both independent and identically distributed (IID) samples and correlated samples generated by MC techniques, such as Markov Chain Monte Carlo or Nested Sampling. Through repeated comparisons, we evaluate test statistics of the metrics, allowing to evaluate the quality of the MC sampling algorithms. Our benchmark suite offers a variety of target functions with different complexities and dimensionalities, providing a versatile platform for testing the capabilities of sampling algorithms. Implemented as a Julia package, MCBench enables users to easily select test cases and metrics from the provided collections, which can be extended as needed. Users can run external sampling algorithms of their choice on these test functions and input the resulting samples to obtain detailed metrics that quantify the quality of their samples compared to the IID samples generated by our package. This approach yields clear, quantitative measures of sampling quality and allows for informed decisions about the effectiveness of different sampling methods. By offering such a standardized method for evaluating MC sampling quality, our benchmark suite provides researchers and practitioners from many scientific fields, such as the natural sciences, engineering, or the social sciences with a valuable tool for developing, validating and refining sampling algorithms.

stat.CO

Bias through time-varying covariates in the analysis of cohort stepped wedge trials: a simulation study

In stepped wedge cluster randomized trials (SW-CRTs), observations collected under the control condition are, on average, from an earlier time than observations collected under the intervention condition. In a cohort design, participants are followed up throughout the study, so correlations between measurements within a participant are dependent of the timing in which the observations are made. Therefore, changes in participants' characteristics over time must be taken into account when estimating intervention effects. For example, participants' age progresses, which may impact the outcome over the study period. Motivated by an SW-CRT of a geriatric care intervention to improve quality of life, we conducted a simulation study to compare model formulations analysing data from an SW-CRT under different scenarios in which time was related to the covariates and the outcome. The aim was to find a model specification that produces reliable estimates of the intervention effect. Six linear mixed effects (LME) models with different specification of fixed effects were fitted. Across 1000 simulations per parameter combination, we computed mean and standard error of the estimated intervention effects. We found that LME models with fixed categorical time effects additional to the fixed intervention effect and two random effects used to account for clustering (within-cluster correlation) and multiple measurements on participants (within-individual correlation) seem to produce unbiased estimates of the intervention effect even if time-varying confounders or their functional influence on outcome were unknown or unmeasured and if secular time trends occurred. Therefore, including (time-varying) covariates describing the study cohort seems to be avoidable.

stat.ME

Multilevel Conditional Autoregressive models for longitudinal and spatially referenced epidemiological data

The classical multilevel model fails to capture the proximity effect in epidemiological studies, where subjects are nested within geographical units. Multilevel Conditional Autoregressive models are alternatives to help explain the spatial effect better. They have been developed for cross-sectional studies but not for longitudinal studies so far. This paper has two goals. Firstly, it further develops the multilevel (growth) models for longitudinal data by adding existing area level random effect terms with CAR prior specification, whose structure is changing over time. We name these models MLM tCARs for longitudinal data. We compare the developed MLM tCARs to the classical multilevel growth model via simulation studies in common spatial data situations. The results indicate the better performance of the MLM tCARs, to retrieve the true regression coefficients and with better fit in general. Secondly, this paper provides a comprehensive decision tree for analysing data in epidemiological studies with spatially nested structure: we also consider the Multilevel Conditional Autoregressive models for cross-sectional studies (MLM CARs). We compare three models (for cross-sectional studies) via simulation studies: the classical multilevel model, the multilevel CAR model and the Restricted CAR model that accounts for spatial confounding. The MLM CARs, particularly the Restricted CAR show better results. We apply the models comparatively on the analysis of the association between greenness and depressive symptoms in the longitudinal Heinz Nixdorf Recall Study. The results show negative association between greenness and depression and a decreasing linear individual time trend for all models. We observe very weak spatial variation and moderate temporal autocorrelation.

stat.ME

Bivariate Analysis of Birth Weight and Gestational Age Depending on Environmental Exposures: Bayesian Distributional Regression with Copulas

In this article, we analyze perinatal data with birth weight (BW) as primarily interesting response variable. Gestational age (GA) is usually an important covariate and included in polynomial form. However, in opposition to this univariate regression, bivariate modeling of BW and GA is recommended to distinguish effects on each, on both, and between them. Rather than a parametric bivariate distribution, we apply conditional copula regression, where marginal distributions of BW and GA (not necessarily of the same form) can be estimated independently, and where the dependence structure is modeled conditional on the covariates separately from these marginals. In the resulting distributional regression models, all parameters of the two marginals and the copula parameter are observation-specific. Besides biometric and obstetric information, data on drinking water contamination and maternal smoking are included as environmental covariates. While the Gaussian distribution is suitable for BW, the skewed GA data are better modeled by the three-parametric Dagum distribution. The Clayton copula performs better than the Gumbel and the symmetric Gaussian copula, indicating lower tail dependence (stronger dependence when both variables are low), although this non-linear dependence between BW and GA is surprisingly weak and only influenced by Cesarean section. A non-linear trend of BW on GA is detected by a classical univariate model that is polynomial with respect to the effect of GA. Linear effects on BW mean are similar in both models, while our distributional copula regression also reveals covariates' effects on all other parameters.

stat.ME

Cross-Leverage Scores for Selecting Subsets of Explanatory Variables

In a standard regression problem, we have a set of explanatory variables whose effect on some response vector is modeled. For wide binary data, such as genetic marker data, we often have two limitations. First, we have more parameters than observations. Second, main effects are not the main focus; instead the primary aim is to uncover interactions between the binary variables that effect the response. Methods such as logic regression are able to find combinations of the explanatory variables that capture higher-order relationships in the response. However, the number of explanatory variables these methods can handle is highly limited. To address these two limitations we need to reduce the number of variables prior to computationally demanding analyses. In this paper, we demonstrate the usefulness of using so-called cross-leverage scores as a means of sampling subsets of explanatory variables while retaining the valuable interactions.

stat.ME

Is there a role for statistics in artificial intelligence?

The research on and application of artificial intelligence (AI) has triggered a comprehensive scientific, economic, social and political discussion. Here we argue that statistics, as an interdisciplinary scientific field, plays a substantial role both for the theoretical and practical understanding of AI and for its future development. Statistics might even be considered a core element of AI. With its specialist knowledge of data evaluation, starting with the precise formulation of the research question and passing through a study design stage on to analysis and interpretation of the results, statistics is a natural partner for other disciplines in teaching, research and practice. This paper aims at contributing to the current discussion by highlighting the relevance of statistical methodology in the context of AI development. In particular, we discuss contributions of statistics to the field of artificial intelligence concerning methodological development, planning and design of studies, assessment of data quality and data collection, differentiation of causality and associations and assessment of uncertainty in results. Moreover, the paper also deals with the equally necessary and meaningful extension of curricula in schools and universities.

cs.CY

Combining heterogeneous subgroups with graph-structured variable selection priors for Cox regression

Important objectives in cancer research are the prediction of a patient's risk based on molecular measurements such as gene expression data and the identification of new prognostic biomarkers (e.g. genes). In clinical practice, this is often challenging because patient cohorts are typically small and can be heterogeneous. In classical subgroup analysis, a separate prediction model is fitted using only the data of one specific cohort. However, this can lead to a loss of power when the sample size is small. Simple pooling of all cohorts, on the other hand, can lead to biased results, especially when the cohorts are heterogeneous. For this situation, we propose a new Bayesian approach suitable for continuous molecular measurements and survival outcome that identifies the important predictors and provides a separate risk prediction model for each cohort. It allows sharing information between cohorts to increase power by assuming a graph linking predictors within and across different cohorts. The graph helps to identify pathways of functionally related genes and genes that are simultaneously prognostic in different cohorts. Results demonstrate that our proposed approach is superior to the standard approaches in terms of prediction performance and increased power in variable selection when the sample size is small.

stat.AP

Identifying treatment effect heterogeneity in dose-finding trials using Bayesian hierarchical models

An important task in drug development is to identify patients, which respond better or worse to an experimental treatment. Identifying predictive covariates, which influence the treatment effect and can be used to define subgroups of patients, is a key aspect of this task. Analyses of treatment effect heterogeneity are however known to be challenging, since the number of possible covariates or subgroups is often large, while samples sizes in earlier phases of drug development are often small. In addition, distinguishing predictive covariates from prognostic covariates, which influence the response independent of the given treatment, can often be difficult. While many approaches for these types of problems have been proposed, most of them focus on the two-arm clinical trial setting, where patients are given either the treatment or a control. In this paper we consider parallel groups dose-finding trials, in which patients are administered different doses of the same treatment. To investigate treatment effect heterogeneity in this setting we propose a Bayesian hierarchical dose-response model with covariate effects on dose-response parameters. We make use of shrinkage priors to prevent overfitting, which can easily occur, when the number of considered covariates is large and sample sizes are small. We compare several such priors in simulations and also investigate dependent modeling of prognostic and predictive effects to better distinguish these two types of effects. We illustrate the use of our proposed approach using a Phase II dose-finding trial and show how it can be used to identify predictive covariates and subgroups of patients with increased treatment effects.

stat.ME

Beyond unimodal regression: modelling multimodality with piecewise unimodal regression or deconvolution models

Shape constraints enable us to reflect prior knowledge in regression settings. A unimodality constraint, for example, can describe the frequent case of a first increasing and then decreasing intensity. Yet, data shapes often exhibit multiple modes. Therefore, we go beyond unimodal regression and propose modelling multimodality with piecewise unimodal regression or with deconvolution models based on unimodal peak shapes. Usefulness of unimodal regression and its multimodal extensions is demonstrated within three applications areas: marine biology, astroparticle physics and breath gas analysis. Despite this diversity, valuable results are obtained in each application. This encourages the use of these methods in other areas as well.

stat.AP

Random projections for Bayesian regression

This article deals with random projections applied as a data reduction technique for Bayesian regression analysis. We show sufficient conditions under which the entire $d$-dimensional distribution is approximately preserved under random projections by reducing the number of data points from $n$ to $k\in O(\operatorname{poly}(d/\varepsilon))$ in the case $n\gg d$. Under mild assumptions, we prove that evaluating a Gaussian likelihood function based on the projected data instead of the original data yields a $(1+O(\varepsilon))$-approximation in terms of the $\ell_2$ Wasserstein distance. Our main result shows that the posterior distribution of Bayesian linear regression is approximated up to a small error depending on only an $\varepsilon$-fraction of its defining parameters. This holds when using arbitrary Gaussian priors or the degenerate case of uniform distributions over $\mathbb{R}^d$ for $β$. Our empirical evaluations involve different simulated settings of Bayesian linear regression. Our experiments underline that the proposed method is able to recover the regression model up to small error while considerably reducing the total running time.

stat.CO