SearcharxivSearch

arXiv subjects

Rainer Schwabe

Publications and source records attributed to Rainer Schwabe.

At least 19 recordsLinked to original sources

Optimal Design in Repeated Testing for Count Data

In this paper, we develop optimal designs for growth curve models with count data based on the Rasch Poisson-Gamma counts (RPGCM) model. This model is often used in educational and psychological testing when test results yield count data. In the RPGCM, the test scores are determined by respondents ability and item difficulty. Locally D-optimal designs are derived for maximum quasi-likelihood estimation to efficiently estimate the mean abilities of the respondents over time. Using the log link, both unstructured, linear and nonlinear growth curves of log mean abilities are taken into account. Finally, the sensitivity of the derived optimal designs due to an imprecise choice of parameter values is analyzed using D-efficiency.

stat.ME

Poisson Regression in one Covariate on Massive Data

The goal of subsampling is to select an informative subset of all observations, when using the full data for statistical analysis is not viable. We construct locally $ D $-optimal subsampling designs under a Poisson regression model with a log link in one covariate. A Representation of the support of locally $ D $-optimal subsampling designs is established. We make statements on scale-location transformations of the covariate that require a simultaneous transformation of the regression parameter. The performance of the methods is demonstrated by illustrating examples. To show the advantage of the optimal subsampling designs, we examine the efficiency of uniform random subsampling as well as of two heuristic designs. Further, the efficiency of locally $ D $-optimal subsampling designs is studied when the parameter is misspecified.

math.ST

A $p$-step-ahead sequential adaptive algorithm for D-optimal nonlinear regression design

Under a nonlinear regression model with univariate response an algorithm for the generation of sequential adaptive designs is studied. At each stage, the current design is augmented by adding $p$ design points where $p$ is the dimension of the parameter of the model. The augmenting $p$ points are such that, at the current parameter estimate, they constitute the locally D-optimal design within the set of all saturated designs. Two relevant subclasses of nonlinear regression models are focused on, which were considered in previous work of the authors on the adaptive Wynn algorithm: firstly, regression models satisfying the `saturated identifiability condition' and, secondly, generalized linear models. Adaptive least squares estimators and adaptive maximum likelihood estimators in the algorithm are shown to be strongly consistent and asymptotically normal, under appropriate assumptions. For both model classes, if a condition of `saturated D-optimality' is satisfied, the almost sure asymptotic D-optimality of the generated design sequence is implied by the strong consistency of the adaptive estimators employed by the algorithm. The condition states that there is a saturated design which is locally D-optimal at the true parameter point (in the class of all designs).

math.ST

D-optimal Subsampling Design for Massive Data Linear Regression

Data reduction is a fundamental challenge of modern technology, where classical statistical methods are not applicable because of computational limitations. We consider multiple linear regression for an extraordinarily large number of observations, but only a few covariates. Subsampling aims at the selection of a given proportion of the existing original data. Under distributional assumptions on the covariates, we derive D-optimal subsampling designs and study their theoretical properties. We make use of fundamental concepts of optimal design theory and an equivalence theorem from constrained convex optimization. The thus obtained subsampling designs provide simple rules for whether to accept or reject a data point, allowing for an easy algorithmic implementation. In addition, we propose a simplified subsampling method with lower computational complexity that deviates from the D-optimal design. We present a simulation study, comparing both subsampling schemes with the IBOSS method in the case of a fixed size of the subsample.

stat.ME

Optimal Subsampling Design for Polynomial Regression in one Covariate

Improvements in technology lead to increasing availability of large data sets which makes the need for data reduction and informative subsamples ever more important. In this paper we construct $ D $-optimal subsampling designs for polynomial regression in one covariate for invariant distributions of the covariate. We study quadratic regression more closely for specific distributions. In particular we make statements on the shape of the resulting optimal subsampling designs and the effect of the subsample size on the design. To illustrate the advantage of the optimal subsampling designs we examine the efficiency of uniform random subsampling.

math.ST

$D$-Optimal and Nearly $D$-Optimal Exact Designs for Binary Response on the Ball

In this paper the results of Radloff and Schwabe (2019a) will be extended for a special class of symmetrical intensity functions. This includes binary response models with logit and probit link. To evaluate the position and the weights of the two non-degenerated orbits on the $k$-dimensional ball usually a system of three equations has to be solved. The symmetry allows to reduce this system to a single equation. As a further result, the number of support points can be reduced to the minimal number. These minimally supported designs are highly efficient. The results can be generalized to arbitrary ellipsoidal design regions.

stat.ME

Optimal designs for discrete choice models via graph Laplacians

In discrete choice experiments, the information matrix depends on the model parameters. Therefore designing optimally informative experiments for arbitrary initial parameters often yields highly nonlinear optimization problems and makes optimal design infeasible. To overcome such challenges, we connect design theory for discrete choice experiments with Laplacian matrices of undirected graphs, resulting in complexity reduction and feasibility of optimal design. We rewrite the $D$-optimality criterion in terms of Laplacians via Kirchhoff's matrix tree theorem, and show that its dual has a simple description via the Cayley-Menger determinant of the Farris transform of the Laplacian matrix. This results in a drastic reduction of complexity and allows us to implement a gradient descent algorithm to find locally $D$-optimal designs. For the subclass of Bradley-Terry paired comparison models, we find a direct link to maximum likelihood estimation for Laplacian-constrained Gaussian graphical models. Finally, we study the performance of our algorithm and demonstrate its application to real and simulated data.

math.ST

Optimal Design for Estimating the Mean Ability over Time in Repeated Item Response Testing

We present general results on D-optimal designs for estimating the mean response in repeated measures growth curve models with metric outcomes. For this situation, we derive a novel equivalence theorem for checking design optimality. The motivation of this work originates from designing a study in psychological item response testing with multiple retests to measure the improvement in ability. Besides introductory linear growth curves for which analytical results can be obtained, we consider two non-linear growth curve models incorporating an increasing mean ability and a saturation effect. For these models, D-optimal designs are determined by computational methods and are validated by means of the equivalence theorem.

math.ST

Experimental Designs for Accelerated Degradation Tests Based on Linear Mixed Effects Models

Accelerated degradation tests are used to provide accurate estimation of lifetime properties of highly reliable products within a relatively short testing time. There data from particular tests at high levels of stress (e.\,g.\ temperature, voltage, or vibration) are extrapolated, through a physically meaningful model, to obtain estimates of lifetime quantiles under normal use conditions. In this work, we consider repeated measures accelerated degradation tests with multiple stress variables, where the degradation paths are assumed to follow a linear mixed effects model which is quite common in settings when repeated measures are made. We derive optimal experimental designs for minimizing the asymptotic variance for estimating the median failure time under normal use conditions when the time points for measurements are either fixed in advance or are also to be optimized.

stat.AP

Optimal Time Plan in Accelerated Degradation Testing

Many highly reliable products are designed to function for years without failure. For such systems accelerated degradation testing may provide significance information about the reliability properties of the system. In this paper, we propose the $c$-optimality criterion for obtaining optimal designs of constant-stress accelerated degradation tests where the degradation path follow a linear mixed effects model. The present work is mainly concerned with developing optimal desings in terms of the time variable rather than considering an optimal design of the stress variable which is usually considered in the majority of the literature. Finally, numerical examples are presented and sensitivity analysis procedures are conducted for evaluating the robustness of the optimal as well as standard designs against misspecifications of the nominal values.

stat.AP

The semi-algebraic geometry of saturated optimal designs for the Bradley-Terry model

Optimal design theory for nonlinear regression studies local optimality on a given design space. We identify designs for the Bradley--Terry paired comparison model with small undirected graphs and prove that every saturated D-optimal design is represented by a path. We discuss the case of four alternatives in detail and derive explicit polynomial inequality descriptions for optimality regions in parameter space. Using these regions, for each point in parameter space we can prescribe a D-optimal design.

math.ST

Optimal Stress Levels in Accelerated Degradation Testing for Various Degradation Models

Accelerated degradation tests are used to provide accurate estimation of lifetime characteristics of highly reliable products within a relatively short testing time. Data from particular tests at high levels of stress (e.g., temperature, voltage, or vibration) are extrapolated, through a physically meaningful statistical model, to attain estimates of lifetime quantiles at normal use conditions. The gamma process is a natural model for estimating the degradation increments over certain degradation paths, which exhibit a monotone and strictly increasing degradation pattern. In this work, we derive first an algorithm-based optimal design for a repeated measures degradation test with single failure mode that corresponds to a single response component. The univariate degradation process is expressed using a gamma model where a generalized linear model is introduced to facilitate the derivation of an optimal design. Consequently, we extend the univariate model and characterize optimal designs for accelerated degradation tests with bivariate degradation processes. The first bivariate model includes two gamma processes as marginal degradation models. The second bivariate models is expressed by a gamma process along with a mixed effects linear model. We derive optimal designs for minimizing the asymptotic variance for estimating some quantile of the failure time distribution at the normal use conditions. Sensitivity analysis is conducted to study the behavior of the resulting optimal designs under misspecifications of adopted nominal values.

stat.AP

In- and Equivariance for Optimal Designs in Generalized Linear Models: The Gamma Model

We give an overview over the usefulness of the concept of equivariance and invariance in the design of experiments for generalized linear models. In contrast to linear models here pairs of transformations have to be considered which act simultaneously on the experimental settings and on the location parameters in the linear component. Given the transformation of the experimental settings the parameter transformations are not unique and may be nonlinear to make further use of the model structure. The general concepts and results are illustrated by models with gamma distributed response. Locally optimal and maximin efficient design are obtained for the common D- and IMSE-criterion.

math.ST

$ D $-optimal designs for Poisson regression with synergetic interaction effect

We characterize $D$-optimal designs in the two-dimensional Poisson regression model with synergetic interaction and provide an explicit proof. The proof is based on the idea of reparameterization of the design region in terms of contours of constant intensity. This approach leads to a substantial reduction of complexity as properties of the sensitivity can be treated along and across the contours separately. Furthermore, some extensions of this result to higher dimensions are presented.

math.ST

Optimal Design for Probit Choice Models with Dependent Utilities

In this paper we derive locally D-optimal designs for discrete choice experiments based on multinomial probit models. These models include several discrete explanatory variables as well as a quantitative one. The commonly used multinomial logit model assumes independent utilities for different choice options. Thus, D-optimal optimal designs for such multinomial logit models may comprise choice sets, e.g., consisting of alternatives which are identical in all discrete attributes but different in the quantitative variable. Obviously such designs are not appropriate for many empirical choice experiments. It will be shown that locally D-optimal designs for multinomial probit models supposing independent utilities consist of counterintuitive choice sets as well. However, locally D-optimal designs for multinomial probit models allowing for dependent utilities turn out to be reasonable for analyzing decisions using discrete choice studies.

stat.ME

Optimality regions for designs in multiple linear regression models with correlated random coefficients

This paper studies optimal designs for linear regression models with correlated effects for single responses. We introduce the concept of rhombic design to reduce the computational complexity and find a semi-algebraic description for the D-optimality of a rhombic design via the Kiefer-Wolfowitz equivalence theorem. Subsequently, we show that the structure of an optimal rhombic design depends directly on the correlation structure of the random coefficients.

math.ST

Convergence of least squares estimators in the adaptive Wynn algorithm for a class of nonlinear regression models

The paper continues the authors' work on the adaptive Wynn algorithm in a nonlinear regression model. In the present paper it is shown that if the mean response function satisfies a condition of `saturated identifiability', which was introduced by Pronzato \cite{Pronzato}, then the adaptive least squares estimators are strongly consistent. The condition states that the regression parameter is identifiable under any saturated design, i.e., the values of the mean response function at any $p$ distinct design points determine the parameter point uniquely where, typically, $p$ is the dimension of the regression parameter vector. Further essential assumptions are compactness of the experimental region and of the parameter space together with some natural continuity assumptions. If the true parameter point is an interior point of the parameter space then under some smoothness assumptions and asymptotic homoscedasticity of random errors the asymptotic normality of adaptive least squares estimators is obtained.

math.ST

The adaptive Wynn-algorithm in generalized linear models with univariate response

For a nonlinear regression model the information matrices of designs depend on the parameter of the model. The adaptive Wynn-algorithm for D-optimal design estimates the parameter at each step on the basis of the employed design points and observed responses so far, and selects the next design point as in the classical Wynn-algorithm for D-optimal design. The name `Wynn-algorithm' is in honor of Henry P. Wynn who established the latter `classical' algorithm in his 1970 paper. The asymptotics of the sequences of designs and maximum likelihood estimates generated by the adaptive algorithm is studied for an important class of nonlinear regression models: generalized linear models whose (univariate) response variables follow a distribution from a one-parameter exponential family. Under the assumptions of compactness of the experimental region and of the parameter space together with some natural continuity assumptions it is shown that the adaptive ML-estimators are strongly consistent and the design sequence is asymptotically locally D-optimal at the true parameter point. If the true parameter point is an interior point of the parameter space then under some smoothness assumptions the asymptotic normality of the adaptive ML-estimators is obtained.

math.ST