SearcharxivSearch

arXiv subjects

Maryna Prus

Publications and source records attributed to Maryna Prus.

10 recordsLinked to original sources

Optimal allocation of trials to sub-regions in crop variety testing with multiple years and correlated genotype effects

Plant breeding and variety trials are usually conducted in multiple environments sampled from a defined target population of environments in order to characterize the performance of breeding lines or varieties. When the population is large and heterogeneous, it may be sub-divided into sub-regions or zones according to administrative and agro-ecological criteria. Analysis then focuses on prediction of performance in the individual sub-regions. Modelling the genotype effect in each sub-region as random, information can be borrowed across sub-regions using best linear unbiased prediction based on a suitable variance-covariance matrix for the genotype-zone effects. Here, we consider the important case where kinship of pedigree information is available for the genotypes under test. This information can be integrated into the variance-covariance matrix for genotype-zone effects. The objective we pursue here is to determine the optimal allocation of a fixed budget of trials to sub-regions. This design problem is solved using a combination of theory and explicit equations on one hand and numerical optimization on the other hand. Our proposed novel approach allows obtaining the optimal allocation when the number of genotypes is in the hundreds, a common setting in large plant breeding programs as well as in variety testing for economically important crops.

stat.AP

A Bayesian Updating Framework for Long-term Multi-Environment Trial Data in Plant Breeding

In variety testing, multi-environment trials (MET) are essential for evaluating the genotypic performance of crop plants. A persistent challenge in the statistical analysis of MET data is the estimation of variance components, which are often still inaccurately estimated or shrunk to exactly zero when using residual (restricted) maximum likelihood (REML) approaches. At the same time, institutions conducting MET typically possess extensive historical data that can, in principle, be leveraged to improve variance component estimation. However, these data are rarely incorporated sufficiently. The purpose of this paper is to address this gap by proposing a Bayesian framework that systematically integrates historical information to stabilize variance component estimation and better quantify uncertainty. Our Bayesian linear mixed model (BLMM) reformulation uses priors and Markov chain Monte Carlo (MCMC) methods to maintain the variance components as positive, yielding more realistic distributional estimates. Furthermore, our model incorporates historical prior information by managing MET data in successive historical data windows. Variance component prior and posterior distributions are shown to be conjugate and belong to the inverse gamma and inverse Wishart families. While Bayesian methodology is increasingly being used for analyzing MET data, to the best of our knowledge, this study comprises one of the first serious attempts to objectively inform priors in the context of MET data. This refers to the proposed Bayesian updating approach. To demonstrate the framework, we consider an application where posterior variance component samples are plugged into an A-optimality experimental design criterion to determine the average optimal allocations of trials to agro-ecological zones in a sub-divided target population of environments (TPE).

stat.ME

Equivalence theorems for compound design problems with application in mixed models

In the present paper we consider design criteria which depend on several designs simultaneously. We formulate equivalence theorems based on moment matrices (if criteria depend on designs via moment matrices) or with respect to the designs themselves (for finite design regions). We apply the obtained optimality conditions to the multiple-group random coefficient regression models and illustrate the results by simple examples.

math.ST

Optimizing the allocation of trials to sub-regions in multi-environment crop variety testing

New crop varieties are extensively tested in multi-environment trials in order to obtain a solid empirical basis for recommendations to farmers. When the target population of environments is large and heterogeneous, a division into sub-regions is often advantageous. When designing such trials, the question arises how to allocate trials to the different subregions. We consider a solution to this problem assuming a linear mixed model. We propose an analytical approach for computation of optimal designs for best linear unbiased prediction of genotype effects and pairwise linear contrasts and illustrate the obtained results by a real data example from Indian nation-wide maize variety trials. It is shown that, except in simple cases such as a compound symmetry model, the optimal allocation depends on the variance-covariance structure for genotypic effects nested within sub-regions.

stat.AP

Optimal Designs for Minimax-Criteria in Random Coefficient Regression Models

We consider minimax-optimal designs for the prediction of individual parameters in random coefficient regression models. We focus on the minimax-criterion, which minimizes the "worst case" for the basic criterion with respect to the covariance matrix of random effects. We discuss particular models: linear and quadratic regression, in detail.

math.ST

Various Optimality Criteria for the Prediction of Individual Response Curves

We consider optimal designs for the Kiefer cirteria, which include the E-criterion as a particular case, and the G-criterion in random coefficients regression (RCR) models. We obtain general the Kiefer criteria for approximate designs and prove the equivalence of the E-criteria in the fixed effects and RCR models. We discuss in detail the G-criterion for ordinary linear regression on specific design regions.

math.ST

Optimal Designs in Multiple Group Random Coefficient Regression Models

The subject of this work is multiple group random coefficients regression models with several treatments and one control group. Such models are often used for studies with cluster randomized trials. We investigate A-, D- and E-optimal designs for estimation and prediction of fixed and random treatment effects, respectively, and illustrate the obtained results by numerical examples.

math.ST

Optimal Design in Hierarchical Models with application in Multi-center Trials

Hierarchical random effect models are used for different purposes in clinical research and other areas. In general, the main focus is on population parameters related to the expected treatment effects or group differences among all units of an upper level (e.g. subjects in many settings). Optimal design for estimation of population parameters are well established for many models. However, optimal designs for the prediction for the individual units may be different. Several settings are identiffed in which individual prediction may be of interest. In this paper we determine optimal designs for the individual predictions, e.g. in multi-center trials, and compare them to a conventional balanced design with respect to treatment allocation. Our investigations show, that balanced designs are far from optimal if the treatment effects vary strongly as compared to the residual error and more subjects should be recruited to the active (new) treatment in multi-center trials. Nevertheless, effciency loss may be limited resulting in a moderate sample size increase when individual predictions are foreseen with a balanced allocation.

stat.AP

Computing optimal experimental designs with respect to a compound Bayes risk criterion

We consider the problem of computing optimal experimental design on a finite design space with respect to a compound Bayes risk criterion, which includes the linear criterion for prediction in a random coefficient regression model. We show that the problem can be restated as constrained A-optimality in an artificial model. This permits using recently developed computational tools, for instance the algorithms based on the second-order cone programming for optimal approximate design, and mixed-integer second-order cone programming for optimal exact designs. We demonstrate the use of the proposed method for the problem of computing optimal designs of a random coefficient regression model with respect to an integrated mean squared error criterion.

stat.CO