SearcharxivSearch

arXiv subjects

Fred van Eeuwijk

Publications and source records attributed to Fred van Eeuwijk.

3 recordsLinked to original sources

Hierarchical Causal Structure Learning

Traditional statistical approaches primarily model associations between variables, however many scientific and practical questions require causal methods instead. These methods typically rely on assumptions about an underlying structure and are often represented by a Directed Acyclic Graph (DAG). While causal structures can be learned for single-level data, hierarchical or multi-level settings, where units (e.g., plants, students or patients) are nested within groups (e.g., environments, schools or hospitals), lack suitable methods. This article addresses such settings. These multi-level structures frequently arise in fields such as agriculture, where plants grow within different environments. Building on nonlinear structural causal models, or additive noise models, we propose an approach that facilitates causal structure learning for hierarchical causal models with additive unobserved group-level effects and group-specific causal functions. We also propose a simulation-based manner to compute hard interventions for the estimated causal model. In a simulation study, we show that the proposed method is able to identify the underlying hierarchical causal mechanism. A winter-wheat example is used to showcase how the proposed method can be used in practice and how to interpret its results.

stat.ME

MixINN: Accelerating Plant Breeding by Combining Mixed Models and Deep Learning for Interaction Prediction

Plant breeding underpins global food security through incremental, accumulating improvements in crop yield, quality and sustainability, achieved via repeated cycles of crop ranking, selection and crossing. Climate change disrupts this process by altering local growing conditions, thereby shifting the relative performance of crop genotypes. Predicting these relative changes in yield is critical for food security. Yet, this problem remains an open challenge in plant breeding, and relatively unexplored within the AI community. We propose MixINN, an approach that first isolates high-quality genotype-environment interaction labels using mixed models, and then predicts these interactions for new crop varieties in future environmental conditions with a deep neural network. We evaluate our method on a corn multi-environment trial across the continental United States and show improved prediction of genotype ranking over current plant breeding methods. MixINN demonstrated superior performance in identifying the 20% most productive corn genotypes, leading to a 5.8% higher average yield, which further improved to 7.2% when targeting specific growing environments. These are competitive results for real-world breeding programs, demonstrating the potential of AI research in accelerating the development of climate-adapted crops, and improving future food security under climate change.

cs.LG

Marker-based estimation of heritability in immortal populations

Heritability is a central parameter in quantitative genetics, both from an evolutionary and a breeding perspective. For plant traits heritability is traditionally estimated by comparing within and between genotype variability. This approach estimates broad-sense heritability, and does not account for different genetic relatedness. With the availability of high-density markers there is growing interest in marker based estimates of narrow-sense heritability, using mixed models in which genetic relatedness is estimated from genetic markers. Such estimates have received much attention in human genetics but are rarely reported for plant traits. A major obstacle is that current methodology and software assume a single phenotypic value per genotype, hence requiring genotypic means. An alternative that we propose here, is to use mixed models at individual plant or plot level. Using statistical arguments, simulations and real data we investigate the feasibility of both approaches, and how these affect genomic prediction with G-BLUP and genome-wide association studies. Heritability estimates obtained from genotypic means had very large standard errors and were sometimes biologically unrealistic. Mixed models at individual plant or plot level produced more realistic estimates, and for simulated traits standard errors were up to 13 times smaller. Genomic prediction was also improved by using these mixed models, with up to a 49% increase in accuracy. For GWAS on simulated traits, the use of individual plant data gave almost no increase in power. The new methodology is applicable to any complex trait where multiple replicates of individual genotypes can be scored. This includes important agronomic crops, as well as bacteria and fungi.

stat.AP