Searcharxiv⌕ Search

arXiv · 2610.03815

GCTAg: scalable mixed-model analysis for biobank-scale agricultural cohorts

Abstract

Genome-wide association studies identify genomic variants associated with traits. Mixed linear model association (MLMA) methods using a whole-genome relationship matrix, such as those implemented in GCTA, are powerful but computationally expensive. Here, we remove key memory and CPU bottlenecks in GCTA, reducing REML memory usage by nearly 75% and substantially accelerating MLMA by orders of magnitude in biobank-scale cohorts while preserving exactness. We further exploit relatedness in the mapping cohort through a reduced-rank Woodbury matrix approach, delivering further orders of magnitude performance gains with controlled genomic inflation. Native on-the-fly dominance recoding also eliminates slow I/O-operations on intermediate files, enabling efficient additive and dominance MLMA analyses in large agricultural cohorts.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Alexander S. Leonard, Qiongyu He, Natasha Watson, Naveen Kumar Kadri, Hubert Pausch. 2026-10-01. GCTAg: scalable mixed-model analysis for biobank-scale agricultural cohorts. https://arxiv.org/abs/2610.03815

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

HRPv2: an automated and enhanced method for full-length homology-based R-gene prediction

Motivation: Plant disease resistance genes, particularly those encoding NB-LRR proteins, are important targets for crop improvement. Proteome-based domain or motif searches can only identify NB-LRRs among existing gene models, meaning they cannot recover loci that have been missed or incorrectly predicted by the reference annotation. The full-length, homology-based R-gene prediction (HRP) method circumvents this issue by reconstructing gene models directly on the genome. However, the original implementation of this method requires several separate phases of classification, comparison and filtering operations. Results: HRPv2 is an automated, enhanced version of the original strategy. The labour-intensive, step-by-step curation process used in HRP has been replaced by a reproducible filtering framework. Enhancing two steps of the homology search process has enabled HRPv2 to better account for the specific NB-LRR variability of the genome. Performance validation confirmed that the number of full-length NB-LRRs annotated in both the automatically predicted gene set of the respective genome assembly and the final NB-LRR repertoire has increased in HRPv2 compared to HRP. Availability and implementation: HRPv2 and its associated documentation and reproducible test data are available at https://github.com/AndolfoG/HRPv2. Detailed installation and dependency information is provided in the repository README. A version-pinned Conda package has also been developed and locally validated to provide a reproducible execution environment.

q-bio.GN↗

A Unified Unsupervised Framework for Genome-Wide Association Studies in Heterogeneous Populations

Genome-wide association studies (GWAS) have greatly advanced the discovery of genetic variants underlying complex traits and diseases. Yet in heterogeneous populations, existing GWAS strategies typically either pool all individuals under an assumption of population homogeneity or perform meta-analysis across predefined subgroups, both of which are limited when latent genetic heterogeneity attenuates subgroup-specific effects and masks true associations or subgroup labels are imprecise. Here we present UCALM, a unified unsupervised framework that infers genetically homogeneous subgroups directly from the data and integrates subgroup-specific GWAS with a novel layered meta-analysis method to capture both shared and subgroup-specific association signals. Through extensive simulations and analyses of large-scale human and livestock cohorts, including the UK Biobank ($n \approx 487{,}000$) and a heterogeneous pig cohort ($n \approx 85{,}000$), we demonstrate that UCALM substantially alleviated the mean genomic inflation across 24 UK Biobank traits to 1.17 compared with 1.43 for GLM and 1.37 for LDAK-KVIK, and further revealed 74 loci in the pig cohort that were previously obscured by conventional approaches. Our results establish a robust and broadly applicable strategy for association mapping in structured populations and improve the resolution of genetic signals across diverse species.

q-bio.GN↗

Tracing model-generated DNA with position-independent watermarking

Genomic language models can write synthetic DNA that carries no record of its origin. A generation-time watermark could provide a provenance signal, if a verifier can detect it using only the DNA sequence, a key, and published detector settings, without knowing where the generated region starts, which strand it lies on, or how it was divided into six-base tokens. We embed the SynthID tournament watermark in two genomic language models, Carbon and GENERator-v2. The verifier searches both strands, every start position, and four window lengths, and corrects its decision for the whole search. In a detection cohort of 1,544 held-out prompts, it found all 3,088 marked sequences per model, unedited and after one substituted, inserted, or deleted base. For ordinary sequences, the one-sided 95% upper confidence bound on the false-positive rate was at most 0.850%; one bound for the wrong-key control reached 1.015% (Carbon, after a deletion). In a development cohort, neither model showed a detectable change in likelihood or in predefined sequence measures. Under random edits placed without reference to the detector, detection stayed complete or nearly complete up to a 2% per-base edit rate. These results show statistical detection, not biological function, secret-key security, or robustness to an editor who sees the detector.

q-bio.GN↗