Searcharxiv⌕ Search

arXiv · 2609.35877

BarcodeMAE+: Rethinking Masked Pretraining and Global Representations for DNA Barcode Foundation Models

Abstract

Many DNA foundation models are pretrained by masking parts of a sequence and asking the model to reconstruct them. Standard masked pretraining exposes the encoder to special [MASK] tokens that are absent at inference, creating a mismatch between training and downstream use. The role of an explicit global sequence representation such as a [CLS] token and how it should be trained also remain poorly understood for DNA barcodes. We introduce BarcodeMAE+ and study model architecture, global [CLS] representation, and auxiliary pretraining objectives across arthropod COI (BIOSCAN-5M) and fungal ITS (UNITE+INSD) barcodes. Across both barcode regions, the encoder-decoder MAE-LM architecture outperforms its matched encoder-only counterpart in nearly all evaluated configurations, supporting MAE-LM as an effective architectural design for DNA barcode foundation models. A trained global [CLS] representation provides substantial additional gains: on BIOSCAN-5M, [CLS] accuracy increases from 47.53% without an auxiliary objective to 80.65% with cross-entropy genus classification. The best auxiliary objective is region-dependent: cross-entropy performs best on BIOSCAN-5M, whereas pairwise same-genus classification performs best on UNITE+INSD, reaching 73.19% on Yeast and 63.07% on Filamentous Fungi. BarcodeMAE+ outperforms published DNA foundation model baselines on BIOSCAN-5M and achieves the highest Yeast accuracy among the evaluated UNITE+INSD baselines using frozen encoder representations. Similarity-weighted softmax KNN voting further stabilizes accuracy as neighbourhood size increases. Overall, encoder-decoder masked pretraining and an explicitly trained global representation are strong design choices for DNA barcode foundation models, while the optimal objective for learning that representation depends on the biological domain.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Monireh Safari, Pablo Millan Arias, Scott C. Lowe, Lila Kari, Angel X. Chang, Graham W. Taylor. 2026-09-26. BarcodeMAE+: Rethinking Masked Pretraining and Global Representations for DNA Barcode Foundation Models. https://arxiv.org/abs/2609.35877

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

CancerZigZag: Iterative Seed-Anchored Diffusion for Generative Modeling of Single-Cell State Transitions

Single-cell cancer datasets are predominantly cross-sectional and rarely provide paired or longitudinal observations linking individual healthy-like cells to tumor-associated states. We introduce CancerZigZag, a seed-initialized diffusion-based framework for exploratory generation of tumor-associated single-cell candidate clouds from unpaired epithelial cell populations. For each cancer context, a diffusion model is trained exclusively on tumor-derived epithelial cells and applied through repeated partial latent-space perturbation and reverse diffusion to held-out healthy-like seeds, generating stochastic candidate clouds without paired measurements or classifier guidance. We applied CancerZigZag to colorectal, breast, lung, and renal cell carcinoma contexts and explored parameter landscapes defined by perturbation depth and the number of ZigZag cycles. Across the reported operating configurations, candidate clouds contained outputs classified toward held-out tumor-derived reference populations for each evaluated seed. Residual seed-dependent organization varied across contexts, with the clearest structure in colorectal cancer, more modest organization in lung cancer, and limited cloud-level structure in breast and renal cell carcinoma. Representative candidates also showed directional concordance with transcriptional shifts observed between held-out healthy-like and tumor-derived reference populations. CancerZigZag is not interpreted as a model of deterministic healthy-to-tumor transformation or cellular progression. Instead, it provides a reference-informed framework for exploring tumor-associated candidate distributions from unpaired healthy-like seeds and quantifying the context-dependent relationship between tumor-associated displacement and residual seed dependence.

q-bio.GN↗

Learning Interpretable Tumor Microenvironment Representations by Fitting Pan-Cancer Cell State-Niche Correlation

In the tumor microenvironment, a cell's state is influenced by cell-cell interactions (CCIs) with neighboring cells in its niches. Identifying dysregulated CCIs that are associated with pathogenic processes pinpoints targets for drug discovery. Imaging-based spatial transcriptomics and single-cell RNA sequencing provide, respectively, single-cell spatial information and transcriptome-wide measurements needed to study CCIs, but neither modality provides both. Existing spatial transcriptomics foundation models also cannot effectively learn from spatially resolved single-cell data with full-transcriptome coverage, explicitly infer the CCI mechanisms driving cell state-niche associations, or be interpretable enough to support direct biological interpretations. Here, we present GITIII-scale, a hierarchical, interpretable pan-cancer spatial transcriptomics foundation model for TME representation learning that investigates cell state-niche associations and their underlying ligand-receptor (LR) signaling pathways. GITIII-scale uses transformers to model interactions between pairs of cells at defined spatial distances, an interpretable single-layer graph transformer without a feed-forward network to decompose how each gene in a receiver cell is influenced by each neighboring sender cell, and a graph transformer to generate cellular-neighborhood embeddings. Trained on our assembled pan-cancer database of specimen-matched scRNA-seq and imaging-based spatial transcriptomics datasets, GITIII-scale generated TME embeddings that recovered niche-associated state changes more accurately than existing spatial transcriptomics foundation models in cancer types unseen during training. A case study of an unseen breast cancer dataset further demonstrated the model's interpretability by identifying potentially drug-targetable LR pathways associated with endothelial overgrowth and tumorigenesis.

q-bio.GN↗

EMMA: an R/Bioconductor package to automate tracking of metadata in functional enrichment analyses

Summary: Functional enrichment analysis (FEA) is a widely used approach for interpreting high-throughput omics data. However, essential methodological details, such as software versions, analysis parameters, and annotation database releases among others, are often incompletely reported, limiting the reproducibility and transparency of enrichment analyses and complicating the assessment of potentially problematic methodological choices. Here we present EMMA, an R/Bioconductor package that integrates with existing FEA tools and automatically captures provenance metadata, such as annotation metadata, software version, and parameters, during the analysis runtime. Our package provides utilities for accessing and exporting the recorded metadata to facilitate transparent reporting and preserve provenance required for reproducible enrichment analyses. This also enables auditing of the results while remaining compatible with existing Bioconductor workflows. Availability and implementation: EMMA is available on Bioconductor under the MIT license (https: //bioconductor.org/packages/EMMA), with its development version also available on GitHub (https: //github.com/imbeimainz/EMMA).

q-bio.GN↗