arXiv · 2603.27519
SPROUT: A Scalable Diffusion Foundation Model for Agricultural Vision
Abstract
Image-based plant phenotyping depends on dense structural understanding of crops, yet pixel-level annotation remains expensive across species, organs, growth stages, and field conditions. General-purpose vision foundation models offer a natural route to label efficiency, but their web-scale pretraining objectives transfer weakly to agricultural imagery, where semantics are often determined by fine organ geometry inside repetitive, texture-dominated scenes. We introduce SPROUT, a diffusion foundation model for multi-crop plant phenotyping. SPROUT learns from 2.6 million unlabeled open-field images (MCD-2.6M) using a pixel-space Diffusion Transformer, and selects transferable features with a label-free effective-rank criterion over denoising timesteps. This design shifts pretraining from crop-based invariance to structure-preserving denoising, making the representation better aligned with dense phenotyping tasks. We evaluate SPROUT across dense phenotyping tasks, including organ segmentation, crop-weed parsing, depth estimation, and counting. SPROUT consistently improves over strong web-pretrained baselines, with the largest gains on dense structural prediction, and shows favorable label and compute efficiency compared with general-purpose and crop-specific foundation models. The source code and MCD-2.6M dataset are publicly available.
Explore related subjects
Keep this discovery
Shuai Xiang, James Burridge, Shouyang Liu, Hao Lu, Tokihiro Fukatsu, Yinqiang Zheng, Wei Guo. 2026-03-29. SPROUT: A Scalable Diffusion Foundation Model for Agricultural Vision. https://arxiv.org/abs/2603.27519
Cite the original work for its findings. Save a collection to share your selection of sources.