arXiv · 2609.13279
The MODA General Attribute Suite: A Four-Track Evaluation Benchmark for Fashion Attribute Extraction
Abstract
Fashion attribute extraction is evaluated inconsistently: results are reported as single aggregate numbers across image types that pose different problems, fields that are not visible in an image are scored as ordinary negatives, and the effect of vocabulary mismatch between datasets is acknowledged but not measured. We release the MODA General Attribute Suite, a four-track benchmark that keeps these problems separate by construction. Each track (localized garment crops, catalogue product images, applicability-aware full-body photographs, and product text) carries its own frozen test set, input contract, metric, and leakage unit, and the tracks are never averaged. The protocol requires label-blind prediction, SHA-256 commitment of prediction files before any label is opened, a fail-closed scorer, and a 10,000-sample paired cluster bootstrap at each track's natural leakage unit; promotion requires a positive interval on every track rather than a favourable mean. We release the scorers, the split builders, our own prediction files with their hashes including the runs we lose, and the MODA_NER(V) model checkpoints for three of the four tracks. The text model is not distributed; its benchmark is. We report baseline results, six interventions that did not improve them, and the limitations of the evidence.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Arkid Mitra. 2026-09-08. The MODA General Attribute Suite: A Four-Track Evaluation Benchmark for Fashion Attribute Extraction. https://arxiv.org/abs/2609.13279
Cite the original work for its findings. Save a collection to share your selection of sources.