arXiv · 2601.13922
Automatic Prompt Optimization for Dataset-Level Feature Discovery
Abstract
Feature extraction from unstructured text is a critical step in many downstream classification pipelines, yet current approaches largely rely on hand-crafted prompts or fixed feature schemas. We formulate feature discovery as a dataset-level prompt optimization problem: given a labelled text corpus, the goal is to induce a global set of interpretable and discriminative feature definitions whose realizations optimize a downstream supervised learning objective. To this end, we propose a multi-agent prompt optimization framework in which language-model agents jointly propose feature definitions, extract feature values, and evaluate feature quality using dataset-level performance and interpretability feedback. Instruction prompts are iteratively refined based on this structured feedback, enabling optimization over prompts that induce shared feature sets rather than per-example predictions. This formulation departs from prior prompt optimization methods that rely on per-sample supervision and provides a principled mechanism for automatic feature discovery from unstructured text.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Adrian Cosma, Oleg Szehr, David Kletz, Alessandro Antonucci, Olivier Pelletier. 2026-01-20. Automatic Prompt Optimization for Dataset-Level Feature Discovery. https://arxiv.org/abs/2601.13922
Cite the original work for its findings. Save a collection to share your selection of sources.