arXiv · 2609.33482
How Synthetic Labels Improve Conformal Prediction: A Perspective on Conditional Coverage
Abstract
Conformal prediction provides distribution-free finite-sample marginal coverage, but post-hoc calibration data may be too scarce to learn how uncertainty varies across inputs. Meanwhile, abundant covariates can often be labeled cheaply by domain models or general-purpose language models. We study whether these synthetic labels can improve conditional coverage when only a small trusted sample is available. Building on score-quantile regression, we introduce prediction-powered quantile learning: a synthetic-labeled pool estimates pinball risk, paired trusted and synthetic outcomes correct its bias, and an independent trusted split performs final conformalization. Profiling pinball risk over scalar corrections reveals that population conditional-coverage error is its functional gradient; the corresponding Hessian removes global shifts and weights remaining shape error by boundary density. Composing this geometry with prediction-powered learning yields a three-resource expansion and a benefit--cost rule for synthetic power. Across eight regression benchmarks, synthetic-powered quantile learning substantially improves downstream conditional coverage while preserving marginal validity and producing more compact prediction sets. A human-rating study finds similar gains from external LLM labels and exposes a quality--quantity--cost tradeoff.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Qianyi Chen, Bo Li. 2026-09-27. How Synthetic Labels Improve Conformal Prediction: A Perspective on Conditional Coverage. https://arxiv.org/abs/2609.33482
Cite the original work for its findings. Save a collection to share your selection of sources.