arXiv · 2307.02719
Understanding Uncertainty Sampling via Equivalent Loss
Abstract
Uncertainty sampling is a classical active-learning strategy, yet the statistical objective induced by its query rule is often implicit. We introduce the equivalent loss, whose gradient is the original loss gradient multiplied by the query probability. This construction places probabilistic, margin-based, and threshold-based uncertainty rules within a common framework. For binary classification, concrete equivalent losses and surrogate link functions show how uncertainty weighting preserves calibration while reshaping optimization geometry. When the equivalent loss is convex, we derive a finite-sample excess-risk bound with a fixed learning rate and an explicit constant controlling the tradeoff between initial error and query-weighted gradient variance. A feasible choice based only on maximal uncertainty yields a leading classification-risk upper bound no larger than the corresponding passive-learning bound at the same expected label budget; knowledge of the average query rate sharpens this comparison. We also analyze pool-based sampling, characterize the integrability obstruction beyond scalar prediction, and examine the query-clock dynamics of momentum methods. Together, these results provide a reusable route from a classical acquisition rule to its induced objective, statistical guarantees, and optimization behavior.
Explore related subjects
Keep this discovery
Shang Liu, Xiaocheng Li. 2023-07-06. Understanding Uncertainty Sampling via Equivalent Loss. https://arxiv.org/abs/2307.02719
Cite the original work for its findings. Save a collection to share your selection of sources.
Discover connections
Connections use source metadata and explicit phrase matches, not verified experimental comparisons.