Searcharxiv⌕ Search

arXiv · 2609.38112

ReCIRC: Rectified Conformal Risk Control

Abstract

Many applications of black-box predictive models require controlling task-relevant error rates, such as missed lesion pixels in segmentation or missed labels in multilabel classification. Conformal risk control (CRC; Angelopoulos et al., arXiv:2208.02814) gives distribution-free guarantees for such losses, but it calibrates a single threshold shared by all inputs. Because conditional risk varies with the input, this marginal guarantee often overprotects easy cases and underprotects hard ones. We propose ReCIRC (Rectified Conformal Risk Control), which inverts each input's estimated local risk curve to reparameterize the calibrated threshold as a risk budget $a$ representing a common target conditional risk, and then applies CRC unchanged to the resulting family. ReCIRC retains CRC's finite-sample marginal guarantee regardless of the accuracy of the estimated curves, while accurate curves yield approximate conditional risk control and, under additional conditions, asymptotically exact conditional risk control; they also support a risk-calibration diagnostic. Across three synthetic and five real-data settings spanning segmentation, multilabel and multiclass classification, and regression, ReCIRC attained the lowest average worst-group risk and mean positive group excess in every setting, while maintaining marginal risk close to the target, whereas changes in prediction size were application-dependent.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Bruno Marcondes e Resende, Helton Graziadei, Thiago Rodrigo Ramos, Rafael Izbicki. 2026-09-29. ReCIRC: Rectified Conformal Risk Control. https://arxiv.org/abs/2609.38112

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

High-Dimensional Partial Least Squares: Spectral Analysis and Fundamental Limitations

Partial Least Squares (PLS) is a widely used method for data integration, designed to extract latent components shared across paired high-dimensional datasets. Despite decades of practical success, a precise theoretical understanding of its behavior in high-dimensional regimes remains limited. In this paper, we study a data integration model in which two high-dimensional data matrices share a low-rank common latent structure while also containing individual-specific components. We analyze the singular vectors of the associated cross-covariance matrix using tools from random matrix theory and derive asymptotic characterizations of the alignment between estimated and true latent directions. These results provide a quantitative explanation of the reconstruction performance of the PLS variant based on Singular Value Decomposition (PLS-SVD) and identify regimes where the method exhibits counter-intuitive or limiting behavior. Building on this analysis, we compare PLS-SVD with principal component analysis applied separately to each dataset and show its asymptotic superiority in detecting the common latent subspace. Overall, our results offer a comprehensive theoretical understanding of high-dimensional PLS-SVD, clarifying both its advantages and fundamental limitations.

stat.ML↗

Empirical Bayes 1-bit matrix completion

The problem of predicting unobserved entries in a binary matrix, known as 1-bit matrix completion, has found diverse applications in fields such as recommendation systems. In this study, we develop an empirical Bayes method for 1-bit matrix completion motivated by the Efron--Morris estimator, a matrix generalization of the James--Stein estimator that shrinks singular values toward zero. The proposed method exploits the underlying low-rank structure of binary matrices, drawing parallels with multidimensional item response theory. Simulation studies and real-data applications demonstrate that the proposed method achieves competitive predictive accuracy and favorable predictive calibration.

stat.ML↗

Certified Adaptive Refresh: Anytime-Valid Monitoring for Federated Conformal RAG

Question-answering services built on retrieval-augmented generation (RAG), in which a language model answers from retrieved documents, are inspected continuously and upgraded repeatedly, so their reliability guarantee must survive both. We study federated conformal RAG: nodes holding private corpora score candidate answers with a shared language model and send compressed scores to a hub that returns an answer set; a miss omits the true answer. We formulate monitoring as a sequential test: alarm when misses exceed a certified bound on the miss rate (the allowance). But a conformal certificate covers one frozen configuration at one look fixed in advance, so it gives no anytime-valid alarm; a threshold calibrated for one model certifies nothing about its replacement; and re-certifying each upgrade at full error level compounds failures. We propose Anytime-FC-RAG with three auditable rules: pick each configuration before its fresh calibration, charge every certificate attempt to one trajectory-wide budget $δ_{\rm cal}$, and fix threshold, allowance and bet before each query. One betting wealth $E_t$, growing in expectation only when misses exceed the allowance, never resets across deployments. Our certified adaptive-refresh theorem shows that, with evidence budget $δ_e$, alarming when $E_t$ reaches $1/δ_e$ has false-alarm probability at most $δ_e + δ_{\rm cal}$ under a random number of history-selected, evidence-driven refreshes of model, retriever, corpus, score or threshold; $δ_{\rm cal}$ pays for wrong certificates. With an exact order-statistic certificate the bound is distribution-free, architecture-agnostic and about misses, not per-input coverage or distribution shift as such. On Qwen2.5/MMLU-Pro the alarm separated audited-harmful from benign shifts completely at a 5% budget; in a model swap, only fresh calibration kept a downgrade's miss rate in check.

stat.ML↗