Searcharxiv⌕ Search

arXiv subjects

Mélissa Tamine

Publications and source records attributed to Mélissa Tamine.

3 recordsLinked to original sources

Shapley-based Data Valuation for LLM Alignment via Sequential Preference Optimization

Data valuation is a natural framework for understanding which data sources matter most when aligning a Large Language Model (LLM) from multiple sources. The standard game-theoretic approach treats each source, or equivalently each preference dataset, as a player in a cooperative game and assigns it a contribution score through the Shapley value. In practice, however, Shapley-based valuation is computationally prohibitive because it requires aligning a separate model for every possible coalition of sources, i.e., an exponential number of alignments. We address this challenge for Direct Alignment Algorithms (DAAs), including IPO, which learn through log-policy ratios with respect to a reference policy. We show that, when a model is aligned sequentially source by source, exact optimization makes each stage contribute additively to the log-probability of a full response, up to a prompt-dependent normalization constant. This allows the log-probability assigned by any coalition to a fixed response to be reconstructed from the base policy and the policies trained on each source individually. This reduces the alignment cost of Shapley-based valuation from exponential to linear, since only one model per source needs to be trained to evaluate coalition scores. We test whether this theoretical property remains approximately valid under finite training across several base models and real-world data sources. We finally compute the Shapley values of these sources under multiple reward models, showing how their estimated contributions vary across evaluation criteria.

cs.LG↗

An Asymptotic Analysis of the Shapley Value for Dataset Valuation

We propose an asymptotic analysis of the Shapley value in a dataset valuation setting in which utilities are modeled as smooth functionals of empirical distributions via reproducing kernel Hilbert space (RKHS) mean embeddings. We prove that, despite its combinatorial definition, the Shapley value of a data source is asymptotically captured by a simple leading term. This term can be interpreted as the first-order contribution of a dataset relative to the surrounding data population. It also identifies the scale of the Shapley value as the number of data sources grows and provides a framework for analyzing existing Shapley value estimators. Moreover, for practitioners working with large numbers of datasets, the leading term becomes a tractable reference against which Shapley value approximations can be benchmarked.

cs.GT↗

On the Impact of the Utility in Semivalue-based Data Valuation

Semivalue-based data valuation uses cooperative-game theory intuitions to assign each data point a value reflecting its contribution to a downstream task. Still, those values depend on the practitioner's choice of utility, raising the question: How robust is semivalue-based data valuation to changes in the utility? This issue is critical when the utility is set as a trade-off between several criteria and when practitioners must select among multiple equally valid utilities. We address this by introducing the notion of a dataset's spatial signature: given a semivalue, we embed each data point into a lower-dimensional space in which any utility becomes a linear functional, making the data valuation framework amenable to a simpler geometric picture. Building on this, we propose a practical methodology centered on an explicit robustness metric that informs practitioners whether and by how much their data valuation results will shift as the utility changes. We validate this approach across diverse datasets and semivalues, demonstrating strong agreement with rank-correlation analyses and offering analytical insight into how choosing a semivalue can amplify or diminish robustness.

cs.AI↗