Shapley-based Data Valuation for LLM Alignment via Sequential Preference Optimization
Data valuation is a natural framework for understanding which data sources matter most when aligning a Large Language Model (LLM) from multiple sources. The standard game-theoretic approach treats each source, or equivalently each preference dataset, as a player in a cooperative game and assigns it a contribution score through the Shapley value. In practice, however, Shapley-based valuation is computationally prohibitive because it requires aligning a separate model for every possible coalition of sources, i.e., an exponential number of alignments. We address this challenge for Direct Alignment Algorithms (DAAs), including IPO, which learn through log-policy ratios with respect to a reference policy. We show that, when a model is aligned sequentially source by source, exact optimization makes each stage contribute additively to the log-probability of a full response, up to a prompt-dependent normalization constant. This allows the log-probability assigned by any coalition to a fixed response to be reconstructed from the base policy and the policies trained on each source individually. This reduces the alignment cost of Shapley-based valuation from exponential to linear, since only one model per source needs to be trained to evaluate coalition scores. We test whether this theoretical property remains approximately valid under finite training across several base models and real-world data sources. We finally compute the Shapley values of these sources under multiple reward models, showing how their estimated contributions vary across evaluation criteria.