arXiv · 2606.31032
Structuring license permissiveness from pairwise comparisons
Abstract
Licenses are legal instruments that inventors rely upon to protect the technologies they build and regulate how they are used---however, the nature of their authorship and selection implies that how they are interpreted, chosen, and enforced is largely unstructured. In practice, this makes it difficult to compare licenses at scale---when is one license considered more permissive than the other, and when are their terms incomparable to each other? Currently, there is a growing list of licenses that are introduced and used, yet no systematic way to study their relationships. This matters for platforms such as Hugging Face, GitHub, and the Python Package Index, where developers publish or build upon technologies that each have their own licenses. Using large language models (LLMs), we introduce methods for comparing licenses at scale: first, in a pairwise fashion to construct and validate a partial ordering based on permissiveness; and by drawing on existing taxonomies of software licenses. Then, we try to recover the structure with the Bradley-Terry model to see if permissiveness can be judged more cheaply and observe a loss of $\sim$20\%---and classify this loss to feature coverage. The former coupled with model rationale allows us to trace restrictiveness, and the latter allows us to understand license selection as a combination of shared provisions.
Explore related subjects
Keep this discovery
Hamidah Oderinwale, David Atkinson, Rachel Hong, Art Abal, Ben Laufer. 2026-06-30. Structuring license permissiveness from pairwise comparisons. https://arxiv.org/abs/2606.31032
Cite the original work for its findings. Save a collection to share your selection of sources.