SearcharxivSearch

arXiv subjects

Mark A. Lemley

Publications and source records attributed to Mark A. Lemley.

10 recordsLinked to original sources

Estimating near-verbatim extraction risk in language models with decoding-constrained beam search

Recent work shows that standard greedy-decoding extraction methods for quantifying memorization in LLMs miss how extraction risk varies across sequences. Probabilistic extraction -- computing the probability of generating a target suffix given a prefix under a decoding scheme -- addresses this, but is tractable only for verbatim memorization, missing near-verbatim instances that pose similar privacy and copyright risks. Quantifying near-verbatim extraction risk is expensive: the set of near-verbatim suffixes is combinatorially large, and reliable Monte Carlo (MC) estimation can require ~100,000 samples per sequence. To mitigate this cost, we introduce decoding-constrained beam search, which yields deterministic lower bounds on near-verbatim extraction risk at a cost comparable to ~20 MC samples per sequence. Across experiments, our approach surfaces information invisible to verbatim methods: many more extractable sequences, substantially larger per-sequence extraction mass, and patterns in how near-verbatim extraction risk manifests across model sizes and types of text.

cs.CL

Extracting memorized pieces of (copyrighted) books from open-weight language models

Plaintiffs and defendants in copyright lawsuits over generative AI often make sweeping, opposing claims about the extent to which large language models (LLMs) memorize protected expression from books in their training data. We show that these polarized positions dramatically oversimplify the relationship between memorization and copyright. To do so, we develop a technique to measure memorization of books, which we apply to 200 books and 14 open-weight LLMs. Through over 3000 experiments, we show that memorization varies both by model and book. With respect to our specific extraction methodology, we find that most LLMs do not memorize most books -- either in whole or in part; however, there are notable exceptions. For instance, Llama 3.1 70B entirely memorizes some books, like Harry Potter and the Sorcerer's Stone; memorization is so extensive that one can deterministically extract the whole book almost verbatim using the book's first few words as an initial prompt. We discuss why our results have significant implications for copyright cases, though not ones that unambiguously favor either side.

cs.CL

Probabilistic "Copies" in Generative AI Models

Recent work shows that it is possible to extract verbatim or near-verbatim text of some copyrighted works from some large language models (LLMs or models). That is evidence that the model weights encode the works in some form - that the model has "memorized" those works from its training data. But LLMs don't store information in the same format as familiar databases. Rather, their weights store statistical relationships between tokens that have been learned from the training data, and those relationships inform a generation process that is often probabilistic rather than deterministic. In the case of memorization, those relationships are strong enough that, in many circumstances, the model might generate a copyrighted work from its training data with some probability. Copyright law has not previously had to decide whether storing information that might or might not produce output similar to a copyrighted work is itself a copy of the work. The answer to the question is important, because it may determine the legality of many LLMs. The statute and case law are largely unhelpful. We argue that copyright law will likely take a functional approach to the question, finding that LLMs contain a copy of a particular work only if it is straightforward to extract that work in outputs. That result is unsatisfying as a policy matter, and we suggest potential changes to the law, but it is the most likely outcome under current law.

cs.CY

Extractable Memorization From First Principles

Recent work on extractable memorization in LLMs suffers from two contrasting validity problems. Some studies overstate extraction, e.g., relying on sequences too short to distinguish memorization from predictability. Others imply that extraction is unreliable evidence of memorization, since models can also reproduce real-world text they weren't explicitly trained on. In different ways, both overlook what makes a valid extraction claim: the model must generate a training sequence with high enough probability to indicate memorization. To determine what's high enough, one has to perform a matched comparison: measuring the generation probabilities of both the training sequences of interest and comparable non-training sequences. Because non-training sequences cannot have been memorized, their probabilities provide a baseline for predictability; a training sequence exceeding this baseline provides evidence of memorization. We formalize matched comparisons in two ways: (1) a conformal test that calibrates a threshold to a chosen FPR when training and non-training sequences are sampled from populations, and (2) a census that calibrates against a matched non-training document when the object is a single document (e.g., a book). We show that matched comparisons enable rigorous, calibrated memorization claims, and reveal where prior setups have validity issues. For instance, on Wikipedia OLMo 2 32B reproduces non-training 10-token suffixes roughly 24% as often as training ones: that share of the training generation rate reflects false positives, not memorization. For Llama 3.1 70B on books, the thresholds we calibrate are as low as 1e-27, supporting memorization claims for sequences that no feasible sampling budget would extract. Based on these results, we refine "extractable memorization" to require a valid memorization claim and near-certain generation within a realistic budget.

cs.LG

Machine Unlearning Doesn't Do What You Think: Lessons for Generative AI Policy and Research

"Machine unlearning" is a popular proposed solution for mitigating the existence of content in an AI model that is problematic for legal or moral reasons, including privacy, copyright, safety, and more. For example, unlearning is often invoked as a solution for removing the effects of specific information from a generative-AI model's parameters, e.g., a particular individual's personal data or the inclusion of copyrighted content in the model's training data. Unlearning is also proposed as a way to prevent a model from generating targeted types of information in its outputs, e.g., generations that closely resemble a particular individual's data or reflect the concept of "Spiderman." Both of these goals--the targeted removal of information from a model and the targeted suppression of information from a model's outputs--present various technical and substantive challenges. We provide a framework for ML researchers and policymakers to think rigorously about these challenges, identifying several mismatches between the goals of unlearning and feasible implementations. These mismatches explain why unlearning is not a general-purpose solution for circumscribing generative-AI model behavior in service of broader positive impact.

cs.LG

The Mirage of Artificial Intelligence Terms of Use Restrictions

Artificial intelligence (AI) model creators commonly attach restrictive terms of use to both their models and their outputs. These terms typically prohibit activities ranging from creating competing AI models to spreading disinformation. Often taken at face value, these terms are positioned by companies as key enforceable tools for preventing misuse, particularly in policy dialogs. But are these terms truly meaningful? There are myriad examples where these broad terms are regularly and repeatedly violated. Yet except for some account suspensions on platforms, no model creator has actually tried to enforce these terms with monetary penalties or injunctive relief. This is likely for good reason: we think that the legal enforceability of these licenses is questionable. This Article systematically assesses of the enforceability of AI model terms of use and offers three contributions. First, we pinpoint a key problem: the artifacts that they protect, namely model weights and model outputs, are largely not copyrightable, making it unclear whether there is even anything to be licensed. Second, we examine the problems this creates for other enforcement. Recent doctrinal trends in copyright preemption may further undermine state-law claims, while other legal frameworks like the DMCA and CFAA offer limited recourse. Anti-competitive provisions likely fare even worse than responsible use provisions. Third, we provide recommendations to policymakers. There are compelling reasons for many provisions to be unenforceable: they chill good faith research, constrain competition, and create quasi-copyright ownership where none should exist. There are, of course, downsides: model creators have fewer tools to prevent harmful misuse. But we think the better approach is for statutory provisions, not private fiat, to distinguish between good and bad uses of AI, restricting the latter.

cs.CY

Foundation Models and Fair Use

Existing foundation models are trained on copyrighted material. Deploying these models can pose both legal and ethical risks when data creators fail to receive appropriate attribution or compensation. In the United States and several other countries, copyrighted content may be used to build foundation models without incurring liability due to the fair use doctrine. However, there is a caveat: If the model produces output that is similar to copyrighted data, particularly in scenarios that affect the market of that data, fair use may no longer apply to the output of the model. In this work, we emphasize that fair use is not guaranteed, and additional work may be necessary to keep model development and deployment squarely in the realm of fair use. First, we survey the potential risks of developing and deploying foundation models based on copyrighted content. We review relevant U.S. case law, drawing parallels to existing and potential applications for generating text, source code, and visual art. Experiments confirm that popular foundation models can generate content considerably similar to copyrighted material. Second, we discuss technical mitigations that can help foundation models stay in line with fair use. We argue that more research is needed to align mitigation strategies with the current state of the law. Lastly, we suggest that the law and technical mitigations should co-evolve. For example, coupled with other policy mechanisms, the law could more explicitly consider safe harbors when strong technical tools are used to mitigate infringement harms. This co-evolution may help strike a balance between intellectual property and innovation, which speaks to the original goal of fair use. But we emphasize that the strategies we describe here are not a panacea and more work is needed to develop policies that address the potential harms of foundation models.

cs.CY

Is Patent Law Technology Specific?

Although patent law purports to cover all manner of technologies, we have noticed recent divergence in the standards applied to biotechnology and to software patents: the Federal Circuit has applied a very permissive standard of obviousness in biotechnology, but a highly restrictive disclosure requirement. The opposite holds true for software patents, which seems to us exactly contrary to sound policy for either industry. These patent standards are grounded in the legal fiction of the "person having ordinary skill in the art" or PHOSITA. We discuss the appropriateness of the PHOSITA standard, concluding that it properly lends flexibility to the patent system. We then discuss the difficulty of applying this standard in different industries, offering suggestions as to how it might be modified to avoid the problems seen in biotechnology and software patents.

cs.CY

ICANN and Antitrust

The Internet Corporation for Assigned Names and Numbers (ICANN) is a private non-profit company which, pursuant to contracts with the US government, acts as the de facto regulator for DNS policy. ICANN decides what TLDs will be made available to users, and which registrars will be permitted to offer those TLDs for sale. In this article we focus on a hitherto-neglected implication of ICANN's assertion that it is a private rather than a public actor: its potential liability under the U.S. antitrust laws, and the liability of those who transact with it. ICANN argues that it is not as closely tied to the government as NSI and IANA were in the days before ICANN was created. If this is correct, it seems likely that ICANN will not benefit from the antitrust immunity those actors enjoyed. Some of ICANN's regulatory actions may restrain competition, e.g. its requirement that applicants for new gTLDs demonstrate that their proposals would not enable competitive (alternate) roots and ICANN's preventing certain types of non-price competition among registrars (requiring the UDRP). ICANN's rule adoption process might be characterized as anticompetitive collusion by existing registrars, who are likely not be subject to the Noerr-Pennington lobbying exemption. Whether ICANN has in fact violated the antitrust laws depends on whether it is an antitrust state actor, whether the DNS is an essential facility, and on whether it can shelter under precedents that protect standard-setting bodies. If (as seems likely) a private ICANN and those who petition it are subject to antitrust law, everyone involved in the process needs to review their conduct with an eye towards legal liability. ICANN should act very differently with respect to both the UDRP and the competitive roots if it is to avoid restraining trade.

cs.CY

Antitrust, Intellectual Property and Standard-Setting Organizations

Standard-setting organizations (SSOs) regularly encounter situations in which one or more companies claim to own proprietary rights that cover a proposed industry standard. The industry cannot adopt the standard without the permission of the intellectual property owner (or owners). How SSOs respond to those who assert intellectual property rights is critically important. Whether or not private companies retain intellectual property rights in group standards will determine whether a standard is "open" or "closed." It will determine who can sell compliant products, and it may well influence whether the standard adopted in the market is one chosen by a group or one offered by a single company. SSO rules governing intellectual property rights will also affect how standards change as technology improves. Given the importance of SSO rules governing intellectual property rights, there has been surprisingly little treatment of SSOs or their intellectual property rules in the legal literature. My aim in this article is to fill that void. To do so, I have surveyed the intellectual property policies of dozens of SSOs, primarily but not exclusively in the computer networking and telecommunications industries.

cs.CY