SearcharxivSearch

arXiv subjects

Gregory M. Dickinson

Publications and source records attributed to Gregory M. Dickinson.

17 recordsLinked to original sources

Stop Removing Stopwords: How an Inherited Preprocessing Default Distorts Legal Text-as-Data

Empirical legal scholarship increasingly treats judicial text as data, and much of it still runs on sparse, interpretable pipelines -- TF-IDF features and linear classifiers -- because the textual feature is often the object of study, not merely a means to a prediction. Yet these pipelines inherit a chain of preprocessing defaults from mid-century information retrieval that were never validated against classification accuracy, the most entrenched being stopword removal. This study introduces an exhaustive single-word ablation that measures a preprocessing step's effect directly against the downstream objective, and applies it to stopword removal as the hardest case to dislodge. Matching Supreme Court Database labels to Caselaw Access Project opinion texts, it examines two binary tasks that bracket F1 headroom, ideological direction (no-removal baseline F1 ~ 0.68) and constitutional versus non-constitutional law type (~ 0.92), across 7,668 and 7,001 opinions. For each task the analysis approximates the best stoplist any expert could build, removing each of roughly 18,500 candidate words and measuring the effect directly. Three findings follow: generic stoplists in common use fall below the no-removal baseline in every test; even optimized stoplists are statistically indistinguishable from removing nothing; and meta-models trained on word-level features cannot predict which removals help, so list curation has nothing to target. The method generalizes to any inherited preprocessing default, and the result is a caution specific to interpretable legal text-as-data: a step that silently reshapes which features a model sees can distort the very doctrinal and ideological signal such research exists to recover. Leaving stopwords in place is a question of measurement validity.

cs.CL

Dark Patterns and Consumer Protection Law for App Makers

Dark patterns in online commerce, especially deceptive user interface designs for apps and websites, undermine consumer autonomy and distort online markets. Although sometimes deception is intentional, the complex app development process can also unintentionally produce manipulative user interfaces. This paper discusses common design pitfalls and proposes strategies for app makers to avoid infringing user autonomy or incurring legal liability under emerging principles of consumer protection law. By focusing on choice architecture and transparent design principles, developers can both facilitate compliance and build user trust and loyalty.

cs.CY

Law Proofing the Future

Gemini said Lawmakers today face continuous calls to "future proof" the legal system against generative artificial intelligence, algorithmic decision-making, targeted advertising, and all manner of emerging technologies. This Article takes a contrarian stance: it is not the law that needs bolstering for the future, but the future that needs protection from the law. From the printing press and the elevator to ChatGPT and online deepfakes, the recurring historical pattern is familiar. Technological breakthroughs provoke wonder, then fear, then legislation. The resulting legal regimes entrench incumbents, suppress experimentation, and displace long-standing legal principles with bespoke but brittle rules. Drawing from history, economics, political science, and legal theory, this Article argues that the most powerful tools for governing technological change--the general-purpose tools of the common law--are in fact already on the books, long predating the technologies they are now called upon to govern, and ready also for whatever the future holds in store. Rather than proposing any new statute or regulatory initiative, this Article offers something far rarer, a defense of doing less. It shows how the law's virtues--generality, stability, and adaptability--are best preserved not through prophylactic regulation, but through accretional judicial decision-making. The epistemic limits that make technological forecasting so unreliable and the hidden costs of early legislative intervention, including biased governmental enforcement and regulatory capture, mean that however fast technology may move, the law must not chase it. The case for legal restraint is thus not a defense of the status quo, but a call to preserve the conditions of freedom and equal justice under which both law and technology can evolve.

cs.CY

Consumer Rights and Algorithms

This article summarizes the field of consumer protection law, from its historical roots to the contemporary challenges of the digital age. It outlines the legal doctrines governing consumer deception and unfair trade practices, highlighting the interplay between common-law, statutory, and private modes of regulation. The article then addresses the impact of artificial intelligence and big data on consumer markets, focusing on digital advertising and new forms of consumer fraud. Finally, it explores regulatory responses to these challenges, including data privacy laws and prohibitions on dark patterns, which illustrate the trade-offs inherent in consumer protection frameworks.

cs.CY

Section 230: A Juridical History

Section 230 of the Communications Decency Act of 1996 is the most important law in the history of the internet. It is also one of the most flawed. Under Section 230, online entities are absolutely immune from lawsuits related to content authored by third parties. The law has been essential to the internet's development over the last twenty years, but it has not kept pace with the times and is now a source of deep consternation to courts and legislatures. Lawmakers and legal scholars from across the political spectrum praise the law for what it has done, while criticizing its protection of bad-actor websites and obstruction of internet law reform. Absent from the fray, however, has been the Supreme Court, which has never issued a decision interpreting Section 230. That is poised to change, as the Court now appears determined to peel back decades of lower court case law and interpret the statute afresh to account for the tremendous technological advances of the last two decades. Rather than offer a proposal for reform, of which there are plenty, this Article acts as a guidebook to reformers by examining how we got to where we are today. It identifies those interpretive steps and missteps by which courts constructed an immunity doctrine insufficiently resilient against technological change, with the aim of aiding lawmakers and scholars in crafting an immunity doctrine better situated to accommodate future innovation.

cs.CY

The Patterns of Digital Deception

Current consumer-protection debates focus on the powerful new data-analysis techniques that have disrupted the balance of power between companies and their customers. Online tracking enables sellers to amass troves of historical data, apply machine-learning tools to construct detailed customer profiles, and target those customers with tailored offers that best suit their interests. It is often a win-win. Sellers avoid pumping dud products and consumers see ads for things they actually want to buy. But the same tools are also used for ill -- to target vulnerable members of the population with scams specially tailored to prey on their weaknesses. The result has been a dramatic rise in online fraud that disproportionately impacts those least able to bear the loss. The law's response has been technology centric. Lawmakers race to identify those technologies that drive consumer deception and target them for regulatory restrictions. But that approach comes at a major cost. General-purpose data-analysis and communications tools have both desirable and undesirable uses, and uniform restrictions on their use impede the good along with the bad. A superior approach would focus not on the technological tools of deception but on what this Article identifies as the legal patterns of digital deception -- those aspects of digital technology that have outflanked the law's existing mechanisms for redressing consumer harm. This Article reorients the discussion from the power of new technologies to the shortcomings in existing regulatory structures that have allowed for their abuse. Focus on these patterns of deception will allow regulators to reallocate resources to offset those shortcomings and thereby enhance efforts to combat online fraud without impeding technological innovation.

econ.GN

Beyond Social Media Analogues

The steady flow of social-media cases toward the Supreme Court shows a nation reworking its fundamental relationship with technology. The cases raise a host of questions ranging from difficult to impossible: how to nurture a vibrant public square when a few tech giants dominate the flow of information, how social media can be at the same time free from conformist groupthink and also protected against harmful disinformation campaigns, and how government and industry can cooperate on such problems without devolving toward censorship. To such profound questions, this Essay offers a comparatively modest contribution -- what not to do. Always the lawyer's instinct is toward analogy, considering what has come before and how it reveals what should come next. Almost invariably, that is the right choice. The law's cautious evolution protects society from disruptive change. But almost is not always, and, with social media, disruptive change is already upon us. Using social-media laws from Texas and Florida as a case study, this Essay shows how social-media's distinct features render it poorly suited to analysis by analogy and argues that courts should instead shift their attention toward crafting legal doctrines targeted to address social media's unique ills.

econ.GN

LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

The advent of large language models (LLMs) and their adoption by the legal community has given rise to the question: what types of legal reasoning can LLMs perform? To enable greater study of this question, we present LegalBench: a collaboratively constructed legal reasoning benchmark consisting of 162 tasks covering six different types of legal reasoning. LegalBench was built through an interdisciplinary process, in which we collected tasks designed and hand-crafted by legal professionals. Because these subject matter experts took a leading role in construction, tasks either measure legal reasoning capabilities that are practically useful, or measure reasoning skills that lawyers find interesting. To enable cross-disciplinary conversations about LLMs in the law, we additionally show how popular legal frameworks for describing legal reasoning -- which distinguish between its many forms -- correspond to LegalBench tasks, thus giving lawyers and LLM developers a common vocabulary. This paper describes LegalBench, presents an empirical evaluation of 20 open-source and commercial LLMs, and illustrates the types of research explorations LegalBench enables.

cs.CL

Privately Policing Dark Patterns

Lawmakers around the country are crafting new laws to target "dark patterns" -- user interface designs that trick or coerce users into enabling cell phone location tracking, sharing browsing data, initiating automatic billing, or making whatever other choices their designers prefer. Dark patterns pose a serious problem. In their most aggressive forms, they interfere with human autonomy, undermine customers' evaluation and selection of products, and distort online markets for goods and services. Yet crafting legislation is a major challenge: Persuasion and deception are difficult to distinguish, and shifting tech trends present an ever-moving target. To address these challenges, this Article proposes leveraging state private law to define and track dark patterns as they evolve. Judge-crafted decisional law can respond quickly to new techniques, flexibly define the boundary between permissible and impermissible designs, and bolster state and federal regulatory enforcement efforts by quickly identifying those designs that most undermine user autonomy.

cs.CY

Big Tech's Tightening Grip on Internet Speech

Online platforms have completely transformed American social life. They have democratized publication, overthrown old gatekeepers, and given ordinary Americans a fresh voice in politics. But the system is beginning to falter. Control over online speech lies in the hands of a select few -- Facebook, Google, and Twitter -- who moderate content for the entire nation. It is an impossible task. Americans cannot even agree among themselves what speech should be permitted. And, more importantly, platforms have their own interests at stake: Fringe theories and ugly name-calling drive away users. Moderation is good for business. But platform beautification has consequences for society's unpopular members, whose unsightly voices are silenced in the process. With control over online speech so centralized, online outcasts are left with few avenues for expression. Concentrated private control over important resources is an old problem. Last century, for example, saw the rise of railroads and telephone networks. To ensure access, such entities are treated as common carriers and required to provide equal service to all comers. Perhaps the same should be true for social media. This Essay responds to recent calls from Congress, the Supreme Court, and academia arguing that, like common carriers, online platforms should be required to carry all lawful content. The Essay studies users' and platforms' competing expressive interests, analyzes problematic trends in platforms' censorship practices, and explores the costs of common-carrier regulation before ultimately proposing market expansion and segmentation as an alternate pathway to avoid the economic and social costs of common-carrier regulation.

econ.GN

Toward Textual Internet Immunity

Internet immunity doctrine is broken. Under Section 230 of the Communications Decency Act of 1996, online entities are absolutely immune from lawsuits related to content authored by third parties. The law has been essential to the internet's development over the last twenty years, but it has not kept pace with the times and is now deeply flawed. Democrats demand accountability for online misinformation. Republicans decry politically motivated censorship. And Congress, President Biden, the Department of Justice, and the Federal Communications Commission all have their own plans for reform. Absent from the fray, however -- until now -- has been the Supreme Court, which has never issued a decision interpreting Section 230. That appears poised to change, however, following Justice Thomas's statement in Malwarebytes v. Enigma in which he urges the Court to prune back decades of lower-court precedent to craft a more limited immunity doctrine. This Essay discusses how courts' zealous enforcement of the early internet's free-information ethos gave birth to an expansive immunity doctrine, warns of potential pitfalls to reform, and explores what a narrower, text-focused doctrine might mean for the tech industry.

econ.GN

Rebooting Internet Immunity

We do everything online. We shop, travel, invest, socialize, and even hold garage sales. Even though we may not care whether a company operates online or in the physical world, however, the question has dramatic consequences for the companies themselves. Online and offline entities are governed by different rules. Under Section 230 of the Communications Decency Act, online entities -- but not physical-world entities -- are immune from lawsuits related to content authored by their users or customers. As a result, online entities have been able to avoid claims for harms caused by their negligence and defective product designs simply because they operate online. The reason for the disparate treatment is the internet's dramatic evolution over the last two decades. The internet of 1996 served as an information repository and communications channel and was well governed by Section 230, which treats internet entities as another form of mass media: Because Facebook, Twitter and other online companies could not possibly review the mass of content that flows through their systems, Section 230 immunizes them from claims related to user content. But content distribution is not the internet's only function, and it is even less so now than it was in 1996. The internet also operates as a platform for the delivery of real-world goods and services and requires a correspondingly diverse immunity doctrine. This Article proposes refining online immunity by limiting it to claims that threaten to impose a content-moderation burden on internet defendants. Where a claim is preventable other than by content moderation -- for example, by redesigning an app or website -- a plaintiff could freely seek relief, just as in the physical world. This approach empowers courts to identify culpable actors in the virtual world and treat like conduct alike wherever it occurs.

econ.GN

An Interpretive Framework for Narrower Immunity Under Section 230 of the Communications Decency Act

Almost all courts to interpret Section 230 of the Communications Decency Act have construed its ambiguously worded immunity provision broadly, shielding Internet intermediaries from tort liability so long as they are not the literal authors of offensive content. Although this broad interpretation effects the basic goals of the statute, it ignores several serious textual difficulties and mistakenly extends protection too far by immunizing even direct participants in tortuous conduct. This analysis, which examines the text and history of Section 230 in light of two strains of pre-Internet vicarious liability defamation doctrine, concludes that the immunity provision of Section 230, though broad, was not intended to abrogate entirely traditional common law notions of vicarious liability. Some bases of vicarious liability remain, and their continuing validity both explains the textual puzzles courts have faced in applying Section 230 and undergirds the push by a small minority of courts to narrow the section's immunity provision.

cs.CY

An Empirical Study of Obstacle Preemption in the Supreme Court

The Supreme Court's federal preemption decisions are notoriously unpredictable. Traditional left-right voting alignments break down in the face of competing ideological pulls. The breakdown of predictable voting blocs leaves the business interests most affected by federal preemption uncertain of the scope of potential liability to injured third parties and unsure even of whether state or federal law will be applied to future claims. This empirical analysis of the Court's decisions over the last fifteen years sheds light on the Court's unique voting alignments in obstacle preemption cases. A surprising anti-obstacle preemption coalition is forming as Justice Thomas gradually positions himself alongside the Court's liberals to form a five-justice voting bloc opposing obstacle preemption.

econ.GN

Calibrating Chevron for Preemption

Now almost three decades since its seminal Chevron decision, the Supreme Court has yet to articulate how that case's doctrine of deference to agency statutory interpretations relates to one of the most compelling federalism issues of our time: regulatory preemption of state law. Should courts defer to preemptive agency interpretations under Chevron, or do preemption's federalism implications demand a less deferential approach? Commentators have provided no shortage of possible solutions, but thus far the Court has resisted all of them. This Article makes two contributions to the debate. First, through a detailed analysis of the Court's recent agency-preemption decisions, I trace its hesitancy to adopt any of the various proposed rules to its high regard for congressional intent where areas of traditional state sovereignty are at risk. Recognizing that congressional intent to delegate preemptive authority varies from case to case, the Court has hesitated to adopt an across-the-board rule. Any such rule would constrain the Court and risk mismatch with congressional intent -- a risk it accepts under Chevron generally but which it finds particularly troublesome in the delicate area of federal preemption. Second, building on this previously underappreciated factor in the Court's analysis, I suggest a novel solution of variable deference that avoids the inflexibility inherent in an across-the-board rule while providing greater predictability than the Court's current haphazard approach. The proposed rule would grant full Chevron-style deference in those cases where congressional delegative intent is most likely -- where Congress has expressly preempted some state law and the agency interpretation merely resolves preemptive scope -- while withholding deference in those cases where Congress has remained completely silent as to preemption and delegative intent is least likely.

econ.GN

Chevron's Sliding Scale in Wyeth v. Levine, 129 S. Ct. 1187 (2009)

In Wyeth v. Levine the Supreme Court once again failed to reconcile the interpretive presumption against preemption with the sometimes competing Chevron doctrine of deference to agencies' reasonable statutory interpretations. Rather than resolve the issue of which principle should govern where the two principles point toward opposite results, the Court continued its recent practice of applying both principles halfheartedly, carving exceptions, and giving neither its proper weight. This analysis situates Wyeth within the larger framework of the Court's recent preemption decisions in an effort to explain the Court's hesitancy to resolve the conflict. The analysis concludes that the Court, motivated by its strong respect for congressional intent and concern to protect federalism, applies both the presumption against preemption and the Chevron doctrine on a sliding scale. Where congressional intent to preempt is clear and vague only as to scope, the Court is usually quite deferential to agency determinations, but where congressional preemptive intent is unclear, agency views are accorded less weight. The Court's variable approach to deference is defensible as necessary to prevent unauthorized incursion into areas of traditional state sovereignty, but its inherent unpredictability sows confusion among regulated parties, and the need for flexibility prevents the Court from adopting any of the more predictable across-the-board approaches to deference proposed by the Court's critics. A superior approach would combine the Court's concern for federalism with the certainty of a bright-line rule by granting deference to agency views where Congress has spoken via a preemption clause of ambiguous scope and no deference where Congress has remained silent.

econ.GN

A Computational Analysis of Oral Argument in the Supreme Court

As the most public component of the Supreme Court's decision-making process, oral argument receives an out-sized share of attention in the popular media. Despite its prominence, however, the basic function and operation of oral argument as an institution remains poorly understood, as political scientists and legal scholars continue to debate even the most fundamental questions about its role. Past study of oral argument has tended to focus on discrete, quantifiable attributes of oral argument, such as the number of questions asked to each advocate, the party of the Justices' appointing president, or the ideological implications of the case on appeal. Such studies allow broad generalizations about oral argument and judicial decision making: Justices tend to vote in accordance with their ideological preferences, and they tend to ask more questions when they are skeptical of a party's position. But they tell us little about the actual goings on at oral argument -- the running dialog between Justice and advocate that is the heart of the institution. This Article fills that void, using machine learning techniques to, for the first time, construct predictive models of judicial decision making based not on oral argument's superficial features or on factors external to oral argument, such as where the case falls on a liberal-conservative spectrum, but on the actual content of the oral argument itself -- the Justices' questions to each side. The resultant models offer an important new window into aspects of oral argument that have long resisted empirical study, including the Justices' individual questioning styles, how each expresses skepticism, and which of the Justices' questions are most central to oral argument dialog.

cs.CY