SearcharxivSearch

arXiv subjects

John Nay

Publications and source records attributed to John Nay.

4 recordsLinked to original sources

Language Models can Subtly Deceive Without Lying: A Case Study on Strategic Phrasing in Legislation

We explore the ability of large language models (LLMs) to engage in subtle deception through strategically phrasing and intentionally manipulating information. This harmful behavior can be hard to detect, unlike blatant lying or unintentional hallucination. We build a simple testbed mimicking a legislative environment where a corporate \textit{lobbyist} module is proposing amendments to bills that benefit a specific company while evading identification of this benefactor. We use real-world legislative bills matched with potentially affected companies to ground these interactions. Our results show that LLM lobbyists can draft subtle phrasing to avoid such identification by strong LLM-based detectors. Further optimization of the phrasing using LLM-based re-planning and re-sampling increases deception rates by up to 40 percentage points. Our human evaluations to verify the quality of deceptive generations and their retention of self-serving intent show significant coherence with our automated metrics and also help in identifying certain strategies of deceptive phrasing. This study highlights the risk of LLMs' capabilities for strategic phrasing through seemingly neutral language to attain self-serving goals. This calls for future research to uncover and protect against such subtle deception.

cs.CL

LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

The advent of large language models (LLMs) and their adoption by the legal community has given rise to the question: what types of legal reasoning can LLMs perform? To enable greater study of this question, we present LegalBench: a collaboratively constructed legal reasoning benchmark consisting of 162 tasks covering six different types of legal reasoning. LegalBench was built through an interdisciplinary process, in which we collected tasks designed and hand-crafted by legal professionals. Because these subject matter experts took a leading role in construction, tasks either measure legal reasoning capabilities that are practically useful, or measure reasoning skills that lawyers find interesting. To enable cross-disciplinary conversations about LLMs in the law, we additionally show how popular legal frameworks for describing legal reasoning -- which distinguish between its many forms -- correspond to LegalBench tasks, thus giving lawyers and LLM developers a common vocabulary. This paper describes LegalBench, presents an empirical evaluation of 20 open-source and commercial LLMs, and illustrates the types of research explorations LegalBench enables.

cs.CL

Climate-Contingent Finance

Climate adaptation could yield significant benefits. However, the uncertainty of which future climate scenarios will occur decreases the feasibility of proactively adapting. Climate adaptation projects could be underwritten by benefits paid for in the climate scenarios that each adaptation project is designed to address because other entities would like to hedge the financial risk of those scenarios. Because the return on investment is a function of the level of climate change, it is optimal for the adapting entity to finance adaptation with repayment as a function of the climate. It is also optimal for entities with more financial downside under a more extreme climate to serve as an investing counterparty because they can obtain higher than market rates of return when they need it most. In this way, parties proactively adapting would reduce the risk they over-prepare, while their investors would reduce the risk they under-prepare. This is superior to typical insurance because, by investing in climate-contingent mechanisms, investors are not merely financially hedging but also outright preventing physical damage, and therefore creating economic value. This coordinates capital through time and place according to parties' risk reduction capabilities and financial profiles, while also providing a diversifying investment return. Climate-contingent finance can be generalized to any situation where entities share exposure to a risk where they lack direct control over whether it occurs (e.g., climate change, or a natural pandemic), and one type of entity can take proactive actions to benefit from addressing the effects of the risk if it occurs (e.g., through innovating on crops that would do well under extreme climate change or vaccination technology that could address particular viruses) with funding from another type of entity that seeks a targeted return to ameliorate the downside.

q-fin.GN

Aligning Artificial Intelligence with Humans through Public Policy

Given that Artificial Intelligence (AI) increasingly permeates our lives, it is critical that we systematically align AI objectives with the goals and values of humans. The human-AI alignment problem stems from the impracticality of explicitly specifying the rewards that AI models should receive for all the actions they could take in all relevant states of the world. One possible solution, then, is to leverage the capabilities of AI models to learn those rewards implicitly from a rich source of data describing human values in a wide range of contexts. The democratic policy-making process produces just such data by developing specific rules, flexible standards, interpretable guidelines, and generalizable precedents that synthesize citizens' preferences over potential actions taken in many states of the world. Therefore, computationally encoding public policies to make them legible to AI systems should be an important part of a socio-technical approach to the broader human-AI alignment puzzle. This Essay outlines research on AI that learn structures in policy data that can be leveraged for downstream tasks. As a demonstration of the ability of AI to comprehend policy, we provide a case study of an AI system that predicts the relevance of proposed legislation to any given publicly traded company and its likely effect on that company. We believe this represents the "comprehension" phase of AI and policy, but leveraging policy as a key source of human values to align AI requires "understanding" policy. Solving the alignment problem is crucial to ensuring that AI is beneficial both individually (to the person or group deploying the AI) and socially. As AI systems are given increasing responsibility in high-stakes contexts, integrating democratically-determined policy into those systems could align their behavior with human goals in a way that is responsive to a constantly evolving society.

cs.CY