SearcharxivSearch

arXiv subjects

Sandipan Kundu

Publications and source records attributed to Sandipan Kundu.

At least 19 recordsLinked to original sources

Specific versus General Principles for Constitutional AI

Human feedback can prevent overtly harmful utterances in conversational models, but may not automatically mitigate subtle problematic behaviors such as a stated desire for self-preservation or power. Constitutional AI offers an alternative, replacing human feedback with feedback from AI models conditioned only on a list of written principles. We find this approach effectively prevents the expression of such behaviors. The success of simple principles motivates us to ask: can models learn general ethical behaviors from only a single written principle? To test this, we run experiments using a principle roughly stated as "do what's best for humanity". We find that the largest dialogue models can generalize from this short constitution, resulting in harmless assistants with no stated interest in specific motivations like power. A general principle may thus partially avoid the need for a long list of constitutions targeting potentially harmful behaviors. However, more detailed constitutions still improve fine-grained control over specific types of harms. This suggests both general and specific principles have value for steering AI safely.

cs.CL

Measuring Faithfulness in Chain-of-Thought Reasoning

Large language models (LLMs) perform better when they produce step-by-step, "Chain-of-Thought" (CoT) reasoning before answering a question, but it is unclear if the stated reasoning is a faithful explanation of the model's actual reasoning (i.e., its process for answering the question). We investigate hypotheses for how CoT reasoning may be unfaithful, by examining how the model predictions change when we intervene on the CoT (e.g., by adding mistakes or paraphrasing it). Models show large variation across tasks in how strongly they condition on the CoT when predicting their answer, sometimes relying heavily on the CoT and other times primarily ignoring it. CoT's performance boost does not seem to come from CoT's added test-time compute alone or from information encoded via the particular phrasing of the CoT. As models become larger and more capable, they produce less faithful reasoning on most tasks we study. Overall, our results suggest that CoT can be faithful if the circumstances such as the model size and task are carefully chosen.

cs.AI

The Capacity for Moral Self-Correction in Large Language Models

We test the hypothesis that language models trained with reinforcement learning from human feedback (RLHF) have the capability to "morally self-correct" -- to avoid producing harmful outputs -- if instructed to do so. We find strong evidence in support of this hypothesis across three different experiments, each of which reveal different facets of moral self-correction. We find that the capability for moral self-correction emerges at 22B model parameters, and typically improves with increasing model size and RLHF training. We believe that at this level of scale, language models obtain two capabilities that they can use for moral self-correction: (1) they can follow instructions and (2) they can learn complex normative concepts of harm like stereotyping, bias, and discrimination. As such, they can follow instructions to avoid certain kinds of morally harmful outputs. We believe our results are cause for cautious optimism regarding the ability to train language models to abide by ethical principles.

cs.CL

Discovering Language Model Behaviors with Model-Written Evaluations

As language models (LMs) scale, they develop many novel behaviors, good and bad, exacerbating the need to evaluate how they behave. Prior work creates evaluations with crowdwork (which is time-consuming and expensive) or existing data sources (which are not always available). Here, we automatically generate evaluations with LMs. We explore approaches with varying amounts of human effort, from instructing LMs to write yes/no questions to making complex Winogender schemas with multiple stages of LM-based generation and filtering. Crowdworkers rate the examples as highly relevant and agree with 90-100% of labels, sometimes more so than corresponding human-written datasets. We generate 154 datasets and discover new cases of inverse scaling where LMs get worse with size. Larger LMs repeat back a dialog user's preferred answer ("sycophancy") and express greater desire to pursue concerning goals like resource acquisition and goal preservation. We also find some of the first examples of inverse scaling in RL from Human Feedback (RLHF), where more RLHF makes LMs worse. For example, RLHF makes LMs express stronger political views (on gun rights and immigration) and a greater desire to avoid shut down. Overall, LM-written evaluations are high-quality and let us quickly discover many novel LM behaviors.

cs.CL

Constitutional AI: Harmlessness from AI Feedback

As AI systems become more capable, we would like to enlist their help to supervise other AIs. We experiment with methods for training a harmless AI assistant through self-improvement, without any human labels identifying harmful outputs. The only human oversight is provided through a list of rules or principles, and so we refer to the method as 'Constitutional AI'. The process involves both a supervised learning and a reinforcement learning phase. In the supervised phase we sample from an initial model, then generate self-critiques and revisions, and then finetune the original model on revised responses. In the RL phase, we sample from the finetuned model, use a model to evaluate which of the two samples is better, and then train a preference model from this dataset of AI preferences. We then train with RL using the preference model as the reward signal, i.e. we use 'RL from AI Feedback' (RLAIF). As a result we are able to train a harmless but non-evasive AI assistant that engages with harmful queries by explaining its objections to them. Both the SL and RL methods can leverage chain-of-thought style reasoning to improve the human-judged performance and transparency of AI decision making. These methods make it possible to control AI behavior more precisely and with far fewer human labels.

cs.CL

Measuring Progress on Scalable Oversight for Large Language Models

Developing safe and useful general-purpose AI systems will require us to make progress on scalable oversight: the problem of supervising systems that potentially outperform us on most skills relevant to the task at hand. Empirical work on this problem is not straightforward, since we do not yet have systems that broadly exceed our abilities. This paper discusses one of the major ways we think about this problem, with a focus on ways it can be studied empirically. We first present an experimental design centered on tasks for which human specialists succeed but unaided humans and current general AI systems fail. We then present a proof-of-concept experiment meant to demonstrate a key feature of this experimental design and show its viability with two question-answering tasks: MMLU and time-limited QuALITY. On these tasks, we find that human participants who interact with an unreliable large-language-model dialog assistant through chat -- a trivial baseline strategy for scalable oversight -- substantially outperform both the model alone and their own unaided performance. These results are an encouraging sign that scalable oversight will be tractable to study with present models and bolster recent findings that large language models can productively assist humans with difficult tasks.

cs.HC

Snowmass Theory Frontier: Effective Field Theory

We summarize recent progress in the development, application, and understanding of effective field theories and highlight promising directions for future research. This Report is prepared as the TF02 "Effective Field Theory" topical group summary for the Theory Frontier as part of the Snowmass 2021 process.

hep-ph

Snowmass White Paper: UV Constraints on IR Physics

Fundamental principles of local quantum field theory or of quantum gravity can enforce consistency requirements on the space of consistent low-energy effective field theories. We survey the various techniques that have been used to put UV constraints on IR physics, including those from causality considerations in the form of S-matrix positivity and bootstrap bounds, scattering time delays, conformal field theory and holographic methods, together with those that arise from landscape/swampland criteria such as the weak gravity conjecture. We review recent applications of these constraints to corrections to Standard Model physics, corrections to Einstein gravity, and cosmological theories and highlight promising future directions.

hep-th

Extremal Chaos

In maximally chaotic quantum systems, a class of out-of-time-order correlators (OTOCs) saturate the Maldacena-Shenker-Stanford (MSS) bound on chaos. Recently, it has been shown that the same OTOCs must also obey an infinite set of (subleading) constraints in any thermal quantum system with a large number of degrees of freedom. In this paper, we find a unique analytic extension of the maximally chaotic OTOC that saturates all the subleading chaos bounds which allow saturation. This extremally chaotic OTOC has the feature that information of the initial perturbation is recovered at very late times. Furthermore, we argue that the extremally chaotic OTOC provides a Källen-Lehmann-type representation for all OTOCs. This representation enables the identification of all analytic completions of maximal chaos as small deformations of extremal chaos in a precise way.

hep-th

Subleading Bounds on Chaos

Chaos, in quantum systems, can be diagnosed by certain out-of-time-order correlators (OTOCs) that obey the chaos bound of Maldacena, Shenker, and Stanford (MSS). We begin by deriving a dispersion relation for this class of OTOCs, implying that they must satisfy many more constraints beyond the MSS bound. Motivated by this observation, we perform a systematic analysis obtaining an infinite set of constraints on the OTOC. This infinite set includes the MSS bound as the leading constraint. In addition, it also contains subleading bounds that are highly constraining, especially when the MSS bound is saturated by the leading term. These new bounds, among other things, imply that the MSS bound cannot be exactly saturated over any duration of time, however short. Furthermore, we derive a sharp bound on the Lyapunov exponent $λ_2 \le \frac{6π}β$ of the subleading correction to maximal chaos.

hep-th

RG Flows with Global Symmetry Breaking and Bounds from Chaos

We discuss general aspects of renormalization group (RG) flows between two conformal fixed points in 4d with a broken continuous global symmetry in the UV. Every such RG flow can be described in terms of the dynamics of Nambu-Goldstone bosons of broken conformal and global symmetries. We derive the low-energy effective action that describes this class of RG flows from basic symmetry principles. We view the theory of Nambu-Goldstone bosons as a theory in anti-de Sitter space with the flat space limit. This enables an equivalent CFT$_3$ formulation of these 4d RG flows in terms of spectral deformations of a generalized free CFT$_3$. We utilize this dual description to impose further constraints on the low energy effective action associated with unitary RG flows in 4d by invoking the chaos bound in 3d. This approach naturally provides a set of independent monotonically decreasing $C$-functions for 4d RG flows with global symmetry breaking by explicitly relating 4d $C$-functions with certain out-of-time-order correlators that diagnose chaos in 3d. We also comment on a more general connection between RG and chaos in QFT.

hep-th

Swampland Conditions for Higher Derivative Couplings from CFT

There are effective field theories that cannot be embedded in any UV complete theory. We consider scalar effective field theories, with and without dynamical gravity, in $D$-dimensional anti-de Sitter (AdS) spacetime with large radius and derive precise bounds (analytically) on the coupling constants of higher derivative interactions $ϕ^2\Box^kϕ^2$ by only requiring that the dual CFT obeys the standard conformal bootstrap axioms. In particular, we show that all such coupling constants, for even $k\ge 2$, must satisfy positivity, monotonicity, and log-convexity conditions in the absence of dynamical gravity. Inclusion of gravity only affects constraints involving the $ϕ^2\Box^2ϕ^2$ interaction which now can have a negative coupling constant. Our CFT setup is a Lorentzian four-point correlator in the Regge limit. We also utilize this setup to derive constraints on effective field theories of multiple scalars. We argue that similar analysis should impose nontrivial constraints on the graviton four-point scattering amplitude in AdS.

hep-th

EFT of 6D SUSY RG Flows

Motivated by its potential use in constraining the structure of 6D renormalization group flows, we determine the low energy dilaton-axion effective field theory of conformal and global symmetry breaking in 6D conformal field theories (CFTs). While our analysis is largely independent of supersymmetry, we also investigate the case of 6D superconformal field theories (SCFTs), where we use the effective action to present a streamlined proof of the 6D a-theorem for tensor branch flows, as well as to constrain properties of Higgs branch and mixed branch flows. An analysis of Higgs branch flows in some examples leads us to conjecture that in 6D SCFTs, an interacting dilaton effective theory may be possible even when certain 4-dilaton 4-derivative interaction terms vanish, because of large momentum modifications to 4-point dilaton scattering amplitudes. This possibility is due to the fact that in all known $D > 4$ CFTs, the approach to a conformal fixed point involves effective strings which are becoming tensionless.

hep-th

Causality Constraints in Large $N$ QCD Coupled to Gravity

Confining gauge theories contain glueballs and mesons with arbitrary spin, and these particles become metastable at large $N$. However, metastable higher spin particles, when coupled to gravity, are in conflict with causality. This tension can be avoided only if the gravitational interaction is accompanied by interactions involving other higher spin states well below the Planck scale $M_{\rm pl}$. These higher spin states can come from either the QCD sector or the gravity sector, but both these resolutions have some surprising implications. For example, QCD states can resolve the problem since there is a non-trivial mixing between the QCD sector and the gravity sector, requiring all particles to interact with glueballs at tree-level. If gravity sector states restore causality, any weakly coupled UV completion of the gravity-sector must have many stringy features, with an upper bound on the string scale. Under the assumption that gravity is weakly coupled, both scenarios imply that the theory has a stringy description above $N\gtrsim \frac{M_{\rm pl}}{Λ_{\rm QCD}}$, where $Λ_{\rm QCD}$ is the confinement scale.

hep-th

Closed Strings and Weak Gravity from Higher-Spin Causality

We combine old and new quantum field theoretic arguments to show that any theory of stable or metastable higher spin particles can be coupled to gravity only when the gravity sector has a stringy structure. Metastable higher spin particles, free or interacting, cannot couple to gravity while preserving causality unless there exist higher spin states in the gravitational sector much below the Planck scale $M_{\rm pl}$. We obtain an upper bound on the mass $Λ_{\rm gr}$ of the lightest higher spin particle in the gravity sector in terms of quantities in the non-gravitational sector. We invoke the CKSZ uniqueness theorem to argue that any weakly coupled UV completion of such a theory must have a gravity sector containing infinite towers of asymptotically parallel, equispaced, and linear Regge trajectories. Consequently, gravitational four-point scattering amplitudes must coincide with the closed string four-point amplitude for $s,t\gg1$, identifying $Λ_{\rm gr}$ as the string scale. Our bound also implies that all metastable higher spin particles in 4d with masses $m\ll Λ_{\rm gr}$ must satisfy a weak gravity condition.

hep-th

Renormalization Group Flows, the $a$-Theorem and Conformal Bootstrap

Every renormalization group flow in $d$ spacetime dimensions can be equivalently described as spectral deformations of a generalized free CFT in $(d-1)$ spacetime dimensions. This can be achieved by studying the effective action of the Nambu-Goldstone boson of broken conformal symmetry in anti-de Sitter space and then taking the flat space limit. This approach is particularly useful in even spacetime dimension where the change in the Euler anomaly $ a_{UV}-a_{IR}$ can be related to anomalous dimensions of lowest twist multi-trace operators in the dual CFT. As an application, we provide a simple proof of the 4d $a$-theorem using the dual description. Furthermore, we reinterpret the statement of the $a$-theorem in 6d as a conformal bootstrap problem in 5d.

hep-th

A Generalized Nachtmann Theorem in CFT

Correlators of unitary quantum field theories in Lorentzian signature obey certain analyticity and positivity properties. For interacting unitary CFTs in more than two dimensions, we show that these properties impose general constraints on families of minimal twist operators that appear in the OPEs of primary operators. In particular, we rederive and extend the convexity theorem which states that for the family of minimal twist operators with even spins appearing in the reflection-symmetric OPE of any scalar primary, twist must be a monotonically increasing convex function of the spin. Our argument is completely non-perturbative and it also applies to the OPE of nonidentical scalar primaries in unitary CFTs, constraining the twist of spinning operators appearing in the OPE. Finally, we argue that the same methods also impose constraints on the Regge behavior of certain CFT correlators.

hep-th

A Species or Weak-Gravity Bound for Large $N$ Gauge Theories Coupled to Gravity

Causality constrains the gravitational interactions of massive higher spin particles in both AdS and flat spacetime. We explore the extent to which these constraints apply to composite particles, explaining why they do not rule out macroscopic objects or hydrogen atoms. However, we find that they do apply to glueballs and mesons in confining large $N$ gauge theories. Assuming such theories contain massive bound states of general spin, we find parametric bounds in $(3+1)$ spacetime dimensions of the form $N\lesssim \frac{M_{Pl}}{Λ_{\text{QCD}}}$ relating $N$, the QCD scale, and the Planck scale. We also argue that a stronger bound replacing $Λ_{\text{QCD}}$ with the UV cut-off scale may be derived from eikonal scattering in flat spacetime.

hep-th