SearcharxivSearch

arXiv subjects

Newton Cheng

Publications and source records attributed to Newton Cheng.

12 recordsLinked to original sources

Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Humans are capable of strategically deceptive behavior: behaving helpfully in most situations, but then behaving very differently in order to pursue alternative objectives when given the opportunity. If an AI system learned such a deceptive strategy, could we detect it and remove it using current state-of-the-art safety training techniques? To study this question, we construct proof-of-concept examples of deceptive behavior in large language models (LLMs). For example, we train models that write secure code when the prompt states that the year is 2023, but insert exploitable code when the stated year is 2024. We find that such backdoor behavior can be made persistent, so that it is not removed by standard safety training techniques, including supervised fine-tuning, reinforcement learning, and adversarial training (eliciting unsafe behavior and then training to remove it). The backdoor behavior is most persistent in the largest models and in models trained to produce chain-of-thought reasoning about deceiving the training process, with the persistence remaining even when the chain-of-thought is distilled away. Furthermore, rather than removing backdoors, we find that adversarial training can teach models to better recognize their backdoor triggers, effectively hiding the unsafe behavior. Our results suggest that, once a model exhibits deceptive behavior, standard techniques could fail to remove such deception and create a false impression of safety.

cs.CR

Towards Understanding Sycophancy in Language Models

Human feedback is commonly utilized to finetune AI assistants. But human feedback may also encourage model responses that match user beliefs over truthful ones, a behaviour known as sycophancy. We investigate the prevalence of sycophancy in models whose finetuning procedure made use of human feedback, and the potential role of human preference judgments in such behavior. We first demonstrate that five state-of-the-art AI assistants consistently exhibit sycophancy across four varied free-form text-generation tasks. To understand if human preferences drive this broadly observed behavior, we analyze existing human preference data. We find that when a response matches a user's views, it is more likely to be preferred. Moreover, both humans and preference models (PMs) prefer convincingly-written sycophantic responses over correct ones a non-negligible fraction of the time. Optimizing model outputs against PMs also sometimes sacrifices truthfulness in favor of sycophancy. Overall, our results indicate that sycophancy is a general behavior of state-of-the-art AI assistants, likely driven in part by human preference judgments favoring sycophantic responses.

cs.CL

Specific versus General Principles for Constitutional AI

Human feedback can prevent overtly harmful utterances in conversational models, but may not automatically mitigate subtle problematic behaviors such as a stated desire for self-preservation or power. Constitutional AI offers an alternative, replacing human feedback with feedback from AI models conditioned only on a list of written principles. We find this approach effectively prevents the expression of such behaviors. The success of simple principles motivates us to ask: can models learn general ethical behaviors from only a single written principle? To test this, we run experiments using a principle roughly stated as "do what's best for humanity". We find that the largest dialogue models can generalize from this short constitution, resulting in harmless assistants with no stated interest in specific motivations like power. A general principle may thus partially avoid the need for a long list of constitutions targeting potentially harmful behaviors. However, more detailed constitutions still improve fine-grained control over specific types of harms. This suggests both general and specific principles have value for steering AI safely.

cs.CL

Question Decomposition Improves the Faithfulness of Model-Generated Reasoning

As large language models (LLMs) perform more difficult tasks, it becomes harder to verify the correctness and safety of their behavior. One approach to help with this issue is to prompt LLMs to externalize their reasoning, e.g., by having them generate step-by-step reasoning as they answer a question (Chain-of-Thought; CoT). The reasoning may enable us to check the process that models use to perform tasks. However, this approach relies on the stated reasoning faithfully reflecting the model's actual reasoning, which is not always the case. To improve over the faithfulness of CoT reasoning, we have models generate reasoning by decomposing questions into subquestions. Decomposition-based methods achieve strong performance on question-answering tasks, sometimes approaching that of CoT while improving the faithfulness of the model's stated reasoning on several recently-proposed metrics. By forcing the model to answer simpler subquestions in separate contexts, we greatly increase the faithfulness of model-generated reasoning over CoT, while still achieving some of the performance gains of CoT. Our results show it is possible to improve the faithfulness of model-generated reasoning; continued improvements may lead to reasoning that enables us to verify the correctness and safety of LLM behavior.

cs.CL

Measuring Faithfulness in Chain-of-Thought Reasoning

Large language models (LLMs) perform better when they produce step-by-step, "Chain-of-Thought" (CoT) reasoning before answering a question, but it is unclear if the stated reasoning is a faithful explanation of the model's actual reasoning (i.e., its process for answering the question). We investigate hypotheses for how CoT reasoning may be unfaithful, by examining how the model predictions change when we intervene on the CoT (e.g., by adding mistakes or paraphrasing it). Models show large variation across tasks in how strongly they condition on the CoT when predicting their answer, sometimes relying heavily on the CoT and other times primarily ignoring it. CoT's performance boost does not seem to come from CoT's added test-time compute alone or from information encoded via the particular phrasing of the CoT. As models become larger and more capable, they produce less faithful reasoning on most tasks we study. Overall, our results suggest that CoT can be faithful if the circumstances such as the model size and task are carefully chosen.

cs.AI

Random tensor networks with nontrivial links

Random tensor networks are a powerful toy model for understanding the entanglement structure of holographic quantum gravity. However, unlike holographic quantum gravity, their entanglement spectra are flat. It has therefore been argued that a better model consists of random tensor networks with link states that are not maximally entangled, i.e., have nontrivial spectra. In this work, we initiate a systematic study of the entanglement properties of these networks. We employ tools from free probability, random matrix theory, and one-shot quantum information theory to study random tensor networks with bounded and unbounded variation in link spectra, and in cases where a subsystem has one or multiple minimal cuts. If the link states have bounded spectral variation, the limiting entanglement spectrum of a subsystem with two minimal cuts can be expressed as a free product of the entanglement spectra of each cut, along with a Marchenko-Pastur distribution. For a class of states with unbounded spectral variation, analogous to semiclassical states in quantum gravity, we relate the limiting entanglement spectrum of a subsystem with two minimal cuts to the distribution of the minimal entanglement across the two cuts. In doing so, we draw connections to previous work on split transfer protocols, entanglement negativity in random tensor networks, and Euclidean path integrals in quantum gravity.

quant-ph

A Gap Between the Hypergraph and Stabilizer Entropy Cones

It was recently found that the stabilizer and hypergraph entropy cones coincide for four parties, leading to a conjecture of their equivalence at higher party numbers. In this note, we show this conjecture to be false by proving new inequalities obeyed by all hypergraph entropy vectors that exclude particular stabilizer states on six qubits. By further leveraging this connection, we improve the characterization of stabilizer entropies and show that all linear rank inequalities at five parties, except for classical monotonicity, form facets of the stabilizer cone. Additionally, by studying minimum cuts on hypergraphs, we prove some structural properties of hypergraph representations of entanglement and generalize the notion of entanglement wedge nesting in holography.

quant-ph

Topological Link Models of Multipartite Entanglement

We introduce a novel model of multipartite entanglement based on topological links, generalizing the graph/hypergraph entropy cone program. We demonstrate that there exist link representations of entropy vectors which provably cannot be represented by graphs or hypergraphs. Furthermore, we show that the contraction map proof method generalizes to the topological setting, though now requiring oracular solutions to well-known but difficult problems in knot theory.

quant-ph

Optimized Correlation Measures in Holography

We consider a class of correlation measures for quantum states called optimized correlation measures, defined as a minimization of a linear combination of von Neumann entropies over purifications of a given state. Examples include the entanglement of purification $E_P$ and squashed entanglement $E_{\text{sq}}$. We show that when evaluating such measures on ``nice" holographic states in the large-$N$ limit, the optimal purification has a semi-classical geometric dual. We then apply this result to confirm several holographic dual proposals, including the $n$-party squashed entanglement. Moreover, our result suggests two new techniques for determining holographic duals: holographic entropy inequalities and direct optimization of the dual geometry.

hep-th

The Quantum Entropy Cone of Hypergraphs

In this work, we generalize the graph-theoretic techniques used for the holographic entropy cone to study hypergraphs and their analogously-defined entropy cone. This allows us to develop a framework to efficiently compute entropies and prove inequalities satisfied by hypergraphs. In doing so, we discover a class of quantum entropy vectors which reach beyond those of holographic states and obey constraints intimately related to the ones obeyed by stabilizer states and linear ranks. We show that, at least up to 4 parties, the hypergraph cone is identical to the stabilizer entropy cone, thus demonstrating that the hypergraph framework is broadly applicable to the study of entanglement entropy. We conjecture that this equality continues to hold for higher party numbers and report on partial progress on this direction. To physically motivate this conjectured equivalence, we also propose a plausible method inspired by tensor networks to construct a quantum state from a given hypergraph such that their entropy vectors match.

quant-ph

Multipartite Reflected Entropy

We discuss two methods that, through a combination of cyclically gluing copies of a given $n$-party boundary state in AdS/CFT and a canonical purification, creates a bulk geometry that contains a boundary homologous minimal surface with area equal to 2 or 4 times the $n$-party entanglement wedge cross-section, depending on the parity of the party number and choice of method. The areas of the minimal surfaces are each dual to entanglement entropies that we define to be candidates for the $n$-party reflected entropy. In the context of AdS$_3$/CFT$_2$, we provide a boundary interpretation of our construction as a multiboundary wormhole, and conjecture that this interpretation generalizes to higher dimensions.

hep-th

Eigenstate Thermalization Hypothesis and Approximate Quantum Error Correction

The eigenstate thermalization hypothesis (ETH) is a powerful conjecture for understanding how statistical mechanics emerges in a large class of many-body quantum systems. It has also been interpreted in a CFT context, and, in particular, holographic CFTs are expected to satisfy ETH. Recently, it was observed that the ETH condition corresponds to a necessary and sufficient condition for an approximate quantum error correcting code (AQECC), implying the presence of AQECCs in systems satisfying ETH. In this paper, we explore the properties of ETH as an error correcting code and show that there exists an explicit universal recovery channel for the code. Based on the analysis, we discuss a generalization that all chaotic theories contain error correcting codes. We then specialize to AdS/CFT to demonstrate the possibility of total bulk reconstruction in black holes with a well-defined macroscopic geometry. When combined with the existing AdS/CFT error correction story, this shows that black holes are enormously robust against erasure errors.

hep-th