SearcharxivSearch

arXiv subjects

Timothy Maxwell

Publications and source records attributed to Timothy Maxwell.

7 recordsLinked to original sources

Towards Understanding Sycophancy in Language Models

Human feedback is commonly utilized to finetune AI assistants. But human feedback may also encourage model responses that match user beliefs over truthful ones, a behaviour known as sycophancy. We investigate the prevalence of sycophancy in models whose finetuning procedure made use of human feedback, and the potential role of human preference judgments in such behavior. We first demonstrate that five state-of-the-art AI assistants consistently exhibit sycophancy across four varied free-form text-generation tasks. To understand if human preferences drive this broadly observed behavior, we analyze existing human preference data. We find that when a response matches a user's views, it is more likely to be preferred. Moreover, both humans and preference models (PMs) prefer convincingly-written sycophantic responses over correct ones a non-negligible fraction of the time. Optimizing model outputs against PMs also sometimes sacrifices truthfulness in favor of sycophancy. Overall, our results indicate that sycophancy is a general behavior of state-of-the-art AI assistants, likely driven in part by human preference judgments favoring sycophantic responses.

cs.CL

Measuring Faithfulness in Chain-of-Thought Reasoning

Large language models (LLMs) perform better when they produce step-by-step, "Chain-of-Thought" (CoT) reasoning before answering a question, but it is unclear if the stated reasoning is a faithful explanation of the model's actual reasoning (i.e., its process for answering the question). We investigate hypotheses for how CoT reasoning may be unfaithful, by examining how the model predictions change when we intervene on the CoT (e.g., by adding mistakes or paraphrasing it). Models show large variation across tasks in how strongly they condition on the CoT when predicting their answer, sometimes relying heavily on the CoT and other times primarily ignoring it. CoT's performance boost does not seem to come from CoT's added test-time compute alone or from information encoded via the particular phrasing of the CoT. As models become larger and more capable, they produce less faithful reasoning on most tasks we study. Overall, our results suggest that CoT can be faithful if the circumstances such as the model size and task are carefully chosen.

cs.AI

Multipoint-BAX: A New Approach for Efficiently Tuning Particle Accelerator Emittance via Virtual Objectives

Although beam emittance is critical for the performance of high-brightness accelerators, optimization is often time limited as emittance calculations, commonly done via quadrupole scans, are typically slow. Such calculations are a type of $\textit{multipoint query}$, i.e. each query requires multiple secondary measurements. Traditional black-box optimizers such as Bayesian optimization are slow and inefficient when dealing with such objectives as they must acquire the full series of measurements, but return only the emittance, with each query. We propose a new information-theoretic algorithm, Multipoint-BAX, for black-box optimization on multipoint queries, which queries and models individual beam-size measurements using techniques from Bayesian Algorithm Execution (BAX). Our method avoids the slow multipoint query on the accelerator by acquiring points through a $\textit{virtual objective}$, i.e. calculating the emittance objective from a fast learned model rather than directly from the accelerator. We use Multipoint-BAX to minimize emittance at the Linac Coherent Light Source (LCLS) and the Facility for Advanced Accelerator Experimental Tests II (FACET-II). In simulation, our method is 20$\times$ faster and more robust to noise compared to existing methods. In live tests, it matched the hand-tuned emittance at FACET-II and achieved a 24% lower emittance than hand-tuning at LCLS. Our method represents a conceptual shift for optimizing multipoint queries, and we anticipate that it can be readily adapted to similar problems in particle accelerators and other scientific instruments.

physics.acc-ph

Site-specific Interrogation of an Ionic Chiral Fragment During Photolysis Using an X-ray Free-Electron Laser

Short-wavelength free-electron lasers with their ultrashort pulses at high intensities have originated new approaches for tracking molecular dynamics from the vista of specific sites. X-ray pump X-ray probe schemes even allow to address individual atomic constituents with a 'trigger'-event that preludes the subsequent molecular dynamics while being able to selectively probe the evolving structure with a time-delayed second X-ray pulse. Here, we use a linearly polarized X-ray photon to trigger the photolysis of a prototypical chiral molecule, namely trifluoromethyloxirane (C$_3$H$_3$F$_3$O), at the fluorine K-edge at around 700 eV. The evolving fluorine-containing fragments are then probed by a second, circularly polarized X-ray pulse of higher photon energy in order to investigate the chemically shifted inner-shell electrons of the ionic motherfragment for their stereochemical sensitivity. We experimentally demonstrate and theoretically support how two-color X-ray pump X-ray probe experiments with polarization control enable XFELs as tools for chiral recognition.

physics.chem-ph

Scientific Opportunities with an X-ray Free-Electron Laser Oscillator

An X-ray free-electron laser oscillator (XFELO) is a new type of hard X-ray source that would produce fully coherent pulses with meV bandwidth and stable intensity. The XFELO complements existing sources based on self-amplified spontaneous emission (SASE) from high-gain X-ray free-electron lasers (XFEL) that produce ultra-short pulses with broad-band chaotic spectra. This report is based on discussions of scientific opportunities enabled by an XFELO during a workshop held at SLAC on June 29 - July 1, 2016

physics.ins-det

Measurements of wake-induced electron beam deflection in a dechirper at the Linac Coherent Light Source

The RadiaBeam/SLAC dechirper, a structure consisting of pairs of flat, metallic, corrugated plates, %a corrugated structure in flat geometry, has been installed just upstream of the undulators in the Linac Coherent Light Source (LCLS). As a dechirper, with the beam passing between the plates on axis, longitudinal wakefields are induced that can remove unwanted energy chirp in the beam. However, with the beam passing off axis, strong transverse wakes are also induced. This mode of operation has already been used for the production of intense, multi-color photon beams using the Fresh-Slice technique, and is being used to develop a diagnostic for attosecond bunch length measurements. Here we measure, as function of offset, the strength of the transverse wakefields that are excited between the two plates, and also for the case of the beam passing near to a single plate. We compare with analytical formulas from the literature, and find good agreement. This report presents the first systematic measurements of the transverse wake strength in a dechirper, one that has been excited by a bunch with the short pulse duration and high energy found in an X-ray free electron laser.

physics.acc-ph

Synchronization and Characterization of an Ultra-Short Laser for Photoemission and Electron-Beam Diagnostics Studies at a Radio Frequency Photoinjector

A commercially-available titanium-sapphire laser system has recently been installed at the Fermilab A0 photoinjector laboratory in support of photoemission and electron beam diagnostics studies. The laser system is synchronized to both the 1.3-GHz master oscillator and a 1-Hz signal use to trigger the radiofrequency system and instrumentation acquisition. The synchronization scheme and performance are detailed. Long-term temporal and intensity drifts are identified and actively suppressed to within 1 ps and 1.5%, respectively. Measurement and optimization of the laser's temporal profile are accomplished using frequency-resolved optical gating.

physics.acc-ph