SearcharxivSearch

arXiv subjects

Nikolas Tezak

Publications and source records attributed to Nikolas Tezak.

13 recordsLinked to original sources

Efficient Training of Language Models to Fill in the Middle

We show that autoregressive language models can learn to infill text after we apply a straightforward transformation to the dataset, which simply moves a span of text from the middle of a document to its end. While this data augmentation has garnered much interest in recent years, we provide extensive evidence that training models with a large fraction of data transformed in this way does not harm the original left-to-right generative capability, as measured by perplexity and sampling evaluations across a wide range of scales. Given the usefulness, simplicity, and efficiency of training models to fill-in-the-middle (FIM), we suggest that future autoregressive language models be trained with FIM by default. To this end, we run a series of ablations on key hyperparameters, such as the data transformation frequency, the structure of the transformation, and the method of selecting the infill span. We use these ablations to prescribe strong default settings and best practices to train FIM models. We have released our best infilling model trained with best practices in our API, and release our infilling benchmarks to aid future research.

cs.CL

Text and Code Embeddings by Contrastive Pre-Training

Text embeddings are useful features in many applications such as semantic search and computing text similarity. Previous work typically trains models customized for different use cases, varying in dataset choice, training objective and model architecture. In this work, we show that contrastive pre-training on unsupervised data at scale leads to high quality vector representations of text and code. The same unsupervised text embeddings that achieve new state-of-the-art results in linear-probe classification also display impressive semantic search capabilities and sometimes even perform competitively with fine-tuned models. On linear-probe classification accuracy averaging over 7 tasks, our best unsupervised model achieves a relative improvement of 4% and 1.8% over previous best unsupervised and supervised text embedding models respectively. The same text embeddings when evaluated on large-scale semantic search attains a relative improvement of 23.4%, 14.7%, and 10.6% over previous best unsupervised methods on MSMARCO, Natural Questions and TriviaQA benchmarks, respectively. Similarly to text embeddings, we train code embedding models on (text, code) pairs, obtaining a 20.8% relative improvement over prior best work on code search.

cs.CL

Evaluating Large Language Models Trained on Code

We introduce Codex, a GPT language model fine-tuned on publicly available code from GitHub, and study its Python code-writing capabilities. A distinct production version of Codex powers GitHub Copilot. On HumanEval, a new evaluation set we release to measure functional correctness for synthesizing programs from docstrings, our model solves 28.8% of the problems, while GPT-3 solves 0% and GPT-J solves 11.4%. Furthermore, we find that repeated sampling from the model is a surprisingly effective strategy for producing working solutions to difficult prompts. Using this method, we solve 70.2% of our problems with 100 samples per problem. Careful investigation of our model reveals its limitations, including difficulty with docstrings describing long chains of operations and with binding operations to variables. Finally, we discuss the potential broader impacts of deploying powerful code generation technologies, covering safety, security, and economics.

cs.LG

Solving Rubik's Cube with a Robot Hand

We demonstrate that models trained only in simulation can be used to solve a manipulation problem of unprecedented complexity on a real robot. This is made possible by two key components: a novel algorithm, which we call automatic domain randomization (ADR) and a robot platform built for machine learning. ADR automatically generates a distribution over randomized environments of ever-increasing difficulty. Control policies and vision state estimators trained with ADR exhibit vastly improved sim2real transfer. For control policies, memory-augmented models trained on an ADR-generated distribution of environments show clear signs of emergent meta-learning at test time. The combination of ADR with our custom robot platform allows us to solve a Rubik's cube with a humanoid robot hand, which involves both control and state estimation problems. Videos summarizing our results are available: https://openai.com/blog/solving-rubiks-cube/

cs.LG

Low dimensional manifolds for exact representation of open quantum systems

Weakly nonlinear degrees of freedom in dissipative quantum systems tend to localize near manifolds of quasi-classical states. We present a family of analytical and computational methods for deriving optimal unitary model transformations based on representations of finite dimensional Lie groups. The transformations are optimal in that they minimize the quantum relative entropy distance between a given state and the quasi-classical manifold. This naturally splits the description of quantum states into quasi-classical coordinates that specify the nearest quasi-classical state and a transformed quantum state that can be represented in fewer basis levels. We derive coupled equations of motion for the coordinates and the transformed state and demonstrate how this can be exploited for efficient numerical simulation. Our optimization objective naturally quantifies the non-classicality of states occurring in some given open system dynamics. This allows us to compare the intrinsic complexity of different open quantum systems.

quant-ph

All-mechanical quantum noise cancellation for accelerometry: broadband with momentum measurements, narrow band without

We show that the ability to make direct measurements of momentum, in addition to the usual direct measurements of position, allows a simple configuration of two identical mechanical oscillators to be used for broadband back-action-free force metrology. This would eliminate the need for an optical reference oscillator in the scheme of Tsang and Caves [Phys. Rev. Lett. 105, 123601 (2010)], along with its associated disadvantages. We also show that if one is restricted to position measurements alone then two copies of the same two-oscillator configuration can be used for narrow-band back-action-free force metrology.

quant-ph

A Coherent Perceptron for All-Optical Learning

We present nonlinear photonic circuit models for constructing programmable linear transformations and use these to realize a coherent Perceptron, i.e., an all-optical linear classifier capable of learning the classification boundary iteratively from training data through a coherent feedback rule. Through extensive semi-classical stochastic simulations we demonstrate that the device nearly attains the theoretical error bound for a model classification problem.

quant-ph

Quantum noise in large-scale coherent nonlinear photonic circuits

A semiclassical simulation approach is presented for studying quantum noise in large-scale photonic circuits incorporating an ideal Kerr nonlinearity. A circuit solver is used to generate matrices defining a set of stochastic differential equations, in which the resonator field variables represent random samplings of the Wigner quasi-probability distributions. Although the semiclassical approach involves making a large-photon-number approximation, tests on one- and two-resonator circuits indicate satisfactory agreement between the semiclassical and full-quantum simulation results in the parameter regime of interest. The semiclassical model is used to simulate random errors in a large-scale circuit that contains 88 resonators and hundreds of components in total, and functions as a 4-bit ripple counter. The error rate as a function of on-state photon number is examined, and it is observed that the quantum fluctuation amplitudes do not increase as signals propagate through the circuit, an important property for scalability.

quant-ph

Squeezed light in an optical parametric oscillator network with coherent feedback quantum control

We present squeezing and anti-squeezing spectra of the output from a degenerate optical parametric oscillator (OPO) network arranged in different coherent quantum feedback configurations. One OPO serves as a quantum plant, the other as a quantum controller. The addition of coherent feedback enables shaping of the output squeezing spectrum of the plant, and is found to be capable of pushing the frequency of maximum squeezing away from the optical driving frequency and broadening the spectrum over a wider frequency band. The experimental results are in excellent agreement with the developed theory, and illustrate the use of coherent quantum feedback to engineer the quantum-optical properties of the plant OPO output.

quant-ph

Transformation of quantum photonic circuit models by term rewriting

The development of practical methods for synthesis and verification of complex photonic circuits presents a grand challenge for the nascent field of quantum engineering. Of course, classical electrical engineering provides essential foundations and serves to illustrate the degree of sophistication that can be achieved in automated circuit design. In this paper we explore the utility of term rewriting approaches to the transformation of quantum circuit models, specifically applying rewrite rules for both reduction/verification and robustness analysis of photonic circuits for autonomous quantum error correction. We outline a workflow for quantum photonic circuit analysis that leverages the Modelica framework for multi-domain physical modeling, which parallels a previously described approach based on VHSIC Hardware Description Language (VHDL).

quant-ph

Specification of photonic circuits using Quantum Hardware Description Language

Following the simple observation that the interconnection of a set of quantum optical input-output devices can be specified using structural mode VHSIC Hardware Description Language (VHDL), we demonstrate a computer-aided schematic capture workflow for modeling and simulating multi-component photonic circuits. We describe an algorithm for parsing circuit descriptions to derive quantum equations of motion, illustrate our approach using simple examples based on linear and cavity-nonlinear optical components, and demonstrate a computational approach to hierarchical model reduction.

quant-ph

Spectral properties of finite laser-driven lattices of ultracold Rydberg atoms

We investigate the spectral properties of a finite laser-driven lattice of ultracold Rydberg atoms exploiting the dipole blockade effect in the frozen Rydberg gas regime. Uniform one-dimensional lattices as well as lattices with variable spacings are considered. In the case of a weak laser coupling, we find a multitude of many-body Rydberg states with well-defined excitation properties which are adiabatically accessible starting from the ground state. A comprehensive analysis of the degeneracies of the spectrum as well as of the single and pair excitations numbers of the eigenstates is performed. In the strong laser regime, analytical solutions for the pseudo-fermionic eigenmodes are derived. Perturbative energy corrections for this approximative approach are provided.

physics.atom-ph

Rydberg-Rydberg interaction profile from the excitation dynamics of ultracold atoms in lattices

We propose a method for the determination of the interaction potential of Rydberg atoms. Specifically, we consider a laser-driven Rydberg gas confined in a one-dimensional lattice and demonstrate that the Rydberg atom number after a laser excitation cycle as a function of the laser detuning provides a measure for the Rydberg interaction coefficient. With the lattice spacing precisely known, the proposed scheme only relies on the measurement of the number of Rydberg atoms and thus circumvents the necessity to map the interaction potential by varying the interparticle separation.

quant-ph