SearcharxivSearch

arXiv subjects

Julian Schulz

Publications and source records attributed to Julian Schulz.

12 recordsLinked to original sources

Demonstrating Generalization Failures via Mixtures of Conditional Policies

Post-training of frontier language models is conducted on curated task suites, and inevitably leaves a distribution shift between training and deployment environments. This exposes developers to generalization failures, which are relatively poorly understood. To better understand such generalization failures, we believe the community should construct clean demonstrations under simplified conditions. To facilitate this, we propose a simple and flexible way to construct language models which fail to generalize in controllable ways when subsequently trained with Reinforcement Learning (RL) on a given distribution of training tasks. Our construction uses Supervised Fine-Tuning on a dataset of a mixture of transcripts corresponding to a collection of 'conditional policies', which can each independently be assigned certain behaviors on each different task distribution, to obtain a model that is then well approximated as a 'mixture of conditional policies.' We observe that RL training then selects for policies that obtain the highest reward on the training distribution. This can produce striking behaviors: in a controlled setting with two distributions containing identical questions prepended with two different 'trigger strings', RL training on either distribution actively degrades performance on the other to zero, even though the underlying task is identical. We also use our construction to illustrate two novel ways in which generalization may fail in future language models, corresponding to distribution shifts of task coverage and temporal context respectively. While our construction is deliberately simple and may not closely resemble 'natural' generalization failures, the resulting 'model organisms' are of interest for alignment stress-testing and generalization science, and can be used as existence proofs that training success and generalization can come apart in structured ways.

cs.AI

3D micro-printing: An enabling technique for arbitrary potential landscapes for photonic quantum-gases

Photonic quantum gases explore the physics of open driven-dissipative quantum systems under ambient conditions and thus open access to thermodynamics and transport phenomena in quantum gases in the weakly interacting regime. Here we introduce the technology of 3D micro-printing to create potential landscapes for photonic quantum gases in dye-filled micro cavities, which surpass the current state of the art in terms of potential size and definition, potential depth, coupling strength, and number of coupled potentials by at least an order of magnitude. We realize as demonstration of the capabilities box potentials with rectangular side walls, anisotropic harmonic potentials, double-well potentials with dimensions on the scale of the wavelength of light as well as potential lattices with topological non-trivial properties. This approach paves the way for experimentally studying the physics of open quantum systems on lattices and might find applications in solving complex ground-state problems like the XY-model.

quant-ph

A Concrete Roadmap towards Safety Cases based on Chain-of-Thought Monitoring

As AI systems approach dangerous capability levels where inability safety cases become insufficient, we need alternative approaches to ensure safety. This paper presents a roadmap for constructing safety cases based on chain-of-thought (CoT) monitoring in reasoning models and outlines our research agenda. We argue that CoT monitoring might support both control and trustworthiness safety cases. We propose a two-part safety case: (1) establishing that models lack dangerous capabilities when operating without their CoT, and (2) ensuring that any dangerous capabilities enabled by a CoT are detectable by CoT monitoring. We systematically examine two threats to monitorability: neuralese and encoded reasoning, which we categorize into three forms (linguistic drift, steganography, and alien reasoning) and analyze their potential drivers. We evaluate existing and novel techniques for maintaining CoT faithfulness. For cases where models produce non-monitorable reasoning, we explore the possibility of extracting a monitorable CoT from a non-monitorable CoT. To assess the viability of CoT monitoring safety cases, we establish prediction markets to aggregate forecasts on key technical milestones influencing their feasibility.

cs.LG

Orbital Frontiers: Harnessing Higher Modes in Photonic Simulators

Photonic platforms have emerged as versatile and powerful classical simulators of quantum dynamics, providing clean, controllable optical analogs of extended structured (i.e., crystalline) electronic systems. While most realizations to date have used only the fundamental mode in each site, recent advances in structured light - particularly the use of higher-order spatial modes, including those with orbital angular momentum - are enabling richer dynamics and new functionalities. These additional degrees of freedom facilitate the emulation of phenomena ranging from topological band structures and synthetic gauge fields to orbitronics. In this perspective, we discuss how exploiting the internal structure of higher-order modes is reshaping the scope and capabilities of photonic platforms for simulating quantum phenomena.

physics.optics

Polarization properties of photon Bose-Einstein condensates

The first experimental realization of a photon Bose-Einstein condensate was demonstrated more than a decade ago. However, the polarization of the condensate has not been fully understood and measured in this weakly driven-dissipative system. In this letter, we experimentally investigate the polarization of thermal and condensed light depending on the power and polarization of the pump beam. With full control over the polarization of the pump, it is possible to create arbitrary states on the surface of the Poincar\'e sphere. We show that, in agreement with previous theoretical work, there is a remarkable increase in the polarization strength of the condensate above the threshold for a fully linearly polarized pump. Above a certain threshold also the degenerate orthogonal polarized state of the cavity is occupied, limiting the degree of polarization to approximately 90%.

quant-ph

Automated Feature Labeling with Token-Space Gradient Descent

We present a novel approach to feature labeling using gradient descent in token-space. While existing methods typically use language models to generate hypotheses about feature meanings, our method directly optimizes label representations by using a language model as a discriminator to predict feature activations. We formulate this as a multi-objective optimization problem in token-space, balancing prediction accuracy, entropy minimization, and linguistic naturalness. Our proof-of-concept experiments demonstrate successful convergence to interpretable single-token labels across diverse domains, including features for detecting animals, mammals, Chinese text, and numbers. Although our current implementation is constrained to single-token labels and relatively simple features, the results suggest that token-space gradient descent could become a valuable addition to the interpretability researcher's toolkit.

cs.LG

Steering Llama 2 via Contrastive Activation Addition

We introduce Contrastive Activation Addition (CAA), an innovative method for steering language models by modifying their activations during forward passes. CAA computes "steering vectors" by averaging the difference in residual stream activations between pairs of positive and negative examples of a particular behavior, such as factual versus hallucinatory responses. During inference, these steering vectors are added at all token positions after the user's prompt with either a positive or negative coefficient, allowing precise control over the degree of the targeted behavior. We evaluate CAA's effectiveness on Llama 2 Chat using multiple-choice behavioral question datasets and open-ended generation tasks. We demonstrate that CAA significantly alters model behavior, is effective over and on top of traditional methods like finetuning and system prompt design, and minimally reduces capabilities. Moreover, we gain deeper insights into CAA's mechanisms by employing various activation space interpretation methods. CAA accurately steers model outputs and sheds light on how high-level concepts are represented in Large Language Models (LLMs).

cs.CL

Dimensional Crossover in a Quantum Gas of Light

The dimensionality of a system profoundly influences its physical behaviour, leading to the emergence of different states of matter in many-body quantum systems. In lower dimensions, fluctuations increase and lead to the suppression of long-range order. For example, in bosonic gases, Bose-Einstein condensation (BEC) in one dimension requires stronger confinement than in two dimensions. We experimentally study the properties of a harmonically trapped photon gas undergoing Bose-Einstein condensation along the dimensional crossover from one to two dimensions. The photons are trapped in a dye microcavity where polymer nanostructures provide the trapping potential for the photon gas. By varying the aspect ratio of the harmonic trap, we tune from an isotropic two-dimensional confinement to an anisotropic, highly elongated one-dimensional trapping potential. Along this transition we determine the caloric properties of the photon gas and find a softening of the second-order Bose-Einstein condensation phase transition observed in two dimensions to a crossover behaviour in one dimension.

cond-mat.quant-gas

Broadband mode division multiplexing of OAM-modes by a micro printed waveguide structure

A light beam carrying orbital angular momentum (OAM) is characterized by a helical phase-front that winds around the center of the beam. These beams have unique properties that have found numerous applications. In the field of data transmission, they represent a degree of freedom that could potentially increase capacity by a factor of several distinct OAM modes. While an efficient method for (de)composing beams based on their OAM exists for free-space optics, a device capable of performing this (de)composition in an integrated, compact fiber application without the use of external active optical elements and for multiple OAM modes simultaneously has not been reported. In this study, a waveguide structure is presented that can serve as a broadband OAM (de)multiplexer. The structure design is based on the adiabatic principle used in photonic lanterns for highly efficient conversion of spatially separated single modes into eigenmodes of a few-mode fiber. In addition, an artificial magnetic field is introduced by twisting the structure during the adiabatic evolution, which removes the degeneracy between modes having the same absolute OAM. This structure can simplify, stabilize, and miniaturize the creation or decomposition of OAM beams, making them useful for various applications.

physics.optics

Photonic quadrupole topological insulator using orbital-induced synthetic flux

The rich physical properties of multiatomic molecules and crystalline structures are determined, to a significant extent, by the underlying geometry and connectivity of atomic orbitals. This orbital degree of freedom has also been used effectively to introduce structural diversity in a few synthetic materials including polariton lattices nonlinear photonic lattices and ultracold atoms in optical lattices. In particular, the mixing of orbitals with distinct parity representations, such as $s$ and $p$ orbitals, has been shown to be especially useful for generating systems that require alternating phase patterns, as with the sign of couplings within a lattice. Here we show that by further breaking the symmetries of such mixed-orbital lattices, it is possible to generate synthetic magnetic flux threading the lattice. This capability allows the generation of multipole higher-order topological phases in synthetic bosonic platforms, in which $π$ flux threading each plaquette of the lattice is required, and which to date have only been implemented using tailored connectivity patterns. We use this insight to experimentally demonstrate a quadrupole photonic topological insulator in a two-dimensional lattice of waveguides that leverage modes with both $s$ and $p$ orbital-type representations. We confirm the nontrivial quadrupole topology of the system by observing the presence of protected zero-dimensional states, which are spatially confined to the corners, and by confirming that these states sit at the band gap. Our approach is also applicable to a broader range of time-reversal-invariant synthetic materials that do not allow for tailored connectivity, e.g. with nanoscale geometries, and in which synthetic fluxes are essential.

physics.optics

Existence of a negative next-nearest-neighbor coupling in evanescently coupled dielectric waveguides

We experimentally demonstrate that the next-nearest-neighbor(NNN)coupling in an array of waveguides can naturally be negative. To do so, dielectric zig-zag shaped waveguide arrays are fabricated with direct laser writing (DLW). By changing the angle of the zig-zag shape it is possible to tune between positive and negative ratios of nearest and next-nearest-neighbor coupling, which also allows to reduce the impact of the NNN-coupling to zero at the correct respective angle. We describe how the correct higher order coupling constants in tight-binding models can be derived, based on non-orthogonal coupled mode theory. We confirm the existence of negative NNN-couplings experimentally and show the improved accuracy of this refined tight-binding model. The negative NNN-coupling has a noticeable impact especially when higher order coupling terms can no longer be neglected. Our results are also of importance for other discrete systems in which the tight-binding model is often used.

physics.optics

Generalized Laws of Refraction and Reflection at Interfaces between Different Photonic Artificial Gauge Fields

Artificial gauge fields enable extending the control over dynamics of uncharged particles, by engineering the potential landscape such that the particles behave as if effective external fields are acting on them. Recent years have witnessed a growing interest in artificial gauge fields that are generated either by geometry or by time-dependent modulation, as they have been the enablers for topological phenomena and synthetic dimensions in many physical settings, e.g., photonics, cold atoms and acoustic waves. Here, we formulate and experimentally demonstrate the generalized laws of refraction and reflection from an interface between two regions with different artificial gauge fields. We use the symmetries in the system to obtain the generalized Snell law for such a gauge interface, and solve for reflection and transmission. We identify total internal reflection (TIR) and complete transmission, and demonstrate the concept in experiments. Additionally, we calculate the artificial magnetic flux at the interface of two regions with different artificial gauge, and present a method to concatenate several gauge interfaces. As an example, we propose a scheme to make a gauge imaging system - a device that is able to reconstruct (image) the shape of an arbitrary wavepacket launched at a certain position to a predesigned location.

physics.optics