SearcharxivSearch

arXiv subjects

Eric Smith

Publications and source records attributed to Eric Smith.

At least 19 recordsLinked to original sources

Large Reasoning Models Learn Better Alignment from Flawed Thinking

Large reasoning models (LRMs) "think" by generating structured chain-of-thought (CoT) before producing a final answer, yet they still lack the ability to reason critically about safety alignment and are easily biased when a flawed premise is injected into their thought process. We propose RECAP (Robust Safety Alignment via Counter-Aligned Prefilling), a principled reinforcement learning (RL) method for post-training that explicitly teaches models to override flawed reasoning trajectories and reroute to safe and helpful responses. RECAP trains on a mixture of synthetically generated counter-aligned CoT prefills and standard prompts, requires no additional training cost or modifications beyond vanilla reinforcement learning from human feedback (RLHF), and substantially improves safety and jailbreak robustness, reduces overrefusal, and preserves core reasoning capability -- all while maintaining inference token budget. Extensive analysis shows that RECAP-trained models engage in self-reflection more frequently and remain robust under adaptive attacks, preserving safety even after repeated attempts to override their reasoning.

cs.LG

Thermodynamic ranking of pathways in reaction networks

One of the puzzles left open by energetic analyses of irreversible stochastic processes is that boundary conditions that prevent the performance of work or the dissipation of heat make no contribution to an entropy-production budget; yet we see ubiquitously in both engineered and living systems that both transient and persistent energy costs are paid to create and maintain such boundaries. We wish to know whether there are inherent limits for the costs of such phenomena, and common units in which those can be traded off against more familiar costs measured in terms of heat dissipation. We give this problem a concrete framing in the context of CRNs, for the problem of extracting a topologically restricted pathway from a larger distributed network, through activation of some reactions and selective elimination of others. We define a thermodynamic cost function for pathways derived from large-deviation theory of stochastic CRNs, which decomposes into two components: an ongoing maintenance cost to sustain a NESS, and a restriction cost, quantifying the ongoing improbability of neutralizing reactions outside the specified pathway. Applying this formalism to detailed-balanced CRNs in the linear response regime, we make use of their formal equivalence to electrical circuits. We prove that the resistance of a CRN decreases as reactions are added that support the throughput current, and that the maintenance cost, the restriction cost, and the thermodynamic cost of nested pathways are bounded below by those of their hosting network. For small CRNs, we show how catalytic and inhibitory mechanisms can drastically alter pathway costs, enabling unfavorable pathways to become favorable and approach the cost of the hosting pathway. Our results provide insights into the thermodynamic principles governing open CRNs and offer a foundation for understanding the evolution of metabolic networks.

q-bio.MN

Llama Guard 3-1B-INT4: Compact and Efficient Safeguard for Human-AI Conversations

This paper presents Llama Guard 3-1B-INT4, a compact and efficient Llama Guard model, which has been open-sourced to the community during Meta Connect 2024. We demonstrate that Llama Guard 3-1B-INT4 can be deployed on resource-constrained devices, achieving a throughput of at least 30 tokens per second and a time-to-first-token of 2.5 seconds or less on a commodity Android mobile CPU. Notably, our experiments show that Llama Guard 3-1B-INT4 attains comparable or superior safety moderation scores to its larger counterpart, Llama Guard 3-1B, despite being approximately 7 times smaller in size (440MB).

cs.DC

Llama Guard 3 Vision: Safeguarding Human-AI Image Understanding Conversations

We introduce Llama Guard 3 Vision, a multimodal LLM-based safeguard for human-AI conversations that involves image understanding: it can be used to safeguard content for both multimodal LLM inputs (prompt classification) and outputs (response classification). Unlike the previous text-only Llama Guard versions (Inan et al., 2023; Llama Team, 2024b,a), it is specifically designed to support image reasoning use cases and is optimized to detect harmful multimodal (text and image) prompts and text responses to these prompts. Llama Guard 3 Vision is fine-tuned on Llama 3.2-Vision and demonstrates strong performance on the internal benchmarks using the MLCommons taxonomy. We also test its robustness against adversarial attacks. We believe that Llama Guard 3 Vision serves as a good starting point to build more capable and robust content moderation tools for human-AI conversation with multimodal capabilities.

cs.CV

Using Counterexample Generation and Theory Exploration to Suggest Missing Hypotheses

Newcomers to ACL2 are sometimes surprised that ACL2 rejects formulas that they believe should be theorems, such as (REVERSE (REVERSE X)) = X. Experienced ACL2 users will recognize that the theorem only holds for intended values of X, and given ACL2's total logic, there are many counterexamples for which this formula is simply not true. Counterexample generation (cgen) is a technique that helps by giving the user a number of counterexamples (and also witnesses) to the formula, e.g., letting the user know that the intended theorem is false when X is equal to 10. In this paper we describe a tool called DrLA that goes further by suggesting additional hypotheses that will make the theorem true. In this case, for example, DrLA may suggest that X needs to be either a TRUE-LIST or a STRING. The suggestions are discovered using the ideas of theory exploration and subsumption from automated theorem proving.

cs.LO

Multilingual Holistic Bias: Extending Descriptors and Patterns to Unveil Demographic Biases in Languages at Scale

We introduce a multilingual extension of the HOLISTICBIAS dataset, the largest English template-based taxonomy of textual people references: MULTILINGUALHOLISTICBIAS. This extension consists of 20,459 sentences in 50 languages distributed across all 13 demographic axes. Source sentences are built from combinations of 118 demographic descriptors and three patterns, excluding nonsensical combinations. Multilingual translations include alternatives for gendered languages that cover gendered translations when there is ambiguity in English. Our benchmark is intended to uncover demographic imbalances and be the tool to quantify mitigations towards them. Our initial findings show that translation quality for EN-to-XX translations is an average of 8 spBLEU better when evaluating with the masculine human reference compared to feminine. In the opposite direction, XX-to-EN, we compare the robustness of the model when the source input only differs in gender (masculine or feminine) and masculine translations are an average of almost 4 spBLEU better than feminine. When embedding sentences to a joint multilingual sentence representations space, we find that for most languages masculine translations are significantly closer to the English neutral sentences when embedded.

cs.CL

Polyhedral geometry and combinatorics of an autocatalytic ecosystem

Developing a mathematical understanding of autocatalysis in reaction networks has both theoretical and practical implications. We review definitions of autocatalytic networks and prove some properties for minimal autocatalytic subnetworks (MASs). We show that it is possible to classify MASs in equivalence classes, and develop mathematical results about their behavior. We also provide linear-programming algorithms to exhaustively enumerate them and a scheme to visualize their polyhedral geometry and combinatorics. We then define cluster chemical reaction networks, a framework for coarse-graining real chemical reactions with positive integer conservation laws. We find that the size of the list of minimal autocatalytic subnetworks in a maximally connected cluster chemical reaction network with one conservation law grows exponentially in the number of species. We end our discussion with open questions concerning an ecosystem of autocatalytic subnetworks and multidisciplinary opportunities for future investigation.

q-bio.MN

Action Functional Gradient Descent algorithm for estimating escape paths in Stochastic Chemical Reaction Networks

We first derive the Hamilton-Jacobi theory underlying continuous-time Markov processes, and then use the construction to develop a variational algorithm for estimating escape (least improbable or first passage) paths for a generic stochastic chemical reaction network that exhibits multiple fixed points. The design of our algorithm is such that it is independent of the underlying dimensionality of the system, the discretization control parameters are updated towards the continuum limit, and there is an easy-to-calculate measure for the correctness of its solution. We consider several applications of the algorithm and verify them against computationally expensive means such as the shooting method and stochastic simulation. While we employ theoretical techniques from mathematical physics, numerical optimization and chemical reaction network theory, we hope that our work finds practical applications with an inter-disciplinary audience including chemists, biologists, optimal control theorists and game theorists.

cond-mat.stat-mech

Toxicity in Multilingual Machine Translation at Scale

Machine Translation systems can produce different types of errors, some of which are characterized as critical or catastrophic due to the specific negative impact that they can have on users. In this paper we focus on one type of critical error: added toxicity. We evaluate and analyze added toxicity when translating a large evaluation dataset (HOLISTICBIAS, over 472k sentences, covering 13 demographic axes) from English into 164 languages. An automatic toxicity evaluation shows that added toxicity across languages varies from 0% to 5%. The output languages with the most added toxicity tend to be low-resource ones, and the demographic axes with the most added toxicity include sexual orientation, gender and sex, and ability. We also perform human evaluation on a subset of 8 translation directions, confirming the prevalence of true added toxicity. We use a measurement of the amount of source contribution to the translation, where a low source contribution implies hallucination, to interpret what causes toxicity. Making use of the input attributions allows us to explain toxicity, because the source contributions significantly correlate with toxicity for 84% of languages studied. Given our findings, our recommendations to reduce added toxicity are to curate training data to avoid mistranslations, mitigate hallucination and check unstable translations.

cs.CL

Supercooling of the A phase of $^3$He

Because of the extreme purity, lack of disorder, and complex order parameter, the first-order superfluid $^3$He A-B transition is the leading model system for first order transitions in the early universe. Here we report on the path dependence of the supercooling of the A phase over a wide range of pressures below 29.3 bar at nearly zero magnetic field. The A phase can be cooled significantly below the thermodynamic A-B transition temperature. While the extent of supercooling is highly reproducible, it depends strongly upon the cooling trajectory: The metastability of the A phase is enhanced by transiting through regions where the A phase is more stable. We provide evidence that some of the additional supercooling is due to the elimination of B phase seeds formed upon passage through the superfluid transition. A greater understanding of the physics is essential before the $^3$He can be exploited to model transitions in the early universe.

cond-mat.other

Perturbation Augmentation for Fairer NLP

Unwanted and often harmful social biases are becoming ever more salient in NLP research, affecting both models and datasets. In this work, we ask whether training on demographically perturbed data leads to fairer language models. We collect a large dataset of human annotated text perturbations and train a neural perturbation model, which we show outperforms heuristic alternatives. We find that (i) language models (LMs) pre-trained on demographically perturbed corpora are typically more fair, and (ii) LMs finetuned on perturbed GLUE datasets exhibit less demographic bias on downstream tasks, and (iii) fairness improvements do not come at the expense of performance on downstream tasks. Lastly, we discuss outstanding questions about how best to evaluate the (un)fairness of large language models. We hope that this exploration of neural demographic perturbation will help drive more improvement towards fairer NLP.

cs.CL

Path-Dependent Supercooling of the $^3$He Superfluid A-B transition

We examine the discontinuous first-order superfluid $^3$He A to B transition in the vicinity of the polycritical point (2.232 mK and 21.22 bar). We find path-dependent transitions: cooling at fixed pressure yields a well defined transition line in the temperature-pressure plane, but this line can be reliably crossed by depressurizing at nearly constant temperature after transiting $T_{\rm c}$ at a higher pressure. This path dependence is not consistent with any of the standard B-phase nucleation mechanisms in the literature. This symmetry breaking transition is a potential simulator for first order transitions in the early universe.

cond-mat.supr-con

Source mass characterization in the ARIADNE axion experiment

The Axion Resonant InterAction Detection Experiment (ARIADNE) is a collaborative effort to search for the QCD axion using nuclear magnetic resonance (NMR), where the axion acts as a mediator of spin-dependent forces between an unpolarized tungsten source mass and a sample of polarized helium-3 gas. Since the experiment involves precision measurement of a small magnetization, it relies on limiting ordinary magnetic noise with superconducting magnetic shielding. In addition to the shielding, proper characterization of the noise level from other sources is crucial. We investigate one such noise source in detail: the magnetic noise due to impurities and Johnson noise in the tungsten source mass.

physics.ins-det

Batch-sequential design and heteroskedastic surrogate modeling for delta smelt conservation

Delta smelt is an endangered fish species in the San Francisco estuary that have shown an overall population decline over the past 30 years. Researchers have developed a stochastic, agent-based simulator to virtualize the system, with the goal of understanding the relative contribution of natural and anthropogenic factors suggested as playing a role in their decline. However, the input configuration space is high-dimensional, running the simulator is time-consuming, and its noisy outputs change nonlinearly in both mean and variance. Getting enough runs to effectively learn input--output dynamics requires both a nimble modeling strategy and parallel supercomputer evaluation. Recent advances in heteroskedastic Gaussian process (HetGP) surrogate modeling helps, but little is known about how to appropriately plan experiments for highly distributed simulator evaluation. We propose a batch sequential design scheme, generalizing one-at-a-time variance-based active learning for HetGP surrogates, as a means of keeping multi-core cluster nodes fully engaged with expensive runs. Our acquisition strategy is carefully engineered to favor selection of replicates which boost statistical and computational efficiencies when training surrogates to isolate signal in high noise regions. Design and modeling performance is illustrated on a range of toy examples before embarking on a large-scale smelt simulation campaign and downstream high-fidelity input sensitivity analysis.

stat.AP

Eikonal solutions for moment hierarchies of Chemical Reaction Networks in the limits of large particle number

Trajectory-based methods are well-developed to approximate steady-state probability distributions for stochastic processes in large-system limits. The trajectories are solutions to equations of motion of Hamiltonian dynamical systems, and are known as eikonals. They also express the leading flow lines along which probability currents balance. The existing eikonal methods for discrete-state processes including chemical reaction networks are based on the Liouville operator that evolves generating functions of the underlying probability distribution. We have previously derived a representation for the generators of such processes that acts directly in the hierarchy of moments of the distribution, rather than on the distribution itself or on its generating function. We show here how in the large-system limit the steady-state condition for that generator reduces to a mapping from eikonals to the ratios of neighboring factorial moments, as a function of the order $k$ of these moments. The construction shows that the boundary values for the moment hierarchy, and thus its whole solution, are anchored in the interior fixed points of the Hamiltonian system, a result familiar from Freidlin-Wenztell theory. The direct derivation of eikonals from the moment representation further illustrates the relation between coherent-state and number fields in Doi-Peliti theory, clarifying the role of canonical transformations in that theory.

cond-mat.stat-mech

Intrinsic and extrinsic thermodynamics for stochastic population processes with multi-level large-deviation structure

A set of core features is set forth as the essence of a thermodynamic description, which derive from large-deviation properties in systems with hierarchies of timescales, but which are \emph{not} dependent upon conservation laws or microscopic reversibility in the substrate hosting the process. The most fundamental elements are the concept of a macrostate in relation to the large-deviation entropy, and the decomposition of contributions to irreversibility among interacting subsystems, which is the origin of the dependence on a concept of heat in both classical and stochastic thermodynamics. A natural decomposition is shown to exist, into a relative entropy and a housekeeping entropy rate, which define respectively the \textit{intensive} thermodynamics of a system and an \textit{extensive} thermodynamic vector embedding the system in its context. Both intensive and extensive components are functions of Hartley information of the momentary system stationary state, which is information \emph{about} the joint effect of system processes on its contribution to irreversibility. Results are derived for stochastic Chemical Reaction Networks, including a Legendre duality for the housekeeping entropy rate to thermodynamically characterize fully-irreversible processes on an equal footing with those at the opposite limit of detailed-balance. The work is meant to encourage development of inherent thermodynamic descriptions for rule-based systems and the living state, which are not conceived as reductive explanations to heat flows.

cond-mat.stat-mech

The information geometry of 2-field functional integrals

2-field functional integrals (2FFI) are an important class of solution methods for generating functions of dissipative processes, including discrete-state stochastic processes, dissipative dynamical systems, and decohering quantum densities. The stationary trajectories of these integrals describe a conserved current by Liouville's theorem, despite the fact that there is no conserved phase space current in the underlying stochastic process. We develop the information geometry of generating functions for discrete-state classical stochastic processes in the Doi-Peliti 2FFI form, showing that the conserved current is a Fisher information between the underlying distribution of the process and the tilting weight of the generating function. To give an interpretation to the time invertibility implied by current conservation, we use generating functions to represent importance sampling protocols, and show that the conserved Fisher information is the differential of a sample volume under deformations of the nominal distribution and the likelihood ratio. We derive a new pair of dual Riemannian connections respecting the symplectic structure of transport along stationary rays that gives rise to Liouville's theorem, and show that dual flatness in the affine coordinates of the coherent-state basis captures the special role played by coherent states in many 2FFI theories. The covariant convective derivative under time translation correctly represents the geometric invariants of generating functions under canonical transformations of the 2FFI field variables of integration.

cond-mat.stat-mech

Path-reversal, Doi-Peliti generating functionals, and dualities between dynamics and inference for stochastic processes

Fluctuation theorems may be partitioned into those that apply the probability measure under the original stochastic process to reversed paths, and those that construct a new, adjoint measure by similarity transform, which locally reverses probability currents. Results that use the original measure have a natural interpretation in terms of time-reversal of the dynamics. Here we develop a general interpretation of fluctuation theorems based on the adjoint process by considering the duality of the Kolmogorov-forward and backward equations, acting on distributions versus observables. The backward propagation of the dependency of observables is related to problems of statistical inference, so we characterize the adjoint construction as a duality between dynamics and inference. The adjoint process corresponds to the Kolmogorov backward equation in a generating functional that erases memory from the dynamics of its underlying distribution. We show how erasure affects general correlation functions by showing that duality under the adjoint fluctuation theorems exchanges the roles of advanced and retarded Green's functions. We derive results for the class of discrete-state stochastic processes corresponding to Chemical Reaction Networks (CRNs), and show that dualization acts on the \emph{finite} representation of the generating event-set, in a manner similar to the usual similarity transform acting on the (potentially infinite) set of state transitions. We construct generating functionals within the Doi-Peliti (DP) functional integral framework, within which duality transformation takes a remarkably simple form as a change of integration variable. Our Green's function analysis recovers the Extended Fluctuation-Dissipation Theorem of Seifert and Speck for non-equilibrium steady states, shows that the causal structure responsible for it applies also to dualization about non-steady states.

cond-mat.stat-mech