Searcharxiv⌕ Search

arXiv subjects

Kevin Yang

Publications and source records attributed to Kevin Yang.

At least 37 records · Page 2Linked to original sources

KPZ equation from ASEP plus general speed-change drift

We derive the KPZ equation as a continuum limit of height functions in asymmetric simple exclusion processes with drift that depends on the local particle configuration. To our knowledge, it is a first such result for a class of particle systems without duality or explicit invariant measures. The tools developed in this paper consist of estimates on the corresponding Kolmogorov equations, giving a more robust proof of the Boltzmann-Gibbs principle. These tools are not exclusive to KPZ.

math.PR↗

A flow-type scaling limit for random growth with memory

We study a stochastic Laplacian growth model, where a set $\mathbf{U}\subseteq\mathbb{R}^{\mathrm{d}}$ grows according to a reflecting Brownian motion in $\mathbf{U}$ stopped at level sets of its boundary local time. We derive a scaling limit for the leading-order behavior of the growing boundary (i.e. "interface"). It is given by a geometric flow-type PDE. It is obtained by an averaging principle for the reflecting Brownian motion. We also show that this geometric flow-type PDE is locally well-posed, and its blow-up times correspond to changes in the diffeomorphism class of the growth model. Our results extend those of Dembo-Groisman-Huang-Sidoravicius '21, which restricts to star-shaped growth domains and radially outwards growth, so that in polar coordinates, the geometric flow transforms into a simple ODE with infinite lifetime. Also, we remove the "separation of scales" assumption that was taken in Dembo-Groisman-Huang-Sidoravicius '21; this forces us to understand the local geometry of the growing interface.

math.PR↗

Learning Personalized Alignment for Evaluating Open-ended Text Generation

Recent research has increasingly focused on evaluating large language models' (LLMs) alignment with diverse human values and preferences, particularly for open-ended tasks like story generation. Traditional evaluation metrics rely heavily on lexical similarity with human-written references, often showing poor correlation with human judgments and failing to account for alignment with the diversity of human preferences. To address these challenges, we introduce PerSE, an interpretable evaluation framework designed to assess alignment with specific human preferences. It is tuned to infer specific preferences from an in-context personal profile and evaluate the alignment between the generated content and personal preferences. PerSE enhances interpretability by providing detailed comments and fine-grained scoring, facilitating more personalized content generation. Our 13B LLaMA-2-based PerSE shows a 15.8% increase in Kendall correlation and a 13.7% rise in accuracy with zero-shot reviewers compared to GPT-4. It also outperforms GPT-4 by 46.01% in Kendall correlation on new domains, indicating its transferability.

cs.CL↗

RLCD: Reinforcement Learning from Contrastive Distillation for Language Model Alignment

We propose Reinforcement Learning from Contrastive Distillation (RLCD), a method for aligning language models to follow principles expressed in natural language (e.g., to be more harmless) without using human feedback. RLCD creates preference pairs from two contrasting model outputs, one using a positive prompt designed to encourage following the given principles, and one using a negative prompt designed to encourage violating them. Using two different prompts causes model outputs to be more differentiated on average, resulting in cleaner preference labels in the absence of human annotations. We then use the preference pairs to train a preference model, which is in turn used to improve a base unaligned language model via reinforcement learning. Empirically, RLCD outperforms RLAIF (Bai et al., 2022b) and context distillation (Huang et al., 2022) baselines across three diverse alignment tasks--harmlessness, helpfulness, and story outline generation--and when using both 7B and 30B model scales for simulating preference data.

cs.CL↗

Improving Pacing in Long-Form Story Planning

Existing LLM-based systems for writing long-form stories or story outlines frequently suffer from unnatural pacing, whether glossing over important events or over-elaborating on insignificant details, resulting in a jarring experience for the reader. We propose a CONCrete Outline ConTrol (CONCOCT) system to improve pacing when automatically generating story outlines. We first train a concreteness evaluator to judge which of two events is more concrete (low-level-detailed). This evaluator can then be used to control pacing in hierarchical outline generation; in this work, we explore a vaguest-first expansion procedure that aims for uniform pacing. We further use the evaluator to filter new outline items based on predicted concreteness. Compared to a baseline hierarchical outline generator, humans judge CONCOCT's pacing to be more consistent over 57% of the time across multiple outline lengths; the gains also translate to downstream stories. All code, data, and models are open-sourced.

cs.CL↗

End-to-end Story Plot Generator

Story plots, while short, carry most of the essential information of a full story that may contain tens of thousands of words. We study the problem of automatic generation of story plots, which includes story premise, character descriptions, plot outlines, etc. To generate a single engaging plot, existing plot generators (e.g., DOC (Yang et al., 2022a)) require hundreds to thousands of calls to LLMs (e.g., OpenAI API) in the planning stage of the story plot, which is costly and takes at least several minutes. Moreover, the hard-wired nature of the method makes the pipeline non-differentiable, blocking fast specialization and personalization of the plot generator. In this paper, we propose three models, $\texttt{OpenPlot}$, $\texttt{E2EPlot}$ and $\texttt{RLPlot}$, to address these challenges. $\texttt{OpenPlot}$ replaces expensive OpenAI API calls with LLaMA2 (Touvron et al., 2023) calls via careful prompt designs, which leads to inexpensive generation of high-quality training datasets of story plots. We then train an end-to-end story plot generator, $\texttt{E2EPlot}$, by supervised fine-tuning (SFT) using approximately 13000 story plots generated by $\texttt{OpenPlot}$. $\texttt{E2EPlot}$ generates story plots of comparable quality to $\texttt{OpenPlot}$, and is > 10$\times$ faster (1k tokens in only 30 seconds on average). Finally, we obtain $\texttt{RLPlot}$ that is further fine-tuned with RLHF on several different reward models for different aspects of story quality, which yields 60.0$\%$ winning rate against $\texttt{E2EPlot}$ along the aspect of suspense and surprise.

cs.CL↗

PREADD: Prefix-Adaptive Decoding for Controlled Text Generation

We propose Prefix-Adaptive Decoding (PREADD), a flexible method for controlled text generation. Unlike existing methods that use auxiliary expert models to control for attributes, PREADD does not require an external model, instead relying on linearly combining output logits from multiple prompts. Specifically, PREADD contrasts the output logits generated using a raw prompt against those generated using a prefix-prepended prompt, enabling both positive and negative control with respect to any attribute encapsulated by the prefix. We evaluate PREADD on three tasks -- toxic output mitigation, gender bias reduction, and sentiment control -- and find that PREADD outperforms not only prompting baselines, but also an auxiliary-expert control method, by 12% or more in relative gain on our main metrics for each task.

cs.CL↗

DOC: Improving Long Story Coherence With Detailed Outline Control

We propose the Detailed Outline Control (DOC) framework for improving long-range plot coherence when automatically generating several-thousand-word-long stories. DOC consists of two complementary components: a detailed outliner and a detailed controller. The detailed outliner creates a more detailed, hierarchically structured outline, shifting creative burden from the main drafting procedure to the planning stage. The detailed controller ensures the more detailed outline is still respected during generation by controlling story passages to align with outline details. In human evaluations of automatically generated stories, DOC substantially outperforms a strong Re3 baseline (Yang et al., 2022) on plot coherence (22.5% absolute gain), outline relevance (28.2%), and interestingness (20.7%). Humans also judged DOC to be much more controllable in an interactive generation setting.

cs.CL↗

Modular Visual Question Answering via Code Generation

We present a framework that formulates visual question answering as modular code generation. In contrast to prior work on modular approaches to VQA, our approach requires no additional training and relies on pre-trained language models (LMs), visual models pre-trained on image-caption pairs, and fifty VQA examples used for in-context learning. The generated Python programs invoke and compose the outputs of the visual models using arithmetic and conditional logic. Our approach improves accuracy on the COVR dataset by at least 3% and on the GQA dataset by roughly 2% compared to the few-shot baseline that does not employ code generation.

cs.CL↗

Fluctuations for some non-stationary interacting particle systems via Boltzmann-Gibbs Principle

Conjecture II.3.6 of Spohn in [Spohn '91] and Lecture 7 of Jensen-Yau in [Jensen-Yau '99] ask for a general derivation of universal fluctuations of hydrodynamic limits in large-scale stochastic interacting particle systems. However, the past few decades have witnessed only minimal progress according to [Goncalves-Landim-Milanes '17]. In this paper, we develop a general method for deriving the so-called Boltzmann-Gibbs principle for a general family of non-integrable and non-stationary interacting particle systems, thereby responding to Spohn and Jensen-Yau. Most importantly, our method depends mostly on local and dynamical, and thus more general/universal, features of the model. This contrasts with previous works, which rely on global and non-universal assumptions on invariant measures or initial measures of the model. As a concrete application of the method, we derive the KPZ equation as a large-scale limit of the height functions for a family of non-stationary and non-integrable exclusion processes with an environment-dependent asymmetry. This establishes a first result to Big Picture Question 1.6 in [AimPLKPZ] for non-stationary and non-integrable "speed-change" models that have also been of interest beyond KPZ [De Masi-Presutti-Spohn-Wick '86, Funaki '18, Funaki-Handa-Uchiyama '91, Komoriya '98].

math.PR↗

Hairer-Quastel universality in non-stationarity via energy solution theory

The paper addresses probabilistic aspects of the KPZ equation and stochastic Burgers equation by providing a solution theory that builds on the energy solution theory Goncalves-Jara '14, Gubinelli-Jara '13, Gubinelli-Perkowski '18, Gubinelli-Perkowski '20. The perspective we adopt is to study the stochastic Burgers equation by writing its solution as a probabilistic solution Gubinelli-Perkowski '17 plus a term that can be studied with deterministic PDE considerations. One motivation is universality of KPZ and stochastic Burgers equations for a certain class of stochastic PDE growth models, first studied in Hairer-Quastel '18. For this, we prove universality for SPDEs with general nonlinearities, thereby extending Hairer-Quastel '18, Hairer-Xu '19, and for many non-stationary initial data, thereby extending Gubinelli-Perkowski '16. Our perspective lets us also prove explicit rates of convergence to white noise invariant measure of stochastic Burgers for non-stationary initial data, in particular extending the spectral gap result of Gubinelli-Perkowski '20 beyond stationary initial data, though for non-stationary data our convergence will be measured in Wasserstein distance and relative entropy, not via the spectral gap as in Gubinelli-Perkowski '20. Actually, we extend the spectral gap in Gubinelli-Perkowski '20 to a log-Sobolev inequality. Our methods can also analyze fractional stochastic Burgers equations; we discuss this briefly. Lastly, we note that our perspective on the KPZ and stochastic Burgers equations provides a first intrinsic notion of solutions for general continuous initial data, in contrast to Holder regular data needed for regularity structures, paracontrolled distributions, and Holder-regular Brownian bridge data for energy solutions.

math.PR↗

Non-Stationary KPZ equation from ASEP with slow bonds

We prove the height functions for a class of non-integrable and non-stationary particle systems converge to the KPZ equation, thereby making progress on the universality of the KPZ equation. The models herein are ASEP [4] with a mesoscopic family of slow bonds, thus we partially extend [16] to non-stationary models and add to the almost empty set of non-integrable, non-stationary interacting particle systems for which universality is established. To do this, we develop further the strategy of [41, 42] introduce a method to establish a novel principle that builds upon the classical hydrodynamic limits of [30] and that we call local hydrodynamics.

math.PR↗

Kardar-Parisi-Zhang Equation from Long-Range Exclusion Processes

We prove here that the height function associated to non-simple exclusion processes with arbitrary jump-length converges to the solution of the Kardar-Parisi-Zhang SPDE under suitable scaling and renormalization. This extends the work of Dembo-Tsai'16 for arbitrary jump-length and Goncalves-Jara '17 for the non-stationary regime. Thus we answer a ``Big Picture Question" from the AIM workshop on KPZ and also expand on the almost empty set of non-integrable and non-stationary particle systems for which weak KPZ universality is proven. We use an approximate microscopic Cole-Hopf transform as in Dembo-Tsai'16 but we develop tools to analyze local statistics of the particle system via local equilibrium and work of Goncalves-Jara '17. Local equilibrium is done via the one-block step in Guo-Papanicolaou-Varadhan '88 for path-space/dynamic statistics.

math.PR↗

Re3: Generating Longer Stories With Recursive Reprompting and Revision

We consider the problem of automatically generating longer stories of over two thousand words. Compared to prior work on shorter stories, long-range plot coherence and relevance are more central challenges here. We propose the Recursive Reprompting and Revision framework (Re3) to address these challenges by (a) prompting a general-purpose language model to construct a structured overarching plan, and (b) generating story passages by repeatedly injecting contextual information from both the plan and current story state into a language model prompt. We then revise by (c) reranking different continuations for plot coherence and premise relevance, and finally (d) editing the best continuation for factual consistency. Compared to similar-length stories generated directly from the same base model, human evaluators judged substantially more of Re3's stories as having a coherent overarching plot (by 14% absolute increase), and relevant to the given initial premise (by 20%).

cs.CL↗

Multi-objective Optimization by Learning Space Partitions

In contrast to single-objective optimization (SOO), multi-objective optimization (MOO) requires an optimizer to find the Pareto frontier, a subset of feasible solutions that are not dominated by other feasible solutions. In this paper, we propose LaMOO, a novel multi-objective optimizer that learns a model from observed samples to partition the search space and then focus on promising regions that are likely to contain a subset of the Pareto frontier. The partitioning is based on the dominance number, which measures "how close" a data point is to the Pareto frontier among existing samples. To account for possible partition errors due to limited samples and model mismatch, we leverage Monte Carlo Tree Search (MCTS) to exploit promising regions while exploring suboptimal regions that may turn out to contain good solutions later. Theoretically, we prove the efficacy of learning space partitioning via LaMOO under certain assumptions. Empirically, on the HyperVolume (HV) benchmark, a popular MOO metric, LaMOO substantially outperforms strong baselines on multiple real-world MOO tasks, by up to 225% in sample efficiency for neural architecture search on Nasbench201, and up to 10% for molecular design.

cs.LG↗

Automated Crossword Solving

We present the Berkeley Crossword Solver, a state-of-the-art approach for automatically solving crossword puzzles. Our system works by generating answer candidates for each crossword clue using neural question answering models and then combines loopy belief propagation with local search to find full puzzle solutions. Compared to existing approaches, our system improves exact puzzle accuracy from 71% to 82% on crosswords from The New York Times and obtains 99.9% letter accuracy on themeless puzzles. Additionally, in 2021, a hybrid of our system and the existing Dr.Fill system outperformed all human competitors for the first time at the American Crossword Puzzle Tournament. To facilitate research on question answering and crossword solving, we analyze our system's remaining errors and release a dataset of over six million question-answer pairs.

cs.CL↗