SearcharxivSearch

arXiv subjects

Dominic Culver

Publications and source records attributed to Dominic Culver.

12 recordsLinked to original sources

Trajectory as the Teacher: Few-Step Discrete Flow Matching via Energy-Navigated Distillation

Discrete flow matching generates text by iteratively transforming noise tokens into coherent language, but may require hundreds of forward passes. Distillation uses the multi-step trajectory to train a student to reproduce the process in a few steps. When the student underperforms, the usual explanation is insufficient capacity. We argue the opposite: the trajectory is the bottleneck, not the student. Each training trajectory is built through a chain of blind stochastic jumps with no evaluation of sequence quality; a single bad decision at an early midpoint propagates through subsequent steps, yet the student must imitate the result. Trajectory-Shaped Discrete Flow Matching (TS-DFM) replaces these blind jumps with guided navigation: a lightweight energy compass evaluates candidate continuations at each midpoint, selecting the most coherent. All shaping is training-only; inference cost is unchanged. On 170M-parameter language modeling, the shaped student at 8 steps achieves 32% lower perplexity than the 1,024-step teacher while being 128x faster, with gains consistent across source distributions and three evaluators of increasing scale. TS-DFM achieves the best perplexity of any discrete-generation baseline we compare against, including methods trained on 6x more data or using 5x larger models.

cs.LG

DACA-GRPO: Denoising-Aware Credit Assignment for Reinforcement Learning in Diffusion Language Models

Diffusion large language models are a compelling alternative to autoregressive models, yet existing RL methods for diffusion treat all denoising steps as equally important and rely on biased, high-variance likelihood estimates. We identify two fundamental weaknesses: the absence of temporal credit assignment across the denoising trajectory, and the systematic bias of mean-field likelihood estimates used for policy optimization. To address these, we propose Denoising-Aware Credit Assignment for GRPO (DACA-GRPO), a lightweight, plug-and-play enhancement for any GRPO-style trainer. DACA-GRPO introduces two complementary mechanisms: Denoising Progress Scores, which extract per-token importance weights from intermediate predictions at no additional forward cost, and Stratified Masking Likelihood, which partitions token positions into strata so that each token is predicted with most of the sequence as context, reducing the mean-field bias. Applied on top of three GRPO base methods, DACA-GRPO achieves consistent improvements across seven benchmarks spanning mathematical reasoning, code generation, constraint satisfaction, and constrained generation, with gains of up to 5.6pp on math reasoning, 7.4pp on code generation, 36.3pp on constraint satisfaction, and 5.9pp on JSON schema adherence.

cs.LG

FS-DFM: Fast and Accurate Long Text Generation with Few-Step Diffusion Language Models

Autoregressive language models (ARMs) deliver strong likelihoods, but are inherently serial: they generate one token per forward pass, which limits throughput and inflates latency for long sequences. Diffusion Language Models (DLMs) parallelize across positions and thus appear promising for language generation, yet standard discrete diffusion typically needs hundreds to thousands of model evaluations to reach high quality, trading serial depth for iterative breadth. We introduce FS-DFM, Few-Step Discrete Flow-Matching. A discrete flow-matching model designed for speed without sacrificing quality. The core idea is simple: make the number of sampling steps an explicit parameter and train the model to be consistent across step budgets, so one big move lands where many small moves would. We pair this with a reliable update rule that moves probability in the right direction without overshooting, and with strong teacher guidance distilled from long-run trajectories. Together, these choices make few-step sampling stable, accurate, and easy to control. On language modeling benchmarks, FS-DFM with 8 sampling steps achieves perplexity parity with a 1,024-step discrete-flow baseline for generating 1,024 tokens using a similar-size model, delivering up to 128 times faster sampling and corresponding latency/throughput gains. Code & pretrained checkpoints: https://github.com/apple/ml-fs-dfm

cs.CL

SaulLM-54B & SaulLM-141B: Scaling Up Domain Adaptation for the Legal Domain

In this paper, we introduce SaulLM-54B and SaulLM-141B, two large language models (LLMs) tailored for the legal sector. These models, which feature architectures of 54 billion and 141 billion parameters, respectively, are based on the Mixtral architecture. The development of SaulLM-54B and SaulLM-141B is guided by large-scale domain adaptation, divided into three strategies: (1) the exploitation of continued pretraining involving a base corpus that includes over 540 billion of legal tokens, (2) the implementation of a specialized legal instruction-following protocol, and (3) the alignment of model outputs with human preferences in legal interpretations. The integration of synthetically generated data in the second and third steps enhances the models' capabilities in interpreting and processing legal texts, effectively reaching state-of-the-art performance and outperforming previous open-source models on LegalBench-Instruct. This work explores the trade-offs involved in domain-specific adaptation at this scale, offering insights that may inform future studies on domain adaptation using strong decoder models. Building upon SaulLM-7B, this study refines the approach to produce an LLM better equipped for legal tasks. We are releasing base, instruct, and aligned versions on top of SaulLM-54B and SaulLM-141B under the MIT License to facilitate reuse and collaborative research.

cs.CL

SaulLM-7B: A pioneering Large Language Model for Law

In this paper, we introduce SaulLM-7B, a large language model (LLM) tailored for the legal domain. With 7 billion parameters, SaulLM-7B is the first LLM designed explicitly for legal text comprehension and generation. Leveraging the Mistral 7B architecture as its foundation, SaulLM-7B is trained on an English legal corpus of over 30 billion tokens. SaulLM-7B exhibits state-of-the-art proficiency in understanding and processing legal documents. Additionally, we present a novel instructional fine-tuning method that leverages legal datasets to further enhance SaulLM-7B's performance in legal tasks. SaulLM-7B is released under the MIT License.

cs.CL

On the E2-term of the bo-Adams spectral sequence

The E_1-term of the (2-local) bo-based Adams spectral sequence for the sphere spectrum decomposes into a direct sum of a v_1-periodic part, and a v_1-torsion part. Lellmann and Mahowald completely computed the d_1-differential on the v_1-periodic part, and the corresponding contribution to the E_2-term. The v_1-torsion part is harder to handle, but with the aid of a computer it was computed through the 20-stem by Davis. Such computer computations are limited by the exponential growth of v_1-torsion in the E_1-term. In this paper, we introduce a new method for computing the contribution of the v_1-torsion part to the E_2-term, whose input is the cohomology of the Steenrod algebra. We demonstrate the efficacy of our technique by computing the bo-Adams spectral sequence beyond the 40-stem.

math.AT

The telescope conjecture at height 2 and the tmf resolution

Mahowald proved the height 1 telescope conjecture at the prime 2 as an application of his seminal work on bo-resolutions. In this paper we study the height 2 telescope conjecture at the prime 2 through the lens of tmf-resolutions. To this end we compute the structure of the tmf-resolution for a specifc type 2 complex Z. We find that, analogous to the height 1 case, the E1-page of the tmf-resolution possesses a decomposition into a v2-periodic summand, and an Eilenberg-MacLane summand which consists of bounded v2-torsion. However, unlike the height 1 case, the E2-page of the tmf-resolution exhibits unbounded v2-torsion. We compare this to the work of Mahowald-Ravenel-Shick, and discuss how the validity of the telescope conjecture is connected to the fate of this unbounded v2-torsion: either the unbounded v2-torsion kills itself off in the spectral sequence, and the telescope conjecture is true, or it persists to form v2-parabolas and the telescope conjecture is false. We also study how to use the tmf-resolution to effectively give low dimensional computations of the homotopy groups of Z. These computations allow us to prove a conjecture of the second author and Egger: the E(2)-local Adams-Novikov spectral sequence for Z collapses.

math.AT

The $BP\langle 2 \rangle$-cooperations at odd primes

In previous work, the author analyzed the co-operations algebra for the second truncated Brown-Peterson spectrum at the prime $p=2$. The purpose of this paper is to carry out the necessary modifications to odd primes.

math.AT

On $BP\langle 2\rangle$-cooperations

In this paper we develop techniques to compute the cooperations algebra for the second truncated Brown-Peterson spectrum $\tBP{2}$. We prove that the cooperations algebra $\tBP{2}_*\tBP{2}$ decomposes as a direct some of a $\F_2$-vector space concentrated in Adams filtration 0 and a $\F_2[v_0,v_1,v_2]$-module which is concentrated in even degrees and $v_2$-torsion free. A recursive procedure is also developed to provide an basis of the $v_2$-torsion free part.

math.AT

A new basis for the complex $K$-theory cooperations algebra

A classical theorem of Adams, Harris, and Switzer states that the 0th grading of complex $K$-theory cooperations, $KU_0ku$ is isomorphic to the space of numerical polynomials. The space of numerical polynomials has a basis provided by the binomial coefficient polynomials, which gives a basis of $KU_0ku$. In this paper, we produce a new $p$-local basis for $KU_0ku_{(p)}$ using the Adams splitting. This basis is established by using well known formulas for the Hazewinkel generators. For $p=2$, we show that this new basis coincides with the classical basis modulo higher Adams filtration.

math.AT