SearcharxivSearch

arXiv subjects

Nate Gillman

Publications and source records attributed to Nate Gillman.

10 recordsLinked to original sources

Goal Force: Teaching Video Models To Accomplish Physics-Conditioned Goals

Recent advancements in video generation have enabled the development of ``world models'' capable of simulating potential futures for robotics and planning. However, specifying precise goals for these models remains a challenge; text instructions are often too abstract to capture physical nuances, while target images are frequently infeasible to specify for dynamic tasks. To address this, we introduce Goal Force, a novel framework that allows users to define goals via explicit force vectors and intermediate dynamics, mirroring how humans conceptualize physical tasks. We train a video generation model on a curated dataset of synthetic causal primitives-such as elastic collisions and falling dominos-teaching it to propagate forces through time and space. Despite being trained on simple physics data, our model exhibits remarkable zero-shot generalization to complex, real-world scenarios, including tool manipulation and multi-object causal chains. Our results suggest that by grounding video generation in fundamental physical interactions, models can emerge as implicit neural physics simulators, enabling precise, physics-aware planning without reliance on external engines. We release all datasets, code, model weights, and interactive video demos at our project page.

cs.CV

Force Prompting: Video Generation Models Can Learn and Generalize Physics-based Control Signals

Recent advances in video generation models have sparked interest in world models capable of simulating realistic environments. While navigation has been well-explored, physically meaningful interactions that mimic real-world forces remain largely understudied. In this work, we investigate using physical forces as a control signal for video generation and propose force prompts which enable users to interact with images through both localized point forces, such as poking a plant, and global wind force fields, such as wind blowing on fabric. We demonstrate that these force prompts can enable videos to respond realistically to physical control signals by leveraging the visual and motion prior in the original pretrained model, without using any 3D asset or physics simulator at inference. The primary challenge of force prompting is the difficulty in obtaining high quality paired force-video training data, both in the real world due to the difficulty of obtaining force signals, and in synthetic data due to limitations in the visual quality and domain diversity of physics simulators. Our key finding is that video generation models can generalize remarkably well when adapted to follow physical force conditioning from videos synthesized by Blender, even with limited demonstrations of few objects. Our method can generate videos which simulate forces across diverse geometries, settings, and materials. We also try to understand the source of this generalization and perform ablations that reveal two key elements: visual diversity and the use of specific text keywords during training. Our approach is trained on only around 15k training examples for a single day on four A100 GPUs, and outperforms existing methods on force adherence and physics realism, bringing world models closer to real-world physics interactions. We release all datasets, code, weights, and interactive video demos at our project page.

cs.CV

Fourier Head: Helping Large Language Models Learn Complex Probability Distributions

As the quality of large language models has improved, there has been increased interest in using them to model non-linguistic tokens. For example, the Decision Transformer recasts agentic decision making as a sequence modeling problem, using a decoder-only LLM to model the distribution over the discrete action space for an Atari agent. However, when adapting LLMs to non-linguistic domains, it remains unclear if softmax over discrete bins captures the continuous structure of the tokens and the potentially complex distributions needed for high quality token generation. We introduce a neural network layer, constructed using Fourier series, which we can easily substitute for any linear layer if we want the outputs to have a more continuous structure. We perform extensive analysis on synthetic datasets, as well as on large-scale decision making and time series forecasting tasks. We also provide theoretical evidence that this layer can better learn signal from data while ignoring high-frequency noise. All of our results support the effectiveness of our proposed Fourier head in scenarios where the underlying data distribution has a natural continuous structure. For example, the Fourier head improves a Decision Transformer agent's returns across four benchmark Atari games by as much as 377%, and increases a state-of-the-art times series foundation model's forecasting performance by 3.5% across 20 benchmarks unseen during training.

cs.LG

Self-Correcting Self-Consuming Loops for Generative Model Training

As synthetic data becomes higher quality and proliferates on the internet, machine learning models are increasingly trained on a mix of human- and machine-generated data. Despite the successful stories of using synthetic data for representation learning, using synthetic data for generative model training creates "self-consuming loops" which may lead to training instability or even collapse, unless certain conditions are met. Our paper aims to stabilize self-consuming generative model training. Our theoretical results demonstrate that by introducing an idealized correction function, which maps a data point to be more likely under the true data distribution, self-consuming loops can be made exponentially more stable. We then propose self-correction functions, which rely on expert knowledge (e.g. the laws of physics programmed in a simulator), and aim to approximate the idealized corrector automatically and at scale. We empirically validate the effectiveness of self-correcting self-consuming loops on the challenging human motion synthesis task, and observe that it successfully avoids model collapse, even when the ratio of synthetic data to real data is as high as 100%.

cs.LG

IsoScore: Measuring the Uniformity of Embedding Space Utilization

The recent success of distributed word representations has led to an increased interest in analyzing the properties of their spatial distribution. Several studies have suggested that contextualized word embedding models do not isotropically project tokens into vector space. However, current methods designed to measure isotropy, such as average random cosine similarity and the partition score, have not been thoroughly analyzed and are not appropriate for measuring isotropy. We propose IsoScore: a novel tool that quantifies the degree to which a point cloud uniformly utilizes the ambient vector space. Using rigorously designed tests, we demonstrate that IsoScore is the only tool available in the literature that accurately measures how uniformly distributed variance is across dimensions in vector space. Additionally, we use IsoScore to challenge a number of recent conclusions in the NLP literature that have been derived using brittle metrics of isotropy. We caution future studies from using existing tools to measure isotropy in contextualized embedding space as resulting conclusions will be misleading or altogether inaccurate.

cs.CL

Large Sets with Small Injective Projections

Let $\ell_1,\ell_2,\dots$ be a countable collection of lines in ${\mathbb R}^d$. For any $t \in [0,1]$ we construct a compact set $Γ\subset{\mathbb R}^d$ with Hausdorff dimension $d-1+t$ which projects injectively into each $\ell_i$, such that the image of each projection has dimension $t$. This immediately implies the existence of homeomorphisms between certain Cantor-type sets whose graphs have large dimensions. As an application, we construct a collection $E$ of disjoint, non-parallel $k$-planes in $\mathbb{R}^d$, for $d \geq k+2$, whose union is a small subset of $\mathbb{R}^d$, either in Hausdorff dimension or Lebesgue measure, while $E$ itself has large dimension. As a second application, for any countable collection of vertical lines $w_i$ in the plane we construct a collection of nonvertical lines $H$, so that $F$, the union of lines in $H$, has positive Lebesgue measure, but each point of each line $w_i$ intersects at most one $h\in H$ and, for each $w_i$, the Hausdorff dimension of $F\cap w_i$ is zero.

math.MG

Patterns of primes in the Sato-Tate conjecture

Fix a non-CM elliptic curve $E/\mathbb{Q}$, and let $a_E(p) = p + 1 - \#E(\mathbb{F}_p)$ denote the trace of Frobenius at $p$. The Sato-Tate conjecture gives the limiting distribution $μ_{ST}$ of $a_E(p)/(2\sqrt{p})$ within $[-1, 1]$. We establish bounded gaps for primes in the context of this distribution. More precisely, given an interval $I\subseteq [-1, 1]$, let $p_{I,n}$ denote the $n$th prime such that $a_E(p)/(2\sqrt{p})\in I$. We show $\liminf_{n\to\infty}(p_{I,n+m}-p_{I,n}) < \infty$ for all $m\ge 1$ for "most" intervals, and in particular, for all $I$ with $μ_{ST}(I)\ge 0.36$. Furthermore, we prove a common generalization of our bounded gap result with the Green-Tao theorem. To obtain these results, we demonstrate a Bombieri-Vinogradov type theorem for Sato-Tate primes.

math.NT

Explicit subconvexity savings for sup-norms of cusp forms on $\mathrm{PGL}_n(\mathbb R)$

Blomer and Maga recently proved that, if $F$ is an $L^2$-normalized Hecke Maass cusp form for $\mathrm{SL}_n(\mathbb Z)$, and $Ω$ is a compact subset of $\mathrm{PGL}_n(\mathbb R)/\mathrm{PO}_n(\mathbb R)$, then we have $\|F|_Ω\|_\infty\ll_Ωλ_F^{n(n-1)/8-δ_n}$ for some $δ_n>0$, where $λ_F$ is the Laplacian eigenvalue of $F$. In the present paper, we prove an explicit version of their result.

math.NT

From partitions to Hodge numbers of Hilbert Schemes of Surfaces

We celebrate the 100th anniversary of Srinivasa Ramanujan's election as a Fellow of the Royal Society, which was largely based on his work with G. H. Hardy on the asymptotic properties of the partition function. After recalling this revolutionary work, marking the birth of the "circle method", we present a contemporary example of its legacy in topology. We deduce the equidistribution of Hodge numbers for Hilbert schemes of suitable smooth projective surfaces.

math.NT

Exact Formulas for Invariants of Hilbert Schemes

A theorem of Göttsche establishes a connection between cohomological invariants of a complex projective surface $S$ and corresponding invariants of the Hilbert scheme of $n$ points on $S.$ This relationship is encoded in certain infinite product $q$-series which are essentially modular forms. Here we make use of the circle method to arrive at exact formulas for certain specializations of these $q$-series, yielding convergent series for the signature and Euler characteristic of these Hilbert schemes. We also analyze the asymptotic and distributional properties of the $q$-series' coefficients.

math.NT