SearcharxivSearch

arXiv subjects

Daniel Levine

Publications and source records attributed to Daniel Levine.

7 recordsLinked to original sources

Adjoint Sampling: Highly Scalable Diffusion Samplers via Adjoint Matching

We introduce Adjoint Sampling, a highly scalable and efficient algorithm for learning diffusion processes that sample from unnormalized densities, or energy functions. It is the first on-policy approach that allows significantly more gradient updates than the number of energy evaluations and model samples, allowing us to scale to much larger problem settings than previously explored by similar methods. Our framework is theoretically grounded in stochastic optimal control and shares the same theoretical guarantees as Adjoint Matching, being able to train without the need for corrective measures that push samples towards the target distribution. We show how to incorporate key symmetries, as well as periodic boundary conditions, for modeling molecules in both cartesian and torsional coordinates. We demonstrate the effectiveness of our approach through extensive experiments on classical energy functions, and further scale up to neural network-based energy models where we perform amortized conformer generation across many molecular systems. To encourage further research in developing highly scalable sampling methods, we plan to open source these challenging benchmarks, where successful methods can directly impact progress in computational chemistry.

cs.LG

Non-Markovian Discrete Diffusion with Causal Language Models

Discrete diffusion models offer a flexible, controllable approach to structured sequence generation, yet they still lag behind causal language models in expressive power. A key limitation lies in their reliance on the Markovian assumption, which restricts each step to condition only on the current state, leading to potential uncorrectable error accumulation. In this paper, we introduce CaDDi (Causal Discrete Diffusion Model), a discrete diffusion model that conditions on the entire generative trajectory, thereby lifting the Markov constraint and allowing the model to revisit and improve past states. By unifying sequential (causal) and temporal (diffusion) reasoning in a single non-Markovian transformer, CaDDi also treats standard causal language models as a special case and permits the direct reuse of pretrained LLM weights with no architectural changes. Empirically, CaDDi outperforms state-of-the-art discrete diffusion baselines on natural-language benchmarks, substantially narrowing the remaining gap to large autoregressive transformers.

cs.LG

CaLMFlow: Volterra Flow Matching using Causal Language Models

We introduce CaLMFlow (Causal Language Models for Flow Matching), a novel framework that casts flow matching as a Volterra integral equation (VIE), leveraging the power of large language models (LLMs) for continuous data generation. CaLMFlow enables the direct application of LLMs to learn complex flows by formulating flow matching as a sequence modeling task, bridging discrete language modeling and continuous generative modeling. Our method implements tokenization across space and time, thereby solving a VIE over these domains. This approach enables efficient handling of high-dimensional data and outperforms ODE solver-dependent methods like conditional flow matching (CFM). We demonstrate CaLMFlow's effectiveness on synthetic and real-world data, including single-cell perturbation response prediction, showcasing its ability to incorporate textual context and generalize to unseen conditions. Our results highlight LLM-driven flow matching as a promising paradigm in generative modeling, offering improved scalability, flexibility, and context-awareness.

cs.LG

Operator Learning Meets Numerical Analysis: Improving Neural Networks through Iterative Methods

Deep neural networks, despite their success in numerous applications, often function without established theoretical foundations. In this paper, we bridge this gap by drawing parallels between deep learning and classical numerical analysis. By framing neural networks as operators with fixed points representing desired solutions, we develop a theoretical framework grounded in iterative methods for operator equations. Under defined conditions, we present convergence proofs based on fixed point theory. We demonstrate that popular architectures, such as diffusion models and AlphaFold, inherently employ iterative operator learning. Empirical assessments highlight that performing iterations through network operators improves performance. We also introduce an iterative graph neural network, PIGN, that further demonstrates benefits of iterations. Our work aims to enhance the understanding of deep learning by merging insights from numerical analysis, potentially guiding the design of future networks with clearer theoretical underpinnings and improved performance.

cs.LG

Brill-Noether and existence of semistable sheaves on del Pezzo surfaces

Let $X$ be a del Pezzo surface. When the degree of $X$ is at least 4, we compute the cohomology of a general sheaf in the moduli space of Gieseker semistable sheaves. We also classify the Chern characters for which the general sheaf in the moduli space is non-special, i.e. has at most one nonzero cohomology group. Our results hold for arbitrary polarizations, slope semistability, and semi-exceptional moduli spaces. When the degree of $X$ is at least 3, we further show our construction of certain vector bundles implies the existence of stable and semistable sheaves with respect to the anti-canonical polarization.

math.AG

Drive for Creativity

We advance a hypothesis that creativity has evolved with evolution of internal representations, possibly from amniotes to primates, and further in human cultural evolution. Representations separated sensing from acting and gave "internal room" for creativity. To see (or perform any sensing), creatures with internal representations had to modify these representations to fit sensor signals. Therefore the knowledge instinct, KI, the drive to fit representations to the world, had to evolve along with internal representations. Until primates, it remained simple, without language internal representations could not evolve from perceptions to abstract representations, and abstract thoughts were not possible. We consider creative vs. non-creative decision making, and compare KI with Kahneman-Tversky's heuristic thinking. We identify higher, conscious levels of KI with the drive for creativity (DC) and discuss the roles of language and music, brain mechanisms involved, and experimental directions for testing the advanced hypotheses.

q-bio.NC

A Re-Evaluation of the Evolved Stars in the Globular Cluster M13

We present photometry for all bright red giant branch (RGB), horizontal branch (HB), and asymptotic giant branch (AGB) stars within 10' of the center of M13. We find support for the idea that the population of HB stars redder than the primary group are noticeably evolved, which resolves a disagreement between distance moduli derived from the tip of the RGB and from stars near the instability strip. The sharp cut at the red end of the HB provides strong evidence that stars from the dominant HB group must still be undergoing blue loops, implying that diffusion is being inhibited. We argue that M13's HB is a somewhat pathological case - the dominant HB population occurs very near the "knee" in optical CMDs, and evolved stars exclusively appear redward of that peak, leading to the incorrect appearance of a continuation of the unevolved HB. M13 has a distinct group of HB stars previously identified with the second U jump, which may be examples of early hot flashers that ignite core helium fusion shortly after leaving the RGB. However, there is not convincing evidence that a large fraction of stars leave the RGB before helium flash. We revisited the helium-sensitive R ratio, and find that M13's ratio is in agreement with theoretical values for primordial helium abundance Y_P = 0.245 and inconsistent with a helium enhancement DY = 0.04. The brightness of the HB (both in comparison to the end of the canonical HB and to the tip of the RGB) also appears to rule out the idea that the envelopes of the reddest HB stars have been significantly enriched in helium. The absolute colors of the turnoffs of M3 and M13 may potentially be used to look for differences in their mean helium abundances.(ABRIDGED)

astro-ph.SR