SearcharxivSearch

arXiv subjects

David Lowry-Duda

Publications and source records attributed to David Lowry-Duda.

At least 19 recordsLinked to original sources

Institutional Books - Visual Elements: An open-source pipeline for extracting, classifying, deduplicating, and captioning visual elements from digital book collections

Historical book collections contain rich visual elements - such as illustrations, photographs, engravings, and decorative art - that are frequently under-explored in large-scale digitization projects. While Optical Character Recognition (OCR) has standardized the extraction of textual content, these visual components offer a layer of nuance and context that remains largely untapped by automated text extraction workflows. This technical report introduces Institutional Books - Visual Elements, an open-source end-to-end pipeline for detecting, classifying, deduplicating, and captioning visual elements from historical book collections. Alongside this pipeline, we release an initial dataset of 22.6 million visual elements extracted from the 983,004 scanned volumes that comprise the Institutional Books: Harvard Library dataset. This work contributes to ongoing, community-wide efforts to enable new use cases for digitized library collections through computational access, from artificial intelligence model training to digital humanities research.

cs.CV

Institutional Books - Enriched Text: A customizable multilingual open-source pipeline for denoising, deduplicating, and annotating OCR text at scale

Released in 2025, Institutional Books: Harvard Library (IB-HL) is a collection of 983,004 volumes (242B o200k_base tokens), originally digitized through Harvard Library's participation in the Google Books Library project. As researchers and developers have begun to use IB-HL, a tension has emerged between standard large-scale preprocessing practices and the goals of careful information stewardship. Many existing pipelines optimize for web text: as a result, they tend to aggressively filter, deduplicate, restrict by language, and sometimes discard meaningful metadata. Meanwhile, researchers seeking to use IB-HL duplicate effort while performing similar processing and analysis. We describe an approach that we call Enriched Text. Instead of producing a single 'complete' stream of tokens, we normalize the text while preserving metadata through annotations. We separate endmatter, detect per-paragraph language, identify clusters of duplicate paragraphs, and compute per-paragraph bits-per-byte scores. We provide this information through HTML-like annotations layered on top of the text. By parsing these annotations, users can tailor the output to their own needs instead of accepting a global editorial decision on content. The pipeline applies to all $\approx$250 languages in the collection. This report describes this project's goals, implementation, and design rationale. The release includes IB-HL-ET (an enriched-text version of IB-HL containing 217B o200k_base tokens across 983,003 volumes, organized into 1.39B annotated subtopic paragraphs) and the pipeline that produced it. These serve to make the collection easier for machines to parse and for humans to study.

cs.CL

Murmurations, Mestre--Nagao sums, and Convolutional Neural Networks for elliptic curves

We apply one-dimensional convolutional neural networks to the Frobenius traces of elliptic curves over $\mathbb{Q}$ and evaluate and interpret their predictive capacity. In keeping with similar experiments by Kazalicki--Vlah, Bujanovi\'{c}--Kazalicki--Novak, and Pozdnyakov, we observe high accuracy predictions for the analytic rank across a range of conductors. We interpret the prediction using saliency curves and explore the interesting interplay between murmurations and Mestre--Nagao sums, the details of which vary with the conductor and the (predicted) rank.

math.NT

On Murmurations and Trace Formulas

In recent work with Bober, Booker, Lee, Seymour-Howell, and Zubrilina, we proved murmuration behavior for Maass forms in the eigenvalue aspect and for modular forms in the weight aspect. Both used an approach based on the Selberg trace formula. But different trace formulas, including those due to Kuznetsov or Petersson, offer different variations. We examine murmurations from the perspective of different trace formulas and outline several families of $L$-functions where one can likely prove additional murmuration behavior.

math.NT

Learning Euler Factors of Elliptic Curves

We apply transformer models and feedforward neural networks to predict Frobenius traces $a_p$ from elliptic curves given other traces $a_q$. We train further models to predict $a_p \bmod 2$ from $a_q \bmod 2$, and cross-analysis such as $a_p \bmod 2$ from $a_q$. Our experiments reveal that these models achieve high accuracy, even in the absence of explicit number-theoretic tools like functional equations of $L$-functions. We also present partial interpretability findings.

math.NT

Machine learning the vanishing order of rational L-functions

In this paper, we study the vanishing order of rational $L$-functions from a data scientific perspective. Each $L$-function is represented in our data by finitely many Dirichlet coefficients, the normalisation of which depends on the context. We observe murmuration-like patterns in averages across our dataset, find that PCA clusters rational $L$-functions by their vanishing order, and record that LDA and neural networks may accurately predict this quantity.

math.NT

The Fibonacci Zeta Function and Modular Forms

We show that a family of Dirichlet series generalizing the Fibonacci zeta function $\sum F(n)^{-s}$ has meromorphic continuation in terms of dihedral $\mathrm{GL}(2)$ Maass forms.

math.NT

A database of rigorous Maass forms

We announce a database of rigorously computed Maass forms on congruence subgroups $\Gamma_0(N)$ and briefly describe the methods of computation.

math.NT

Learning Fricke signs from Maass form Coefficients

In this paper, we conduct a data-scientific investigation of Maass forms. We find that averaging the Fourier coefficients of Maass forms with the same Fricke sign reveals patterns analogous to the recently discovered "murmuration" phenomenon, and that these patterns become more pronounced when parity is incorporated as an additional feature. Approximately 43% of the forms in our dataset have an unknown Fricke sign. For the remaining forms, we employ Linear Discriminant Analysis (LDA) to machine learn their Fricke sign, achieving 96% (resp. 94%) accuracy for forms with even (resp. odd) parity. We apply the trained LDA model to forms with unknown Fricke signs to make predictions. The average values based on the predicted Fricke signs are computed and compared to those for forms with known signs to verify the reasonableness of the predictions. Additionally, a subset of these predictions is evaluated against heuristic guesses provided by Hejhal's algorithm, showing a match approximately 95% of the time. We also use neural networks to obtain results comparable to those from the LDA model.

math.NT

The Fibonacci Zeta Function and Continuation

We introduce a family of Dirichlet series associated to real quadratic number fields that generalize the ordinary Fibonacci zeta function $\sum F(n)^{-s}$, where $F(n)$ denotes the $n$th Fibonacci number. We then give three different methods of meromorphic continuation to $\mathbb{C}$. Two are purely analytic and classical, while the third uses shifted convolutions and modular forms.

math.NT

Murmurations of Maass forms

We prove the existence of murmurations in the family of Maass forms of weight 0 and level 1 with their Laplace eigenvalue parameter going to infinity (i.e., correlations between the parity and Hecke eigenvalues at primes growing in proportion to the analytic conductor).

math.NT

Towards a classification of isolated $j$-invariants

We develop an algorithm to test whether a non-CM elliptic curve $E/\mathbb{Q}$ gives rise to an isolated point of any degree on any modular curve of the form $X_1(N)$. This builds on prior work of Zywina which gives a method for computing the image of the adelic Galois representation associated to $E$. Running this algorithm on all elliptic curves presently in the $L$-functions and Modular Forms Database and the Stein-Watkins Database gives strong evidence for the conjecture that $E$ gives rise to an isolated point on $X_1(N)$ if and only if $j(E)=-140625/8, -9317,$ $351/4$, or $-162677523113838677$.

math.NT

Counting Divisors in the Outputs of a Binary Quadratic Form

For a fixed natural number $h$, we prove meromorphic continuation of the two-variable Dirichlet series $\sum_m r_2(m) σ_w(m + h) (m + h)^{-s + w}$ to $\mathbb{C}^2$ and use this to obtain asymptotics for $\sum_{m^2 + n^2 \leq X} σ_w(m^2 + n^2 + h)$. We approach this continuation through spectral theory. Our results are comparable to earlier work of Bykovskii, who used different methods to study the sums $\sum_{n^2 \leq X} σ_w(n^2 + h)$.

math.NT

Murmurations of modular forms in the weight aspect

We prove the existence of "murmurations" in the family of holomorphic modular forms of level $1$ and weight $k\to\infty$, that is, correlations between their root numbers and Hecke eigenvalues at primes growing in proportion to the analytic conductor. This is the first demonstration of murmurations in an archimedean family.

math.NT

Sums of Cusp Form Coefficients Along Quadratic Sequences

Let $f(z) = \sum A(n) n^{(k-1)/2} e(nz)$ be a cusp form of weight $k \geq 3$ on $Γ_0(N)$ with character $χ$. By studying a certain shifted convolution sum, we prove that $\sum_{n \leq X} A(n^2+h) = c_{f,h} X + O_{f,h,ε}(X^{\frac{3}{4}+ε})$ for $ε>0$, which improves a result of Blomer from 2008 with error $X^{\frac{6}{7}+ε}$. This includes an appendix due to Raphael S. Steiner, proving stronger bounds for certain spectral averages.

math.NT

Computing classical modular forms

We discuss practical and some theoretical aspects of computing a database of classical modular forms in the L-functions and Modular Forms Database (LMFDB).

math.NT

Improved bounds on number fields of small degree

We study the number of degree $n$ number fields with discriminant bounded by $X$. In this article, we improve an upper bound due to Schmidt on the number of such fields that was previously the best known upper bound for $6 \leq n \leq 94$.

math.NT