SearcharxivSearch

arXiv subjects

Grayson Jorgenson

Publications and source records attributed to Grayson Jorgenson.

9 recordsLinked to original sources

Haldane Bundles: A Dataset for Learning to Predict the Chern Number of Line Bundles on the Torus

Characteristic classes, which are abstract topological invariants associated with vector bundles, have become an important notion in modern physics with surprising real-world consequences. As a representative example, the incredible properties of topological insulators, which are insulators in their bulk but conductors on their surface, can be completely characterized by a specific characteristic class associated with their electronic band structure, the first Chern class. Given their importance to next generation computing and the computational challenge of calculating them using first-principles approaches, there is a need to develop machine learning approaches to predict the characteristic classes associated with a material system. To aid in this program we introduce the {\emph{Haldane bundle dataset}}, which consists of synthetically generated complex line bundles on the $2$-torus. We envision this dataset, which is not as challenging as noisy and sparsely measured real-world datasets but (as we show) still difficult for off-the-shelf architectures, to be a testing ground for architectures that incorporate the rich topological and geometric priors underlying characteristic classes.

cond-mat.mes-hall

ColMix -- A Simple Data Augmentation Framework to Improve Object Detector Performance and Robustness in Aerial Images

In the last decade, Convolutional Neural Network (CNN) and transformer based object detectors have achieved high performance on a large variety of datasets. Though the majority of detection literature has developed this capability on datasets such as MS COCO, these detectors have still proven effective for remote sensing applications. Challenges in this particular domain, such as small numbers of annotated objects and low object density, hinder overall performance. In this work, we present a novel augmentation method, called collage pasting, for increasing the object density without a need for segmentation masks, thereby improving the detector performance. We demonstrate that collage pasting improves precision and recall beyond related methods, such as mosaic augmentation, and enables greater control of object density. However, we find that collage pasting is vulnerable to certain out-of-distribution shifts, such as image corruptions. To address this, we introduce two simple approaches for combining collage pasting with PixMix augmentation method, and refer to our combined techniques as ColMix. Through extensive experiments, we show that employing ColMix results in detectors with superior performance on aerial imagery datasets and robust to various corruptions.

cs.CV

Internal Representations of Vision Models Through the Lens of Frames on Data Manifolds

While the last five years have seen considerable progress in understanding the internal representations of deep learning models, many questions remain. This is especially true when trying to understand the impact of model design choices, such as model architecture or training algorithm, on hidden representation geometry and dynamics. In this work we present a new approach to studying such representations inspired by the idea of a frame on the tangent bundle of a manifold. Our construction, which we call a neural frame, is formed by assembling a set of vectors representing specific types of perturbations of a data point, for example infinitesimal augmentations, noise perturbations, or perturbations produced by a generative model, and studying how these change as they pass through a network. Using neural frames, we make observations about the way that models process, layer-by-layer, specific modes of variation within a small neighborhood of a datapoint. Our results provide new perspectives on a number of phenomena, such as the manner in which training with augmentation produces model invariance or the proposed trade-off between adversarial training and model generalization.

cs.LG

In What Ways Are Deep Neural Networks Invariant and How Should We Measure This?

It is often said that a deep learning model is "invariant" to some specific type of transformation. However, what is meant by this statement strongly depends on the context in which it is made. In this paper we explore the nature of invariance and equivariance of deep learning models with the goal of better understanding the ways in which they actually capture these concepts on a formal level. We introduce a family of invariance and equivariance metrics that allows us to quantify these properties in a way that disentangles them from other metrics such as loss or accuracy. We use our metrics to better understand the two most popular methods used to build invariance into networks: data augmentation and equivariant layers. We draw a range of conclusions about invariance and equivariance in deep learning models, ranging from whether initializing a model with pretrained weights has an effect on a trained model's invariance, to the extent to which invariance learned via training can generalize to out-of-distribution data.

cs.LG

Testing predictions of representation cost theory with CNNs

It is widely acknowledged that trained convolutional neural networks (CNNs) have different levels of sensitivity to signals of different frequency. In particular, a number of empirical studies have documented CNNs sensitivity to low-frequency signals. In this work we show with theory and experiments that this observed sensitivity is a consequence of the frequency distribution of natural images, which is known to have most of its power concentrated in low-to-mid frequencies. Our theoretical analysis relies on representations of the layers of a CNN in frequency space, an idea that has previously been used to accelerate computations and study implicit bias of network training algorithms, but to the best of our knowledge has not been applied in the domain of model robustness.

cs.LG

Automorphism loci for degree 3 and degree 4 endomorphisms of the projective line

Let $f$ be an endomorphism of the projective line. There is a natural conjugation action on the space of such morphisms by elements of the projective linear group. The group of automorphisms, or stabilizer group, of a given $f$ for this action is known to be a finite group. We determine explicit families that parameterize all endomorphisms defined over $\bar{\mathbb{Q}}$ of degree $3$ and degree $4$ that have a nontrivial automorphism, the \textit{automorphism locus} of the moduli space of dynamical systems. We analyze the geometry of these loci in the appropriate moduli space of dynamical systems. Further, for each family of maps, we study the possible structures of $\mathbb{Q}$-rational preperiodic points which occur under specialization.

math.DS

Secant indices of projective varieties

To each subvariety $X$ in projective $n$-space of codimension $m$ we associate an integer sequence of length $m + 1$ from $1$ to the degree of $X$ recording the maximal cardinalities of finite, reduced intersections of $X$ with linear subvarieties. We call this the sequence of secant indices of $X$. Similar numbers have been studied independently with the aim of classifying subvarieties with extremal secant spaces. Our focus in this note is the study of the combinatorial properties that the secant indices satisfy collectively. We show these sequences are strictly increasing for nondegenerate smooth subvarieties, develop a method to compute term-wise lower bounds for the secant indices, and compute these lower bounds for Veronese and Segre varieties. In the case of Veronese varieties, the truth of the Eisenbud-Green-Harris conjecture would imply the lower bounds we find are in fact equal to the secant indices. Along the way we state several relevant questions and additional conjectures which to our knowledge are open.

math.AG

A relative Segre zeta function

The choice of a homogeneous ideal in a polynomial ring defines a closed subscheme $Z$ in a projective space as well as an infinite sequence of cones over $Z$ in progressively higher dimension projective spaces. Recent work of Aluffi introduces the Segre zeta function, a rational power series with integer coefficients which captures the relationship between the Segre class of $Z$ and those of its cones. The goal of this note is to define a relative version of this construction for closed subschemes of projective bundles over a smooth variety. If $Z$ is a closed subscheme of such a projective bundle $P(E)$, this relative Segre zeta function will be a rational power series which describes the Segre class of the cone over $Z$ in every projective bundle "dominating" $P(E)$. When the base variety is a point we recover the absolute Segre zeta function for projective spaces. Part of our construction requires $Z$ to be the zero scheme of a section of a bundle on $P(E)$ of rank smaller than that of $E$ that is able to extend to larger projective bundles. The question of what bundles may extend in this sense seems independently interesting and we discuss some related results, showing that at a minimum one can always count on direct sums of line bundles to extend. Furthermore, the relative Segre zeta function depends only on the Segre class of $Z$ and the total Chern class of the bundle defining $Z$, and the basic forms of the numerator and denominator can be described. As an application of our work we derive a Segre zeta function for products of projective spaces and prove its key properties.

math.AG

Linear recurrence sequences and the duality defect conjecture

It is conjectured that the dual variety of every smooth nonlinear subvariety of dimension $> \frac{2N}{3}$ in projective $N$-space is a hypersurface, an expectation known as the duality defect conjecture. This would follow from the truth of Hartshorne's complete intersection conjecture but nevertheless remains open for the case of subvarieties of codimension $> 2$. A combinatorial approach to proving the conjecture in the codimension $2$ case was developed by Holme, and following this approach Oaland devised an algorithm for proving the conjecture in the codimension $3$ case for particular $N$. This combinatorial approach gives a potential method of proving the duality defect conjecture in many of the cases by studying the positivity of certain homogeneous integer linear recurrence sequences. We give a generalization of the algorithm of Oaland to the higher codimension cases, obtaining with this bounds the degrees of counterexamples would have to satisfy, and using the relationship with recurrence sequences we prove that the conjecture holds in the codimension $3$ case when $N$ is odd.

math.AG