SearcharxivSearch

arXiv subjects

Tyler J. Jarvis

Publications and source records attributed to Tyler J. Jarvis.

At least 19 recordsLinked to original sources

eGAD! double descent is explained by Generalized Aliasing Decomposition

A central problem in data science is to use potentially noisy samples of an unknown function to predict values for unseen inputs. In classical statistics, predictive error is understood as a trade-off between the bias and the variance that balances model simplicity with its ability to fit complex functions. However, over-parameterized models exhibit counterintuitive behaviors, such as "double descent" in which models of increasing complexity exhibit decreasing generalization error. Others may exhibit more complicated patterns of predictive error with multiple peaks and valleys. Neither double descent nor multiple descent phenomena are well explained by the bias-variance decomposition. We introduce a novel decomposition that we call the generalized aliasing decomposition (GAD) to explain the relationship between predictive performance and model complexity. The GAD decomposes the predictive error into three parts: 1) model insufficiency, which dominates when the number of parameters is much smaller than the number of data points, 2) data insufficiency, which dominates when the number of parameters is much greater than the number of data points, and 3) generalized aliasing, which dominates between these two extremes. We demonstrate the applicability of the GAD to diverse applications, including random feature models from machine learning, Fourier transforms from signal processing, solution methods for differential equations, and predictive formation enthalpy in materials discovery. Because key components of the GAD can be explicitly calculated from the relationship between model class and samples without seeing any data labels, it can answer questions related to experimental design and model selection before collecting data or performing experiments. We further demonstrate this approach on several examples and discuss implications for predictive modeling and data science.

math.ST

Chebyshev Subdivision and Reduction Methods for Solving Multivariable Systems of Equations

We present a new algorithm for finding isolated zeros of a system of real-valued functions in a bounded interval in $\mathbb{R}^n$. It uses the Chebyshev proxy method combined with a mixture of subdivision, reduction methods, and elimination checks that leverage special properties of Chebyshev polynomials. We prove the method has R-quadratic convergence locally near simple zeros of the system. We also analyze the temporal complexity and the numerical stability of the algorithm and provide numerical evidence in dimensions up to three that the method is both fast and accurate on a wide range of problems. The algorithm should also work well in higher dimensions. Our tests show that the algorithm outperforms other standard methods on this problem of finding all real zeros in a bounded domain. Our Python implementation of the algorithm is publicly available on GitHub.

math.NA

Mathematical Analysis of Redistricting in Utah

We discuss difficulties of evaluating partisan gerrymandering in the congressional districts in Utah and the failure of many common metrics in Utah. We explain why the Republican vote share in the least-Republican district (LRVS) is a good indicator of the advantage or disadvantage each party has in the Utah congressional districts. Although the LRVS only makes sense in settings with at most one competitive district, in that setting it directly captures the extent to which a given redistricting plan gives advantage or disadvantage to the Republican and Democratic parties. We use the LRVS to evaluate the most common measures of partisan gerrymandering in the context of Utah's 2011 congressional districts. We do this by generating large ensembles of alternative redistricting plans using Markov chain Monte Carlo methods. We also discuss the implications of this new metric and our results on the question of whether the 2011 Utah congressional plan was gerrymandered.

cs.CY

A general algorithm for calculating irreducible Brillouin zones

Calculations of properties of materials require performing numerical integrals over the Brillouin zone (BZ). Integration points in density functional theory codes are uniformly spread over the BZ (despite integration error being concentrated in small regions of the BZ) and preserve symmetry to improve computational efficiency. Integration points over an irreducible Brillouin zone (IBZ), a rotationally distinct region of the BZ, do not have to preserve crystal symmetry for greater efficiency. This freedom allows the use of adaptive meshes with higher concentrations of points at locations of large error, resulting in improved algorithmic efficiency. We have created an algorithm for constructing an IBZ of any crystal structure in 2D and 3D. The algorithm uses convex hull and half-space representations for the BZ and IBZ to make many aspects of construction and symmetry reduction of the BZ trivial. The algorithm is simple, general, and available as open-source software.

cond-mat.mtrl-sci

Analysis of Normal-Form Algorithms for Solving Systems of Polynomial Equations

We examine several of the normal-form multivariate polynomial rootfinding methods of Telen, Mourrain, and Van Barel and some variants of those methods. We analyze the performance of these variants in terms of their asymptotic temporal complexity as well as speed and accuracy on a wide range of numerical experiments. All variants of the algorithm are problematic for systems in which many roots are very close together. We analyze performance on one such system in detail, namely the 'devastating example' that Noferini and Townsend used to demonstrate instability of resultant-based methods.

math.NA

Tandem Blocks in Deep Convolutional Neural Networks

Due to the success of residual networks (resnets) and related architectures, shortcut connections have quickly become standard tools for building convolutional neural networks. The explanations in the literature for the apparent effectiveness of shortcuts are varied and often contradictory. We hypothesize that shortcuts work primarily because they act as linear counterparts to nonlinear layers. We test this hypothesis by using several variations on the standard residual block, with different types of linear connections, to build small image classification networks. Our experiments show that other kinds of linear connections can be even more effective than the identity shortcuts. Our results also suggest that the best type of linear connection for a given application may depend on both network width and depth.

stat.ML

A Brief Survey of FJRW Theory

In this paper we describe some of the constructions of FJRW theory. We also briefly describe its relation to Saito-Givental theory via Landau-Ginzburg mirror symmetry and its relation to Gromov-Witten theory via the Landau-Ginzburg/Calabi-Yau correspondence. We conclude with a discussion of some of the recent results in the field.

math.AG

Chern Classes and Compatible Power Operations in Inertial K-theory

Let [X/G] be a smooth Deligne-Mumford quotient stack. In a previous paper the authors constructed a class of exotic products called inertial products on K(I[X/G]), the Grothendieck group of vector bundles on the inertia stack I[X/G]. In this paper we develop a theory of Chern classes and compatible power operations for inertial products. When G is diagonalizable these give rise to an augmented $λ$-ring structure on inertial K-theory. One well-known inertial product is the virtual product. Our results show that for toric Deligne-Mumford stacks there is a $λ$-ring structure on inertial K-theory. As an example, we compute the $λ$-ring structure on the virtual K-theory of the weighted projective lines P(1,2) and P(1,3). We prove that after tensoring with C, the augmentation completion of this $λ$-ring is isomorphic as a $λ$-ring to the classical K-theory of the crepant resolutions of singularities of the coarse moduli spaces of the cotangent bundles $T^*P(1,2)$ and $T^*P(1,3)$, respectively. We interpret this as a manifestation of mirror symmetry in the spirit of the Hyper-Kaehler Resolution Conjecture.

math.AG

A plethora of inertial products

For a smooth Deligne-Mumford stack X we describe a large number of inertial products on K(IX) and A*(IX) and corresponding inertial Chern characters. We do this by developing a theory of inertial pairs. Each inertial pair determines an inertial product on K(IX) and an inertial product on A*(IX) and Chern character ring homomorphisms between them. We show that there are many inertial pairs; indeed, every vector bundle V on X defines two new inertial pairs. We recover, as special cases, the orbifold products of Chen-Run, Abramovich-Graber-Vistoli, Jarvis-Kaufmann-Kimura, and Edidin-Jarvis-Kimura and the virtual product of Gonzalez-Lupercio-Segovia-Uribe-Xicotencatl. We also introduce an entirely new product we call the localized orbifold product, which is defined on the complexification of K(IX). The inertial products developed in this paper are used in a subsequent paper to describe a theory of inertial Chern classes and power operations in inertial K-theory. These constructions provide new manifestations of mirror symmetry, in the spirit of the Hyper-Kaehler Resolution Conjecture.

math.AG

Witten's D_4 Integrable Hierarchies Conjecture

We prove that the total descendant potential functions of the theory of Fan-Jarvis-Ruan-Witten for D_4 with symmetry group and D_4^T with symmetry group G_{max}, respectively, are both tau-functions of the D_4 Kac-Wakimoto/Drinfeld-Sokolov hierarchy. This completes the proof, begun in [FJR], of the Witten Integrable Hierarchies Conjecture for all simple (ADE) singularities.

math.AG

The Witten equation, mirror symmetry and quantum singularity theory

For any non-degenerate, quasi-homogeneous hypersurface singularity, we describe a family of moduli spaces, a virtual cycle, and a corresponding cohomological field theory associated to the singularity. This theory is analogous to Gromov-Witten theory and generalizes the theory of r-spin curves, which corresponds to the simple singularity A_{r-1}. We also resolve two outstanding conjectures of Witten. The first conjecture is that ADE-singularities are self-dual; and the second conjecture is that the total potential functions of ADE-singularities satisfy corresponding ADE-integrable hierarchies. Other cases of integrable hierarchies are also discussed.

math.AG

Integral Models of Extremal Rational Elliptic Surfaces

Miranda and Persson classified all extremal rational elliptic surfaces in characteristic zero. We show that each surface in Miranda and Persson's classification has an integral model with good reduction everywhere (except for those of type X_{11}(j), which is an exceptional case), and that every extremal rational elliptic surface over an algebraically closed field of characteristic p > 0 can be obtained by reducing one of these integral models mod p.

math.AG

The Witten equation and its virtual fundamental cycle

We study a system of nonlinear elliptic PDEs associated with a quasi-homogeneous polynomial. These equations were proposed by Witten as the replacement for the Cauchy-Riemann equation in the singularity (Landau-Ginzburg) setting. We introduce a perturbation to the equation and construct a virtual cycle for the moduli space of its solutions. Then, we study the wall-crossing of the deformation of the virtual cycle under perturbation and match it to classical Picard-Lefschetz theory. An extended virtual cycle is obtained for the original equation. Finally, we prove that the extended virtual cycle satisfies a set of axioms similar to those of Gromov-Witten theory and r-spin theory.

math.AG

Quantum Singularity Theory for A_{r-1} and r-Spin Theory

We give a review of the quantum singularity theory of Fan-Jarvis-Ruan and the r-spin theory of Jarvis-Kimura-Vaintrob and describe the work of Abramovich-Jarvis showing that for the singularity A_{r-1} = x^r the stack of A_{r-1}-curves of is canonically isomorphic to the stack of r-spin curves. We prove that the A_{r-1}-theory satisfies all the axioms of Jarvis-Kimura-Vaintrob for an r-spin virtual class. Therefore, the results of Lee, Faber-Shadrin-Zovonkine, and Givental all apply to the A_{r-1}-theory. In particular, this shows that the Witten Integrable Hierarchies Conjecture is true for the A_{r-1}-theory; that is, the total descendant potential function of the A_{r-1}-theory satisfies the r-th Gelfand-Dikii hierarchy.

math.AG

A representation-valued relative Riemann-Hurwitz theorem and the Hurwitz-Hodge bundle

We provide a formula describing the G-module structure of the Hurwitz-Hodge bundle for admissible G-covers in terms of the Hodge bundle of the base curve, and more generally, for describing the G-module structure of the push-forward to the base of any sheaf on a family of admissible G-covers. This formula can be interpreted as a representation-ring-valued relative Riemann-Hurwitz formula for families of admissible G-covers.

math.AG

Logarithmic trace and orbifold products

We give a purely equivariant construction of orbifold products for quotient Deligne-Mumford stacks [X/G] where G is an arbitrary linear algebraic group (not necessarily finite). The key to our construction is the definition of the "logarithmic trace" of an equivariant vector bundle. We also prove that there is an orbifold Chern character homomorphism which induces an isomorphism of a canonical summand in the orbifold Grothendieck ring with the orbifold Chow ring. As an application we obtain an associative orbifold product on the Grothendieck ring of [X/G] (as opposed to its inerita stack) taken with complex coefficients.

math.AG

Geometry and analysis of spin equations

We introduce W-spin structures on a Riemann surface and give a precise definition to the corresponding W-spin equations for any quasi-homogeneous polynomial W. Then, we construct examples of nonzero solutions of spin equations in the presence of Ramond marked points. The main result of the paper is a compactness theorem for the moduli space of the solutions of W-spin equations when W is a non-degenerate, quasi-homogeneous polynomial whose variables all have weight (or fractional degree) wt(x_i) < 1/2. In particular, the compactness theorem holds for the A,D, and E superpotentials.

math.DG