SearcharxivSearch

arXiv subjects

Alexander Strang

Publications and source records attributed to Alexander Strang.

18 recordsLinked to original sources

A fast diagonalization algorithm to enable singular value decomposition of large matrices for efficient template matching

Many computational problems possess the following three features: (1) the problem could be solved efficiently if a linear operator could be diagonalized, (2) direct diagonalization is infeasible due to scale, but, (3) the operator is known to commute with a permutation since the problem is (a) symmetric with respect to a change of coordinates, or (b) the operator is block circulant. For example, high-resolution template matching demands repeated convolution of large, high resolution images with an exhaustive list of related templates. The associated linear operator could be compressed, via a low rank approximation, if diagonalized. The required decomposition expensive, but, the entire problem is symmetric to in-plane rotations. In these cases, the eigenfunctions are constrained by the symmetry. These constraints allow efficient decomposition. We illustrate a parallelized algorithm that allows fast diagonalization of any block-circulant matrix. When the index space can be partitioned into $l$ classes of $m$ interchangeable elements, the algorithm reduces storage costs by a factor of $m$ and, if $w$ workers are available, the algorithm demands $\mathcal{O}(l^2m\log(m)/w) + \mathcal{O}(ml^3/w)$ floating point operations per worker yielding a $m^2$ speedup. We use this procedure to decompose a high precision template matching matrix. We compare runtime on a reduced problem, where the fast procedure ran 205 times faster per recovered feature, for 22.5 times as many features. On a typical template matching matrix, the decomposition required 25 times less time than computing the matrix entries. We also demonstrate decomposition of a $\sim$30 times larger matrix that covers the complete set of possible projections represented in a cryo-EM image at 2~{\AA} resolution in 14 minutes. This procedure is more stable and allows recovery of in-plane rotations to arbitrary precision.

q-bio.QM

Sample continuation in Bayesian hierarchical model via variational inference

Posterior distributions arising in ill-posed Bayesian inverse problems are often both analytically intractable and highly sensitive to parameters of the chosen prior family. We aim to understand the sensitivity of intractable posterior distributions to changes in prior assumptions by tracking how a sample representation of the posterior changes as the prior parameters change. This enables sensitivity analysis for small perturbations in the prior, providing insights into the robustness of the posterior estimates under minor changes in assumptions. It also allows solution continuation when dealing with significant alterations in prior beliefs, facilitating a comprehensive understanding of how large shifts in assumptions affect the posterior distribution. We focus on a class of non-conjugate hierarchical models tailored to encourage sparsity in linear inverse problems. The specific hierarchical model of interest is chosen since it is parameterized by a small number of shape parameters, and includes most classical sparsity promoting priors as special cases. As the shape parameters change, the posterior can transition continuously from a tractable unimodal distribution to an intractable multimodal distribution. To track the change in the posterior, we adopt particle based variational inference methods, specifically Stein Variational Gradient Descent (SVGD). SVGD iteratively updates a set of samples to minimize the KL-divergence away from a desired target distribution. We augment SVGD by Birth-Death sampling, which can efficiently exchange mass between separated modes, while simultaneously optimizing the kernel bandwidth used to derive the SVGD update. This method enables the discovery of new modes by tracing the modes as they branch out of a simpler, unimodal posterior, derived within the same family of priors.

stat.ME

Stochastic Gradient Descent in the Saddle-to-Saddle Regime of Deep Linear Networks

Deep linear networks (DLNs) are used as an analytically tractable model of the training dynamics of deep neural networks. While gradient descent in DLNs is known to exhibit saddle-to-saddle dynamics, the impact of stochastic gradient descent (SGD) noise on this regime remains poorly understood. We investigate the dynamics of SGD during training of DLNs in the saddle-to-saddle regime. We model the training dynamics as stochastic Langevin dynamics with anisotropic, state-dependent noise. Under the assumption of aligned and balanced weights, we derive an exact decomposition of the dynamics into a system of one-dimensional per-mode stochastic differential equations. This establishes that the maximal diffusion along a mode precedes the corresponding feature being completely learned. We also derive the stationary distribution of SGD for each mode: in the absence of label noise, its marginal distribution along specific features coincides with the stationary distribution of gradient flow, while in the presence of label noise it approximates a Boltzmann distribution. Finally, we confirm experimentally that the theoretical results hold qualitatively even without aligned or balanced weights. These results establish that SGD noise encodes information about the progression of feature learning but does not fundamentally alter the saddle-to-saddle dynamics.

cs.LG

Do Generalized-Gamma Scale Mixtures of Normals Fit Large Image Datasets?

A scale mixture of normals is a distribution formed by mixing a collection of normal distributions with fixed mean but different variances. A generalized gamma scale mixture draws the variances from a generalized gamma distribution. Generalized gamma scale mixtures of normals have been proposed as an attractive class of parametric priors for Bayesian inference in inverse imaging problems. Generalized gamma scale mixtures have two shape parameters, one that controls the behavior of the distribution about its mode, and the other that controls its tail decay. In this paper, we provide the first demonstration that the prior model is realistic for multiple large imaging data sets. We draw data from remote sensing, medical imaging, and image classification applications. We study the realism of the prior when applied to Fourier and wavelet (Haar and Gabor) transformations of the images, as well as to the coefficients produced by convolving the images against the filters used in the first layer of AlexNet, a popular convolutional neural network trained for image classification. We discuss data augmentation procedures that improve the fit of the model, procedures for identifying approximately exchangeable coefficients, and characterize the parameter regions that best describe the observed data sets. These regions are significantly broader than the region of primary focus in computational work. We show that this prior family provides a substantially better fit to each data set than any of the standard priors it contains. These include Gaussian, Laplace, $\ell_p$, and Student's $t$ priors. Finally, we identify cases where the prior is unrealistic and highlight characteristic features of images that suggest the model will fit poorly.

stat.AP

On the Local Structure and Approximation Stability of Block Isotropic Gaussian Fields

Skew-symmetric functions are a class of functions defined on a product space $M \times M$ that are antisymmetric with respect to the order of their inputs. In [13], the authors proved that non-deterministic skew-symmetric Gaussian fields cannot be stationary or isotropic and proposed an alternative notion: stationarity (isotropy) in each component space. Our work focuses on local quadratic approximations of the associated Gaussian fields. Local quadratic approximations to random fields are random polynomials parametrized by a jointly sampled gradient vector and Hessian matrix. We characterize the distribution of the corresponding random vectors and random matrices. Then, we study the error in the quadratic approximation, which is also a Gaussian field. We investigate the error induced by the quadratic approximation in three senses: the pointwise error, the maximal error over an ellipsoidal region, and the worst-case error for multivariate Gaussian inputs at a given confidence level. Next, we explore the limiting behavior of the worst-case error as the distance between an expansion point and evaluation points approaches zero and infinity. Finally, we study how, as the input dimension increases, the variance of multivariate Gaussian distributions must be restricted to keep the worst-case error bound constant.

math.PR

Disc Game Dynamics: A Latent Space Perspective on Selection and Learning in Games

Evolutionary game theory studies populations that change in response to an underlying game. Often, the functional form relating outcome to player attributes or strategy is complex, preventing mathematical progress. In this work, we axiomatically derive a latent space representation for pairwise, symmetric, zero-sum games by seeking a coordinate space in which the optimal training direction for an agent responding to an opponent depends only on their opponent's coordinates. The associated embedding represents the original game as a linear combination of copies of a simple game, the disc game, in a new coordinate space. In this article, we show that disc-game embedding is useful for studying learning dynamics. We demonstrate that a series of classical evolutionary processes simplify to constrained oscillator equations in the latent space. In particular, the continuous replicator equation reduces to a Hamiltonian system of coupled oscillators that exhibit Poincar\'e recurrence. This reduction allows exact, finite-dimensional closure when the underlying game is finite-rank, and optimal approximation otherwise. It also establishes an exact equivalence between the continuous replicator equation and adaptive dynamics in the transformed coordinates. By identifying a minimal rank representation, the disc game embedding offers numerical methods that could decouple the cost of simulation from the number of attributes used to define agents. These results generalize to metapopulation models that mix inhomogeneously, and to any time-differentiable dynamic where the rate of growth of a type, relative to its expected payout, is a nonnegative function of its frequency. We recommend disc-game embedding as an organizing paradigm for learning and selection in response to symmetric two-player zero-sum games.

cs.GT

A Theoretical Review of Area Production Rates as Test Statistics for Detecting Nonequilibrium Dynamics in Ornstein-Uhlenbeck Processes

A stochastic process is at thermodynamic equilibrium if it obeys time-reversal symmetry; forward and reverse time are statistically indistinguishable at steady state. Non-equilibrium processes break time-reversal symmetry by maintaining circulating probability currents. In physical processes, these currents require a continual use and exchange of energy. Accordingly, signatures of non-equilibrium behavior are important markers of energy use in biophysical systems. In this article we consider a particular signature of nonequilibrium behavior: area production rates. These are, the average rate at which a stochastic process traces out signed area in its projections onto coordinate planes. Area production is an example of a linear observable: a path integral over an observed trajectory against a linear vector field. We provide a summary review of area production rates in Ornstein-Uhlenbeck (OU) processes. Then, we show that, given an OU process, a weighted Frobenius norm of the area production rate matrix is the optimal test statistic for detecting nonequilibrium behavior in the sense that its coefficient of variation decays faster in the length of time observed than the coefficient of variation of any other linear observable. We conclude by showing that this test statistic estimates the entropy production rate of the process.

math-ph

Shared-Endpoint Correlations and Hierarchy in Random Flows on Graphs

We analyze the correlation between randomly chosen edge weights on neighboring edges in a directed graph. This shared-endpoint correlation controls the expected organization of randomly drawn edge flows when the flow on each edge is conditionally independent of the flows on other edges given its endpoints. To model different relationships between endpoints and flow, we draw edge weights in two stages. First, assign a random description to the vertices by sampling random attributes at each vertex. Then, sample a Gaussian process (GP) and evaluate it on the pair of endpoints connected by each edge. We model different relationships between endpoint attributes and flow by varying the kernel associated with the GP. We then relate the expected flow structure to the smoothness class containing functions generated by the GP. We compute the exact shared-endpoint correlation for the squared exponential kernel and provide accurate approximations for Mat\'ern kernels. In addition, we provide asymptotics in both smooth and rough limits and isolate three distinct domains distinguished by the regularity of the ensemble of sampled functions. Taken together, these results demonstrate a consistent effect; smoother functions relating attributes to flow produce more organized flows.

math.ST

Decomposing force fields as flows on graphs reconstructed from stochastic trajectories

Disentangling irreversible and reversible forces from random fluctuations is a challenging problem in the analysis of stochastic trajectories measured from real-world dynamical systems. We present an approach to approximate the dynamics of a stationary Langevin process as a discrete-state Markov process evolving over a graph-representation of phase-space, reconstructed from stochastic trajectories. Next, we utilise the analogy of the Helmholtz-Hodge decomposition of an edge-flow on a contractible simplicial complex with the associated decomposition of a stochastic process into its irreversible and reversible parts. This allows us to decompose our reconstructed flow and to differentiate between the irreversible currents and reversible gradient flows underlying the stochastic trajectories. We validate our approach on a range of solvable and nonlinear systems and apply it to derive insight into the dynamics of flickering red-blood cells and healthy and arrhythmic heartbeats. In particular, we capture the difference in irreversible circulating currents between healthy and passive cells and healthy and arrhythmic heartbeats. Our method breaks new ground at the interface of data-driven approaches to stochastic dynamics and graph signal processing, with the potential for further applications in the analysis of biological experiments and physiological recordings. Finally, it prompts future analysis of the convergence of the Helmholtz-Hodge decomposition in discrete and continuous spaces.

cond-mat.stat-mech

The Almost Sure Evolution of Hierarchy Among Similar Competitors

While generic competitive systems exhibit mixtures of hierarchy and cycles, real-world systems are predominantly hierarchical. We demonstrate and extend a mechanism for hierarchy; systems with similar agents approach perfect hierarchy in expectation. A variety of evolutionary mechanisms plausibly select for nearly homogeneous populations, however, extant work does not explicitly link selection dynamics to hierarchy formation via population concentration. Moreover, previous work lacked numerical demonstration. This paper contributes in four ways. First, populations that converge to perfect hierarchy in expectation converge to hierarchy in probability. Second, we analyze hierarchy formation in populations subject to the continuous replicator dynamic with diffusive exploration, linking population dynamics to emergent structure. Third, we show how to predict the degree of cyclicity sustained by concentrated populations at internal equilibria. This theory can differentiate learning rules and random payout models. Finally, we provide direct numerical evidence by simulating finite populations of agents subject to a modified Moran process with Gaussian exploration. As examples, we consider three bimatrix games and an ensemble of games with random payouts. Through this analysis, we explicitly link the temporal dynamics of a population undergoing selection to the development of hierarchy.

q-bio.PE

Machine learning in and out of equilibrium

The algorithms used to train neural networks, like stochastic gradient descent (SGD), have close parallels to natural processes that navigate a high-dimensional parameter space -- for example protein folding or evolution. Our study uses a Fokker-Planck approach, adapted from statistical physics, to explore these parallels in a single, unified framework. We focus in particular on the stationary state of the system in the long-time limit, which in conventional SGD is out of equilibrium, exhibiting persistent currents in the space of network parameters. As in its physical analogues, the current is associated with an entropy production rate for any given training trajectory. The stationary distribution of these rates obeys the integral and detailed fluctuation theorems -- nonequilibrium generalizations of the second law of thermodynamics. We validate these relations in two numerical examples, a nonlinear regression network and MNIST digit classification. While the fluctuation theorems are universal, there are other aspects of the stationary state that are highly sensitive to the training details. Surprisingly, the effective loss landscape and diffusion matrix that determine the shape of the stationary distribution vary depending on the simple choice of minibatching done with or without replacement. We can take advantage of this nonequilibrium sensitivity to engineer an equilibrium stationary state for a particular application: sampling from a posterior distribution of network weights in Bayesian machine learning. We propose a new variation of stochastic gradient Langevin dynamics (SGLD) that harnesses without replacement minibatching. In an example system where the posterior is exactly known, this SGWORLD algorithm outperforms SGLD, converging to the posterior orders of magnitude faster as a function of the learning rate.

cs.LG

Path-following methods for Maximum a Posteriori estimators in Bayesian hierarchical models: How estimates depend on hyperparameters

Maximum a posteriori (MAP) estimation, like all Bayesian methods, depends on prior assumptions. These assumptions are often chosen to promote specific features in the recovered estimate. The form of the chosen prior determines the shape of the posterior distribution, thus the behavior of the estimator and complexity of the associated optimization problem. Here, we consider a family of Gaussian hierarchical models with generalized gamma hyperpriors designed to promote sparsity in linear inverse problems. By varying the hyperparameters, we move continuously between priors that act as smoothed $\ell_p$ penalties with flexible $p$, smoothing, and scale. We then introduce a predictor-corrector method that tracks MAP solution paths as the hyperparameters vary. Path following allows a user to explore the space of possible MAP solutions and to test the sensitivity of solutions to changes in the prior assumptions. By tracing paths from a convex region to a non-convex region, the user can find local minimizers in strongly sparsity promoting regimes that are consistent with a convex relaxation derived using related prior assumptions. We show experimentally that these solutions. are less error prone than direct optimization of the non-convex problem.

stat.ME

Principal Trade-off Analysis

How are the advantage relations between a set of agents playing a game organized and how do they reflect the structure of the game? In this paper, we illustrate "Principal Trade-off Analysis" (PTA), a decomposition method that embeds games into a low-dimensional feature space. We argue that the embeddings are more revealing than previously demonstrated by developing an analogy to Principal Component Analysis (PCA). PTA represents an arbitrary two-player zero-sum game as the weighted sum of pairs of orthogonal 2D feature planes. We show that the feature planes represent unique strategic trade-offs and truncation of the sequence provides insightful model reduction. We demonstrate the validity of PTA on a quartet of games (Kuhn poker, RPS+2, Blotto, and Pokemon). In Kuhn poker, PTA clearly identifies the trade-off between bluffing and calling. In Blotto, PTA identifies game symmetries, and specifies strategic trade-offs associated with distinct win conditions. These symmetries reveal limitations of PTA unaddressed in previous work. For Pokemon, PTA recovers clusters that naturally correspond to Pokemon types, correctly identifies the designed trade-off between those types, and discovers a rock-paper-scissor (RPS) cycle in the Pokemon generation type - all absent any specific information except game outcomes.

cs.GT

Hierarchical Ensemble Kalman Methods with Sparsity-Promoting Generalized Gamma Hyperpriors

This paper introduces a computational framework to incorporate flexible regularization techniques in ensemble Kalman methods for nonlinear inverse problems. The proposed methodology approximates the maximum a posteriori (MAP) estimate of a hierarchical Bayesian model characterized by a conditionally Gaussian prior and generalized gamma hyperpriors. Suitable choices of hyperparameters yield sparsity-promoting regularization. We propose an iterative algorithm for MAP estimation, which alternates between updating the unknown with an ensemble Kalman method and updating the hyperparameters in the regularization to promote sparsity. The effectiveness of our methodology is demonstrated in several computed examples, including compressed sensing and subsurface flow inverse problems.

stat.CO

Similarity Suppresses Cyclicity: Why Similar Competitors Form Hierarchies

Competitive systems can exhibit both hierarchical (transitive) and cyclic (intransitive) structures. Despite theoretical interest in cyclic competition, which offers richer dynamics, and occupies a larger subset of the space of possible competitive systems, most real-world systems are predominantly transitive. Why? Here, we introduce a generic mechanism which promotes transitivity, even when there is ample room for cyclicity. Consider a competitive system where outcomes are mediated by competitor attributes via a performance function. We demonstrate that, if competitive outcomes depend smoothly on competitor attributes, then similar competitors compete transitively. We quantify the rate of convergence to transitivity given the similarity of the competitors and the smoothness of the performance function. Thus, we prove the adage regarding apples and oranges. Similar objects admit well ordered comparisons. Diverse objects may not. To test that theory, we run a series of evolution experiments designed to mimic genetic training algorithms. We consider a series of canonical bimatrix games and an ensemble of random performance functions that demonstrate the generality of our mechanism, even when faced with highly cyclic games. We vary the training parameters controlling the evolution process, and the shape parameters controlling the performance function, to evaluate the robustness of our results. These experiments illustrate that, if competitors evolve to optimize performance, then their traits may converge, leading to transitivity.

q-bio.PE

A Variational Inference Approach to Inverse Problems with Gamma Hyperpriors

Hierarchical models with gamma hyperpriors provide a flexible, sparse-promoting framework to bridge $L^1$ and $L^2$ regularizations in Bayesian formulations to inverse problems. Despite the Bayesian motivation for these models, existing methodologies are limited to \textit{maximum a posteriori} estimation. The potential to perform uncertainty quantification has not yet been realized. This paper introduces a variational iterative alternating scheme for hierarchical inverse problems with gamma hyperpriors. The proposed variational inference approach yields accurate reconstruction, provides meaningful uncertainty quantification, and is easy to implement. In addition, it lends itself naturally to conduct model selection for the choice of hyperparameters. We illustrate the performance of our methodology in several computed examples, including a deconvolution problem and sparse identification of dynamical systems from time series data.

stat.ME

The Network HHD: Quantifying Cyclic Competition in Trait-Performance Models of Tournaments

Competitive tournaments appear in sports, politics, population ecology, and animal behavior. All of these fields have developed methods for rating competitors and ranking them accordingly. A tournament is intransitive if it is not consistent with any ranking. Intransitive tournaments contain rock-paper-scissor type cycles. The discrete Helmholtz-Hodge decomposition (HHD) is well adapted to describing intransitive tournaments. It separates a tournament into perfectly transitive and perfectly cyclic components, where the perfectly transitive component is associated with a set of ratings. The size of the cyclic component can be used as a measure of intransitivity. Here we show that the HHD arises naturally from two classes of tournaments with simple statistical interpretations. We then discuss six different sets of assumptions that define equivalent decompositions. This analysis motivates the choice to use the HHD among other existing methods. Success in competition is typically mediated by the traits of the competitors. A trait-performance model assumes that the probability that one competitor beats another can be expressed as a function of their traits. We show that, if the traits of each competitor are drawn independently and identically from a trait distribution then the expected degree of intransitivity in the network can be computed explicitly. Using this result we show that increasing the number of pairs of competitors who could compete promotes cyclic competition, and that increasing the correlation in the performance of $A$ against $B$ with the performance of $A$ against $C$ promotes transitive competition. The expected size of cyclic competition can thus be understood by analyzing this correlation. An illustrative example is provided.

physics.soc-ph

Relationships Between Characteristic Path Length, Efficiency, Clustering Coefficients, and Graph Density

The graph theoretic properties of the clustering coefficient, characteristic (or average) path length, global and local efficiency, provide valuable information regarding the structure of a graph. These four properties have applications to biological and social networks and have dominated much of the the literature in these fields. While much work has done in applied settings, there has yet to be a mathematical comparison of these metrics from a theoretical standpoint. Motivated by networks appearing in neuroscience, we show in this paper that these properties can be linked together using a single property - graph density.

math.CO