SearcharxivSearch

arXiv subjects

Vipul Periwal

Publications and source records attributed to Vipul Periwal.

At least 19 recordsLinked to original sources

Variational Garrote for Statistical Physics-based Sparse and Robust Variable Selection

Selecting key variables from high-dimensional data is increasingly important in the era of big data. Sparse regression serves as a powerful tool for this purpose by promoting model simplicity and explainability. In this work, we revisit a valuable yet underutilized method, the statistical physics-based Variational Garrote (VG), which introduces explicit feature selection spin variables and leverages variational inference to derive a tractable loss function. We enhance VG by incorporating modern automatic differentiation techniques, enabling scalable and efficient optimization. We evaluate VG on both fully controllable synthetic datasets and complex real-world datasets. Our results demonstrate that VG performs especially well in highly sparse regimes, offering more consistent and robust variable selection than Ridge and LASSO regression across varying levels of sparsity. We also uncover a sharp transition: as superfluous variables are admitted, generalization degrades abruptly and the uncertainty of the selection variables increases. This transition point provides a practical signal for estimating the correct number of relevant variables, an insight we successfully apply to identify key predictors in real-world data. We expect that VG offers strong potential for sparse modeling across a wide range of applications, including compressed sensing and model pruning in machine learning.

cs.LG

Multiplicative learning from observation-prediction ratios

Additive parameter updates, as used in gradient descent and its adaptive extensions, underpin most modern machine-learning optimization. Yet, such additive schemes often demand numerous iterations and intricate learning-rate schedules to cope with scale and curvature of loss functions. Here we introduce Expectation Reflection (ER), a multiplicative learning paradigm that updates parameters based on the ratio of observed to predicted outputs, rather than their differences. ER eliminates the need for ad hoc loss functions or learning-rate tuning while maintaining internal consistency. Extending ER to multilayer networks, we demonstrate its efficacy in image classification, achieving optimal weight determination in a single iteration. We further show that ER can be interpreted as a modified gradient descent incorporating an inverse target-propagation mapping. Together, these results position ER as a fast and scalable alternative to conventional optimization methods for neural-network training.

cs.LG

Efficient and stable derivative-free Steffensen algorithm for root finding

We explore a family of numerical methods, based on the Steffensen divided difference iterative algorithm, that do not evaluate the derivative of the objective functions. The family of methods achieves second-order convergence with two function evaluations per iteration with marginal additional computational cost. An important side benefit of the method is the improvement in stability for different initial conditions compared to the vanilla Steffensen method. We present numerical results for scalar functions, fields, and scalar fields. This family of methods outperforms the Steffensen method with respect to standard quantitative metrics in most cases.

math.NA

Topological approach to void finding applied to the SDSS galaxy map

The structure of the low redshift Universe is dominated by a multi-scale void distribution delineated by filaments and walls of galaxies. The characteristics of voids; such as morphology, average density profile, and correlation function, can be used as cosmological probes. However, their physical properties are difficult to infer due to shot noise and the general lack of tracer particles used to define them. In this work, we construct a robust, topology-based void finding algorithm that utilizes Persistent Homology (PH) to detect persistent features in the data. We apply this approach to a volume limited sub-sample of galaxies in the SDSS I/II Main Galaxy catalog with the $r$-band absolute magnitude brighter than $M_r=-20.19$, and a set of mock catalogs constructed using the Horizon Run 4 cosmological $N$-body simulation. We measure the size distribution of voids, their averaged radial profile, sphericity, and the centroid nearest neighbor separation, using conservative values for the threshold and persistence. We find $32$ topologically robust voids in the SDSS data over the redshift range $0.02 \leq z \leq 0.116$, with effective radii in the range $21 - 56 \, h^{-1} \, {\rm Mpc}$. The median nearest neighbor void separation is found to be $\sim 57 \, h^{-1} \, {\rm Mpc}$, and the median radial void profile is consistent with the expected shape from the mock data.

astro-ph.CO

Annealing approach to root-finding

The Newton-Raphson method is a fundamental root-finding technique with numerous applications in physics. In this study, we propose a parameterized variant of the Newton-Raphson method, inspired by principles from physics. Through analytical and empirical validation, we demonstrate that this novel approach offers increased robustness and faster convergence during root-finding iterations. Furthermore, we establish connections to the Adomian series method and provide a natural interpretation within a series framework. Remarkably, the introduced parameter, akin to a temperature variable, enables an annealing approach. This advancement sets the stage for a fresh exploration of numerical iterative root-finding methodologies.

math.NA

Tight basis cycle representatives for persistent homology of large data sets

Persistent homology (PH) is a popular tool for topological data analysis that has found applications across diverse areas of research. It provides a rigorous method to compute robust topological features in discrete experimental observations that often contain various sources of uncertainties. Although powerful in theory, PH suffers from high computation cost that precludes its application to large data sets. Additionally, most analyses using PH are limited to computing the existence of nontrivial features. Precise localization of these features is not generally attempted because, by definition, localized representations are not unique and because of even higher computation cost. For scientific applications, such a precise location is a sine qua non for determining functional significance. Here, we provide a strategy and algorithms to compute tight representative boundaries around nontrivial robust features in large data sets. To showcase the efficiency of our algorithms and the precision of computed boundaries, we analyze three data sets from different scientific fields. In the human genome, we found an unexpected effect on loops through chromosome 13 and the sex chromosomes, upon impairment of chromatin loop formation. In a distribution of galaxies in the universe, we found statistically significant voids. In protein homologs with significantly different topology, we found voids attributable to ligand-interaction, mutation, and differences between species.

cs.LG

Inference of stochastic time series with missing data

Inferring dynamics from time series is an important objective in data analysis. In particular, it is challenging to infer stochastic dynamics given incomplete data. We propose an expectation maximization (EM) algorithm that iterates between alternating two steps: E-step restores missing data points, while M-step infers an underlying network model of restored data. Using synthetic data generated by a kinetic Ising model, we confirm that the algorithm works for restoring missing data points as well as inferring the underlying model. At the initial iteration of the EM algorithm, the model inference shows better model-data consistency with observed data points than with missing data points. As we keep iterating, however, missing data points show better model-data consistency. We find that demanding equal consistency of observed and missing data points provides an effective stopping criterion for the iteration to prevent overshooting the most accurate model inference. Armed with this EM algorithm with this stopping criterion, we infer missing data points and an underlying network from a time-series data of real neuronal activities. Our method recovers collective properties of neuronal activities, such as time correlations and firing statistics, which have previously never been optimized to fit.

physics.data-an

Dory: Overcoming Barriers to Computing Persistent Homology

Persistent homology (PH) is an approach to topological data analysis (TDA) that computes multi-scale topologically invariant properties of high-dimensional data that are robust to noise. While PH has revealed useful patterns across various applications, computational requirements have limited applications to small data sets of a few thousand points. We present Dory, an efficient and scalable algorithm that can compute the persistent homology of large data sets. Dory uses significantly less memory than published algorithms and also provides significant reductions in the computation time compared to most algorithms. It scales to process data sets with millions of points. As an application, we compute the PH of the human genome at high resolution as revealed by a genome-wide Hi-C data set. Results show that the topology of the human genome changes significantly upon treatment with auxin, a molecule that degrades cohesin, corroborating the hypothesis that cohesin plays a crucial role in loop formation in DNA.

cs.LG

Statistical Physics of Epidemic on Network Predictions for SARS-CoV-2 Parameters

The SARS-CoV-2 pandemic has necessitated mitigation efforts around the world. We use only reported deaths in the two weeks after the first death to determine infection parameters, in order to make predictions of hidden variables such as the time dependence of the number of infections. Early deaths are sporadic and discrete so the use of network models of epidemic spread is imperative, with the network itself a crucial random variable. Location-specific population age distributions and population densities must be taken into account when attempting to fit these events with parametrized models. These characteristics render naive Bayesian model comparison impractical as the networks have to be large enough to avoid finite-size effects. We reformulated this problem as the statistical physics of independent location-specific `balls' attached to every model in a six-dimensional lattice of 56448 parametrized models by elastic springs, with model-specific `spring constants' determined by the stochasticity of network epidemic simulations for that model. The distribution of balls then determines all Bayes posterior expectations. Important characteristics of the contagion are determinable: the fraction of infected patients that die ($0.017\pm 0.009$), the expected period an infected person is contagious ($22 \pm 6$ days) and the expected time between the first infection and the first death ($25 \pm 8$ days) in the US. The rate of exponential increase in the number of infected individuals is $0.18\pm 0.03$ per day, corresponding to 65 million infected individuals in one hundred days from a single initial infection, which fell to 166000 with even imperfect social distancing effectuated two weeks after the first recorded death. The fraction of compliant socially-distancing individuals matters less than their fraction of social contact reduction for altering the cumulative number of infections.

q-bio.PE

Inverse Ising inference from high-temperature re-weighting of observations

Maximum Likelihood Estimation (MLE) is the bread and butter of system inference for stochastic systems. In some generality, MLE will converge to the correct model in the infinite data limit. In the context of physical approaches to system inference, such as Boltzmann machines, MLE requires the arduous computation of partition functions summing over all configurations, both observed and unobserved. We present here a conceptually and computationally transparent data-driven approach to system inference that is based on the simple question: How should the Boltzmann weights of observed configurations be modified to make the probability distribution of observed configurations close to a flat distribution? This algorithm gives accurate inference by using only observed configurations for systems with a large number of degrees of freedom where other approaches are intractable.

stat.ML

Data-driven inference of hidden nodes in networks

The explosion of activity in finding interactions in complex systems is driven by availability of copious observations of complex natural systems. However, such systems, e.g. the human brain, are rarely completely observable. Interaction network inference must then contend with hidden variables affecting the behavior of the observed parts of the system. We present a novel data-driven approach for model inference with hidden variables. From configurations of observed variables, we identify the observed-to-observed, hidden-to-observed, observed-to-hidden, and hidden-to-hidden interactions, the configurations of hidden variables, and the number of hidden variables. We demonstrate the performance of our method by simulating a kinetic Ising model, and show that our method outperforms existing methods. Turning to real data, we infer the hidden nodes in a neuronal network in the salamander retina and a stock market network. We show that predictive modeling with hidden variables is significantly more accurate than that without hidden variables. Finally, an important hidden variable problem is to find the number of clusters in a dataset. We apply our method to classify MNIST handwritten digits. We find that there are about 60 clusters which are roughly equally distributed amongst the digits.

physics.data-an

Causality inference in stochastic systems from neurons to currencies: Profiting from small sample size

Success in modeling complex phenomena such as human perception hinges critically on the availability of data and computational power. Significant progress has been made in modeling such phenomena using probabilistic methods, particularly in image analysis and speech recognition. Maximum Likelihood Estimation (MLE) combined with Bayesian model selection is the basis of much of this progress, as MLE converges to the true model with copious data. In the sciences, large enough datasets are rarae aves, so alternatives to MLE must be developed for small sample size. We introduce a data-driven statistical physics approach to model inference based on minimizing a free energy of data and show superior model recovery for small sample sizes. We demonstrate coupling strength inference in non-equilibrium kinetic Ising models, including in the difficult large coupling variability regime, and show scaling to systems of arbitrary size. As applications, we infer a functional connectivity network in the salamander retina and a currency exchange rate network from time-series data of neuronal spiking and currency exchange rates, respectively. Accurate small sample size inference is critical for devising a profitable currency hedging strategy.

physics.data-an

The jigsaw puzzle of sequence phenotype inference: Piecing together Shannon entropy, importance sampling, and Empirical Bayes

A nucleotide sequence 35 base pairs long can take 1,180,591,620,717,411,303,424 possible values. An example of systems biology datasets, protein binding microarrays, contain activity data from about 40000 such sequences. The discrepancy between the number of possible configurations and the available activities is enormous. Thus, albeit that systems biology datasets are large in absolute terms, they oftentimes require methods developed for rare events due to the combinatorial increase in the number of possible configurations of biological systems. A plethora of techniques for handling large datasets, such as Empirical Bayes, or rare events, such as importance sampling, have been developed in the literature, but these cannot always be simultaneously utilized. Here we introduce a principled approach to Empirical Bayes based on importance sampling, information theory, and theoretical physics in the general context of sequence phenotype model induction. We present the analytical calculations that underlie our approach. We demonstrate the computational efficiency of the approach on concrete examples, and demonstrate its efficacy by applying the theory to publicly available protein binding microarray transcription factor datasets and to data on synthetic cAMP-regulated enhancer sequences. As further demonstrations, we find transcription factor binding motifs, predict the activity of new sequences and extract the locations of transcription factor binding sites. In summary, we present a novel method that is efficient (requiring minimal computational time and reasonable amounts of memory), has high predictive power that is comparable with that of models with hundreds of parameters, and has a limited number of optimized parameters, proportional to the sequence length.

q-bio.QM

The Universality of Cancer

Cancer has been characterized as a constellation of hundreds of diseases differing in underlying mutations and depending on cellular environments. Carcinogenesis as a stochastic physical process has been studied for over sixty years, but there is no accepted standard model. We show that the hazard rates of all cancers are characterized by a simple dynamic stochastic process on a half-line, with a universal linear restoring force balancing a universal simple Brownian motion starting from a universal initial distribution. Only a critical radius defining the transition from normal to tumorigenic genomes distinguishes between different cancer types when time is measured in cell--cycle units. Reparametrizing to chronological time units introduces two additional parameters: the onset of cellular senescence with age and the time interval over which this cessation in replication takes place. This universality implies that there may exist a finite separation between normal cells and tumorigenic cells in all tissue types that may be a viable target for both early detection and preventive therapy.

q-bio.TO

Zen and the Science of Pattern Identification: An Inquiry into Bayesian Skepticism

Finding patterns in data is one of the most challenging open questions in information science. The number of possible relationships scales combinatorially with the size of the dataset, overwhelming the exponential increase in availability of computational resources. Physical insights have been instrumental in developing efficient computational heuristics. Using quantum field theory methods and rethinking three centuries of Bayesian inference, we formulated the problem in terms of finding landscapes of patterns and solved this problem exactly. The generality of our calculus is illustrated by applying it to handwritten digit images and to finding structural features in proteins from sequence alignments without any presumptions about model priors suited to specific datasets. Landscapes of patterns can be uncovered on a desktop computer in minutes.

q-bio.QM

Black Hole Attractor Varieties and Complex Multiplication

Black holes in string theory compactified on Calabi-Yau varieties a priori might be expected to have moduli dependent features. For example the entropy of the black hole might be expected to depend on the complex structure of the manifold. This would be inconsistent with known properties of black holes. Supersymmetric black holes appear to evade this inconsistency by having moduli fields that flow to fixed points in the moduli space that depend only on the charges of the black hole. Moore observed in the case of compactifications with elliptic curve factors that these fixed points are arithmetic, corresponding to curves with complex multiplication. The main goal of this talk is to explore the possibility of generalizing such a characterization to Calabi-Yau varieties with finite fundamental groups.

math.AG

Complex Multiplication Symmetry of Black Hole Attractors

We show how Moore's observation, in the context of toroidal compactifications in type IIB string theory, concerning the complex multiplication structure of black hole attractor varieties, can be generalized to Calabi-Yau compactifications with finite fundamental groups. This generalization leads to an alternative general framework in terms of motives associated to a Calabi-Yau variety in which it is possible to address the arithmetic nature of the attractor varieties in a universal way via Deligne's period conjecture.

hep-th

String field theory, non-commutative Chern-Simons theory and Lie algebra cohomology

Motivated by noncommutative Chern-Simons theory, we construct an infinite class of field theories that satisfy the axioms of Witten's string field theory. These constructions have no propagating open string degrees of freedom. We demonstrate the existence of non-trivial classical solutions. We find Wilson loop-like observables in these examples.

hep-th