SearcharxivSearch

arXiv subjects

Alex Rodriguez

Publications and source records attributed to Alex Rodriguez.

33 records · Page 2Linked to original sources

Non-parametric learning critical behavior in Ising partition functions: PCA entropy and intrinsic dimension

We provide and critically analyze a framework to learn critical behavior in classical partition functions through the application of non-parametric methods to data sets of thermal configurations. We illustrate our approach in phase transitions in 2D and 3D Ising models. First, we extend previous studies on the intrinsic dimension of 2D partition function data sets, by exploring the effect of volume in 3D Ising data. We find that as opposed to 2D systems for which this quantity has been successfully used in unsupervised characterizations of critical phenomena, in the 3D case its estimation is far more challenging. To circumvent this limitation, we then use the principal component analysis (PCA) entropy, a "Shannon entropy" of the normalized spectrum of the covariance matrix. We find a striking qualitative similarity to the thermodynamic entropy, which the PCA entropy approaches asymptotically. The latter allows us to extract -- through a conventional finite-size scaling analysis with modest lattice sizes -- the critical temperature with less than $1\%$ error for both 2D and 3D models while being computationally efficient. The PCA entropy can readily be applied to characterize correlations and critical phenomena in a huge variety of many-body problems and suggests a (direct) link between easy-to-compute quantities and entropies.

cond-mat.stat-mech

Machine Learning Catalysis of Quantum Tunneling

Optimizing the probability of quantum tunneling between two states, while keeping the resources of the underlying physical system constant, is a task of key importance due to its critical role in various applications. We show that, by applying Machine Learning techniques when the system is coupled to an ancilla, one optimizes the parameters of both the ancillary component and the coupling, ultimately resulting in the maximization of the tunneling probability. We provide illustrative examples for the paradigmatic scenario involving a two-mode system and a two-mode ancilla in the presence of several interacting particles. Physically, the increase of the tunneling probability is rooted in the decrease of the two-well asymmetry due to the coherent oscillations induced by the coupling to the ancilla. We also argue that the enhancement of the tunneling probability is not hampered by weak coupling to noisy environments.

quant-ph

ZundEig: The Structure of the Proton in Liquid Water From Unsupervised Learning

The structure of the excess proton in liquid water has been the subject of lively debate from both experimental and theoretical fronts for the last century. Fluctuations of the proton are typically interpreted in terms of limiting states referred to as the Eigen and Zundel species. Here we put these ideas under the microscope taking advantage of recent advances in unsupervised learning that use local atomic descriptors to characterize environments of acidic water combined with advanced clustering techniques. Our agnostic approach leads to the observation of only a single charged cluster and two neutral ones. We demonstrate that the charged cluster involving the excess proton, is best seen as an ionic topological defect in water's hydrogen bond network forming a single local minimum on the global free-energy landscape. This charged defect is a highly fluxional moiety where the idealized Eigen and Zundel species are neither limiting configurations nor distinct thermodynamic states. Instead, the ionic defect enhances the presence of neutral water defects through strong interactions with the network. We dub the combination of the charged and neutral defect clusters as ZundEig demonstrating that the fluctuations between these local environments provide a general framework for rationalizing more descriptive notions of the proton in the existing literature.

cond-mat.mtrl-sci

Catalysis of quantum tunneling by ancillary system learning

Given the key role that quantum tunneling plays in a wide range of applications, a crucial objective is to maximize the probability of tunneling from one quantum state/level to another, while keeping the resources of the underlying physical system fixed. In this work, we demonstrate that an effective solution to this challenge can be achieved by coupling the tunneling system with an ancillary system of the same kind. By utilizing machine learning techniques, the parameters of both the ancillary system and the coupling can be optimized, leading to the maximization of the tunneling probability. We provide illustrative examples for the paradigmatic scenario involving a two-mode system and a two-mode ancilla with arbitrary couplings and in the presence of several interacting particles. Importantly, the enhancement of the tunneling probability appears to be minimally affected by noise and decoherence in both the system and the ancilla.

quant-ph

Complexity of spin configurations dynamics due to unitary evolution and periodic projective measurements

We study the Hamiltonian dynamics of a many-body quantum system subjected to periodic projective measurements which leads to probabilistic cellular automata dynamics. Given a sequence of measured values, we characterize their dynamics by performing a principal component analysis. The number of principal components required for an almost complete description of the system, which is a measure of complexity we refer to as PCA complexity, is studied as a function of the Hamiltonian parameters and measurement intervals. We consider different Hamiltonians that describe interacting, non-interacting, integrable, and non-integrable systems, including random local Hamiltonians and translational invariant random local Hamiltonians. In all these scenarios, we find that the PCA complexity grows rapidly in time before approaching a plateau. The dynamics of the PCA complexity can vary quantitatively and qualitatively as a function of the Hamiltonian parameters and measurement protocol. Importantly, the dynamics of PCA complexity present behavior that is considerably less sensitive to the specific system parameters for models which lack simple local dynamics, as is often the case in non-integrable models. In particular, we point out a figure of merit that considers the local dynamics and the measurement direction to predict the sensitivity of the PCA complexity dynamics to the system parameters.

cond-mat.stat-mech

Building Flexible, Low-Cost Wireless Access Networks With Magma

Billions of people remain without Internet access due to availability or affordability of service. In this paper, we present Magma, an open and flexible system for building low-cost wireless access networks. Magma aims to connect users where operator economics are difficult due to issues such as low population density or income levels, while preserving features expected in cellular networks such as authentication and billing policies. To achieve this, and in contrast to traditional cellular networks, Magma adopts an approach that extensively leverages Internet design patterns, terminating access network-specific protocols at the edge and abstracting the access network from the core architecture. This decision allows Magma to refactor the wireless core using SDN (software-defined networking) principles and leverage other techniques from modern distributed systems. In doing so, Magma lowers cost and operational complexity for network operators while achieving resilience, scalability, and rich policy support.

cs.NI

DADApy: Distance-based Analysis of DAta-manifolds in Python

DADApy is a python software package for analysing and characterising high-dimensional data manifolds. It provides methods for estimating the intrinsic dimension and the probability density, for performing density-based clustering and for comparing different distance metrics. We review the main functionalities of the package and exemplify its usage in toy cases and in a real-world application. DADApy is freely available under the open-source Apache 2.0 license.

cs.LG

The Collective Burst Mechanism of Angular Jumps in Liquid Water

Understanding the microscopic origins of collective reorientational motions in aqueous systems requires techniques that allow us to reach beyond our chemical imagination. Herein, we elucidate a mechanism using unsupervised learning, showing that large angular jumps in liquid water involve highly cooperative orchestrated motions. Our automatized detection of angular fluctuations, unravels a heterogeneity in the type of angular jumps occurring concertedly in the system. We show that large orientational motions require a highly collective dynamic process involving correlated motion of up to 10% of water molecules in the hydrogen-bond network that form spatially connected clusters. This phenomenon is rooted in the collective fluctuations of the network topology which results in the creation of defects in waves on the ThZ timescale. The mechanism we propose involves a cascade of hydrogen-bond fluctuations underlying angular jumps and provides new insights into the current localized picture of angular jumps, and in its wide use in the interpretations of numerous spectroscopies as well in reorientational dynamics of water near biological and inorganic systems.

cond-mat.soft

High Dimensional Fluctuations in Liquid Water: Combining Chemical Intuition with Unsupervised Learning

The microscopic description of the local structure of water remains an open challenge. Here, we adopt an agnostic approach to understanding water's hydrogen bond network using data harvested from molecular dynamics simulations of an empirical water model. A battery of state-of-the-art unsupervised data-science techniques are used to characterize the free energy landscape of water starting from encoding the water environment using local-atomic descriptors, through dimensionality reduction and finally the use of advanced clustering techniques. Analysis of the free energy at ambient conditions was found to be consistent with a rough single basin and independent of the choice of the water model. We find that the fluctuations of the water network occur in a high-dimensional space which we characterize using a combination of both atomic descriptors and chemical-intuition based coordinates. We demonstrate that a combination of both types of variables are needed in order to adequately capture the complexity of the fluctuations in the hydrogen bond network at different length-scales both at room temperature and also close to the critical point of water. Our results provide a general framework for examining fluctuations in water under different conditions.

cond-mat.soft

Intrinsic dimension of path integrals: data mining quantum criticality and emergent simplicity

Quantum many-body systems are characterized by patterns of correlations that define highly-non trivial manifolds when interpreted as data structures. Physical properties of phases and phase transitions are typically retrieved via simple correlation functions, that are related to observable response functions. Recent experiments have demonstrated capabilities to fully characterize quantum many-body systems via wave-function snapshots, opening new possibilities to analyze quantum phenomena. Here, we introduce a method to data mine the correlation structure of quantum partition functions via their path integral (or equivalently, stochastic series expansion) manifold. We characterize path-integral manifolds generated via state-of-the-art Quantum Monte Carlo methods utilizing the intrinsic dimension (ID) and the variance of distances from nearest neighbors (NN): the former is related to dataset complexity, while the latter is able to diagnose connectivity features of points in configuration space. We show how these properties feature universal patterns in the vicinity of quantum criticality, that reveal how data structures {\it simplify} systematically at quantum phase transitions. This is further reflected by the fact that both ID and variance of NN-distances exhibit universal scaling behavior in the vicinity of second-order and Berezinskii-Kosterlitz-Thouless critical points. Finally, we show how non-Abelian symmetries dramatically influence quantum data sets, due to the nature of (non-commuting) conserved charges in the quantum case. Complementary to neural network representations, our approach represents a first, elementary step towards a systematic characterization of path integral manifolds before any dimensional reduction is taken, that is informative about universal behavior and complexity, and can find immediate application to both experiments and Monte Carlo simulations.

cond-mat.stat-mech

Unsupervised learning universal critical behavior via the intrinsic dimension

The identification of universal properties from minimally processed data sets is one goal of machine learning techniques applied to statistical physics. Here, we study how the minimum number of variables needed to accurately describe the important features of a data set - the intrinsic dimension ($I_d$) - behaves in the vicinity of phase transitions. We employ state-of-the-art nearest neighbors-based $I_d$-estimators to compute the $I_d$ of raw Monte Carlo thermal configurations across different phase transitions: first-, second-order and Berezinskii-Kosterlitz-Thouless. For all the considered cases, we find that the $I_d$ uniquely characterizes the transition regime. The finite-size analysis of the $I_d$ allows not just to identify critical points with an accuracy comparable with methods that rely on {\it a priori} identification of order parameters, but also to determine the corresponding (critical) exponent $ν$ in case of continuous transitions. For the case of topological transitions, this analysis overcomes the reported limitations affecting other unsupervised learning methods. Our work reveals how raw data sets display unique signatures of universal behavior in the absence of any dimensional reduction scheme, and suggest direct parallelism between conventional order parameters in real space, and the intrinsic dimension in the data space.

cond-mat.stat-mech

Automatic topography of high-dimensional data sets by non-parametric Density Peak clustering

Data analysis in high-dimensional spaces aims at obtaining a synthetic description of a data set, revealing its main structure and its salient features. We here introduce an approach providing this description in the form of a topography of the data, namely a human-readable chart of the probability density from which the data are harvested. The approach is based on an unsupervised extension of Density Peak clustering and a non-parametric density estimator that measures the probability density in the manifold containing the data. This allows finding automatically the number and the height of the peaks of the probability density, and the depth of the "valleys" separating them. Importantly, the density estimator provides a measure of the error, which allows distinguishing genuine density peaks from density fluctuations due to finite sampling. The approach thus provides robust and visual information about the density peaks' height, their statistical reliability, and their hierarchical organization, offering a conceptually powerful extension of the standard clustering partitions. We show that this framework is particularly useful in the analysis of complex data sets.

stat.ML

The mechanism of RNA base fraying: molecular dynamics simulations analyzed with core-set Markov state models

The process of RNA base fraying (i.e. the transient opening of the termini of a helix) is involved in many aspects of RNA dynamics. We here use molecular dynamics simulations and Markov state models to characterize the kinetics of RNA fraying and its sequence and direction dependence. In particular, we first introduce a method for determining biomolecular dynamics employing core-set Markov state models constructed using an advanced clustering technique. The method is validated on previously reported simulations. We then use the method to analyze extensive trajectories for four different RNA model duplexes. Results obtained using D. E. Shaw research and AMBER force fields are compared and discussed in detail, and show a non-trivial interplay between the stability of intermediate states and the overall fraying kinetics.

physics.comp-ph

Estimating the intrinsic dimension of datasets by a minimal neighborhood information

Analyzing large volumes of high-dimensional data is an issue of fundamental importance in data science, molecular simulations and beyond. Several approaches work on the assumption that the important content of a dataset belongs to a manifold whose Intrinsic Dimension (ID) is much lower than the crude large number of coordinates. Such manifold is generally twisted and curved, in addition points on it will be non-uniformly distributed: two factors that make the identification of the ID and its exploitation really hard. Here we propose a new ID estimator using only the distance of the first and the second nearest neighbor of each point in the sample. This extreme minimality enables us to reduce the effects of curvature, of density variation, and the resulting computational cost. The ID estimator is theoretically exact in uniformly distributed datasets, and provides consistent measures in general. When used in combination with block analysis, it allows discriminating the relevant dimensions as a function of the block size. This allows estimating the ID even when the data lie on a manifold perturbed by a high-dimensional noise, a situation often encountered in real world data sets. We demonstrate the usefulness of the approach on molecular simulations and image analysis.

stat.ML

METAGUI 3: a graphical user interface for choosing the collective variables in molecular dynamics simulations

Molecular dynamics (MD) simulations allow the exploration of the phase space of biopolymers through the integration of equations of motion of their constituent atoms. The analysis of MD trajectories often relies on the choice of collective variables (CVs) along which the dynamics of the system is projected. We developed a graphical user interface (GUI) for facilitating the interactive choice of the appropriate CVs. The GUI allows: defining interactively new CVs; partitioning the configurations into microstates characterized by similar values of the CVs; calculating the free energies of the microstates for both unbiased and biased (metadynamics) simulations; clustering the microstates in kinetic basins; visualizing the free energy landscape as a function of a subset of the CVs used for the analysis. A simple mouse click allows one to quickly inspect structures corresponding to specific points in the landscape.

physics.comp-ph