SearcharxivSearch

arXiv subjects

Simone Martino

Publications and source records attributed to Simone Martino.

4 recordsLinked to original sources

dynsight: an Open Python Platform for Simulation and Experimental Trajectory Data Analysis

The study of complex many-body systems via analysis of the trajectories of the units that dynamically move and interact within them is a non-trivial task. The workflow for extracting meaningful information from the raw trajectory data is often composed of a series of interconnected steps, such as, (i) identifying and tracking the constitutive objects/particles, resolving their trajectories (e.g., in experimental cases, where these are not automatically available as in typical molecular simulations), (ii) translating the trajectories into data that are easier to handle/analyze by using well suited descriptors, and (iii) extracting meaningful information from such data. Each of these different tasks often requires non-negligible programming skills, the use of various types of representations or methods, and the availability/development of an interface between them. Despite the considerable potential that new tools contributed to each of these individual steps, their integration under a common framework would decrease the barrier to usage (especially by diverse communities of users), avoid fragmentation, and ultimately facilitate the development of new approaches in data analysis. To this end, here we introduce dynsight, an open Python platform that streamlines the extraction and analysis of time-series data from simulation- or experimentally-resolved trajectories. dynsight simplifies workflows, enhances accessibility, and facilitates time-series and trajectories data analysis offering a useful tool to unraveling the dynamic complexity of a variety of systems (or signals) across different scales. dynsight is open source (https://github.com/GMPavanLab/dynsight) and can be easily installed using pip.

cond-mat.mtrl-sci

Data-driven assessment of optimal spatiotemporal resolutions for information extraction in noisy time series data

In general, comprehension of any type of complex system depends on the resolution used to examine the phenomena occurring within it. However, identifying a priori, for example, the best time frequencies/scales to study a certain system over-time, or the spatial distances at which correlations, symmetries, and fluctuations are, most often non-trivial. Here we describe an unsupervised approach that, starting solely from the data of a system, allows learning the characteristic length scales of the dominant key events/processes and the optimal spatiotemporal resolutions to characterize them. We tested this approach on time series data obtained from simulation or experimental trajectories of various example many-body complex systems ranging from the atomic to the macroscopic scale and having diverse internal dynamic complexities. Our method automatically analyzes the system data by analyzing correlations at all relevant inter-particle distances and at all possible inter-frame intervals in which their time series can be subdivided, namely, at all space and time resolutions. The optimal spatiotemporal resolution for studying a certain system thus maximizes information extraction and classification from the system's data, which we prove to be related to the characteristic spatiotemporal length scales of the local/collective physical events dominating it. This approach is broadly applicable and can be used to optimize the study of different types of data (static distributions, time series, or signals). The concept of 'optimal resolution' has a general character and provides a robust basis for characterizing any type of system based on its data, as well as to guide data analysis in general.

physics.data-an

Relevant, hidden, and frustrated information in high-dimensional analyses of complex dynamical systems with internal noise

Extracting from trajectory data meaningful information to understand complex molecular systems might be non-trivial. High-dimensional analyses are typically assumed to be desirable, if not required, to prevent losing important information. But to what extent such high-dimensionality is really needed/beneficial often remains unclear. Here we challenge such a fundamental general problem. As a representative case of a system with internal dynamical complexity, we study atomistic molecular dynamics trajectories of liquid water and ice coexisting in dynamical equilibrium at the solid/liquid transition temperature. To attain an intrinsically high-dimensional analysis, we use as an example the Smooth Overlap of Atomic Positions (SOAP) descriptor, obtaining a large dataset containing 2.56e6 576-dimensional SOAP vectors that we analyze in various ways. Our results demonstrate how the time-series data contained in one single SOAP dimension accounting only <0.001% of the total dataset's variance (neglected and discarded in typical variance-based dimensionality-reduction approaches) allows resolving a remarkable amount of information, classifying/discriminating the bulk of water and ice phases, as well as two solid-interface and liquid-interface layers as four statistically distinct dynamical molecular environments. Adding more dimensions to this one is found not only ineffective but even detrimental to the analysis due to recurrent negligible-information/non-negligible-noise additions and "frustrated information" phenomena leading to information loss. Such effects are proven general and are observed also in completely different systems and descriptors' combinations. This shows how high-dimensional analyses are not necessarily better than low-dimensional ones to elucidate the internal complexity of physical/chemical systems, especially when these are characterized by non-negligible internal noise.

physics.chem-ph

A data driven approach to classify descriptors based on their efficiency in translating noisy trajectories into physically-relevant information

Reconstructing the physical complexity of many-body dynamical systems can be challenging. Starting from the trajectories of their constitutive units (raw data), typical approaches require selecting appropriate descriptors to convert them into time-series, which are then analyzed to extract interpretable information. However, identifying the most effective descriptor is often non-trivial. Here, we report a data-driven approach to compare the efficiency of various descriptors in extracting information from noisy trajectories and translating it into physically relevant insights. As a prototypical system with non-trivial internal complexity, we analyze molecular dynamics trajectories of an atomistic system where ice and water coexist in equilibrium near the solid/liquid transition temperature. We compare general and specific descriptors often used in aqueous systems: number of neighbors, molecular velocities, Smooth Overlap of Atomic Positions (SOAP), Local Environments and Neighbors Shuffling (LENS), Orientational Tetrahedral Order, and distance from the fifth neighbor ($d_5$). Using Onion Clustering -- an efficient unsupervised method for single-point time-series analysis -- we assess the maximum extractable information for each descriptor and rank them via a high-dimensional metric. Our results show that advanced descriptors like SOAP and LENS outperform classical ones due to higher signal-to-noise ratios. Nonetheless, even simple descriptors can rival or exceed advanced ones after local signal denoising. For example, $d_5$, initially among the weakest, becomes the most effective at resolving the system's non-local dynamical complexity after denoising. This work highlights the critical role of noise in information extraction from molecular trajectories and offers a data-driven approach to identify optimal descriptors for systems with characteristic internal complexity.

cond-mat.mtrl-sci