SearcharxivSearch

arXiv subjects

Matteo Rucco

Publications and source records attributed to Matteo Rucco.

15 recordsLinked to original sources

PesTwin: A modular agent-based framework for pest and vector population control

Species-specific pest and vector control strategies, including the sterile insect technique, Wolbachia-based interventions, and genetic control technologies, offer powerful alternatives to broad-spectrum chemical control, with applications ranging from targeted crop protection to large-scale disease control. Among these, genetic control technologies are advancing rapidly, but the pace of technological development is outstripping the modelling tools needed to predict outcomes, guide technology design and its implementation, compare alternative strategies across different use settings, and support regulatory and operational decision-making. Here we present PesTwin, an agent-based modelling framework for simulating genetic control technologies across species, ecological settings, and deployment strategies within a common computational environment. PesTwin captures stochastic demographic effects, species-specific life-history traits, heterogeneous dispersal, and temporal variation in resource availability and infestation pressure. We validate PesTwin against published laboratory cage data from four genetic control systems, drawn from three studies, in two insect species, showing close agreement between predicted and observed population trajectories, including their replicate-to-replicate variability. We then illustrate how the same validated models extend beyond the cage to spatially explicit, field-scale scenarios, using PesTwin to explore how the timing, density and spatial placement of releases shape suppression and spread across heterogeneous landscapes. By making genetic control systems testable in silico before they are built or released, PesTwin can shorten the path from laboratory construct to field intervention: informing which constructs to prioritise, how to design the experiments that test them, where and when to release, and what evidence is needed to evaluate them.

q-bio.PE

The Shape of Chocolate: A Topological Perspective on Food Microstructure

We present a computational framework for characterizing the molecular self-organization of cocoa butter (Theobroma cacao) during dark chocolate tempering through the lens of Topological Data Analysis (TDA). A physics-inspired particle simulation models N=100 triglyceride molecules across the full temperature range 15--60 degrees C, spanning all six crystalline polymorphs of cocoa butter (Forms I--VI) as well as the melt and superheating regimes. At each temperature tick, we construct a Vietoris-Rips filtration and compute the persistent homology groups H0 (connected components), H1 (independent cycles), and H2 (3D voids). The resulting persistence diagrams are analyzed via persistent entropy E = -sum_i p_i log2(p_i), where p_i = l_i / sum_j l_j and l_i = death_i - birth_i denotes feature lifetime; essential classes are assigned death = m+1 (m = eps_max) following the standard persistent entropy convention (Rucco 2026, arXiv:2602.09058). Our results demonstrate that Form V (the optimal tempering polymorph, 29.5--34 degrees C) is characterized by a distinctive topological signature: a local minimum in the H0 persistent entropy (E0 = 5.74 +/- 0.04 bits), a pronounced depression in the first Betti number beta_1 (1562 +/- 35), and a global minimum in the H2 entropy (E2 = 12.29 +/- 0.25 bits) reflecting coherent inter-bilayer lamellar cavities. Via Theorem 1 and Corollary 1 of Rucco (2026), persistent entropy is proven to separate the ordered and disordered phases by an asymptotically non-vanishing gap whenever a phase transition induces the creation or destruction of topological mass at macroscopic scales -- a condition we verify empirically across all eight cocoa butter regimes. These findings suggest that TDA-based metrics could serve as non-invasive quality indicators for industrial chocolate tempering processes.

cond-mat.mtrl-sci

PesTwin: a biology-informed Digital Twin for enabling precision farming

In a context of growing agricultural demand and new challenges related to food security and accessibility, boosting agricultural productivity is more important than ever. Reducing the damage caused by invasive insect species is a crucial lever to achieve this objective. In support of these challenges, and in line with the principles of precision agriculture and Integrated Pest Management (IPM), an innovative simulation framework is presented, aiming to become the digital twin of a pest invasion. Through a flexible rule-based approach of the Agent-Based Modeling (ABM) paradigm, the framework supports the fine-tuning of the main ecological interactions of the pest with its crop host and the environment. Forecasting insect infestation in realistic scenarios, considering both spatial and temporal dimensions, is made possible by integrating heterogeneous data sources: pest biodata collected in the laboratory, environmental data from weather stations, and GIS data of a real crop field. In this study, an application to the global pest of soft fruit, the invasive fruit fly Drosophila suzukii, also known as Spotted Wing Drosophila (SWD), is presented.

q-bio.QM

Persistent Entropy as a Detector of Phase Transitions

Persistent entropy is a scalar summary of persistence barcodes widely used to detect regime changes, yet there is no account of when a structural change in a barcode must produce a detectable change in entropy. We establish a model-agnostic theorem supplying such conditions. Treating persistence diagrams as random objects indexed by a control parameter, we identify a dispersion-condensation mechanism in the normalized persistence weights and derive an explicit lower bound on the entropy difference between the two regimes, valid with high probability at finite sample size and insensitive to the absolute scale of bar lifetimes. We also give a procedure for verifying the hypotheses on empirical barcodes. Applied to convolutional networks, the criterion shows that the circular organization of learned filters reported by Gabrielsson and Carlsson emerges through a sharp topological phase transition, and locates its onset: within a few hundred iterations on MNIST, but an order of magnitude later on CIFAR-10. The same criterion detects the Kuramoto synchronization and Vicsek order-disorder transitions.

stat.ML

Application of the representative measure approach to assess the reliability of decision trees in dealing with unseen vehicle collision data

Machine learning algorithms are fundamental components of novel data-informed Artificial Intelligence architecture. In this domain, the imperative role of representative datasets is a cornerstone in shaping the trajectory of artificial intelligence (AI) development. Representative datasets are needed to train machine learning components properly. Proper training has multiple impacts: it reduces the final model's complexity, power, and uncertainties. In this paper, we investigate the reliability of the $\varepsilon$-representativeness method to assess the dataset similarity from a theoretical perspective for decision trees. We decided to focus on the family of decision trees because it includes a wide variety of models known to be explainable. Thus, in this paper, we provide a result guaranteeing that if two datasets are related by $\varepsilon$-representativeness, i.e., both of them have points closer than $\varepsilon$, then the predictions by the classic decision tree are similar. Experimentally, we have also tested that $\varepsilon$-representativeness presents a significant correlation with the ordering of the feature importance. Moreover, we extend the results experimentally in the context of unseen vehicle collision data for XGboost, a machine-learning component widely adopted for dealing with tabular data.

cs.LG

An In-Depth Analysis of Data Reduction Methods for Sustainable Deep Learning

In recent years, Deep Learning has gained popularity for its ability to solve complex classification tasks, increasingly delivering better results thanks to the development of more accurate models, the availability of huge volumes of data and the improved computational capabilities of modern computers. However, these improvements in performance also bring efficiency problems, related to the storage of datasets and models, and to the waste of energy and time involved in both the training and inference processes. In this context, data reduction can help reduce energy consumption when training a deep learning model. In this paper, we present up to eight different methods to reduce the size of a tabular training dataset, and we develop a Python package to apply them. We also introduce a representativeness metric based on topology to measure how similar are the reduced datasets and the full training dataset. Additionally, we develop a methodology to apply these data reduction methods to image datasets for object detection tasks. Finally, we experimentally compare how these data reduction methods affect the representativeness of the reduced dataset, the energy consumption and the predictive performance of the model.

cs.LG

Towards personalized diagnosis of Glioblastoma in Fluid-attenuated inversion recovery (FLAIR) by topological interpretable machine learning

Glioblastoma multiforme (GBM) is a fast-growing and highly invasive brain tumour, it tends to occur in adults between the ages of 45 and 70 and it accounts for 52 percent of all primary brain tumours. Usually, GBMs are detected by magnetic resonance images (MRI). Among MRI, Fluid-attenuated inversion recovery (FLAIR) sequence produces high quality digital tumour representation. Fast detection and segmentation techniques are needed for overcoming subjective medical doctors (MDs) judgment. In the present investigation, we intend to demonstrate by means of numerical experiments that topological features combined with textural features can be enrolled for GBM analysis and morphological characterization on FLAIR. To this extent, we have performed three numerical experiments. In the first experiment, Topological Data Analysis (TDA) of a simplified 2D tumour growth mathematical model had allowed to understand the bio-chemical conditions that facilitate tumour growth: the higher the concentration of chemical nutrients the more virulent the process. In the second experiment topological data analysis was used for evaluating GBM temporal progression on FLAIR recorded within 90 days following treatment (e.g., chemo-radiation therapy - CRT) completion and at progression. The experiment had confirmed that persistent entropy is a viable statistics for monitoring GBM evolution during the follow-up period. In the third experiment we had developed a novel methodology based on topological and textural features and automatic interpretable machine learning for automatic GBM classification on FLAIR. The algorithm reached a classification accuracy up to the 97%.

eess.IV

Topological Run-time Monitoring for Complex Systems

In this paper we introduce a new data-driven run-time monitoring system for analysing the behaviour of time evolving complex systems. The monitor controls the evolution of the whole system but it is mined from the data produced by its single interacting components. Relevant behavioural changes happening at the component level and that are responsible for global system evolution are captured by the monitor. Topological Data Analysis is used for shaping and analysing the data for mining an automaton mimicking the global system dynamics, the so-called Persistent Entropy Automaton (PEA). A slight augmented PEA, the monitor, can be used to run current or past executions of the system to mine temporal invariants, for instance through statistical reasoning. Such invariants can be formulated as properties of a temporal logic, e.g. bounded LTL, that can be run-time model-checked. We have performed a feasibility assessment of the PEA and the associated monitoring system by analysing a simulated biological complex system, namely the human immune system. The application of the monitor to simulated traces reveals temporal properties that should be satisfied in order to reach immunization memory.

cs.LO

Persistent Entropy for Separating Topological Features from Noise in Vietoris-Rips Complexes

Persistent homology studies the evolution of k-dimensional holes along a nested sequence of simplicial complexes (called a filtration). The set of bars (i.e. intervals) representing birth and death times of k-dimensional holes along such sequence is called the persistence barcode. k-Dimensional holes with short lifetimes are informally considered to be "topological noise", and those with long lifetimes are considered to be "topological features" associated to the filtration. Persistent entropy is defined as the Shannon entropy of the persistence barcode of a given filtration. In this paper we present new important properties of persistent entropy of Cech and Vietoris-Rips filtrations. Among the properties, we put a focus on the stability theorem that allows to use persistent entropy for comparing persistence barcodes. Later, we derive a simple method for separating topological noise from features in Vietoris-Rips filtrations.

cs.OH

Topological classifier for detecting the emergence of epileptic seizures

In this work we study how to apply topological data analysis to create a method suitable to classify EEGs of patients affected by epilepsy. The topological space constructed from the collection of EEGs signals is analyzed by Persistent Entropy acting as a global topological feature for discriminating between healthy and epileptic signals. The Physionet data-set has been used for testing the classifier.

q-bio.NC

A new topological entropy-based approach for measuring similarities among piecewise linear functions

In this paper we present a novel methodology based on a topological entropy, the so-called persistent entropy, for addressing the comparison between discrete piecewise linear functions. The comparison is certified by the stability theorem for persistent entropy. The theorem is used in the implementation of a new algorithm. The algorithm transforms a discrete piecewise linear function into a filtered simplicial complex that is analyzed with persistent homology and persistent entropy. Persistent entropy is used as discriminant feature for solving the supervised classification problem of real long length noisy signals of DC electrical motors. The quality of classification is stated in terms of the area under receiver operating characteristic curve (AUC=94.52%).

cs.DM

Topological characterization of S[B] systems: From data to models of complexity

In this paper we propose a methodology for deriving a model of a complex system by exploiting the information extracted from Topological Data Analysis. Central to our approach is the S[B] paradigm in which a complex system is represented by a two-level model. One level, the structural S one, is derived using the newly introduced quantitative concept of Persistent Entropy. The other level, the behavioral B one, is characterized by a network of interacting computational agents described by a Higher Dimensional Automaton. The methodology yields also a representation of the evolution of the derived two-level model as a Persistent Entropy Automaton. The presented methodology is applied to a real case study, the Idiotypic Network of the mammal immune system.

cs.ET

Neural Hypernetwork Approach for Pulmonary Embolism diagnosis

This work introduces an integrative approach based on Q-analysis with machine learning. The new approach, called Neural Hypernetwork, has been applied to a case study of pulmonary embolism diagnosis. The objective of the application of neural hyper-network to pulmonary embolism (PE) is to improve diagnose for reducing the number of CT-angiography needed. Hypernetworks, based on topological simplicial complex, generalize the concept of two-relation to many-body relation. Furthermore, Hypernetworks provide a significant generalization of network theory, enabling the integration of relational structure, logic and analytic dynamics. Another important results is that Q-analysis stays close to the data, while other approaches manipulate data, projecting them into metric spaces or applying some filtering functions to highlight the intrinsic relations. A pulmonary embolism (PE) is a blockage of the main artery of the lung or one of its branches, frequently fatal. Our study uses data on 28 diagnostic features of 1,427 people considered to be at risk of PE. The resulting neural hypernetwork correctly recognized 94% of those developing a PE. This is better than previous results that have been obtained with other methods (statistical selection of features, partial least squares regression, topological data analysis in a metric space).

physics.med-ph

Using Topological Data Analysis for diagnosis pulmonary embolism

Pulmonary Embolism (PE) is a common and potentially lethal condition. Most patients die within the first few hours from the event. Despite diagnostic advances, delays and underdiagnosis in PE are common.To increase the diagnostic performance in PE, current diagnostic work-up of patients with suspected acute pulmonary embolism usually starts with the assessment of clinical pretest probability using plasma d-Dimer measurement and clinical prediction rules. The most validated and widely used clinical decision rules are the Wells and Geneva Revised scores. We aimed to develop a new clinical prediction rule (CPR) for PE based on topological data analysis and artificial neural network. Filter or wrapper methods for features reduction cannot be applied to our dataset: the application of these algorithms can only be performed on datasets without missing data. Instead, we applied Topological data analysis (TDA) to overcome the hurdle of processing datasets with null values missing data. A topological network was developed using the Iris software (Ayasdi, Inc., Palo Alto). The PE patient topology identified two ares in the pathological group and hence two distinct clusters of PE patient populations. Additionally, the topological netowrk detected several sub-groups among healthy patients that likely are affected with non-PE diseases. TDA was further utilized to identify key features which are best associated as diagnostic factors for PE and used this information to define the input space for a back-propagation artificial neural network (BP-ANN). It is shown that the area under curve (AUC) of BP-ANN is greater than the AUCs of the scores (Wells and revised Geneva) used among physicians. The results demonstrate topological data analysis and the BP-ANN, when used in combination, can produce better predictive models than Wells or revised Geneva scores system for the analyzed cohort

physics.med-ph