SearcharxivSearch

arXiv subjects

Francisco A. Rodrigues

Publications and source records attributed to Francisco A. Rodrigues.

At least 19 recordsLinked to original sources

Noniterative Likelihood-Derived Estimation through Auxiliary Estimating Equations

Closed-form estimators are useful when parametric models are repeatedly refitted but likelihood maximization is iterative. We study a likelihood-derived construction of auxiliary estimating equations obtained by differentiating a positive auxiliary function and centering the derivatives under a baseline model. The resulting estimators are just-identified Z-estimators, or equivalently GMM estimators; the contribution is a constructive route to explicitly invertible equations rather than a replacement for general estimating-equation theory. We establish local existence, uniqueness and asymptotic normality of the selected root, characterize optimal linear combinations through Godambe information, and give constrained variants and influence-function diagnostics. Transformed exponential families, Beta, Weibull AFT regression, Wishart covariance estimation, zero-inflated counts, tail modelling and copula dependence illustrate the construction. Simulations and survival-tree split screening quantify the statistical-computational trade-off.

stat.ME

Epidemic spreading on multigraphs

Multigraphs are graphs in which multiple links between pairs of nodes are allowed, whereas they are forbidden in simple graphs, the latter being widely used in network science. Simple graphs generated by the configuration model have served as a benchmark for validating theoretical approaches to dynamical processes on networks. However, generating large scale-free networks with degree exponent $γ<3$ introduces uncontrolled disassortative correlations and severe computational limitations due to the prohibition of reconnecting hubs. These constraints do not exist in multigraphs. We investigate how multiple connections affect epidemic spreading by comparing several epidemic models exhibiting an active steady state on simple graphs and multigraphs sharing the same degree sequence and natural upper cutoff. By analyzing epidemic thresholds, finite-size scaling, and localization, we show that differences between simple graphs and multigraphs emerge only when epidemic activity can persist on isolated hubs (star subgraphs) for times exponentially long in the hub degree. Our results remove a methodological barrier to the study of dynamical processes on large scale-free networks.

physics.soc-ph

Impacts of bridging nodes on the epidemic activation mechanisms

Bridging nodes, which connect critical components of a network, play an important role in maintaining structural integrity and facilitating communication within the network, representing indirect yet relevant connections. Epidemic triggering mechanisms in networks often involve long-range mutual activation of hubs, mediated by paths composed of low-degree nodes. While low-degree nodes are abundant in networks, their role in bridging central nodes in epidemic activation mechanisms has not been thoroughly analyzed. Starting with a backbone network with a power-law degree distribution, we investigate the role of adding degree-2 bridging nodes that are preferentially attached to hubs. Our findings reveal that bridging nodes can mediate an indirect feedback interaction between hubs that modifies the epidemic localization and activation mechanisms of the epidemic processes with recurrent infections. In particular, the collective activation observed in the presence of waning immunity, which produces a finite epidemic threshold in power-law networks with degree exponent $γ>3$, is altered to a localized activation with a vanishing threshold. Our numerical results are analytically supported by the non-backtracking matrix properties that emerge in the recurrent dynamical message-passing theory.

physics.soc-ph

Discovering equations from data: symbolic regression in dynamical systems

The process of discovering equations from data lies at the heart of physics and in many other areas of research, including mathematical ecology and epidemiology. Recently, machine learning methods known as symbolic regression emerged as a way to automate this task. This study presents an overview of the current literature on symbolic regression, while also comparing the efficiency of five state-of-the-art methods in recovering the governing equations from nine processes, including chaotic dynamics and epidemic models. Benchmark results demonstrate the PySR method as the most suitable for inferring equations, with some estimates being indistinguishable from the original analytical forms. These results highlight the potential of symbolic regression as a robust tool for inferring and modeling real-world phenomena.

cs.LG

Neural networks for dengue forecasting: a systematic review

Background: Early forecasts of dengue are an important tool for disease mitigation. Neural networks are powerful predictive models that have made contributions to many areas of public health. In this study, we reviewed the application of neural networks in the dengue forecasting literature, with the objective of informing model design for future work. Methods: Following PRISMA guidelines, we conducted a systematic search of studies that use neural networks to forecast dengue in human populations. We summarized the relative performance of neural networks and comparator models, architectures and hyper-parameters, choices of input features, geographic spread, and model transparency. Results: Sixty two papers were included. Most studies implemented shallow feed-forward neural networks, using historical dengue incidence and climate variables. Prediction horizons varied greatly, as did the model selection and evaluation approach. Building on the strengths of neural networks, most studies used granular observations at the city level, or on its subdivisions, while also commonly employing weekly data. Performance of neural networks relative to comparators, such as tree-based supervised models, varied across study contexts, and we found that 63% of all studies do include at least one such model as a baseline, and in those cases about half of the studies report neural networks as the best performing model. Conclusions: The studies suggest that neural networks can provide competitive forecasts for dengue, and can reliably be included in the set of candidate models for future dengue prediction efforts. The use of deep networks is relatively unexplored but offers promising avenues for further research, as does the use of a broader set of input features and prediction in light of structural changes in the data generation mechanism.

cs.LG

Prediction and inference in complex networks: a brief review and perspectives

Inference and prediction are fundamental to the study of complex systems, where network data are often incomplete, inaccurate or obtained indirectly. In this paper, we review recent advances in network sampling and comparison, as well as in link prediction and network reconstruction from time series. We summarise key methodological developments and emerging approaches that integrate statistical and machine learning perspectives. We also outline promising research directions for enhancing the inference and prediction of complex networked systems.

cond-mat.stat-mech

The systemic impact of edges in financial networks

In this paper, we assess how the stability of financial networks is affected by interconnectedness considering its tiniest variation: the edge. We compute the impact of edges as the percentage difference in the systemic risk (SR) of the whole network caused by the inclusion of that edge. We apply this framework to a thorough Brazilian dataset to compute the impact of bank-firm edges. After observing that (i) edges are heterogeneous regarding their impact on the SR, and (ii) the fraction of edges whose impact on the SR is non-positive increases with the level of the initial shock, we use machine learning techniques to try to predict two variables: the criticality of the edges (defining as critical an edge whose impact on SR is significantly greater than that of the others) and the sign of the edge impact. The level of accuracy obtained in these prediction exercises was very high. These results have important implications for the development macroprudential policies aimed at financial stability. Our framework allows to identify, based on features related to the origin and destination nodes of the edge (i.e., the lending bank and the borrowing firm), whether an additional loan will have a significant and positive impact on the SR.

cond-mat.stat-mech

Beyond the Power Law: Estimation, Goodness-of-Fit, and a Semiparametric Extension in Complex Networks

Scale-free networks play a fundamental role in the study of complex networks and various applied fields due to their ability to model a wide range of real-world systems. A key characteristic of these networks is their degree distribution, which often follows a power-law distribution, where the probability mass function is proportional to $x^{-α}$, with $α$ typically ranging between $2 < α< 3$. In this paper, we introduce Bayesian inference methods to obtain more accurate estimates than those obtained using traditional methods, which often yield biased estimates, and precise credible intervals. Through a simulation study, we demonstrate that our approach provides nearly unbiased estimates for the scaling parameter, enhancing the reliability of inferences. We also evaluate new goodness-of-fit tests to improve the effectiveness of the Kolmogorov-Smirnov test, commonly used for this purpose. Our findings show that the Watson test offers superior power while maintaining a controlled type I error rate, enabling us to better determine whether data adheres to a power-law distribution. Finally, we propose a piecewise extension of this model to provide greater flexibility, evaluating the estimation and its goodness-of-fit features as well. In the complex networks field, this extension allows us to model the full degree distribution, instead of just focusing on the tail, as is commonly done. We demonstrate the utility of these novel methods through applications to two real-world datasets, showcasing their practical relevance and potential to advance the analysis of power-law behavior.

physics.soc-ph

Consequences of non-Markovian healing processes on epidemic models with recurrent infections on networks

Infections diseases are marked by recovering time distributions which can be far from the exponential one associated with Markovian/Poisson processes, broadly applied in epidemic compartmental models. In the present work, we tackled this problem by investigating a susceptible-infected-recovered-susceptible model on networks with $η$ independent infectious compartments (SI$_η$RS), each one with a Markovian dynamics, that leads to a Gamma-distributed recovering times. We analytically develop a theory for the epidemic lifespan on star graphs with a center and $K$ leaves showing that the epidemic lifespan scales with a non-universal power-law $τ_{K}\sim K^{α/μη}$ plus logarithm corrections, where $α^{-1}$ and $μ^{-1}$ are the mean waning immunity and recovering times, respectively. Compared with standard SIRS dynamics with $η=1$ and the same mean recovering time, the epidemic lifespan on star graphs is severely reduced as the number of stages increases. In particular, the case $η\rightarrow\infty$ leads to a finite lifespan. Numerical simulations support the approximated analytical calculations. For the SIS dynamics, numerical simulations show that the lifespan increases exponentially with the number of leaves, with a nonuniversal rate that decays with the number of infectious compartments. We investigated the SI$_η$RS dynamics on power-law networks with degree distribution $P(K)\sim k^{-γ}$. When $γ<5/2$, the epidemic spreading is ruled by a maximum $k$-core activation, the alteration of the hub activity time does not alter either the epidemic threshold or the localization pattern. For $γ>3$, where hub mutual activation is at work, the localization is reduced but not sufficiently to alter the threshold scaling with the network size. Therefore, the activation mechanisms remain the same as in the case of Markovian healing.

physics.soc-ph

Sampling with censored data: a practical guide

In this review, we present a simple guide for researchers to obtain pseudo-random samples with censored data. We focus our attention on the most common types of censored data, such as type I, type II, and random censoring. We discussed the necessary steps to sample pseudo-random values from long-term survival models where an additional cure fraction is informed. For illustrative purposes, these techniques are applied in the Weibull distribution. The algorithms and codes in R are presented, enabling the reproducibility of our study. Finally, we developed an R package that encapsulates these methodologies, providing researchers with practical tools for implementation.

stat.CO

Integrating socioeconomic and geographic data to enhance infectious disease prediction in Brazilian cities

Supervised machine learning models and public surveillance data has been employed for infectious disease forecasting in many settings. These models leverage various data sources capturing drivers of disease spread, such as climate conditions or human behavior. However, few models have incorporated the organizational structure of different geographic locations for forecasting. Traveling waves of seasonal outbreaks have been reported for dengue, influenza, and other infectious diseases, and many of the drivers of infectious disease dynamics may be shared across different cities, either due to their geographic or socioeconomic proximity. In this study, we developed a machine learning model to predict case counts of four infectious diseases across Brazilian cities one week ahead by incorporating information from related cities. We compared selecting related cities using both geographic distance and GDP per capita. Incorporating information from geographically proximate cities improved predictive performance for two of the four diseases, specifically COVID-19 and Zika. We also discuss the impact on forecasts in the presence of anomalous contagion patterns and the limitations of the proposed methodology.

stat.AP

Accuracy of discrete- and continuous-time mean-field theories for epidemic processes on complex networks

Discrete- and continuous-time approaches are frequently used to model the role of heterogeneity on dynamical interacting agents on the top of complex networks. While, on the one hand, one does not expect drastic differences between these approaches, and the choice is usually based on one's expertise or methodological convenience, on the other hand, a detailed analysis of the differences is necessary to guide the proper choice of one or another approach. We tackle this problem, by comparing both discrete- and continuous-time mean-field theories for the susceptible-infected-susceptible (SIS) epidemic model on random networks with power-law degree distributions. We compare the discrete epidemic link equations (ELE) and continuous pair quenched mean-field (PQMF) theories with the corresponding stochastic simulations, both theories that reckon pairwise interactions explicitly. We show that ELE converges to PQMF theory when the time step goes to zero. The epidemic localization analysis performed reckoning the inverse participation ratio (IPR) indicates that both theories present the same localization depence on the network degree exponent $γ$: for $γ<5/2$ the epidemic is localized on the maximum k-core of network with a vanishing in the infinite-size limit while for $γ>5/2$, the localization happens on hubs what leads to a finite value of IPR. However, the IPR and epidemic threshold of ELE depend on the time-step discretization such that a larger time-step leads to more localized epidemics. A remarkable difference between discrete and continuous time approaches is revealed in the epidemic prevalence near the epidemic threshold, in which the discrete-time stochastic simulations indicate a mean-field critical exponent $θ=1$ instead of the value $θ=1/(3-γ)$ obtained rigorously and verified numerically for the continuous-time SIS on the same networks.

physics.soc-ph

Objective Bayesian Analysis for the Differential Entropy of the Gamma Distribution

The present paper introduces a fully objective Bayesian analysis to obtain the posterior distribution of an entropy measure. Notably, we consider the gamma distribution, which describes many natural phenomena in physics, engineering, and biology. We reparametrize the model in terms of entropy, and different objective priors are derived, such as Jeffreys prior, reference prior, and matching priors. Since the obtained priors are improper, we prove that the obtained posterior distributions are proper and that their respective posterior means are finite. An intensive simulation study is conducted to select the prior that returns better results regarding bias, mean square error, and coverage probabilities. The proposed approach is illustrated in two datasets: the first relates to the Achaemenid dynasty reign period, and the second describes the time to failure of an electronic component in a sugarcane harvest machine.

math.ST

Machine learning in physics: a short guide

Machine learning is a rapidly growing field with the potential to revolutionize many areas of science, including physics. This review provides a brief overview of machine learning in physics, covering the main concepts of supervised, unsupervised, and reinforcement learning, as well as more specialized topics such as causal inference, symbolic regression, and deep learning. We present some of the principal applications of machine learning in physics and discuss the associated challenges and perspectives.

cs.LG

Machine learning-based prediction of Q-voter model in complex networks

In this article, we consider machine learning algorithms to accurately predict two variables associated with the $Q$-voter model in complex networks, i.e., (i) the consensus time and (ii) the frequency of opinion changes. Leveraging nine topological measures of the underlying networks, we verify that the clustering coefficient (C) and information centrality (IC) emerge as the most important predictors for these outcomes. Notably, the machine learning algorithms demonstrate accuracy across three distinct initialization methods of the $Q$-voter model, including random selection and the involvement of high- and low-degree agents with positive opinions. By unraveling the intricate interplay between network structure and dynamics, this research sheds light on the underlying mechanisms responsible for polarization effects and other dynamic patterns in social systems. Adopting a holistic approach that comprehends the complexity of network systems, this study offers insights into the intricate dynamics associated with polarization effects and paves the way for investigating the structure and dynamics of complex systems through modern machine learning methods.

physics.soc-ph

Cultural heterogeneity constrains diffusion of innovations

Rogers' diffusion of innovations theory asserts that cultural similarity among individuals plays a crucial role in the acceptance of an innovation in a community. However, most studies on the diffusion of innovations have relied on epidemic-like models where the individuals have no preference on whom they interact with. Here, we use an agent-based model to study the diffusion of innovations in a community of synthetic heterogeneous agents whose interaction preferences depend on their cultural similarity. The community heterogeneity and the agents' interaction preferences are described by Axelrod's model, whereas the diffusion of innovations is described by a variant of the Daley and Kendall model of rumour propagation. The interplay between the social dynamics and the spreading of the innovation is controlled by the parameter $p \in [0,1]$, which yields the probability that the agent engages in social interaction or attempts to spread the innovation. Our findings support Roger's empirical observations that cultural heterogeneity curbs the diffusion of innovations.

physics.soc-ph

An informational approach to uncover the age group interactions in epidemic spreading from macro analysis

We investigate the use of transfer entropy (TE) as a proxy to detect the contact patterns of the population in epidemic processes. We first apply the measure to a classical age-stratified SIR model and observe that the recovered patterns are consistent with the age-mixing matrix that encodes the interaction of the population. We then apply the TE analysis to real data from the COVID-19 pandemic in Spain and show that it can provide information on how the behavior of individuals changed through time. We also demonstrate how the underlying dynamics of the process allow us to build a coarse-grained representation of the time series that provides more information than raw time series. The macro-level representation is a more effective scale for analysis, which is an interesting result within the context of causal analysis across different scales. These results open the path for more research on the potential use of informational approaches to extract retrospective information on how individuals change and adapt their behavior during a pandemic, which is essential for devising adequate strategies for an efficient control of the spreading.

physics.soc-ph

Group polarization, influence, and domination in online interaction networks: A case study of the 2022 Brazilian elections

In this work, we investigate the evolution of polarization, influence, and domination in online interaction networks. Twitter data collected before and during the 2022 Brazilian elections is used as a case study. From a theoretical perspective, we develop a methodology called d-modularity that allows discovering the contribution of specific groups to network polarization using the well-known modularity measure. While the overall network modularity (somewhat unexpectedly) decreased, the proposed group-oriented approach allows concluding that the contribution of the right-leaning community to this modularity increased, remaining very high during the analyzed period. Our methodology is general enough to be used in any situation when the contribution of specific groups to overall network modularity and polarization is needed to investigate. Moreover, using the concept of partial domination, we are able to compare the reach of sets of influential profiles from different groups and their ability to accomplish coordinated communication inside their groups and across segments of the entire network during some specific time window. We show that in the whole network, the left-leaning high-influential information spreaders dominated, reaching a substantial fraction of users with fewer spreaders. However, when comparing domination inside the groups, the results are inverse. Right-leaning spreaders dominate their communities using few nodes, showing as the most capable of accomplishing coordinated communication. The results bring evidence of extreme isolation and the ease of accomplishing coordinated communication that characterized right-leaning communities during the 2022 Brazilian elections.

cs.SI