SearcharxivSearch

arXiv subjects

Stefano Iacus

Publications and source records attributed to Stefano Iacus.

13 recordsLinked to original sources

Ergodic Network Stochastic Differential Equations

We propose a novel framework for Network Stochastic Differential Equations (N-SDE), where each node in a network is governed by an SDE influenced by interactions with its neighbors. The evolution of each node is driven by the interplay of three key components: the node's intrinsic dynamics (\emph{momentum effect}), feedback from neighboring nodes (\emph{network effect}), and a \emph{stochastic volatility} term modeled by Brownian motion. Our primary objective is to estimate the parameters of the N-SDE system from high-frequency discrete-time observations. The motivation behind this model lies in its ability to analyze very high-dimensional time series by leveraging the inherent sparsity of the underlying network graph. We consider two distinct scenarios: \textit{i) known network structure}: the graph is fully specified, and we establish conditions under which the parameters can be identified, considering the linear growth of the parameter space with the number of edges. \textit{ii) unknown network structure}: the graph must be inferred from the data. For this, we develop an iterative procedure using adaptive Lasso, tailored to a specific subclass of N-SDE models. In this work, we assume the network graph is oriented, paving the way for novel applications of SDEs in causal inference, enabling the study of cause-effect relationships in dynamic systems. Through extensive simulation studies, we demonstrate the performance of our estimators across various graph topologies in high-dimensional settings. We also showcase the framework's applicability to real-world datasets, highlighting its potential for advancing the analysis of complex networked systems.

stat.ME

Adaptive Elastic-Net estimation for sparse diffusion processes

Penalized estimation methods for diffusion processes and dependent data have recently gained significant attention due to their effectiveness in handling high-dimensional stochastic systems. In this work, we introduce an adaptive Elastic-Net estimator for ergodic diffusion processes observed under high-frequency sampling schemes. Our method combines the least squares approximation of the quasi-likelihood with adaptive $\ell_1$ and $\ell_2$ regularization. This approach allows to enhance prediction accuracy and interpretability while effectively recovering the sparse underlying structure of the model. In the spirit of analyzing high-dimensional scenarios, we provide finite-sample guarantees for the (block-diagonal) estimator's performance by deriving high-probability non-asymptotic bounds for the $\ell_2$ estimation error. These results complement the established oracle properties in the high-frequency asymptotic regime with mixed convergence rates, ensuring consistent selection of the relevant interactions and achieving optimal rates of convergence. Furthermore, we utilize our results to analyze one-step-ahead predictions, offering non-asymptotic control over the $\ell_1$ prediction error. The performance of our method is evaluated through simulations and real data applications, demonstrating its effectiveness, particularly in scenarios with strongly correlated variables.

math.ST

GREI Data Repository AI Taxonomy

The Generalist Repository Ecosystem Initiative (GREI), funded by the NIH, developed an AI taxonomy tailored to data repository roles to guide AI integration across repository management. It categorizes the roles into stages, including acquisition, validation, organization, enhancement, analysis, sharing, and user support, providing a structured framework for implementing AI in repository workflows.

cs.DL

Migration patterns, friendship networks, and the diaspora: the potential of Facebook Social Connectedness Index to anticipate displacement patterns induced by Russia invasion of Ukraine in the European Union

The conflict in Ukraine is causing large-scale displacement in Europe and in the World. Based on the United Nations High Commissioner for Refugees (UNHCR) estimates, more than 7 million people fled the country as of 5 September 2022. In this context, it is extremely important to anticipate where these people are moving so that national to local authorities can better manage challenges related to their reception and integration. This work shows how innovative data from social media can provide useful insights on conflict-induced migration flows. In particular, we explore the potential of Facebook's Social Connectedness Index (SCI) for predicting migration flows in the context of the war in Ukraine, building on previous research findings that the presence of a diaspora network is one of the major migration drivers. To do so, we first evaluate the relationship between the Ukrainian diaspora and the number of refugees from Ukraine registered for Temporary Protection or similar national schemes as a proxy of migratory flows into the EU. We find a very strong correlation between the two (Pearson's r=0.94, p<0.0001), which indicates that the diaspora is attracting the people fleeing the war, who tend to reach their compatriots, in particular in the countries where the Ukrainian immigration was more a recent phenomenon. Second, we compare Facebook's SCI with available official data on diaspora at regional level in Europe. Our results suggest that the index, along with other readily available covariates, is a strong predictor of the Ukrainian diaspora at regional scale. Finally, we discuss the potential of Facebook's SCI to provide timely and spatially detailed information on human diaspora for those countries where this information might be missing or outdated, and to complement official statistics for fast policy response during conflicts.

physics.soc-ph

Data Innovation in Demography, Migration and Human Mobility

With the consolidation of the culture of evidence-based policymaking, the availability of data has become central to policymakers. Nowadays, innovative data sources offer an opportunity to describe demographic, mobility, and migratory phenomena more accurately by making available large volumes of real-time and spatially detailed data. At the same time, however, data innovation has led to new challenges (ethics, privacy, data governance models, data quality) for citizens, statistical offices, policymakers and the private sector. Focusing on the fields of demography, mobility, and migration studies, the aim of this report is to assess the current state of data innovation in the scientific literature as well as to identify areas in which data innovation has the most concrete potential for policymaking. Consequently, this study has reviewed more than 300 articles and scientific reports, as well as numerous tools, that employed non-traditional data sources to measure vital population events (mortality, fertility), migration and human mobility, and the population change and population distribution. The specific findings of our report form the basis of a discussion on a) how innovative data is used compared to traditional data sources; b) domains in which innovative data have the greatest potential to contribute to policymaking; c) the prospects of innovative data transition towards systematically contributing to official statistics and policymaking.

cs.CY

The potential of Facebook advertising data for understanding flows of people from Ukraine to the European Union

This work contributes to the discussion on how innovative data can support a fast crisis response. We use operational data from Facebook to gain useful insights on where people fleeing Ukraine following the Russian invasion are likely to be displaced, focusing on the European Union. In this context, it is extremely important to anticipate where these people are moving so that local and national authorities can better manage challenges related to their reception and integration. By means of the Ukrainian-speaking Monthly Active Users estimates provided by Facebook advertising platform, we analyse the flows of people fleeing the country towards the European Union. At the fifth week since the beginning of the war, our results indicate an increase in the number of Ukrainian-speaking Facebook users in all the EU countries, with Poland registering the highest percentage share ($33\%$) of the overall increase, followed by Germany ($17\%$), and Czechia ($15\%$). We assess the reliability of prewar Facebook estimates by comparison with official statistics on the Ukrainian diaspora, finding a strong correlation between the two data sources (Pearson's $r=0.93$, $p<0.0001$). We then compare our results with data on arrivals in Poland and Hungary reported by the UNHCR, and we observe a similarity in their trend. In conclusion, we show how Facebook advertising data could offer timely insights on international mobility during crisis, supporting initiatives aimed at providing humanitarian assistance to the displaced people, as well as local and national authorities to better manage their reception and integration.

stat.AP

On penalized estimation for dynamical systems with small noise

We consider a dynamical system with small noise for which the drift is parametrized by a finite dimensional parameter. For this model we consider minimum distance estimation from continuous time observations under $l^p$-penalty imposed on the parameters in the spirit of the Lasso approach with the aim of simultaneous estimation and model selection. We study the consistency and the asymptotic distribution of these Lasso-type estimators for different values of $p$. For $p=1$ we also consider the adaptive version of the Lasso estimator and establish its oracle properties.

math.ST

Measuring Social Well Being in The Big Data Era: Asking or Listening?

The literature on well being measurement seems to suggest that "asking" for a self-evaluation is the only way to estimate a complete and reliable measure of well being. At the same time "not asking" is the only way to avoid biased evaluations due to self-reporting. Here we propose a method for estimating the welfare perception of a community simply "listening" to the conversations on Social Network Sites. The Social Well Being Index (SWBI) and its components are proposed through to an innovative technique of supervised sentiment analysis called iSA which scales to any language and big data. As main methodological advantages, this approach can estimate several aspects of social well being directly from self-declared perceptions, instead of approximating it through objective (but partial) quantitative variables like GDP; moreover self-perceptions of welfare are spontaneous and not obtained as answers to explicit questions that are proved to bias the result. As an application we evaluate the SWBI in Italy through the period 2012-2015 through the analysis of more than 143 millions of tweets.

cs.CY

On a family of test statistics for discretely observed diffusion processes

We consider parametric hypotheses testing for multidimensional ergodic diffusion processes observed at discrete time. We propose a family of test statistics, related to the so called $ϕ$-divergence measures. By taking into account the quasi-likelihood approach developed for studying the stochastic differential equations, it is proved that the tests in this family are all asymptotically distribution free. In other words, our test statistics weakly converge to the chi squared distribution. Furthermore, our test statistic is compared with the quasi likelihood ratio test. In the case of contiguous alternatives, it is also possible to study in detail the power function of the tests. Although all the tests in this family are asymptotically equivalent, we show by Monte Carlo analysis that, in the small sample case, the performance of the test strictly depends on the choice of the function $ϕ$. Furthermore, in this framework, the simulations show that there are not uniformly most powerful tests.

math.ST

Divergences Test Statistics for Discretely Observed Diffusion Processes

In this paper we propose the use of $ϕ$-divergences as test statistics to verify simple hypotheses about a one-dimensional parametric diffusion process $\de X_t = b(X_t, θ)\de t + σ(X_t, θ)\de W_t$, from discrete observations $\{X_{t_i}, i=0, ..., n\}$ with $t_i = iΔ_n$, $i=0, 1, >..., n$, under the asymptotic scheme $Δ_n\to0$, $nΔ_n\to\infty$ and $nΔ_n^2\to 0$. The class of $ϕ$-divergences is wide and includes several special members like Kullback-Leibler, Rényi, power and $α$-divergences. We derive the asymptotic distribution of the test statistics based on $ϕ$-divergences. The limiting law takes different forms depending on the regularity of $ϕ$. These convergence differ from the classical results for independent and identically distributed random variables. Numerical analysis is used to show the small sample properties of the test statistics in terms of estimated level and power of the test.

math.ST

Parametric estimation for partially hidden diffusion processes sampled at discrete times

For a one dimensional diffusion process $X=\{X(t) ; 0\leq t \leq T \}$, we suppose that $X(t)$ is hidden if it is below some fixed and known threshold $τ$, but otherwise it is visible. This means a partially hidden diffusion process. The problem treated in this paper is to estimate finite dimensional parameter in both drift and diffusion coefficients under a partially hidden diffusion process obtained by a discrete sampling scheme. It is assumed that the sampling occurs at regularly spaced time intervals of length $h_n$ such that $n h_n=T$. The asymptotic is when $h_n\to0$, $T\to\infty$ and $n h_n^2\to 0$ as $n\to\infty$. Consistency and asymptotic normality for estimators of parameters in both drift and diffusion coefficients are proved.

math.ST

Renyi information for ergodic diffusion processes

In this paper we derive explicit formulas of the Rényi information, Shannon entropy and Song measure for the invariant density of one dimensional ergodic diffusion processes. In particular, the diffusion models considered include the hyperbolic, the generalized inverse Gaussian, the Pearson, the exponential familiy and a new class of skew-$t$ diffusions.

math.PR

Average treatment effect estimation via random recursive partitioning

A new matching method is proposed for the estimation of the average treatment effect of social policy interventions (e.g., training programs or health care measures). Given an outcome variable, a treatment and a set of pre-treatment covariates, the method is based on the examination of random recursive partitions of the space of covariates using regression trees. A regression tree is grown either on the treated or on the untreated individuals {\it only} using as response variable a random permutation of the indexes 1...$n$ ($n$ being the number of units involved), while the indexes for the other group are predicted using this tree. The procedure is replicated in order to rule out the effect of specific permutations. The average treatment effect is estimated in each tree by matching treated and untreated in the same terminal nodes. The final estimator of the average treatment effect is obtained by averaging on all the trees grown. The method does not require any specific model assumption apart from the tree's complexity, which does not affect the estimator though. We show that this method is either an instrument to check whether two samples can be matched (by any method) and, when this is feasible, to obtain reliable estimates of the average treatment effect. We further propose a graphical tool to inspect the quality of the match. The method has been applied to the National Supported Work Demonstration data, previously analyzed by Lalonde (1986) and others.

math.ST