SearcharxivSearch

arXiv subjects

Stefano M. Iacus

Publications and source records attributed to Stefano M. Iacus.

15 recordsLinked to original sources

Dynamic Attention (DynAttn): Interpretable High-Dimensional Spatio-Temporal Forecasting (with Application to Conflict Fatalities)

Forecasting conflict-related fatalities remains a central challenge in political science and policy analysis due to the sparse, bursty, and highly non-stationary nature of violence data. We introduce DynAttn, an interpretable dynamic-attention forecasting framework for high-dimensional spatio-temporal count processes. DynAttn combines rolling-window estimation, shared elastic-net feature gating, a compact weight-tied self-attention encoder, and a zero-inflated negative binomial (ZINB) likelihood. This architecture produces calibrated multi-horizon forecasts of expected casualties and exceedance probabilities, while retaining transparent diagnostics through feature gates, ablation analysis, and elasticity measures. We evaluate DynAttn using global country-level and high-resolution PRIO-grid-level conflict data from the VIEWS forecasting system, benchmarking it against established statistical and machine-learning approaches, including DynENet, LSTM, Prophet, PatchTST, and the official VIEWS baseline. Across forecast horizons from one to twelve months, DynAttn consistently achieves substantially higher predictive accuracy, with particularly large gains in sparse grid-level settings where competing models often become unstable or degrade sharply. Beyond predictive performance, DynAttn enables structured interpretation of regional conflict dynamics. In our application, cross-regional analyses show that short-run conflict persistence and spatial diffusion form the core predictive backbone, while climate stress acts either as a conditional amplifier or a primary driver depending on the conflict theater.

stat.AP

The Human Flourishing Geographic Index: A County-Level Dataset for the United States, 2013--2023

Quantifying human flourishing, a multidimensional construct including happiness, health, purpose, virtue, relationships, and financial stability, is critical for understanding societal well-being beyond economic indicators. Existing measures often lack fine spatial and temporal resolution. Here we introduce the Human Flourishing Geographic Index (HFGI), derived from analyzing approximately 2.6 billion geolocated U.S. tweets (2013-2023) using fine-tuned large language models to classify expressions across 48 indicators aligned with Harvard's Global Flourishing Study framework plus attitudes towards migration and perception of corruption. The dataset offers monthly and yearly county- and state-level indicators of flourishing-related discourse, validated to confirm that the measures accurately represent the underlying constructs and show expected correlations with established indicators. This resource enables multidisciplinary analyses of well-being, inequality, and social change at unprecedented resolution, offering insights into the dynamics of human flourishing as reflected in social media discourse across the United States over the past decade.

cs.CL

Deep literature reviews: an application of fine-tuned language models to migration research

This paper presents a hybrid framework for literature reviews that augments traditional bibliometric methods with large language models (LLMs). By fine-tuning open-source LLMs, our approach enables scalable extraction of qualitative insights from large volumes of research content, enhancing both the breadth and depth of knowledge synthesis. To improve annotation efficiency and consistency, we introduce an error-focused validation process in which LLMs generate initial labels and human reviewers correct misclassifications. Applying this framework to over 20000 scientific articles about human migration, we demonstrate that a domain-adapted LLM can serve as a "specialist" model - capable of accurately selecting relevant studies, detecting emerging trends, and identifying critical research gaps. Notably, the LLM-assisted review reveals a growing scholarly interest in climate-induced migration. However, existing literature disproportionately centers on a narrow set of environmental hazards (e.g., floods, droughts, sea-level rise, and land degradation), while overlooking others that more directly affect human health and well-being, such as air and water pollution or infectious diseases. This imbalance highlights the need for more comprehensive research that goes beyond physical environmental changes to examine their ecological and societal consequences, particularly in shaping migration as an adaptive response. Overall, our proposed framework demonstrates the potential of fine-tuned LLMs to conduct more efficient, consistent, and insightful literature reviews across disciplines, ultimately accelerating knowledge synthesis and scientific discovery.

cs.CL

Empirical $L^2$-distance test statistics for ergodic diffusions

The aim of this paper is to introduce a new type of test statistic for simple null hypothesis on one-dimensional ergodic diffusion processes sampled at discrete times. We deal with a quasi-likelihood approach for stochastic differential equations (i.e. local gaussian approximation of the transition functions) and define a test statistic by means of the empirical $L^2$-distance between quasi-likelihoods. We prove that the introduced test statistic is asymptotically distribution free; namely it weakly converges to a $χ^2$ random variable. Furthermore, we study the power under local alternatives of the parametric test. We show by the Monte Carlo analysis that, in the small sample case, the introduced test seems to perform better than other tests proposed in literature.

math.PR

Discrete time approximation of a COGARCH(p,q) model and its estimation

In this paper, we construct a sequence of discrete time stochastic processes that converges in probability and in the Skorokhod metric to a COGARCH(p,q) model. The result is useful for the estimation of the continuous model defined for irregularly spaced time series data. The estimation procedure is based on the maximization of a pseudo log-likelihood function and is implemented in the yuima package.

math.ST

Estimation and Simulation of a COGARCH(p,q) model in the YUIMA project

In this paper we show how to simulate and estimate a COGARCH(p,q) model in the R package yuima. Several routines for simulation and estimation are available. Indeed for the generation of a COGARCH(p,q) trajectory, the user can choose between two alternative schemes. The first is based on the Euler discretization of the stochastic differential equations that identifies a COGARCH(p,q) model while the second one considers the explicit solution of the variance process. Estimation is based on the matching of the empirical with the theoretical autocorrelation function. In this case three different approaches are implemented: minimization of the mean square error, minimization of the absolute mean error and the generalized method of moments where the weighting matrix is continuously updated. Numerical examples are given in order to explain methods and classes used in the yuima package.

stat.CO

Implementation of Lévy CARMA model in Yuima package

The paper shows how to use the R package yuima available on CRAN for the simulation and the estimation of a general Lévy Continuous Autoregressive Moving Average (CARMA) model. The flexibility of the package is due to the fact that the user is allowed to choose several parametric Lévy distribution for the increments. Some numerical examples are given in order to explain the main classes and the corresponding methods implemented in yuima package for the CARMA model.

stat.CO

Parameter estimation for the discretely observed fractional Ornstein-Uhlenbeck process and the Yuima R package

This paper proposes consistent and asymptotically Gaussian estimators for the drift, the diffusion coefficient and the Hurst exponent of the discretely observed fractional Ornstein-Uhlenbeck process. For the estimation of the drift, the results are obtained only in the case when 1/2 < H < 3/4. This paper also provides ready-to-use software for the R statistical environment based on the YUIMA package.

stat.CO

Estimation for the change point of the volatility in a stochastic differential equation

We consider a multidimensional Itô process $Y=(Y_t)_{t\in[0,T]}$ with some unknown drift coefficient process $b_t$ and volatility coefficient $σ(X_t,θ)$ with covariate process $X=(X_t)_{t\in[0,T]}$, the function $σ(x,θ)$ being known up to $θ\inΘ$. For this model we consider a change point problem for the parameter $θ$ in the volatility component. The change is supposed to occur at some point $t^*\in (0,T)$. Given discrete time observations from the process $(X,Y)$, we propose quasi-maximum likelihood estimation of the change point. We present the rate of convergence of the change point estimator and the limit thereoms of aymptotically mixed type.

math.ST

Change point estimation for the telegraph process observed at discrete times

The telegraph process models a random motion with finite velocity and it is usually proposed as an alternative to diffusion models. The process describes the position of a particle moving on the real line, alternatively with constant velocity $+ v$ or $-v$. The changes of direction are governed by an homogeneous Poisson process with rate $λ>0.$ In this paper, we consider a change point estimation problem for the rate of the underlying Poisson process by means of least squares method. The consistency and the rate of convergence for the change point estimator are obtained and its asymptotic distribution is derived. Applications to real data are also presented.

math.ST

Parametric estimation for the standard and geometric telegraph process observed at discrete times

The telegraph process $X(t)$, $t>0$, (Goldstein, 1951) and the geometric telegraph process $S(t) = s_0 \exp\{(μ-\frac12σ^2)t + σX(t)\}$ with $μ$ a known constant and $σ>0$ a parameter are supposed to be observed at $n+1$ equidistant time points $t_i=iΔ_n,i=0,1,..., n$. For both models $λ$, the underlying rate of the Poisson process, is a parameter to be estimated. In the geometric case, also $σ>0$ has to be estimated. We propose different estimators of the parameters and we investigate their performance under the high frequency asymptotics, i.e. $Δ_n \to 0$, $nΔ= T<\infty$ as $n \to \infty$, with $T>0$ fixed. The process $X(t)$ in non markovian, non stationary and not ergodic thus we use approximation arguments to derive estimators. Given the complexity of the equations involved only estimators on the first model can be studied analytically. Therefore, we run an extensive Monte Carlo analysis to study the performance of the proposed estimators also for small sample size $n$.

math.ST

Nonparametric estimation of distribution and density functions in presence of missing data: an IFS approach

In this paper we consider a class of nonparametric estimators of a distribution function F, with compact support, based on the theory of IFSs. The estimator of F is tought as the fixed point of a contractive operator T defined in terms of a vector of parameters p and a family of affine maps W which can be both depend of the sample (X_1, X_2, ...., X_n). Given W, the problem consists in finding a vector p such that the fixed point of T is ``sufficiently near'' to F. It turns out that this is a quadratic constrained optimization problem that we propose to solve by penalization techniques. If F has a density f, we can also provide an estimator of f based on Fourier techniques. IFS estimators for F are asymptotically equivalent to the empirical distribution function (e.d.f.) estimator. We will study relative efficiency of the IFS estimators with respect to the e.d.f. for small samples via Monte Carlo approach. For well behaved distribution functions F and for a particular family of so-called wavelet maps the IFS estimators can be dramatically better than the e.d.f. (or the kernel estimator for density estimation) in presence of missing data, i.e. when it is only possibile to observe data on subsets of the whole support of F. This research has also produced a free package for the R statistical environment which is ready to be used in applications.

math.ST

Approximating distribution functions by iterated function systems

In this paper an iterated function system on the space of distribution functions is built. The inverse problem is introduced and studied by convex optimization problems. Some applications of this method to approximation of distribution functions and to estimation theory are given.

math.ST

Statistical analysis of stochastic resonance with ergodic diffusion noise

A subthreshold signal is transmitted through a channel and may be detected when some noise -- with known structure and proportional to some level -- is added to the data. There is an optimal noise level, called stochastic resonance, that corresponds to the highest Fisher information in the problem of estimation of the signal. As noise we consider an ergodic diffusion process and the asymptotic is considered as time goes to infinity. We propose consistent estimators of the subthreshold signal and we solve further a problem of hypotheses testing. We also discuss evidence of stochastic resonance for both estimation and hypotheses testing problems via examples.

math.ST

Statistical analysis of the inhomogeneous telegrapher's process

We consider a problem of estimation for the telegrapher's process on the line, say X(t), driven by a Poisson process with non constant rate. It turns out that the finite-dimensional law of the process X(t) is a solution to the telegraph equation with non constant coefficients. We give the explicit law P(theta) of the process X(t) for a parametric class of intensity functions for the Poisson process. We propose an estimator for the parameter theta of P(theta) and we discuss its properties as a first attempt to apply statistics to these models.

math.PR