SearcharxivSearch

arXiv subjects

Guilherme Pumi

Publications and source records attributed to Guilherme Pumi.

At least 19 recordsLinked to original sources

PRTree: An R Package for Probabilistic Regression Trees with Built-in Missing Data Handling

PRTree is an R package for fitting Probabilistic Regression Trees (PRTrees), a class of regression trees that replaces deterministic splits with probabilistic associations to produce smooth prediction functions. The package implements both the original methodology and its recent extension for handling missing predictor values, allowing model fitting and prediction directly from incomplete datasets without prior imputation. It provides a unified framework for model fitting, prediction, visualization, smoothing-parameter selection, cross-validation, and model diagnostics through a standard R interface. Computationally intensive routines are implemented in FORTRAN, while the high-level interface follows the usual R workflow based on S3 classes and generic methods. This paper reviews the underlying methodology, describes the software architecture and main package components, and illustrates their use through reproducible examples.

stat.ME

Handling Missing Data in Probabilistic Regression Trees

Probabilistic Regression Trees (PRTrees) are a smooth and consistent alternative to classical regression trees, producing continuous predictions through probabilistic split assignments. This paper extends the PRTree framework to accommodate missing predictor values directly during tree construction, eliminating the need for prior imputation. Three strategies are proposed, each exploiting the available information differently: a uniform-probability approach, a partial-observation approach, and a dimension-reduced smoothing approach. These modifications are defined to preserve the fundamental probabilistic properties of the original methodology, including probability conservation and marginal compatibility, under arbitrary patterns of missing covariate values. The proposed methods are evaluated on several real-world datasets exhibiting different levels of missingness and are compared with classical regression trees. The results show that the effectiveness of probabilistic tree construction depends strongly on the treatment of missing observations. Across the considered datasets, the fill strategy emerged as the dominant modeling component, often exerting a larger influence on predictive performance than either the smoothing distribution or the proxy-selection criterion. In datasets where a substantial proportion of observations contained missing predictor values, the proposed methods frequently outperformed CART, while maintaining the interpretability and flexibility of tree-based models.

stat.ML

GARTFIMA Models: A Class of Observation-Driven Models with Tempered Fractional Dynamics

This paper introduces a class of observation-driven models whose systematic component includes a tempered fractional differencing term. This specification generalizes long-range dependent models based on the fractional differencing operator, enabling a more general and robust model specification while offering theoretical advantages. We propose a partial maximum likelihood approach for parameter estimation and address hypothesis testing, confidence intervals, goodness-of-fit assessment, and both in-sample and out-of-sample forecasting. A Monte Carlo simulation study evaluates the finite-sample performance of the proposed estimation method, and an empirical application illustrates the model's practical utility.

stat.ME

Bayes Estimation of GLARMA Models With Applications

This work presents a Bayesian approach for parameter estimation in the class of Generalized Linear Autoregressive Moving Average (GLARMA) models, extending the methodology beyond the common exponential family setting. The proposed framework accommodates positive, double-bounded, and count time series through a unified MCMC-based estimation procedure implemented in \texttt{nimble}. We discuss prior specifications for the model parameters and conduct an extensive Monte Carlo simulation study to evaluate the finite-sample performance of the approach under three distinct data-generating mechanisms: Negative Binomial-GLARMA for count data, Beta-GLARMA for double-bounded outcomes, and Gamma-GLARMA for positive continuous time series. The simulation study assesses point and interval estimation, prior sensitivity, and the behaviour of the estimators under varying levels of temporal persistence, including challenging scenarios near the boundaries of the stationarity region. The practical utility of the proposed framework is illustrated through two empirical applications: analysing monthly net electricity generation by nuclear plants in the United States using a Gamma-GLARMA model with harmonic seasonal components; and modelling monthly hospital admissions due to chronic obstructive pulmonary disease in Belo Horizonte, Brazil, using a Negative Binomial-GLARMA model with principal components derived from air pollution covariates. The results demonstrate that the proposed Bayesian framework provides reliable and stable estimation across all settings, offering a flexible and practical tool for analysing non-Gaussian time series in a wide range of applications.

stat.ME

Estimation of Long-Range Dependent Models with Missing Data: to Impute or not to Impute?

Among the most important models for long-range dependent time series is the class of ARFIMA$(p,d,q)$ (Autoregressive Fractionally Integrated Moving Average) models. Estimating the long-range dependence parameter $d$ in ARFIMA models is a well-studied problem, but the literature regarding the estimation of $d$ in the presence of missing data is very sparse. There are two basic approaches to dealing with the problem: missing data can be imputed using some plausible method, and then the estimation can proceed as if no data were missing, or we can use a specially tailored methodology to estimate $d$ in the presence of missing data. In this work, we review some of the methods available for both approaches and compare them through a Monte Carlo simulation study. We present a comparison among 35 different setups to estimate $d$, under tenths of different scenarios, considering percentages of missing data ranging from as few as 10\% up to 70\% and several levels of dependence.

stat.ME

A Novel Multiple Imputation Approach For Parameter Estimation in Observation-Driven Time Series Models With Missing Data

Handling missing data in time series is a complex problem due to the presence of temporal dependence. General-purpose imputation methods, while widely used, often distort key statistical properties of the data, such as variance and dependence structure, leading to biased estimation and misleading inference. These issues become more pronounced in models that explicitly rely on capturing serial dependence, as standard imputation techniques fail to preserve the underlying dynamics. This paper proposes a novel multiple imputation method specifically designed for parameter estimation in observation-driven models (ODM). The approach takes advantage of the iterative nature of the systematic component in ODM to propagate the dependence structure through missing data, minimizing its impact on estimation. Unlike traditional imputation techniques, the proposed method accommodates continuous, discrete, and mixed-type data while preserving key distributional and dependence properties. We evaluate its performance through Monte Carlo simulations in the context of GARMA models, considering time series with up to 70\% missing data. An application to the proportion of stocked energy stored in South Brazil further demonstrates its practical utility.

stat.ME

A two-step approach to production frontier estimation and the Matsuoka's distribution

In this work, we introduce a deterministic frontier model in which efficiency is governed by the Matsuoka distribution, a parsimonious one-parameter specification on $(0,1)$ designed to reflect patterns typically observed in efficiency data. Based on this formulation, we develop a two-step semiparametric estimation procedure: a nonparametric smoothing for the regression component, followed by a feasible method of moments estimation for the efficiency parameter with plug-in reconstruction of the frontier. Theoretical results establish convergence rates, asymptotic normality, and an oracle property for the parametric estimator of the efficiency parameter. A Monte Carlo study demonstrates that the procedure performs consistently with the theoretical results and improves upon a fully nonparametric alternative. Applying the method to Brazilian temporary crops with land and agrochemicals as inputs, we find that both regions exhibit isoquants close to the constant elasticity substitution form, but differ in the relative productivity of inputs. Most notably, statistical tests provide evidence that the South is relatively more efficient than the Center-West, highlighting the empirical relevance of the proposed approach.

stat.ME

A GARMA Framework for Unit-Bounded Time Series Based on the Unit-Lindley Distribution with Application to Renewable Energy Data

The Unit-Lindley is a one-parameter family of distributions in $(0,1)$ obtained from an appropriate transformation of the Lindley distribution. In this work, we introduce a class of dynamical time series models for continuous random variables taking values in $(0,1)$ based on the Unit-Lindley distribution. The models pertaining to the proposed class are observation-driven ones for which, conditionally on a set of covariates, the random component is modeled by a Unit-Lindley distribution. The systematic component aims at modeling the conditional mean through a dynamical structure resembling the classical ARMA models. Parameter estimation in conducted using partial maximum likelihood, for which an asymptotic theory is available. Based on asymptotic results, the construction of confidence intervals, hypotheses testing, model selection, and forecasting can be carried on. A Monte Carlo simulation study is conducted to assess the finite sample performance of the proposed partial maximum likelihood approach. Finally, an application considering forecasting of the proportion of net electricity generated by conventional hydroelectric power in the United States is presented. The application show the versatility of the proposed method compared to other benchmarks models in the literature.

math.ST

A Matsuoka-Based GARMA Model for Hydrological Forecasting: Theory, Estimation, and Applications

Time series in natural sciences, such as hydrology and climatology, and other environmental applications, often consist of continuous observations constrained to the unit interval (0,1). Traditional Gaussian-based models fail to capture these bounds, requiring more flexible approaches. This paper introduces the Matsuoka Autoregressive Moving Average (MARMA) model, extending the GARMA framework by assuming a Matsuoka-distributed random component taking values in (0,1) and an ARMA-like systematic structure allowing for random time-dependent covariates. Parameter estimation is performed via partial maximum likelihood (PMLE), for which we present the asymptotic theory. It enables statistical inference, including confidence intervals and model selection. To construct prediction intervals, we propose a novel bootstrap-based method that accounts for dependence structure uncertainty. A comprehensive Monte Carlo simulation study assesses the finite sample performance of the proposed methodologies, while an application to forecasting the useful water volume of the Guarapiranga Reservoir in Brazil showcases their practical usefulness.

stat.ME

Order selection in GARMA models for count time series: a Bayesian perspective

Estimation in GARMA models has traditionally been carried out under the frequentist approach. To date, Bayesian approaches for such estimation have been relatively limited. In the context of GARMA models for count time series, Bayesian estimation achieves satisfactory results in terms of point estimation. Model selection in this context often relies on the use of information criteria. Despite its prominence in the literature, the use of information criteria for model selection in GARMA models for count time series have been shown to present poor performance in simulations, especially in terms of their ability to correctly identify models, even under large sample sizes. In this study, we study the problem of order selection in GARMA models for count time series, adopting a Bayesian perspective through the application of the Reversible Jump Markov Chain Monte Carlo approach. Monte Carlo simulation studies are conducted to assess the finite sample performance of the developed ideas, including point and interval inference, sensitivity analysis, effects of burn-in and thinning, as well as the choice of related priors and hyperparameters. Two real-data applications are presented, one considering automobile production in Brazil and the other considering bus exportation in Brazil before and after the COVID-19 pandemic, showcasing the method's capabilities and further exploring its flexibility.

stat.ME

Mitigating the choice of the duration in DDMS models through a parametric link

One of the most important hyper-parameters in duration-dependent Markov-switching (DDMS) models is the duration of the hidden states. Because there is currently no procedure for estimating this duration or testing whether a given duration is appropriate for a given data set, an ad hoc duration choice must be heuristically justified. In this paper, we propose and examine a methodology that mitigates the choice of duration in DDMS models when forecasting is the goal. Two Monte Carlo simulations, based on classical applications of DDMS models, are employed to evaluate the methodology. In addition, an empirical investigation is carried out to forecast the volatility of the S\&P 500, which showcases the capabilities of the proposed model.

stat.AP

Bayesian Analysis of Beta Autoregressive Moving Average Models

This work presents a Bayesian approach for the estimation of Beta Autoregressive Moving Average ($β$ARMA) models. We discuss standard choice for the prior distributions and employ a Hamiltonian Monte Carlo algorithm to sample from the posterior. We propose a method to approach the problem of unit roots in the model's systematic component. We then present a series of Monte Carlo simulations to evaluate the performance of this Bayesian approach. In addition to parameter estimation, we evaluate the proposed approach to verify the presence of unit roots in the model's systematic component and study prior sensitivity. An empirical application is presented to exemplify the usefulness of the method. In the application, we compare the fitted Bayesian and frequentist approaches in terms of their out-of-sample forecasting capabilities.

stat.ME

Parameterization of Copulas and Covariance Decay of Stochastic Processes

In this work we study the problem of constructing stochastic processes with a predetermined covariance decay by parameterizing its marginals and a given family of copulas. We show that the proposed methodology is compatibility-free and present several examples to illustrate the theory, including the important Gaussian and Euclidean families of copulas. We associate the theory to common applied time series models.

math.ST

Unit-Weibull Autoregressive Moving Average Models

In this work we introduce the class of unit-Weibull Autoregressive Moving Average models for continuous random variables taking values in $(0,1)$. The proposed model is an observation driven one, for which, conditionally on a set of covariates and the process' history, the random component is assumed to follow a unit-Weibull distribution parameterized through its $ρ$th quantile. The systematic component prescribes an ARMA-like structure to model the conditional $ρ$th quantile by means of a link. Parameter estimation in the proposed model is performed using partial maximum likelihood, for which we provide closed formulas for the score vector and partial information matrix. We also discuss some inferential tools, such as the construction of confidence intervals, hypotheses testing, model selection, and forecasting. A Monte Carlo simulation study is conducted to assess the finite sample performance of the proposed partial maximum likelihood approach. Finally, we examine the prediction power by contrasting our method with others in the literature using the Manufacturing Capacity Utilization from the US.

math.ST

Positive Time Series Regression Models

In this paper we discuss dynamic ARMA-type regression models for time series taking values in $(0,\infty)$. In the proposed model, the conditional mean is modeled by a dynamic structure containing autoregressive and moving average terms, time-varying regressors, unknown parameters and link functions. We introduce the new class of models and discuss partial maximum likelihood estimation, hypothesis testing inference, diagnostic analysis and forecasting.

stat.ME

On the behavior of the DFA and DCCA in trend-stationary processes

In this work, we develop the asymptotic theory of the Detrended Fluctuation Analysis (DFA) and Detrended Cross-Correlation Analysis (DCCA) for trend-stationary stochastic processes without any assumption on the specific form of the underlying distribution. All results are presented and derived under the general framework of potentially overlapping boxes for the polynomial fit. We prove the stationarity of the DFA and DCCA, viewed as stochastic processes, obtain closed forms for moments up to second order, including the covariance structure for DFA and DCCA and a miscellany of law of large number related results. Our results generalize and improve several results presented in the literature. To verify the behavior of our theoretical results in small samples, we present a Monte Carlo simulation study and an empirical application to econometric time series.

math.ST

Kumaraswamy autoregressive moving average models for double bounded environmental data

In this paper we introduce the Kumaraswamy autoregressive moving average models (KARMA), which is a dynamic class of models for time series taking values in the double bounded interval $(a,b)$ following the Kumaraswamy distribution. The Kumaraswamy family of distribution is widely applied in many areas, especially hydrology and related fields. Classical examples are time series representing rates and proportions observed over time. In the proposed KARMA model, the median is modeled by a dynamic structure containing autoregressive and moving average terms, time-varying regressors, unknown parameters and a link function. We introduce the new class of models and discuss conditional maximum likelihood estimation, hypothesis testing inference, diagnostic analysis and forecasting. In particular, we provide closed-form expressions for the conditional score vector and conditional Fisher information matrix. An application to environmental real data is presented and discussed.

stat.ME

A Dynamic Model for Double Bounded Time Series With Chaotic Driven Conditional Averages

In this work we introduce a class of dynamic models for time series taking values on the unit interval. The proposed model follows a generalized linear model approach where the random component, conditioned on the past information, follows a beta distribution, while the conditional mean specification may include covariates and also an extra additive term given by the iteration of a map that can present chaotic behavior. The resulting model is very flexible and its systematic component can accommodate short and long range dependence, periodic behavior, laminar phases, etc. We derive easily verifiable conditions for the stationarity of the proposed model, as well as conditions for the law of large numbers and a Birkhoff-type theorem to hold. A Monte Carlo simulation study is performed to assess the finite sample behavior of the partial maximum likelihood approach for parameter estimation in the proposed model. Finally, an application to the proportion of stored hydroelectrical energy in Southern Brazil is presented.

math.ST