SearcharxivSearch

arXiv subjects

Marcelo Bourguignon

Publications and source records attributed to Marcelo Bourguignon.

17 recordsLinked to original sources

Modeling double bounded data based on correlated gamma random variables

Many types of bounded data defined on the unit interval arise naturally as ratios of the form $X/(X + Y)$. In the existing literature, the main statistical models proposed for this type of bounded data typically based on the assumption that the random variables $X$ and $Y$ are independent. However, this assumption is often unrealistic in practical applications, where $X$ and $Y$ tend to be correlated due to shared underlying mechanisms or common sources of variability. In this paper, we overcome such limitations and propose a model in which the marginal distributions of the two components are linked by a copula, leading to a more flexible and realistic representation of unit-interval data. In particular, in the proposed model, $X$ and $Y$ are dependent gamma random variables whose joint distribution is specified via Morgenstern's bivariate distribution}, allowing for positive and negative correlations between the components. The mathematical properties and practical applications are rigorously investigated. The resulting distribution exhibits a wide range of shapes, accommodating different degrees of skewness and, for some parameter configurations, more complex density structures. A Monte Carlo simulation study is carried out that shows the good performance of the maximum likelihood estimator in several scenarios of parameter choices. The potential and limitations of efficient likelihood-based computations are also discussed. We evaluate the effectiveness of the new model and its estimates in modeling real-world datasets related to economics.

stat.ME

On the distribution of a random variable involved in an independent ratio

In this paper, using inverse integral transforms, we derive the exact distribution of the random variable $X$ that is involved in the ratio $Z \stackrel{d}{=} X/(X+Y)$ where $X$ and $Y$ are independent random variables having the same support, and $Z$ and $Y$ have known distributions. We introduce new distributions this way. As applications of the obtained results, several examples are presented.

math.PR

Parametric quantile regression for income data

Univariate normal regression models are statistical tools widely applied in many areas of economics. Nevertheless, income data have asymmetric behavior and are best modeled by non-normal distributions. The modeling of income plays an important role in determining workers' earnings, as well as being an important research topic in labor economics. Thus, the objective of this work is to propose parametric quantile regression models based on two important asymmetric income distributions, namely, Dagum and Singh-Maddala distributions. The proposed quantile models are based on reparameterizations of the original distributions by inserting a quantile parameter. We present the reparameterizations, some properties of the distributions, and the quantile regression models with their inferential aspects. We proceed with Monte Carlo simulation studies, considering the maximum likelihood estimation performance evaluation and an analysis of the empirical distribution of two residuals. The Monte Carlo results show that both models meet the expected outcomes. We apply the proposed quantile regression models to a household income data set provided by the National Institute of Statistics of Chile. We showed that both proposed models had a good performance both in terms of model fitting. Thus, we conclude that results were favorable to the use of Singh-Maddala and Dagum quantile regression models for positive asymmetric data, such as income data.

stat.ME

The shared weighted Lindley frailty model for cluster failure time data

The primary goal of this paper is to introduce a novel frailty model based on the weighted Lindley (WL) distribution for modeling clustered survival data. We study the statistical properties of the proposed model. In particular, the amount of unobserved heterogeneity is directly parameterized on the variance of the frailty distribution such as gamma and inverse Gaussian frailty models. Parametric and semiparametric versions of the WL frailty model are studied. A simple expectation-maximization (EM) algorithm is proposed for parameter estimation. Simulation studies are conducted to evaluate its finite sample performance. Finally, we apply the proposed model to a real data set to analyze times after surgery in patients diagnosed with colorectal cancer and compare our results with classical frailty models carried out in this application, which shows the superiority of the proposed model. We implement an R package that includes estimation for fitting the proposed model based on the EM-algorithm.

stat.ME

A parametric quantile beta regression for modeling case fatality rates of COVID-19

Motivated by the case fatality rate (CFR) of COVID-19, in this paper, we develop a fully parametric quantile regression model based on the generalized three-parameter beta (GB3) distribution. Beta regression models are primarily used to model rates and proportions. However, these models are usually specified in terms of a conditional mean. Therefore, they may be inadequate if the observed response variable follows an asymmetrical distribution, such as CFR data. In addition, beta regression models do not consider the effect of the covariates across the spectrum of the dependent variable, which is possible through the conditional quantile approach. In order to introduce the proposed GB3 regression model, we first reparameterize the GB3 distribution by inserting a quantile parameter and then we develop the new proposed quantile model. We also propose a simple interpretation of the predictor-response relationship in terms of percentage increases/decreases of the quantile. A Monte Carlo study is carried out for evaluating the performance of the maximum likelihood estimates and the choice of the link functions. Finally, a real COVID-19 dataset from Chile is analyzed and discussed to illustrate the proposed approach.

stat.ME

A Bimodal Model for Extremes Data

In extreme values theory, for a sufficiently large block size, the maxima distribution is approximated by the generalized extreme value (GEV) distribution. The GEV distribution is a family of continuous probability distributions, which has wide applicability in several areas including hydrology, engineering, science, ecology and finance. However, the GEV distribution is not suitable to model extreme bimodal data. In this paper, we propose an extension of the GEV distribution that incorporate an additional parameter. The additional parameter introduces bimodality and to vary tail weight, i.e., this proposed extension is more flexible than the GEV distribution. Inference for the proposed distribution were performed under the likelihood paradigm. A Monte Carlo experiment is conducted to evaluate the performances of these estimators in finite samples with a discussion of the results. Finally, the proposed distribution is applied to environmental data sets, illustrating their capabilities in challenging cases in extreme value theory.

stat.ME

A Model for Bimodal Rates and Proportions

The beta model is the most important distribution for fitting data with the unit interval. However, the beta distribution is not suitable to model bimodal unit interval data. In this paper, we propose a bimodal beta distribution constructed by using an approach based on the alpha-skew-normal model. We discuss several properties of this distribution such as bimodality, real moments, entropy measures and identifiability. Furthermore, we propose a new regression model based on the proposed model and discuss residuals. Estimation is performed by maximum likelihood. A Monte Carlo experiment is conducted to evaluate the performances of these estimators in finite samples with a discussion of the results. An application is provided to show the modelling competence of the proposed distribution when the data sets show bimodality.

stat.ME

On the bimodal Gumbel model with application to environmental data

The Gumbel model is a very popular statistical model due to its wide applicability for instance in the course of certain survival, environmental, financial or reliability studies. In this work, we have introduced a bimodal generalization of the Gumbel distribution that can be an alternative to model bimodal data. We derive the analytical shapes of the corresponding probability density function and the hazard rate function and provide graphical illustrations. Furthermore, We have discussed the properties of this density such as mode, bimodality, moment generating function and moments. Our results were verified using the Markov chain Monte Carlo simulation method. The maximum likelihood method is used for parameters estimation. Finally, we also carry out an application to real data that demonstrates the usefulness of the proposed distribution.

stat.ME

Parametric quantile regression models for fitting double bounded response with application to COVID-19 mortality rate data

In this paper, we develop two fully parametric quantile regression models, based on power Johnson SB distribution Cancho et al. (2020), for modeling unit interval response at different quantiles. In particular, the conditional distribution is modelled by the power Johnson SB distribution. The maximum likelihood method is employed to estimate the model parameters. Simulation studies are conducted to evaluate the performance of the maximum likelihood estimators in finite samples. Furthermore, we discuss residuals and influence diagnostic tools. The effectiveness of our proposals is illustrated with two data set given by the mortality rate of COVID-19 in different countries.

stat.ME

Improved estimators in beta prime regression models

In this paper, we consider the beta prime regression model recently proposed by \cite{bour18}, which is tailored to situations where the response is continuous and restricted to the positive real line with skewed and long tails and the regression structure involves regressors and unknown parameters. We consider two different strategies of bias correction of the maximum-likelihood estimators for the parameters that index the model. In particular, we discuss bias-corrected estimators for the mean and the dispersion parameters of the model. Furthermore, as an alternative to the two analytically bias-corrected estimators discussed, we consider a bias correction mechanism based on the parametric bootstrap. The numerical results show that the bias correction scheme yields nearly unbiased estimates. An example with real data is presented and discussed.

stat.ME

The negative binomial beta prime regression model with cure rate

This paper introduces a cure rate survival model by assuming that the time to the event of interest follows a beta prime distribution and that the number of competing causes of the event of interest follows a negative binomial distribution. This model provides a novel alternative to the existing cure rate regression models due to its flexibility, as the beta prime model can exhibit greater levels of skewness and kurtosis than those of the gamma and inverse Gaussian distributions. Moreover, the hazard rate of this model can have an upside-down bathtub or an increasing shape. We approach both parameter estimation and local influence based on likelihood methods. In special, three perturbation schemes are considered for local influence. Numerical evaluation of the proposed model is performed by Monte Carlo simulations. In order to illustrate the potential for practice of our model we apply it to a real data set.

stat.ME

Zero-Modified Poisson-Lindley distribution with applications in zero-inflated and zero-deflated count data

The main object of this article is to present an extension of the zero-inflated Poisson-Lindley distribution, called of zero-modified Poisson-Lindley. The additional parameter $π$ of the zero-modified Poisson-Lindley has a natural interpretation in terms of either zero-deflated/inflated proportion. Inference is dealt with by using the likelihood approach. In particular the maximum likelihood estimators of the distribution's parameter are compared in small and large samples. We also consider an alternative bias-correction mechanism based on Efron's bootstrap resampling. The model is applied to real data sets and found to perform better than other competing models.

stat.ME

A new regression model for positive data

In this paper, we propose a regression model where the response variable is beta prime distributed using a new parameterization of this distribution that is indexed by mean and precision parameters. The proposed regression model is useful for situations where the variable of interest is continuous and restricted to the positive real line and is related to other variables through the mean and precision parameters. The variance function of the proposed model has a quadratic form. In addition, the beta prime model has properties that its competitor distributions of the exponential family do not have. Estimation is performed by maximum likelihood. Furthermore, we discuss residuals and influence diagnostic tools. Finally, we also carry out an application to real data that demonstrates the usefulness of the proposed model.

stat.ME

Extended Poisson INAR(1) processes with equidispersion, underdispersion and overdispersion

Real count data time series often show the phenomenon of the underdispersion and overdispersion. In this paper, we develop two extensions of the first-order integer-valued autoregressive process with Poisson innovations, based on binomial thinning, for modeling integer-valued time series with equidispersion, underdispersion and overdispersion. The main properties of the models are derived. The methods of conditional maximum likelihood, Yule-Walker and conditional least squares are used for estimating the parameters, and their asymptotic properties are established. We also use a test based on our processes for checking if the count time series considered is overdispersed or underdispersed. The proposed models are fitted to time series of number of weekly sales and of cases of family violence illustrating its capabilities in challenging cases of overdispersed and underdispersed count data.

stat.ME

Fractional approaches for the distribution of innovation sequence of INAR(1) processes

In this paper, we present a fractional decomposition of the probability generating function of the innovation process of the first-order non-negative integer-valued autoregressive [INAR(1)] process to obtain the corresponding probability mass function. We also provide a comprehensive review of integer-valued time series models, based on the concept of thinning operators, with geometric-type marginals. In particular, we develop four fractional approaches to obtain the distribution of innovation processes of the INAR(1) model and show that the distribution of the innovations sequence has geometric-type distribution. These approaches are discussed in detail and illustrated through a few examples. Finally, using the methods presented here, we develop four new first-order non-negative integer-valued autoregressive process for autocorrelated counts with overdispersion with known marginals, and derive some properties of these models.

stat.ME

A skew true INAR(1) process with application

Integer-valued time series models have been a recurrent theme considered in many papers in the last three decades, but only a few of them have dealt with models on $\mathbb Z$ (that is, including both negative and positive integers). Our aim in this paper is to introduce a first-order integer-valued autoregressive process on $\mathbb Z$ with skew discrete Laplace marginals (Kozubowski and Inusah, 2006). For this, we define a new operator that acts on two independent latent processes, similarly as made by Freeland (2010). We derive some joint and conditional basic properties of the proposed process such as characteristic function, moments, higher-order moments and jumps. Estimators for the parameters of our model are proposed and their asymptotic normality are established. We run a Monte Carlo simulation to evaluate the finite-sample performance of these estimators. In order to illustrate the potentiality of our process, we apply it to a real data set about population increase rates.

stat.ME

A new class of fatigue life distributions

In this paper, we introduce the Birnbaum-Saunders power series class of distributions which is obtained by compounding Birnbaum-Saunders and power series distributions. The new class of distributions has as a particular case the two-parameter Birnbaum-Saunders distribution. The hazard rate function of the proposed class can be increasing and upside-down bathtub shaped. We provide important mathematical properties such as moments, order statistics, estimation of the parameters and inference for large sample. Three special cases of the new class are investigated with some details. We illustrate the usefulness of the new distributions by means of two applications to real data sets.

stat.ME