SearcharxivSearch

arXiv subjects

Francisco Louzada

Publications and source records attributed to Francisco Louzada.

At least 19 recordsLinked to original sources

Noniterative Likelihood-Derived Estimation through Auxiliary Estimating Equations

Closed-form estimators are useful when parametric models are repeatedly refitted but likelihood maximization is iterative. We study a likelihood-derived construction of auxiliary estimating equations obtained by differentiating a positive auxiliary function and centering the derivatives under a baseline model. The resulting estimators are just-identified Z-estimators, or equivalently GMM estimators; the contribution is a constructive route to explicitly invertible equations rather than a replacement for general estimating-equation theory. We establish local existence, uniqueness and asymptotic normality of the selected root, characterize optimal linear combinations through Godambe information, and give constrained variants and influence-function diagnostics. Transformed exponential families, Beta, Weibull AFT regression, Wishart covariance estimation, zero-inflated counts, tail modelling and copula dependence illustrate the construction. Simulations and survival-tree split screening quantify the statistical-computational trade-off.

stat.ME

Violent event-related fatality patterns in Ethiopia: a Bayesian spatiotemporal perspective

Fatalities resulting from violence in armed conflict have long been a significant public health issue in Ethiopia. Despite the severity of this problem, more comprehensive quantitative scientific studies need to be conducted to elucidate the sequence and dynamics of these occurrences. In response, this study introduces a spatio-temporal statistical method designed to uncover the patterns of fatalities associated with violent events in Ethiopia. The research employs a two-part zero-inflated Bayesian generalized additive mixed model, which integrates a spatio-temporal component to map the fatality patterns across Ethiopian regions. The dataset utilized originates from the Armed Conflict Location and Event Data Project, covering fatality counts related to violent events from 1997 to 2022. The analysis revealed that nine out of thirteen administrative regions exhibited a probability greater than 0.6 for fatality occurrence due to violent events, with five regions surpassing a 0.7 probability threshold. These five regions include Benishangul Gumz, Gambela, Oromia, Somali, and the South West Ethiopian People's Region. Notably, the Tigray region displayed the highest probability (0.558) of experiencing more than 20 deaths per violent event, followed by the Benishangul Gumz region with a probability of 0.306. Encouragingly, the findings also indicate an average decline in fatalities per violent event over time. Specifically, the probability of more than 20 deaths per event was 0.401 in 2020, which decreased to 0.148 by 2022. These insights are invaluable for the government, policymakers, political leaders, and traditional or religious authorities in Ethiopia, enabling them to make informed, strategic decisions to mitigate and ultimately prevent violence-related fatalities in the country.

stat.AP

Dynamic Skewness in Stochastic Volatility Models: A Penalized Prior Approach

Financial time series often exhibit skewness and heavy tails, making it essential to use models that incorporate these characteristics to ensure greater reliability in the results. Furthermore, allowing temporal variation in the skewness parameter can bring significant gains in the analysis of this type of series. However, for more robustness, it is crucial to develop models that balance flexibility and parsimony. In this paper, we propose dynamic skewness stochastic volatility models in the SMSN family (DynSSV-SMSN), using priors that penalize model complexity. Parameter estimation was carried out using the Hamiltonian Monte Carlo (HMC) method via the \texttt{RStan} package. Simulation results demonstrated that penalizing priors present superior performance in several scenarios compared to the classical choices. In the empirical application to returns of cryptocurrencies, models with heavy tails and dynamic skewness provided a better fit to the data according to the DIC, WAIC, and LOO-CV information criteria.

q-fin.ST

Chilean Avian flu and its marine impacts: an online Statistical Process Control task

The rapid spread of the HPAI H5N1 virus, responsible for the Avian Flu, is causing a great catastrophe on the South American Pacific coast (especially in the south of Peru and north of Chile). Although very little attention has been delivered to this pandemic, it presents a tremendous lethal rate, though the number of infected humans is relatively low. Towards monitoring and statistical control, this work shows the Chilean national statistics from the year 2023, and presents the developed online tool for supporting the government's decision-making. Additionally, a Bayesian hierarchical spatiotemporal model was used to model the joint analysis of the weekly registered animal including spatial covariates as well as specific and shared spatial effects that take into account the potential autocorrelation between the HPAI H5N1 virus per Region. Our findings allow us to identify the hot-spot areas with high amounts of dead bodies (mostly pinnipeds and penguins) and their evolution over time.

stat.AP

On the unification of zero-adjusted cure survival models

This paper proposes a unified version of survival models that accounts for both zero-adjustment and cure proportions in various latent competing causes, useful in data where survival times may be zero or cure proportions are present. These models are particularly relevant in scenarios like childbirth duration in sub-Saharan Africa. Different competing cause distributions were considered, including Binomial, Geometric, Poisson, and Negative Binomial. The model's maximum likelihood point estimators and asymptotic confidence intervals were evaluated through simulation, demonstrating improved accuracy with larger sample sizes. The model best fits real obstetric data when assuming geometrically distributed causes. This flexible model, capable of considering different distributions for the lifetime of susceptible individuals and competing causes, is an effective tool for adjusting survival data, indicating broad application potential.

math.ST

Ablation Studies for Novel Treatment Effect Estimation Models

Ablation studies are essential for understanding the contribution of individual components within complex models, yet their application in nonparametric treatment effect estimation remains limited. This paper emphasizes the importance of ablation studies by examining the Bayesian Causal Forest (BCF) model, particularly the inclusion of the estimated propensity score $\hatπ(x_i)$ intended to mitigate regularization-induced confounding (RIC). Through a partial ablation study utilizing a total of nine synthetic, we demonstrate that excluding $\hatπ(x_i)$ does not diminish the model's performance in estimating average and conditional average treatment effects or in uncertainty quantification. Moreover, omitting $\hatπ(x_i)$ reduces computational time by approximately 21%. These findings could suggest that the BCF model's inherent flexibility suffices in adjusting for confounding without explicitly incorporating the propensity score. The study advocates for the routine use of ablation studies in treatment effect estimation to ensure model components are essential and to prevent unnecessary complexity.

stat.ME

Sampling with censored data: a practical guide

In this review, we present a simple guide for researchers to obtain pseudo-random samples with censored data. We focus our attention on the most common types of censored data, such as type I, type II, and random censoring. We discussed the necessary steps to sample pseudo-random values from long-term survival models where an additional cure fraction is informed. For illustrative purposes, these techniques are applied in the Weibull distribution. The algorithms and codes in R are presented, enabling the reproducibility of our study. Finally, we developed an R package that encapsulates these methodologies, providing researchers with practical tools for implementation.

stat.CO

Objective Bayesian Analysis for the Differential Entropy of the Gamma Distribution

The present paper introduces a fully objective Bayesian analysis to obtain the posterior distribution of an entropy measure. Notably, we consider the gamma distribution, which describes many natural phenomena in physics, engineering, and biology. We reparametrize the model in terms of entropy, and different objective priors are derived, such as Jeffreys prior, reference prior, and matching priors. Since the obtained priors are improper, we prove that the obtained posterior distributions are proper and that their respective posterior means are finite. An intensive simulation study is conducted to select the prior that returns better results regarding bias, mean square error, and coverage probabilities. The proposed approach is illustrated in two datasets: the first relates to the Achaemenid dynasty reign period, and the second describes the time to failure of an electronic component in a sugarcane harvest machine.

math.ST

fakenewsbr: A Fake News Detection Platform for Brazilian Portuguese

The proliferation of fake news has become a significant concern in recent times due to its potential to spread misinformation and manipulate public opinion. This paper presents a comprehensive study on detecting fake news in Brazilian Portuguese, focusing on journalistic-type news. We propose a machine learning-based approach that leverages natural language processing techniques, including TF-IDF and Word2Vec, to extract features from textual data. We evaluate the performance of various classification algorithms, such as logistic regression, support vector machine, random forest, AdaBoost, and LightGBM, on a dataset containing both true and fake news articles. The proposed approach achieves high accuracy and F1-Score, demonstrating its effectiveness in identifying fake news. Additionally, we developed a user-friendly web platform, fakenewsbr.com, to facilitate the verification of news articles' veracity. Our platform provides real-time analysis, allowing users to assess the likelihood of fake news articles. Through empirical analysis and comparative studies, we demonstrate the potential of our approach to contribute to the fight against the spread of fake news and promote more informed media consumption.

cs.CL

Generalizing the normality: a novel towards different estimation methods for skewed information

Normality is the most often mathematical supposition used in data modeling. Nonetheless, even based on the law of large numbers (LLN), normality is a strong presumption given that the presence of asymmetry and multi-modality in real-world problems is expected. Thus, a flexible modification in the Normal distribution proposed by Elal-Olivero [12] adds a skewness parameter, called Alpha-skew Normal (ASN) distribution, enabling bimodality and fat-tail, if needed, although sometimes not trivial to estimate this third parameter (regardless of the location and scale). This work analyzed seven different statistical inferential methods towards the ASNdistribution on synthetic data and historical data of water flux from 21 rivers (channels) in the Atacama region. Moreover, the contribution of this paper is related to the probability estimation surrounding the rivers' flux level in Copiapo city neighborhood, the most important economic city of the third Chilean region, and known to be located in one of the driest areas on Earth, besides the North and the South Pole

stat.ME

Incorporation of frailties into a non-proportional hazard regression model and its diagnostics for reliability modeling of downhole safety valves

In this paper, our proposal consists of incorporating frailty into a statistical methodology for modeling time-to-event data, based on non-proportional hazards regression model. Specifically, we use the generalized time-dependent logistic (GTDL) model with a frailty term introduced in the hazard function to control for unobservable heterogeneity among the sampling units. We also add a regression in the parameter that measures the effect of time, since it can directly reflect the influence of covariates on the effect of time-to-failure. The practical relevance of the proposed model is illustrated in a real problem based on a data set for downhole safety valves (DHSVs) used in offshore oil and gas production wells. The reliability estimation of DHSVs can be used, among others, to predict the blowout occurrence, assess the workover demand and aid decision-making actions.

stat.AP

Spatial Statistical Models: an overview under the Bayesian Approach

Spatial documentation is exponentially increasing given the availability of Big IoT Data, enabled by the devices miniaturization and data storage capacity. Bayesian spatial statistics is a useful statistical tool to determine the dependence structure and hidden patterns over space through prior knowledge and data likelihood. Nevertheless, this modeling class is not well explored as the classification and regression machine learning models given their simplicity and often weak (data) independence supposition. In this manner, this systematic review aimed to unravel the main models presented in the literature in the past 20 years, identify gaps, and research opportunities. Elements such as random fields, spatial domains, prior specification, covariance function, and numerical approximations were discussed. This work explored the two subclasses of spatial smoothing global and local.

stat.ME

Power laws in the Roman Empire: a survival analysis

The Roman Empire shaped Western civilization, and many Roman principles are embodied in modern institutions. Although its political institutions proved both resilient and adaptable, allowing it to incorporate diverse populations, the Empire suffered from many internal conflicts. Indeed, most emperors died violently, from assassination, suicide, or in battle. These internal conflicts produced patterns in the length of time that can be identified by statistical analysis. In this paper, we study the underlying patterns associated with the reign of the Roman emperors by using statistical tools of survival data analysis. We consider all the 175 Roman emperors and propose a new power-law model with change points to predict the time-to-violent-death of the Roman emperors. This model encompasses data in the presence of censoring and long-term survivors, providing more accurate predictions than previous models. Our results show that power-law distributions can also occur in survival data, as verified in other data types from natural and artificial systems, reinforcing the ubiquity of power law distributions. The generality of our approach paves the way to further related investigations not only in other ancient civilizations but also in applications in engineering and medicine.

stat.AP

Power laws distributions in objective priors

The use of objective prior in Bayesian applications has become a common practice to analyze data without subjective information. Formal rules usually obtain these priors distributions, and the data provide the dominant information in the posterior distribution. However, these priors are typically improper and may lead to improper posterior. Here, we show, for a general family of distributions, that the obtained objective priors for the parameters either follow a power-law distribution or has an asymptotic power-law behavior. As a result, we observed that the exponents of the model are between 0.5 and 1. Understand these behaviors allow us to easily verify if such priors lead to proper or improper posteriors directly from the exponent of the power-law. The general family considered in our study includes essential models such as Exponential, Gamma, Weibull, Nakagami-m, Haf-Normal, Rayleigh, Erlang, and Maxwell Boltzmann distributions, to list a few. In summary, we show that comprehending the mechanisms describing the shapes of the priors provides essential information that can be used in situations where additional complexity is presented.

math.ST

Multiple repairable systems under dependent competing risks with nonparametric Frailty

The aim of this article is to analyze data from multiple repairable systems under the presence of dependent competing risks. In order to model this dependence structure, we adopted the well-known shared frailty model. This model provides a suitable theoretical basis for generating dependence between the components failure times in the dependent competing risks model. It is known that the dependence effect in this scenario influences the estimates of the model parameters. Hence, under the assumption that the cause-specific intensities follow a PLP, we propose a frailty-induced dependence approach to incorporate the dependence among the cause-specific recurrent processes. Moreover, the misspecification of the frailty distribution may lead to errors when estimating the parameters of interest. Because of this, we considered a Bayesian nonparametric approach to model the frailty density in order to offer more flexibility and to provide consistent estimates for the PLP model, as well as insights about heterogeneity among the systems. Both simulation studies and real case studies are provided to illustrate the proposed approaches and demonstrate their validity.

stat.AP

Random Machines Regression Approach: an ensemble support vector regression model with free kernel choice

Machine learning techniques always aim to reduce the generalized prediction error. In order to reduce it, ensemble methods present a good approach combining several models that results in a greater forecasting capacity. The Random Machines already have been demonstrated as strong technique, i.e: high predictive power, to classification tasks, in this article we propose an procedure to use the bagged-weighted support vector model to regression problems. Simulation studies were realized over artificial datasets, and over real data benchmarks. The results exhibited a good performance of Regression Random Machines through lower generalization error without needing to choose the best kernel function during tuning process.

stat.ML

Random Machines: A bagged-weighted support vector model with free kernel choice

Improvement of statistical learning models in order to increase efficiency in solving classification or regression problems is still a goal pursued by the scientific community. In this way, the support vector machine model is one of the most successful and powerful algorithms for those tasks. However, its performance depends directly from the choice of the kernel function and their hyperparameters. The traditional choice of them, actually, can be computationally expensive to do the kernel choice and the tuning processes. In this article, it is proposed a novel framework to deal with the kernel function selection called Random Machines. The results improved accuracy and reduced computational time. The data study was performed in simulated data and over 27 real benchmarking datasets.

stat.ML

An Extended Poisson Family of Life Distribution: A Unified Approach in Competitive and Complementary Risks

In this paper, we introduce a new approach to generate flexible parametric families of distributions. These models arise on competitive and complementary risks scenario, in which the lifetime associated with a particular risk is not observable, rather, we observe only the minimum/maximum lifetime value among all risks. The latent variables have a zero truncated Poisson distribution. For the proposed family of distribution, the extra shape parameter has an important physical interpretation in the competing and complementary risks scenario. The mathematical properties and inferential procedures are discussed. The proposed approach is applied in some existing distributions in which it is fully illustrated by an important data set.

stat.AP