SearcharxivSearch

arXiv subjects

Ishapathik Das

Publications and source records attributed to Ishapathik Das.

9 recordsLinked to original sources

Goodness-of-fit testing for the Pareto type-I distribution based on a mean residual life characterization

The statistical analysis of heavy-tailed data has received considerable attention because extreme observations frequently arise in many practical applications. The Pareto type-I distribution is a fundamental heavy-tailed model used in economics, finance, actuarial science, insurance, reliability, and extreme value analysis. In this paper, we propose novel goodness-of-fit tests for the Pareto distribution using a mean residual life characterization. The test statistic is constructed using U-statistic theory, and its asymptotic behaviour is established under both the null and alternative hypotheses. Its finite-sample performance is evaluated through Monte Carlo simulations using maximum-likelihood and method-of-moments estimation and compared with existing tests. The results show that the proposed test controls the nominal significance level and performs competitively in terms of power across a broad range of alternatives. Finally, the proposed methodology is illustrated using the Danish fire insurance loss and pollution datasets.

stat.ME

Robust Modeling of Extremes in the Presence of Inliers with Enhanced Tail Estimation

Extreme value theory provides a fundamental framework for modeling rare and extreme events; however, threshold selection remains a persistent challenge, particularly in the presence of inliers such as instantaneous or early failures. Such observations commonly arise in applications including reliability studies and environmental data, where clusters of observations near the origin or at the origin can substantially distort classical threshold selection procedures and tail inference. In this paper, we propose a robust modeling framework that accounts for inliers, extremes, and the tail proportion. Parameter estimation is carried out using maximum likelihood. The proposed methodology is compared with classical numerical and graphical diagnostic tools, including the mean excess plot, parameter stability plot, Hill plot, and Pickands plot, as well as existing extreme value mixture models. The theoretical properties of the proposed model are established, and its performance is evaluated through extensive Monte Carlo simulations and real-data applications. The results demonstrate that the proposed methodology provides more accurate threshold estimation and more reliable extreme-value inference in the presence of inliers compared with existing classical approaches. Overall, the proposed methodology provides more accurate threshold estimation and tail inference in the presence of inliers, addressing key limitations of existing methods.

stat.ME

A Flexible Modeling of Extremes in the Presence of Inliers

Many random phenomena, including life-testing and environmental data, show positive values and excess zeros, which pose modeling challenges. In life testing, immediate failures result in zero lifetimes, often due to defects or poor quality, especially in electronics and clinical trials. These failures, called inliers at zero, are difficult to model using standard approaches. The presence and proportion of inliers may influence the accuracy of extreme value analysis, bias parameter estimates, or even lead to severe events or extreme effects, such as drought or crop failure. In such scenarios, a key issue in extreme value analysis is determining a suitable threshold to capture tail behaviour accurately. Although some extreme value mixture models address threshold and tail estimation, they often inadequately handle inliers, resulting in suboptimal results. Bulk model misspecification can affect the threshold, extreme value estimates, and, in particular, the tail proportion. There is no unified framework for defining extreme value mixture models, especially the tail proportion. This paper proposes a flexible model that handles extremes, inliers, and the tail proportion. Parameters are estimated using maximum likelihood estimation. Compared the proposed model estimates with the classical mean excess plot, parameter stability plot, and Pickands plot estimates. Theoretical results are established, and the proposed model outperforms traditional methods in both simulation studies and real data analysis.

stat.ME

Modeling Extreme Events in the Presence of Inlier: A Mixture Approach

In many random phenomena, such as life-testing experiments and environmental data (like rainfall data), there are often positive values and an excess of zeros, which create modeling challenges. In life testing, immediate failures result in zero lifetimes, often due to defects or poor quality, especially in electronics and clinical trials. These failures, called zero inliers, are difficult to model using standard approaches. When studying extreme values in the above scenarios, a key issue is selecting an appropriate threshold for accurate tail approximation of the population using asymptotic models. While some extreme value mixture models address threshold estimation and tail approximation, conventional parametric and non-parametric bulk and generalised Pareto distribution (GPD) approaches often neglect inliers, leading to suboptimal results. This paper introduces a framework for modeling extreme events and inliers using the GPD, addressing threshold uncertainty and effectively capturing inliers at zero. The model's parameters are estimated using the maximum likelihood estimation (MLE) method, ensuring optimal precision. Through simulation studies and real-world applications, we demonstrate that the proposed model significantly outperforms the traditional methods, which typically neglect inliers at the origin.

stat.ME

Mean Residual Life Ageing Intensity Function

The ageing intensity function is a powerful analytical tool that provides valuable insights into the ageing process across diverse domains such as reliability engineering, actuarial science, and healthcare. Its applications continue to expand as researchers delve deeper into understanding the complex dynamics of ageing and its implications for society. One common approach to defining the ageing intensity function is through the hazard rate or failure rate function, extensively explored in scholarly literature. Equally significant to the hazard rate function is the mean residual life function, which plays a crucial role in analyzing the ageing patterns exhibited by units or components. This article introduces the mean residual life ageing intensity (MRLAI) function to delve into component ageing behaviours across various distributions. Additionally, we scrutinize the closure properties of the MRLAI function across different reliability operations. Furthermore, a new order termed the mean residual life ageing intensity order is defined to analyze the ageing behaviour of a system, and the closure property of this order under various reliability operations is discussed.

math.ST

Geo-Spatial Cluster based Hybrid Spatio-Temporal Copula Interpolation

In the absence of Gaussianity assumptions without disturbing spatial continuity interpolating along the whole spatial surface for different time lags is challenging. The past researchers pay enough attention to Spatio-temporal interpolation ignoring the dynamic behavior of a spatial mean function, threshold distance, and direction of maintaining spatial continuity. Therefore, we employ hierarchical spatial clustering (HSC) to preserve local spatial stationarity. This research work introduces a hybrid extreme valued copula-based Spatio-temporal interpolation algorithm. Spatial dependence is captured by a blended extreme valued probability distribution (BEVD). Temporal dependency is modeled by the Bi-directional long short-time memory (BLSTM) at different temporal granularities, 1 month, 2 months, and 3 months. Spatio-temporal dependence is modeled by the Gumbel-Hougaard copula (GH). We apply the proposed Spatio-temporal interpolation approach to the air pollution data (Outdoor Particulate Matter (PM) concentration) of Delhi, collected from the website of the Central Pollution Control Board, India as a crucial circumstantial study. This article describes a probabilistic-recurrent neural networking algorithm for Spatio-temporal interpolation. This Spatio-temporal hybrid copula interpolation algorithm outperforms and is efficient enough to detect spatial trends and temporal influence. From the entire research, we notice that PM concentration in a year reaches a maximum, generally in November and December. The northern and central part of Del-hi is the most sensitive regarding air pollution.

stat.AP

Modified Bivariate Weibull Distribution Allowing Instantaneous and Early Failures

In reliability and life data analysis, the Weibull distribution is widely used to accommodate more data characteristics by changing the values of the parameters. We frequently observe many zeros or close to zero data points in reliability and life testing experiments. We call this phenomenon a nearly instantaneous failure. Many researchers modified the commonly used univariate parametric models such as exponential, gamma, Weibull, and log-normal distributions to appropriately fit such data having instantaneous failure observations. Researchers also find bivariate correlated life testing data having many observations near a particular point while the remaining observations follow some continuous distribution. This situation defines as responses having early failures for such bivariate responses. If the point is the origin, then we call the situation a nearly instantaneous failure for the responses. Here, we propose a modified bivariate Weibull distribution that allows early failure by combining bivariate uniform distribution and bivariate Weibull distribution. The bivariate Weibull distribution is constructed using a 2-dimensional copula, assuming the marginal distributions as two parametric Weibull distributions. We derive some properties of that modified bivariate Weibull distribution, mainly the joint probability density function, the survival (reliability) function, and the hazard (failure rate) function. The model's unknown parameters are estimated using the Maximum Likelihood Estimation (MLE) technique combined with a machine learning clustering algorithm. Numerical examples are provided using simulated data to illustrate and test the performance of the proposed methodologies. The method is also applied to real data and compared with existing approaches to model such data in the literature.

stat.ME

Spatial Cluster-based Copula Model to Interpolate Skewed Conditional Spatial Random Field

Interpolating a skewed conditional spatial random field with missing data is cumbersome in the absence of Gaussianity assumptions. Maintaining spatial homogeneity and continuity around the observed random spatial point is also challenging, especially when interpolating along a spatial surface, focusing on the boundary points as a neighborhood. Otherwise, the point far away from one may appear the closest to another. As a result, importing the hierarchical clustering concept on the spatial random field is as convenient as developing the copula with the interface of the Expectation-Maximization algorithm and concurrently utilizing the idea of the Bayesian framework. This paper introduces a spatial cluster-based C-vine copula and a modified Gaussian kernel to derive a novel spatial probability distribution. Another investigation in this paper uses an algorithm in conjunction with a different parameter estimation technique to make spatial-based copula interpolation more compatible and efficient. We apply the proposed spatial interpolation approach to the air pollution of Delhi as a crucial circumstantial study to demonstrate this newly developed novel spatial estimation technique.

stat.ME

Statistical assessment of spatio-temporal impact of lockdown on air pollution using different modelling approaches in India

One of the main contributors to air pollution is particulate matter (PMxy), which causes several COVID-19 related diseases such as respiratory problems and cardiovascular disorders. Therefore, the spatial and temporal trend analysis of particulate matter and the mass concentration of all aerosol particles less than 2.5 m in diameter (PM2.5) has become critical to control the risk factors of co-morbidity of a patient. Lockdown plays a significant role in maintaining COVID-19 cases as well as air pollution, including particulate matter. This study aims to analyse the effect of the lockdown on controlling air pollution in metropolitan cities in India through various statistical modelling approaches. Most research articles in the literature assume a linear relationship between responses and covariates and take independent and identically distributed error terms in the model, which may not be appropriate for analysing such air pollution data. In this study, we performed a pattern analysis of daily PM2.5 emissions in various major activity zones during 2019 and 2020. By measuring the lockdown effect, we also considered seasonal influence.

stat.AP