SearcharxivSearch

arXiv subjects

Aniket Biswas

Publications and source records attributed to Aniket Biswas.

11 recordsLinked to original sources

Improving the adjusted Benjamini--Hochberg method using e-values in knockoff-assisted variable selection

Considering the knockoff-based multiple testing framework of Barber and Candès [2015], we revisit the method of Sarkar and Tang [2022] and identify it as a specific case of an un-normalized e-value weighted Benjamini-Hochberg procedure. Building on this insight, we extend the method to use bounded p-to-e calibrators that enable more refined and flexible weight assignments. Our approach generalizes the method of Sarkar and Tang [2022], which emerges as a special case corresponding to an extreme calibrator. Within this framework, we propose three procedures: an e-value weighted Benjamini-Hochberg method, its adaptive extension using an estimate of the proportion of true null hypotheses, and an adaptive weighted Benjamini-Hochberg method. We establish control of the false discovery rate (FDR) for the proposed methods. While we do not formally prove that the proposed methods outperform those of Barber and Candès [2015] and Sarkar and Tang [2022], simulation studies and real-data analysis demonstrate large and consistent improvement over the latter in all cases, and better performance than the knockoff method in scenarios with low target FDR, a small number of signals, and weak signal strength. Simulation studies and a real-data application in HIV-1 drug resistance analysis demonstrate strong finite sample FDR control and exhibit improved, or at least competitive, power relative to the aforementioned methods.

stat.ME

A new count model based on Poisson-Transmuted Geometric convolution

A novel over-dispersed discrete distribution, namely the PoiTG distribution is derived by the convolution of a Poisson variate and an independently distributed transmuted geometric random variable. This distribution generalizes the geometric, transmuted geometric, and PoiG distributions. Various important statistical properties of this count model, such as the probability generating function, the moment generating function, the moments, the survival function, and the hazard rate function are investigated. Stochastic ordering for the proposed model are also studied in details. The maximum likelihood estimators of the parameters are obtained using general optimization approach and the EM algorithm approach. It is envisaged that the proposed distribution may prove to be useful for the practitioners for modelling over-dispersed count data compared to its closest competitors.

math.ST

A new over-dispersed count model

A new two-parameter discrete distribution, namely the PoiG distribution is derived by the convolution of a Poisson variate and an independently distributed geometric random variable. This distribution generalizes both the Poisson and geometric distributions and can be used for modelling over-dispersed as well as equi-dispersed count data. A number of important statistical properties of the proposed count model, such as the probability generating function, the moment generating function, the moments, the survival function and the hazard rate function. Monotonic properties are studied, such as the log concavity and the stochastic ordering are also investigated in detail. Method of moment and the maximum likelihood estimators of the parameters of the proposed model are presented. It is envisaged that the proposed distribution may prove to be useful for the practitioners for modelling over-dispersed count data compared to its closest competitors.

stat.ME

A new method for constructing continuous distributions on the unit interval

A novel approach towards construction of absolutely continuous distributions over the unit interval is proposed. Considering two absolutely continuous random variables with positive support, this method conditions on their convolution to generate a new random variable in the unit interval. This approach is demonstrated using some popular choices of the positive random variables such as the exponential, Lindley, gamma. Some existing distributions like the uniform and the beta are formulated with this method. Several new structures of density functions having potential for future application in real life problems are also provided. One of the new distributions having one parameter is considered for parameter estimation and real life modelling application and shown to provide better fit than the popular one parameter Topp-Leone model.

math.ST

A new generalization of the geometric distribution using Azzalini's mechanism: properties and application

The skewing mechanism of Azzalini for continuous distributions is used for the first time to derive a new generalization of the geometric distribution. Various structural properties of the proposed distribution are investigated. Characterizations, including a new result for the geometric distribution, in terms of the proposed model are established. Extensive simulation experiment is done to evaluate performance of the maximum likelihood estimation method. Likelihood ratio test for the necessity of additional skewing parameter is derived and corresponding simulation based power study is also reported. Two real life count datasets are analyzed with the proposed model and compared with some recently introduced two-parameter count models. The findings clearly indicate the superiority of the proposed model over the existing ones in modelling real life count data.

math.ST

Modified estimator for the proportion of true null hypotheses under discrete setup with proven FDR control by the adaptive Benjamini-Hochberg procedure

Some crucial issues about a recently proposed estimator for the proportion of true null hypotheses ($π_0$) under discrete setup are discussed. An estimator for $π_0$ is introduced under the same setup. The estimator may be seen as a modification of a very popular estimator for $π_0$, originally proposed under the assumption of continuous test statistics. It is shown that adaptive Benjamini-Hochberg procedure remains conservative with the new estimator for $π_0$ being plugged in.

math.ST

Bias corrected estimators for proportion of true null hypotheses under exponential model: Application of adaptive FDR-controlling in segmented failure data

Two recently introduced model based bias corrected estimators for proportion of true null hypotheses ($π_0$) under multiple hypotheses testing scenario have been restructured for exponentially distributed random observations available for each of the common hypotheses. Based on stochastic ordering, a new motivation behind formulation of some related estimators for $π_0$ is given. The reduction of bias for the model based estimators are theoretically justified and algorithms for computing the estimators are also presented. The estimators are also used to formulate a popular adaptive multiple testing procedure. Extensive numerical study supports superiority of the bias corrected estimators. We also point out the adverse effect of using the model based bias correction method without proper assessment of the underlying distribution. A case-study is done with a synthetic dataset in connection with reliability and warranty studies to demonstrate the applicability of the procedure, under a non-Gaussian set up. The results obtained are in line with the intuition and experience of the subject expert. An intriguing discussion has been attempted to conclude the article that also indicates the future scope of study.

math.ST

Estimating the proportion of true null hypotheses with application in microarray data

A new formulation for the proportion of true null hypotheses $(π_0)$, based on the sum of all $p$-values and the average of expected $p$-value under the false null hypotheses has been proposed in the current work. This formulation of the parameter of interest $π_0$ has also been used to construct a new estimator for the same. The proposed estimator removes the problem of choosing tuning parameters in the existing estimators. Though the formulation is quite general, computation of the new estimator demands use of an initial estimate of $π_0$. The issue of choosing an appropriate initial estimator is also discussed in this work. The current work assumes normality of each gene expression level and also assumes similar tests for all the hypotheses. Extensive simulation study shows that, the proposed estimator performs better than its closest competitor, the estimator proposed in Cheng et al., 2015 over a substantial continuous subinterval of the parameter space, under independence and weak dependence among the gene expression levels. The proposed method of estimation is applied to two real gene expression level data-sets and the results are in line with what is obtained by the competing method.

math.ST

Selection of link function in binary regression: A case-study with world happiness report on immigration

Selection of appropriate link function for binary regression remains an important issue for data analysis and its influence on related inference. We prescribe a new data-driven methodology to search for the same, considering some popular classification assessment metrics. A case-study with World Happiness report,2018 with special reference to immigration is presented for demonstrating utility of the prescribed routine.

stat.AP

Further study on inferential aspects of log-Lindley distribution with an application of stress-strength reliability in insurance

The log-Lindley distribution was recently introduced in the literature as a viable alternative to the Beta distribution. This distribution has a simple structure and possesses useful theoretical properties relevant in insurance. Classical estimation methods have been well studied. We introduce estimation of parameters from Bayesian point of view for this distribution. Explicit structure of stress-strength reliability and its inference under both classical and Bayesian set-up is addressed. Extensive simulation studies show marked improvement with Bayesian approach over classical given reasonable prior information. An application of a useful metric of discrepancy derived from stress-strength reliability is considered and computed for two categories of firm with respect to a certain financial indicator.

math.ST

R = P(Y < X) for unit-Lindley distribution: inference with an application in public health

The unit-Lindley distribution was recently introduced in the literature as a viable alternative to the Beta and the Kumaraswamy distributions with support in (0; 1). This distribution enjoys many virtuous properties over the named distributions. In this article, we address the issue of parameter estimation from a Bayesian perspective and study relative performance of different estimators through extensive simulation studies. Significant emphasis is given to the estimation of stress-strength reliability employing classical as well as Bayesian approach. A non-trivial useful application in the public health domain is presented proposing a simple metric of discrepancy.

stat.ME