Searcharxiv⌕ Search

arXiv subjects

Kaushik Jana

Publications and source records attributed to Kaushik Jana.

6 recordsLinked to original sources

Causal Analysis at Extreme Quantiles with Application to London Traffic Flow Data

Transport engineers employ various interventions to enhance traffic-network performance. Quantifying the impacts of Cycle Superhighways is complicated due to the non-random assignment of such an intervention over the transport network. Treatment effects on asymmetric and heavy-tailed distributions are better reflected at extreme tails rather than at the median. We propose a novel method to estimate the treatment effect at extreme tails incorporating heavy-tailed features in the outcome distribution. The analysis of London transport data using the proposed method indicates that the extreme traffic flow increased substantially after Cycle Superhighways came into operation.

stat.ME↗

Nonparametric quantile regression for time series with replicated observations and its application to climate data

This paper proposes a model-free nonparametric estimator of conditional quantile of a time series regression model where the covariate vector is repeated many times for different values of the response. This type of data is abound in climate studies. To tackle such problems, our proposed method exploits the replicated nature of the data and improves on restrictive linear model structure of conventional quantile regression. Relevant asymptotic theory for the nonparametric estimators of the mean and variance function of the model are derived under a very general framework. We provide a detailed simulation study which clearly demonstrates the gain in efficiency of the proposed method over other benchmark models, especially when the true data generating process entails nonlinear mean function and heteroskedastic pattern with time dependent covariates. The predictive accuracy of the non-parametric method is remarkably high compared to other methods when attention is on the higher quantiles of the variable of interest. Usefulness of the proposed method is then illustrated with two climatological applications, one with a well-known tropical cyclone wind-speed data and the other with an air pollution data.

stat.ME↗

Scoring Predictions at Extreme Quantiles

Prediction of quantiles at extreme tails is of interest in numerous applications. Extreme value modelling provides various competing predictors for this point prediction problem. A common method of assessment of a set of competing predictors is to evaluate their predictive performance in a given situation. However, due to the extreme nature of this inference problem, it can be possible that the predicted quantiles are not seen in the historical records, particularly when the sample size is small. This situation poses a problem to the validation of the prediction with its realisation. In this article, we propose two non-parametric scoring approaches to assess extreme quantile prediction mechanisms. The proposed assessment methods are based on predicting a sequence of equally extreme quantiles on different parts of the data. We then use the quantile scoring function to evaluate the competing predictors. The performance of the scoring methods is compared with the conventional scoring method and the superiority of the former methods are demonstrated in a simulation study. The methods are then applied to reanalyse cyber Netflow data from Los Alamos National Laboratory and daily precipitation data at a station in California available from Global Historical Climatology Network.

stat.AP↗

Improving linear quantile regression for replicated data

This paper deals with improvement of linear quantile regression, when there are a few distinct values of the covariates but many replicates. On can improve asymptotic efficiency of the estimated regression coefficients by using suitable weights in quantile regression, or simply by using weighted least squares regression on the conditional sample quantiles. The asymptotic variances of the unweighted and weighted estimators coincide only in some restrictive special cases, e.g., when the density of the conditional response has identical values at the quantile of interest over the support of the covariate. The dominance of the weighted estimators is demonstrated in a simulation study, and through the analysis of a data set on tropical cyclones.

stat.AP↗

Space-efficient estimation of empirical tail dependence coefficients for bivariate data streams

This article proposes a space-efficient approximation to empirical tail dependence coefficients of an indefinite bivariate stream of data. The approximation, which has stream-length invariant error bounds, utilises recent work on the development of a summary for bivariate empirical copula functions. The work in this paper accurately approximates a bivariate empirical copula in the tails of each marginal distribution, therefore modelling the tail dependence between the two variables observed in the data stream. Copulas evaluated at these marginal tails can be used to estimate the tail dependence coefficients. Modifications to the space-efficient bivariate copula approximation, presented in this paper, allow the error of approximations to the tail dependence coefficients to remain stream-length invariant. Theoretical and numerical evidence of this, including a case-study using the Los Alamos National Laboratory netflow data-set, is provided within this article.

stat.CO↗

The Statistical Face of a Region under Monsoon Rainfall in Eastern India

A region under rainfall is a contiguous spatial area receiving positive precipitation at a particular time. The probabilistic behavior of such a region is an issue of interest in meteorological studies. A region under rainfall can be viewed as a shape object of a special kind, where scale and rotational invariance are not necessarily desirable attributes of a mathematical representation. For modeling variation in objects of this type, we propose an approximation of the boundary that can be represented as a real valued function, and arrive at further approximation through functional principal component analysis, after suitable adjustment for asymmetry and incompleteness in the data. The analysis of an open access satellite data set on monsoon precipitation over Eastern India leads to explanation of most of the variation in shapes of the regions under rainfall through a handful of interpretable functions that can be further approximated parametrically. The most important aspect of shape is found to be the size followed by contraction/elongation, mostly along two pairs of orthogonal axes. The different modes of variation are remarkably stable across calendar years and across different thresholds for minimum size of the region.

stat.AP↗