SearcharxivSearch

arXiv subjects

Subhra Sankar Dhar

Publications and source records attributed to Subhra Sankar Dhar.

At least 19 recordsLinked to original sources

A novel characterization of structures in smooth regression curves: from a viewpoint of persistent homology

We characterize structures such as monotonicity, convexity, and modality in smooth regression curves using persistent homology. Persistent homology is a key tool in topological data analysis that detects higher-dimensional topological features such as connected components and holes (cycles or loops) in the data. In other words, persistent homology is a multiscale version of homology that characterizes sets based on the connected components and holes. We use super-level sets of functions to extract geometric features via persistent homology. In particular, we explore structures in regression curves via the persistent homology of super-level sets of a function, where the function of interest is - the first derivative of the regression function. In the course of this study, we extend an existing procedure of estimating the persistent homology for the first derivative of a regression function and establish its consistency. Moreover, as an application of the proposed methodology, we demonstrate that the persistent homology of the derivative of a function can reveal hidden structures in the function that are not visible from the persistent homology of the function itself. In particular, we characterize structures such as monotonicity, convexity, and modality, and propose a measure of statistical significance to infer these structures in practice. Finally, we conduct an empirical study to implement the proposed methodology on simulated and real data sets and compare the derived results with an existing methodology.

math.AT

Modeling of Pneumococcal and Respiratory Syncytial Virus Pneumonia: An Epidemiological Review, with Statistical Inference

Infectious diseases continue to pose significant public health challenges worldwide, requiring effective prevention and control strategies to mitigate their negative impact. Infectious diseases can be broadly classified into two groups: vaccine-preventable diseases (e.g., measles, polio, influenza, hepatitis B, pneumonia) and vaccine-non-preventable diseases (e.g., HIV/AIDS). Vaccine-preventable disease models are one of the essential tools for understanding infectious disease dynamics, evaluating intervention strategies, and guiding public health policies. In this review article, we explore the recent advancements in modeling two particular vaccine-preventable infectious diseases. Here, we consider both deterministic and stochastic models to comprehensively capture the complexity of disease transmission, vaccine efficacy, and population-level immunity. We highlight the application of these models to the infectious diseases, namely, bacterial and viral pneumonia caused by the bacteria Streptococcus pneumoniae (S. pneumoniae) and the respiratory syncytial virus (RSV). Pneumonia carry a substantial global burden, where modeling has played a crucial role in assessing vaccine impacts and optimizing immunization strategies to minimize the disease burden. By synthesizing recent methodologies and findings, this review provides valuable insights for future research and policy decisions aimed at improving vaccine-preventable disease control for pneumonia caused by S. pneumoniae and RSV.

q-bio.PE

Identifying Topological Differences in Two Populations of Random Geometric Objects

We propose a statistical framework to identify topological differences in two populations of random geometric objects. The proposed framework involves first associating a topological signature with random geometric objects and then performing a two-sample test using the observed topological signatures. We associate persistence barcodes, a topological signature from topological data analysis, with each observed random geometric object. This, in turn, yields a two-sample problem on the space of persistence barcodes. As the space of persistence barcodes is not suitable for standard statistical analysis, we translate the two-sample problem on a suitable subset of a Euclidean space. In the course of this study, we embed the topological signatures in an ordered convex cone in a Euclidean space using functions from tropical geometry. We show that the embedding is a sufficient statistic for the persistence barcodes. This fact leads to the proposal of a two-sample test based on this sufficient statistic, and its equivalence to the two-sample problem on the barcode space is established. Finally, the consistency of the proposed test is studied.

stat.ME

Least Square Estimation: SDEs Perturbed by Lévy Noise with Sparse Sample Paths

This article investigates the least squares estimators (LSE) for the unknown parameters in stochastic differential equations (SDEs) that are affected by Lévy noise, particularly when the sample paths are sparse. Specifically, given $n$ sparsely observed curves related to this model, we derive the least squares estimators for the unknown parameters: the drift coefficient, the diffusion coefficient, and the jump-diffusion coefficient. We also establish the asymptotic rate of convergence for the proposed LSE estimators. Additionally, in the supplementary materials, the proposed methodology is applied to a benchmark dataset of functional data/curves, and a small simulation study is conducted to illustrate the findings.

stat.ME

A Robust Persistent Homology : Trimming Approach

This article studies the robust version of persistent homology based on trimming methodology to capture the geometric feature through support of the data in presence of outliers. Precisely speaking, the proposed methodology works when the outliers lie outside the main data cloud as well as inside the data cloud. In the course of theoretical study, it is established that the Bottleneck distance between the proposed robust version of persistent homology and its population analogue can be made arbitrary small with a certain rate for a sufficiently large sample size. The practicability of the methodology is shown for various simulated data and bench mark real data associated with cellular biology.

stat.ME

Identifying arbitrary transformation between the slopes in scalar-on-function regression

In this article, we study whether the slope functions of two scalar-on-function regression models in two samples are associated with any arbitrary transformation along the vertical axis. The problem is formally stated as a statistical hypothesis test, and corresponding test statistic is formed based on the estimated second derivative of the unknown transformation. The asymptotic properties of the test statistic are investigated using some advanced techniques related to the empirical process. Moreover, to implement the test for small sample size data, a bootstrap algorithm is proposed, and it is shown that the bootstrap version of the test is as good as the original test for sufficiently large sample size. Furthermore, the utility of the proposed methodology is shown for simulated datasets, and DTI data is analyzed using the proposed methodology.

stat.ME

Statistical estimation of $π$: varying choices over dimensions

This article studies statistical estimation of $π$ based on the fact that the ratio of the volumes of a $d$-dimensional hypersphere and a $d$-dimensional hypercube is a certain function of $π$, and the function depends on the dimension $d$. The estimation of $π$ is carried out for various choices of $d$ (strictly speaking, $d\in\{1, 2, \ldots, 20\}$) using the idea of Monte Carlo simulations. Various intriguing facts are observed, and the estimation of $π$ using infinite dimensional observations is outlined. Moreover, the R codes associated with relevant numerical studies are provided.

stat.OT

L-Estimation Approach to Tobit Models with Endogeneity and Weakly Dependent Errors

This article introduces an L-estimator for the semiparametric Tobit model with endogenous regressors. The estimation procedure follows a two-stage approach: the first stage employs least squares, while the second stage utilizes the L-estimation technique. We establish the large-sample properties of the proposed estimators under weakly dependent data. The utility of the proposed methodology is demonstrated for various simulated data and a benchmark real data set.

stat.ME

Generalized M-Estimation in Censored Regression Model under Endogeneity

We propose and study M-estimation to estimate the parameters in the censored regression model in the presence of endogeneity, i.e., the Tobit model. In the course of this study, we follow two-stage procedures: the first stage consists of applying control function procedures to address the issue of endogeneity using instrumental variables, and the second stage applies the M-estimation technique to estimate the unknown parameters involved in the model. The large sample properties of the proposed estimators are derived and analyzed. The finite sample properties of the estimators are studied through Monte Carlo simulation and a real data application related to women's labor force participation.

stat.ME

Testing Independence of Infinite Dimensional Random Elements: A Sup-norm Approach

In this article, we study the test for independence of two random elements $X$ and $Y$ lying in an infinite dimensional space ${\cal{H}}$ (specifically, a real separable Hilbert space equipped with the inner product $\langle ., .\rangle_{\cal{H}}$). In the course of this study, a measure of association is proposed based on the sup-norm difference between the joint probability density function of the bivariate random vector $(\langle l_{1}, X \rangle_{\cal{H}}, \langle l_{2}, Y \rangle_{\cal{H}})$ and the product of marginal probability density functions of the random variables $\langle l_{1}, X \rangle_{\cal{H}}$ and $\langle l_{2}, Y \rangle_{\cal{H}}$, where $l_{1}\in{\cal{H}}$ and $l_{2}\in{\cal{H}}$ are two arbitrary elements. It is established that the proposed measure of association equals zero if and only if the random elements are independent. In order to carry out the test whether $X$ and $Y$ are independent or not, the sample version of the proposed measure of association is considered as the test statistic after appropriate normalization, and the asymptotic distributions of the test statistic under the null and the local alternatives are derived. The performance of the new test is investigated for simulated data sets and the practicability of the test is shown for three real data sets related to climatology, biological science and chemical science.

math.ST

Estimation of time-varying recovery and death rates from epidemiological data: A new approach

The time-to-recovery or time-to-death for various infectious diseases can vary significantly among individuals, influenced by several factors such as demographic differences, immune strength, medical history, age, pre-existing conditions, and infection severity. To capture these variations, time-since-infection dependent recovery and death rates offer a detailed description of the epidemic. However, obtaining individual-level data to estimate these rates is challenging, while aggregate epidemiological data (such as the number of new infections, number of active cases, number of new recoveries, and number of new deaths) are more readily available. In this article, a new methodology is proposed to estimate time-since-infection dependent recovery and death rates using easily available data sources, accommodating irregular data collection timings reflective of real-world reporting practices. The Nadaraya-Watson estimator is utilized to derive the number of new infections. This model improves the accuracy of epidemic progression descriptions and provides clear insights into recovery and death distributions. The proposed methodology is validated using COVID-19 data and its general applicability is demonstrated by applying it to some other diseases like measles and typhoid.

stat.AP

Inspecting discrepancy between multivariate distributions using half-space depth based information criteria

This article inspects whether a multivariate distribution is different from a specified distribution or not, and it also tests the equality of two multivariate distributions. In the course of this study, a graphical tool-kit using well-known half-space depth based information criteria is proposed, which is a two-dimensional plot, regardless of the dimension of the data, and it is even useful in comparing high-dimensional distributions. The simple interpretability of the proposed graphical tool-kit motivates us to formulate test statistics to carry out the corresponding testing of hypothesis problems. It is established that the proposed tests based on the same information criteria are consistent, and moreover, the asymptotic distributions of the test statistics under contiguous/local alternatives are derived, which enable us to compute the asymptotic power of these tests. Furthermore, it is observed that the computations associated with the proposed tests are unburdensome. Besides, these tests perform better than many other tests available in the literature when data are generated from various distributions such as heavy tailed distributions, which indicates that the proposed methodology is robust as well. Finally, the usefulness of the proposed graphical tool-kit and tests is shown on two benchmark real data sets.

stat.ME

Operator on Operator Regression in Quantum Probability

This article introduces operator on operator regression in quantum probability. Here in the regression model, the response and the independent variables are certain operator valued observables, and they are linearly associated with unknown scalar coefficient (denoted by $β$), and the error is a random operator. In the course of this study, we propose a quantum version of a class of estimators (denoted by $M$ estimator) of $β$, and the large sample behaviour of those quantum version of the estimators are derived, given the fact that the true model is also linear and the samples are observed eigenvalue pairs of the operator valued observables.

stat.ME

Testing Homological Equivalence Using Betti Numbers

In this article, we propose a one-sample test to check whether the support of the unknown distribution generating the data is homologically equivalent to the support of some specified distribution or not OR using the corresponding two-sample test, one can test whether the supports of two unknown distributions are homologically equivalent or not. In the course of this study, test statistics based on the Betti numbers are formulated, and the consistency of the tests is established under the critical and the supercritical regimes. Moreover, some simulation studies are conducted and results are compared with the existing methodologies such as Robinson's permutation test and test based on mean persistent landscape functions. Furthermore, the practicability of the tests is shown on two well-known real data sets also.

stat.ME

Co-variance Operator of Banach Valued Random Elements: U-Statistic Approach

This article proposes a co-variance operator for Banach valued random elements using the concept of $U$-statistic. We then study the asymptotic distribution of the proposed co-variance operator along with related large sample properties. Moreover, specifically for Hilbert space valued random elements, the asymptotic distribution of the proposed estimator is derived even for dependent data under some mixing conditions.

math.ST

Shift identification in time varying regression quantiles

This article investigates whether time-varying quantile regression curves are the same up to the horizontal shift or not. The errors and the covariates involved in the regression model are allowed to be locally stationary. We formalize this issue in a corresponding non-parametric hypothesis testing problem, and develop an integrated-squared-norm based test (SIT) as well as a simultaneous confidence band (SCB) approach. The asymptotic properties of SIT and SCB under null and local alternatives are derived. Moreover, the asymptotic properties of these tests are also studied when the compared data sets are dependent. We then propose valid wild bootstrap algorithms to implement SIT and SCB. Furthermore, the usefulness of the proposed methodology is illustrated via analysing simulated and real data related to COVID-19 outbreak and climate science.

stat.ME

A New Graphical Device and Related Tests for the Shape of Non-parametric Regression Function

We consider a non-parametric regression model $y = m(x) + ε$ and propose a novel graphical device to check whether the $r$-th ($r \geqslant 1$) derivative of the regression function $m(x)$ is positive or otherwise. Since the shape of the regression function can be completely characterized by its derivatives, the graphical device can correctly identify the shape of the regression function. The proposed device includes the check for monotonicity and convexity of the function as special cases. We also present an example to elucidate the practical utility of the graphical device. In addition, we employ the graphical device to formulate a class of test statistics and derive its asymptotic distribution. The tests are exhibited in various simulated and real data examples.

stat.ME

On Variable Screening in Multiple Nonparametric Regression Model

In this article, we study the problem of variable screening in multiple nonparametric regression model. The proposed methodology is based on the fact that the partial derivative of the regression function with respect to the irrelevant variable should be negligible. The Statistical property of the proposed methodology is investigated under both cases : (i) when the variance of the error term is known, and (ii) when the variance of the error term is unknown. Moreover, we establish the practicality of our proposed methodology for various simulated and real data related to interdisciplinary sciences such as Economics, Finance and other sciences.

stat.ME