SearcharxivSearch

arXiv subjects

Hanna Jankowski

Publications and source records attributed to Hanna Jankowski.

7 recordsLinked to original sources

Bandwidth-free nonparametric density estimation for grouped data

In some situations, data is collected under systematical and technical constraints due to uncertainty in experimental reports, intermittent measurements, confidentiality, and non-detects. For this reason, it might not be possible to retrieve or receive the data in a conventional format but rather in a grouped form where only the number of occurrences is known within intervals. The challenge is to estimate the density of the underlying ungrouped data based on the observed grouped data with no information regarding the underlying distribution. To overcome this problem, this study introduces a mean-adjusted log-concave (MALC) density estimation method for univariate grouped data, aiming to provide a bandwidth-free non-parametric approach that does not rely on specific distributional assumptions. The performance of the MALC method is evaluated through simulations across various distributions with different sample sizes and grid widths. The results demonstrate the robustness and effectiveness of the MALC approach in grouped data analysis, offering a broader range of applications over traditional methods.

stat.ME

Joint estimation of the basic reproduction number and serial interval using Sequential Bayes

Early in an infectious disease outbreak, timely and accurate estimation of the basic reproduction number ($R_0$) and the serial interval (SI) is critical for understanding transmission dynamics and informing public health responses. While many methods estimate these quantities separately, and a small number jointly estimate them from incidence data, existing joint approaches are largely likelihood-based and do not fully exploit prior information. We propose a novel Bayesian framework for the joint estimation of $R_0$ and the serial interval using only case count data, implemented through a sequential Bayes approach. Our method assumes an SIR model and employs a mildly informative joint prior constructed by linking log-Gamma marginal distributions for $R_0$ and the SI via a Gaussian copula, explicitly accounting for their dependence. The prior is updated sequentially as new incidence data become available, allowing for real-time inference. We assess the performance of the proposed estimator through extensive simulation studies under correct model specification as well as under model misspecification, including when the true data come from an SEIR or SEAIR model, and under varying degrees of prior misspecification. Comparisons with the widely used White and Pagano likelihood-based joint estimator show that our approach yields substantially more precise and stable estimates of $R_0$, with comparable or improved bias, particularly in the early stages of an outbreak. Estimation of the SI is more sensitive to prior misspecification; however, when prior information is reasonably accurate, our method provides reliable SI estimates and remains more stable than the competing approach. We illustrate the practical utility of the proposed method using Canadian COVID-19 incidence data at both national and provincial levels.

stat.ME

Multiple Comparisons using Composite Likelihood in Clustered Data

We study the problem of multiple hypothesis testing for multidimensional data when inter-correlations are present. The problem of multiple comparisons is common in many applications. When the data is multivariate and correlated, existing multiple comparisons procedures based on maximum likelihood estimation could be prohibitively computationally intensive. We propose to construct multiple comparisons procedures based on composite likelihood statistics. We focus on data arising in three ubiquitous cases: multivariate Gaussian, probit, and quadratic exponential models. To help practitioners assess the quality of our proposed methods, we assess their empirical performance via Monte Carlo simulations. It is shown that composite likelihood based procedures maintain good control of the familywise type I error rate in the presence of intra-cluster correlation, whereas ignoring the correlation leads to erratic performance. Using data arising from a diabetic nephropathy study, we show how our composite likelihood approach makes an otherwise intractable analysis possible.

math.ST

Convergence of linear functionals of the Grenander estimator under misspecification

Under the assumption that the true density is decreasing, it is well known that the Grenander estimator converges at rate $n^{1/3}$ if the true density is curved [Sankhyā Ser. A 31 (1969) 23-36] and at rate $n^{1/2}$ if the density is flat [Ann. Probab. 11 (1983) 328-345; Canad. J. Statist. 27 (1999) 557-566]. In the case that the true density is misspecified, the results of Patilea [Ann. Statist. 29 (2001) 94-123] tell us that the global convergence rate is of order $n^{1/3}$ in Hellinger distance. Here, we show that the local convergence rate is $n^{1/2}$ at a point where the density is misspecified. This is not in contradiction with the results of Patilea [Ann. Statist. 29 (2001) 94-123]: the global convergence rate simply comes from locally curved well-specified regions. Furthermore, we study global convergence under misspecification by considering linear functionals. The rate of convergence is $n^{1/2}$ and we show that the limit is made up of two independent terms: a mean-zero Gaussian term and a second term (with nonzero mean) which is present only if the density has well-specified locally flat regions.

math.ST

Asymptotics of the discrete log-concave maximum likelihood estimator and related applications

The assumption of log-concavity is a flexible and appealing nonparametric shape constraint in distribution modelling. In this work, we study the log-concave maximum likelihood estimator (MLE) of a probability mass function (pmf). We show that the MLE is strongly consistent and derive its pointwise asymptotic theory under both the well- and misspecified setting. Our asymptotic results are used to calculate confidence intervals for the true log-concave pmf. Both the MLE and the associated confidence intervals may be easily computed using the R package logcondiscr. We illustrate our theoretical results using recent data from the H1N1 pandemic in Ontario, Canada.

stat.ME

Nonequilibrium density fluctuations for the zero range process with colour

We examine the fluctuations of the empirical density measure for the colour version of the symmetric nearest neighbour zero range particle systems in dimension one. We show that the weak limit of these fluctuations is the solution of a system of coupled generalized Ornstein-Uhlenbeck processes. We also discuss how this result may be used to prove a central limit theorem for the tagged particle on the level of finite dimensional distributions, and identify the limiting variance. This is the central limit theorem associated to propagation of chaos for this interacting particle system.

math.PR

Logarithmic Sobolev inequality for the inhomogeneous zero range process

We prove that the logarithmic Sobolev constant for the inhomogeneous symmetric nearest neighbour zero range process on a cube of size N^d grows as N^2. We apply this result to the inhomogeneous process which arises in the study of the homogeneous version of the zero range interacting particle system with colours.

math.PR