SearcharxivSearch

arXiv subjects

Angelika Rohde

Publications and source records attributed to Angelika Rohde.

At least 19 recordsLinked to original sources

Concentration of additive functionals of Stratonovich-type

Additive functionals $\overline{J}_t=\frac{1}{t}\int_0^tU(X_s)\circ dX_s$ of Stratonovich-type recently attracted much attention in the context of inference of thermodynamic properties of complex systems from observations $U$ of individual fluctuating paths $(X_s)_{0\le s\le t}$, whereby $X_0$ is initiated from some general measure. Concentration results on $\overline{J}_t$, albeit desirable, are virtually nonexistent. They turn out to be significantly more challenging to prove than for classical Lebesgue-type functionals $\overline{\rho}_t=\frac{1}{t}\int_0^t V(X_s)ds$ because the tilt deforms the full second-order structure of the Feynman-Kac generator instead of contributing an additive potential. This renders the generator generally non-self-adjoint even under detailed balance. We overcome this by working with a symmetrized Dirichlet form with a new effective potential that now couples the observable to the non-equilibrium character of the dynamics. We prove concentration inequalities for $\overline{J}_t$ for any bounded, sufficiently smooth vector-valued function $U$ of a general geometrically ergodic diffusion process $X_s$, including explicit sub-gamma and Bernstein-type inequalities, and we obtain explicit upper bounds on ${\rm Var}(\overline{J}_t)$. Strikingly, under detailed balance the concentration of $\overline{J}_t$ is distinctively sub-Gaussian at all times and all deviations, with a variance proxy fixed by the noise alone and independent of the spectral gap, which has no analog for $\overline{\rho}_t$.

math.PR

Distinguishing case-mix from context heterogeneity in prognostic regression model synthesis settings

Prognostic regression models often synthesize data from multiple sites, whether within a multi-site study, across federated settings, or in individual participant data meta-analysis. Here, a site is any data source, such as a hospital, registry, trial, or study, and need not be a physical center. Analysts must then decide whether one regression model represents all sites or whether site-specific models are needed. Established measures such as coefficient-level tau^2 quantify heterogeneity but do not distinguish its source. We focus on diagnosing whether coefficient heterogeneity reflects case-mix or site-specific context effects. Case-mix heterogeneity can arise when linear regression terms approximate multivariable non-linear relationships in populations with different covariate distributions. Contextual heterogeneity arises when comparable patients require different regression relationships across sites. We do this by fitting site-specific local regressions in a dimension-reduced space and partitioning the smoothed coefficient surfaces into a cross-site reference and site-specific deviations. An autoencoder and custom loss structure the latent space around local prognostic relationships. We then project this partition onto the outcome scale to derive observation- and site-level summaries. We demonstrate the approach on a COPD trial with two sites. In the three leading latent slope coordinates, coefficient-surface variation was predominantly contextual. The derived observation-level outcome-scale variance partition was case-mix-leading, whereas its between-site aggregation was concentrated in contextual differences rather than case-mix shifts. A permuted-site negative control assesses whether the contextual summary can arise when site labels carry no signal. This diagnostic distinction can inform whether joint or site-specific regression models should be evaluated.

stat.ME

Before and beyond the mixing time: New approximations for additive functionals of stationary Gauss-Markov processes

Whereas classical invariance principles for ergodic Markov chains address the situation in which the time horizon of observations is much larger than the mixing time, the quality of approximation is questionable when this is not the case anymore -- even when starting the Markov chain in the invariant law. In this article, we prove quantitative and functional limit theorems for additive functionals along triangular arrays of stationary Gaussian Markov processes when the mixing time $t_{\text{mix}}$ scales sub-, super- and proportionately to the number of observations $n$. Our major finding is a phase-transition at $t_{\text{mix}}\asymp n$, together with the identification and interrelation properties of the emerging new limit processes at and before the mixing time.

math.PR

Self-organized regime switching in null-recurrent dynamics

Based on discrete observations $X_0,X_{\Delta},\dots, X_{n\Delta}$ for $\Delta=n^{-\gamma}$ with $\gamma\in [0,1)$ of the null-recurrent dynamic $dX_t = \sigma(X_t)dW_t$ with a Brownian motion $W$ and $\sigma(x)=\alpha\mathbb{1}\{x<\rho\} + \beta\mathbb{1}\{x\geq \rho\}$, we derive rate of convergence and limiting distribution of the profile MLE for $\rho$. This includes low-frequency asymptotics ($\gamma=0$) for which the observations form a null-recurrent Markov chain. The derived non-standard limit is the argsup over a doubly stochastic drifted Poisson process explicitly involving the local time of oscillating Brownian motion. Its dependence on $\rho$ as well as the unknown volatility levels $\alpha$ and $\beta$ is shown to be continuous w.r.t. the topology of weak convergence, enabling statistical inference. Whereas this limit is independent of the sampling frequency, the profile MLE's rate of convergence equals $n^{-(1+\gamma)/2}$ and is proven to be minimax optimal. The surprising idea of the proof of the limit theorem is to relate the long-term behavior of the null-recurrent Markov chain to the infill asymptotics on a fixed time interval. Indeed, in the very special case that $(X_t)_{t\geq 0}$ is started in the true parameter $X_0=\rho_0$, the process $(X_t-\rho_0)_{t\geq 0}$ is shown to possess a desirable distributional self-similarity. On basis of the strong Markov property, the artificial constallation of starting in $\rho_0$ is finally overcome by a coupling argument.

math.ST

Asymptotic equivalence for nonparametric additive regression

We prove asymptotic equivalence of nonparametric additive regression and an appropriate Gaussian white noise experiment in which a multidimensional shifted Wiener process is observed, whose dimension equals the number of additive components. The shift depends on the additive components of the regression function and solely the one- and two-dimensional marginal distributions of the covariates via an explicitly specified bounded but non-compact linear operator~$\Gamma$. The number of additive components $d$ is allowed to increase moderately with respect to the sample size. In the special case of pairwise independent components of the covariates, the white noise model decomposes into $d$ independent univariate processes. Moreover, we study approximation in some semiparametric setting where $\Gamma$ splits into a multiplication operator and an asymptotically negligible Hilbert-Schmidt operator.

math.ST

Contrasting Global and Patient-Specific Regression Models via a Neural Network Representation

When developing clinical prediction models, it can be challenging to balance between global models that are valid for all patients and personalized models tailored to individuals or potentially unknown subgroups. To aid such decisions, we propose a diagnostic tool for contrasting global regression models and patient-specific (local) regression models. The core utility of this tool is to identify where and for whom a global model may be inadequate. We focus on regression models and specifically suggest a localized regression approach that identifies regions in the predictor space where patients are not well represented by the global model. As localization becomes challenging when dealing with many predictors, we propose modeling in a dimension-reduced latent representation obtained from an autoencoder. Using such a neural network architecture for dimension reduction enables learning a latent representation simultaneously optimized for both good data reconstruction and for revealing local outcome-related associations suitable for robust localized regression. We illustrate the proposed approach with a clinical study involving patients with chronic obstructive pulmonary disease. Our findings indicate that the global model is adequate for most patients but that indeed specific subgroups benefit from personalized models. We also demonstrate how to map these subgroup models back to the original predictors, providing insight into why the global model falls short for these groups. Thus, the principal application and diagnostic yield of our tool is the identification and characterization of patients or subgroups whose outcome associations deviate from the global model.

stat.ME

Small Data Explainer -- The impact of small data methods in everyday life

The emergence of breakthrough artificial intelligence (AI) techniques has led to a renewed focus on how small data settings, i.e., settings with limited information, can benefit from such developments. This includes societal issues such as how best to include under-represented groups in data-driven policy and decision making, or the health benefits of assistive technologies. We provide a conceptual overview, clarify the relationship between small data and big data, and identify common themes from exemplary case studies and application areas. Potential solutions are described in a more detailed technical overview of current data analysis and modelling techniques, highlighting contributions from different disciplines, such as knowledge-driven modelling from statistics and data-driven modelling from computer science. By linking application settings, conceptual contributions and specific techniques, we highlight what is already feasible and suggest what an agenda for fully leveraging small data might look like.

cs.CY

The weak-feature-impact phase transition of the NPMLE in monotone binary regression

Statistical literature provides pointwise limiting distributions of the nonparametric maximum likelihood estimator (NPMLE) in monotone binary regression for the two extremal cases: If the feature-label relation is strictly monotone and sufficiently smooth, it converges at a nonparametric rate with scaled Chernoff-type limiting distribution, and it converges at the parametric $\sqrt{n}$-rate if the underlying relation is flat. In this article, we provide the complete picture of the distributional metamorphosis of the NPMLE, revealing a new limiting distribution which provides a significantly better distributional approximation for small samples in case of a weak feature-label relationship. It is shown to continuously interpolate between the two extremal cases. The innovative way to determine this distribution is to generate it as a limit of the NPMLE in the newly introduced weak-feature-impact triangular array for a particular parameter-sample-size constellation. Moreover, a phase transition is likewise observed for the suitably rescaled $L^{1}$-error in this weak-feature-impact scenario. As a by-product, its limiting distribution for flat regression functions is obtained, which was unknown before. The proof develops a completely new strategy, notably not based on the switch relation. A novel type of local minimax lower bounds accompanies these results.

math.ST

The level of self-organized criticality in oscillating Brownian motion: $n$-consistency and stable Poisson-type convergence of the MLE

For some discretely observed path of oscillating Brownian motion with level of self-organized criticality $\rho_0$, we prove in the infill asymptotics that the MLE is $n$-consistent, where $n$ denotes the sample size, and derive its limit distribution with respect to stable convergence. As the transition density of this homogeneous Markov process is not even continuous in $\rho_0$, the analysis is highly non-standard. Therefore, interesting and somewhat unexpected phenomena occur: The likelihood function splits into several components, each of them contributing very differently depending on how close the argument $\rho$ is to $\rho_0$. Correspondingly, the MLE is successively excluded to lay outside a compact set, a $1/\sqrt{n}$-neighborhood and finally a $1/n$-neighborhood of $\rho_0$ asymptotically. The crucial argument to derive the stable convergence is to exploit the semimartingale structure of the sequential suitably rescaled local log-likelihood function (as a process in time). Both sequentially and as a process in $\rho$, it exhibits a bivariate Poissonian behavior in the stable limit with its intensity being a multiple of the local time at $\rho_0$.

math.ST

Asymptotic Equivalence of Locally Stationary Processes and Bivariate White Noise

We consider a general class of statistical experiments, in which an $n$-dimensional centered Gaussian random variable is observed and its covariance matrix is the parameter of interest. The covariance matrix is assumed to be well-approximable in a linear space of lower dimension $K_n$ with eigenvalues uniformly bounded away from zero and infinity. We prove asymptotic equivalence of this experiment and a class of $K_n$-dimensional Gaussian models with informative expectation in Le Cam's sense when $n$ tends to infinity and $K_n$ is allowed to increase moderately in $n$ at a polynomial rate. For this purpose we derive a new localization technique for non-i.i.d. data and a novel high-dimensional Central Limit Law in total variation distance. These results are key ingredients to show asymptotic equivalence between the experiments of locally stationary Gaussian time series and a bivariate Wiener process with the log spectral density as its drift. Therein a novel class of matrices is introduced which generalizes circulant Toeplitz matrices traditionally used for strictly stationary time series.

math.ST

Computationally tractable nonparametric bootstrap of high-dimensional sample covariance matrices

We introduce a new ``$(m,mp/n)$ out of $(n,p)$'' sampling-with-replace\-ment bootstrap for eigenvalue statistics of high-dimensional sample covariance matrices based on $n$ independent $p$-dimensional random vectors. As it only uses $q=\lfloor mp/n\rfloor $ coordinates of the observations in a subsample of size $m \ll n $ from the original data, it is computationally tractable for large scale data. In the high-dimensional scenario $p/n\rightarrow c\in (0,\infty)$, this fully nonparametric bootstrap is shown to consistently reproduce the empirical spectral measure if $m/n\rightarrow 0$. If $m^2/n\rightarrow 0$, it approximates correctly the distribution of linear spectral statistics. The crucial component is a suitably defined Representative Subpopulation Condition which is shown to be verified in a large variety of situations. Our proofs are conducted under minimal moment requirements and incorporate delicate results on non-centered quadratic forms, combinatorial trace moments estimates as well as a conditional bootstrap martingale CLT which may be of independent interest.

math.ST

Improving prediction models by incorporating external data with weights based on similarity

In clinical settings, we often face the challenge of building prediction models based on small observational data sets. For example, such a data set might be from a medical center in a multi-center study. Differences between centers might be large, thus requiring specific models based on the data set from the target center. Still, we want to borrow information from the external centers, to deal with small sample sizes. There are approaches that either assign weights to each external data set or each external observation. To incorporate information on differences between data sets and observations, we propose an approach that combines both into weights that can be incorporated into a likelihood for fitting regression models. Specifically, we suggest weights at the data set level that incorporate information on how well the models that provide the observation weights distinguish between data sets. Technically, this takes the form of inverse probability weighting. We explore different scenarios where covariates and outcomes differ among data sets, informing our simulation design for method evaluation. The concept of effective sample size is used for understanding the effectiveness of our subgroup modeling approach. We demonstrate our approach through a clinical application, predicting applied radiotherapy doses for cancer patients. Generally, the proposed approach provides improved prediction performance when external data sets are similar. We thus provide a method for quantifying similarity of external data sets to the target data set and use this similarity to include external observations for improving performance in a target data set prediction modeling task with small data.

stat.ME

Sharp adaptive and pathwise stable similarity testing for scalar ergodic diffusions

Within the nonparametric diffusion model, we develop a multiple test to infer about similarity of an unknown drift $b$ to some reference drift $b_0$: At prescribed significance, we simultaneously identify those regions where violation from similiarity occurs, without a priori knowledge of their number, size and location. This test is shown to be minimax-optimal and adaptive. At the same time, the procedure is robust under small deviation from Brownian motion as the driving noise process. A detailed investigation for fractional driving noise, which is neither a semimartingale nor a Markov process, is provided for Hurst indices close to the Brownian motion case.

math.ST

Non-uniform bounds and Edgeworth expansions in self-normalized limit theorems

We study Edgeworth expansions in limit theorems for self-normalized sums. Non-uniform bounds for expansions in the central limit theorem are established while only imposing minimal moment conditions. Within this result, we address the case of non-integer moments leading to a reduced remainder. Furthermore, we provide non-uniform bounds for expansions in local limit theorems. The enhanced tail-accuracy of our non-uniform bounds allows for deriving an Edgeworth-type expansion in the entropic central limit theorem as well as a central limit theorem in total variation distance for self-normalized sums.

math.PR

A central limit theorem concerning uncertainty in estimates of individual admixture

The concept of individual admixture (IA) assumes that the genome of individuals is composed of alleles inherited from $K$ ancestral populations. Each copy of each allele has the same chance $q_k$ to originate from population $k$, and together with the allele frequencies $p$ in all populations at all $M$ markers, comprises the admixture model. Here, we assume a supervised scheme, i.e.\ allele frequencies $p$ are given through a reference database of size $N$, and $q$ is estimated via maximum likelihood for a single sample. We study laws of large numbers and central limit theorems describing effects of finiteness of both, $M$ and $N$, on the estimate of $q$. We recall results for the effect of finite $M$, and provide a central limit theorem for the effect of finite $N$, introduce a new way to express the uncertainty in estimates in standard barplots, give simulation results, and discuss applications in forensic genetics.

q-bio.PE

Interactive versus non-interactive locally differentially private estimation: Two elbows for the quadratic functional

Local differential privacy has recently received increasing attention from the statistics community as a valuable tool to protect the privacy of individual data owners without the need of a trusted third party. Similar to the classical notion of randomized response, the idea is that data owners randomize their true information locally and only release the perturbed data. Many different protocols for such local perturbation procedures can be designed. In most estimation problems studied in the literature so far, however, no significant difference in terms of minimax risk between purely non-interactive protocols and protocols that allow for some amount of interaction between individual data providers could be observed. In this paper we show that for estimating the integrated square of a density, sequentially interactive procedures improve substantially over the best possible non-interactive procedure in terms of minimax rate of estimation. In particular, in the non-interactive scenario we identify an elbow in the minimax rate at $s=\frac34$, whereas in the sequentially interactive scenario the elbow is at $s=\frac12$. This is markedly different from both, the case of direct observations, where the elbow is well known to be at $s=\frac14$, as well as from the case where Laplace noise is added to the original data, where an elbow at $s= \frac94$ is obtained. We also provide adaptive estimators that achieve the optimal rate up to log-factors, we draw connections to non-parametric goodness-of-fit testing and estimation of more general integral functionals and conduct a series of numerical experiments. The fact that a particular locally differentially private, but interactive, mechanism improves over the simple non-interactive one is also of great importance for practical implementations of local differential privacy.

math.ST

Geometrizing rates of convergence under local differential privacy constraints

We study the problem of estimating a functional $θ(\mathbb P)$ of an unknown probability distribution $\mathbb P \in\mathcal P$ in which the original iid sample $X_1,\dots, X_n$ is kept private even from the statistician via an $α$-local differential privacy constraint. Let $ω_{TV}$ denote the modulus of continuity of the functional $θ$ over $\mathcal P$, with respect to total variation distance. For a large class of loss functions $l$ and a fixed privacy level $α$, we prove that the privatized minimax risk is equivalent to $l(ω_{TV}(n^{-1/2}))$ to within constants, under regularity conditions that are satisfied, in particular, if $θ$ is linear and $\mathcal P$ is convex. Our results complement the theory developed by Donoho and Liu (1991) with the nowadays highly relevant case of privatized data. Somewhat surprisingly, the difficulty of the estimation problem in the private case is characterized by $ω_{TV}$, whereas, it is characterized by the Hellinger modulus of continuity if the original data $X_1,\dots, X_n$ are available. We also find that for locally private estimation of linear functionals over a convex model a simple sample mean estimator, based on independently and binary privatized observations, always achieves the minimax rate. We further provide a general recipe for choosing the functional parameter in the optimal binary privatization mechanisms and illustrate the general theory in numerous examples. Our theory allows to quantify the price to be paid for local differential privacy in a large class of estimation problems. This price appears to be highly problem specific.

math.ST

Locally Adaptive Confidence Bands

We develop honest and locally adaptive confidence bands for probability densities. They provide substantially improved confidence statements in case of inhomogeneous smoothness, and are easily implemented and visualized. The article contributes conceptual work on locally adaptive inference as a straightforward modification of the global setting imposes severe obstacles for statistical purposes. Among others, we introduce a statistical notion of local Hölder regularity and prove a correspondingly strong version of local adaptivity. We substantially relax the straightforward localization of the self-similarity condition in order not to rule out prototypical densities. The set of densities permanently excluded from the consideration is shown to be pathological in a mathematically rigorous sense. On a technical level, the crucial component for the verification of honesty is the identification of an asymptotically least favorable stationary case by means of Slepian's comparison inequality.

math.ST