SearcharxivSearch

arXiv subjects

Maurilio Gutzeit

Publications and source records attributed to Maurilio Gutzeit.

6 recordsLinked to original sources

Modelling volume-outcome relationships in health care

Despite the ongoing strong interest in associations between quality of care and the volume of health care providers, a unified statistical framework for analyzing them is missing, and many studies suffer from poor statistical modelling choices. We propose a flexible, additive mixed model for studying volume-outcome associations in health care that takes into account individual patient characteristics as well as provider-specific effects through a multi-level approach. More specifically, we treat volume as a continuous variable, and its effect on the considered outcome is modelled as a smooth function. We take account of different case-mixes by including patient-specific risk factors and of clustering on the provider level through random intercepts. This strategy enables us to extract a smooth volume effect as well as volume-independent provider effects. These two quantities can be compared directly in terms of their magnitude, which gives insight into the sources of variability of quality of care. Based on a causal DAG, we derive conditions under which the volume-effect can be interpreted as a causal effect. The paper provides confidence sets for each of the estimated quantities relying on joint estimation of all effects and parameters. Our approach is illustrated through simulation studies and an application to German health care data. Keywords: health care quality measurement, volume-outcome analysis, minimum provider volume, additive regression models, random intercept

stat.ME

Minimax $L_2$-Separation Rate in Testing the Sobolev-Type Regularity of a function

In this paper we study the problem of testing if an $L_2-$function $f$ belonging to a certain $l_2$-Sobolev-ball $B_t(R)$ of radius $R>0$ with smoothness level $t>0$ indeed exhibits a higher smoothness level $s>t$, that is, belongs to $B_s(R)$. We assume that only a perturbed version of $f$ is available, where the noise is governed by a standard Brownian motion scaled by $\frac{1}{\sqrt{n}}$. More precisely, considering a testing problem of the form $$H_0:~f\in B_s(R)~~\mathrm{vs.}~~H_1:~f\in B_t(R),~\inf_{h\in B_s}\Vert f-h\Vert_{L_2}>ρ$$ for some $ρ>0$, we approach the task of identifying the smallest value for $ρ$, denoted $ρ^\ast$, enabling the existence of a test $φ$ with small error probability in a minimax sense. By deriving lower and upper bounds on $ρ^\ast$, we expose its precise dependence on $n$: $$ρ^\ast\sim n^{-\frac{t}{2t+1/2}}.$$ As a remarkable aspect of this composite-composite testing problem, it turns out that the rate does not depend on $s$ and is equal to the rate in signal-detection, i.e. the case of a simple null hypothesis.

math.ST

Two-sample Hypothesis Testing for Inhomogeneous Random Graphs

The study of networks leads to a wide range of high dimensional inference problems. In many practical applications, one needs to draw inference from one or few large sparse networks. The present paper studies hypothesis testing of graphs in this high-dimensional regime, where the goal is to test between two populations of inhomogeneous random graphs defined on the same set of $n$ vertices. The size of each population $m$ is much smaller than $n$, and can even be a constant as small as 1. The critical question in this context is whether the problem is solvable for small $m$. We answer this question from a minimax testing perspective. Let $P,Q$ be the population adjacencies of two sparse inhomogeneous random graph models, and $d$ be a suitably defined distance function. Given a population of $m$ graphs from each model, we derive minimax separation rates for the problem of testing $P=Q$ against $d(P,Q)>ρ$. We observe that if $m$ is small, then the minimax separation is too large for some popular choices of $d$, including total variation distance between corresponding distributions. This implies that some models that are widely separated in $d$ cannot be distinguished for small $m$, and hence, the testing problem is generally not solvable in these cases. We also show that if $m>1$, then the minimax separation is relatively small if $d$ is the Frobenius norm or operator norm distance between $P$ and $Q$. For $m=1$, only the latter distance provides small minimax separation. Thus, for these distances, the problem is solvable for small $m$. We also present near-optimal two-sample tests in both cases, where tests are adaptive with respect to sparsity level of the graphs.

stat.ME

Minimax Euclidean Separation Rates for Testing Convex Hypotheses in $\mathbb{R}^d$

We consider composite-composite testing problems for the expectation in the Gaussian sequence model where the null hypothesis corresponds to a convex subset $\mathcal{C}$ of $\mathbb{R}^d$. We adopt a minimax point of view and our primary objective is to describe the smallest Euclidean distance between the null and alternative hypotheses such that there is a test with small total error probability. In particular, we focus on the dependence of this distance on the dimension $d$ and the sample size/variance parameter $n$ giving rise to the minimax separation rate. In this paper we discuss lower and upper bounds on this rate for different smooth and non- smooth choices for $\mathcal{C}$.

math.ST

Two-Sample Tests for Large Random Graphs Using Network Statistics

We consider a two-sample hypothesis testing problem, where the distributions are defined on the space of undirected graphs, and one has access to only one observation from each model. A motivating example for this problem is comparing the friendship networks on Facebook and LinkedIn. The practical approach to such problems is to compare the networks based on certain network statistics. In this paper, we present a general principle for two-sample hypothesis testing in such scenarios without making any assumption about the network generation process. The main contribution of the paper is a general formulation of the problem based on concentration of network statistics, and consequently, a consistent two-sample test that arises as the natural solution for this problem. We also show that the proposed test is minimax optimal for certain network statistics.

stat.ME

An optimal algorithm for the Thresholding Bandit Problem

We study a specific \textit{combinatorial pure exploration stochastic bandit problem} where the learner aims at finding the set of arms whose means are above a given threshold, up to a given precision, and \textit{for a fixed time horizon}. We propose a parameter-free algorithm based on an original heuristic, and prove that it is optimal for this problem by deriving matching upper and lower bounds. To the best of our knowledge, this is the first non-trivial pure exploration setting with \textit{fixed budget} for which optimal strategies are constructed.

stat.ML