SearcharxivSearch

arXiv subjects

Micha Mandel

Publications and source records attributed to Micha Mandel.

11 recordsLinked to original sources

Bounding Causal Effects for Ordinal Outcomes Under Positive Dependence

Defining and estimating causal effects for ordinal data is challenging. Standard average treatment effects are not appropriate for ordinal scales, and alternative estimands, such as the probabilities that the treatment outcome exceeds or does not worsen the control outcome, are generally not identifiable. Existing work provides sharp bounds for these quantities based only on marginal distributions. Motivated by a previous observation showing that bounds obtained under an independence working assumption can be substantially tighter, we investigate conditions under which such bounds are valid. We show that commonly used notions of positive dependence, including positive quadrant dependence and positive regression dependence, are not sufficient to justify these bounds. We then propose a new dependence condition, diagonal tail dominance (DTD), under which the independence-based bounds are guaranteed to hold. We explain why this condition is quite strong and may not be appropriate in many settings, limiting the justification for using the independence-based bounds. However, local DTD may be plausible in many applications, and we derive improved bounds that exploit an independence working assumption on selected parts of the probability table. Through theoretical results, numerical examples, and an analysis of data from a clinical trial of a new treatment for acute ischemic stroke, we illustrate the properties of the bounds and the role of the proposed conditions.

stat.ME

Cox Regression on the Plane

The Cox proportional hazards model is the most widely used regression model in univariate survival analysis, yet extensions to bivariate survival data remain scarce. We propose two novel extensions based on a Lehmann-type representation of the survival function. The first, the simple Lehmann model, is a direct extension that retains a straightforward structure. The second, the generalized Lehmann model, allows greater flexibility by incorporating three distinct regression parameters and includes the simple Lehmann model as a special case. The models admit a direct interpretation in terms of survival probabilities, providing a transparent, fully semiparametric framework for assessing covariate effects on both marginal survival probabilities and their dependence, without requiring specification of a copula or frailty distribution. To estimate the regression parameters, we build on a pseudo-observation-based approach for bivariate survival data and extend it to the generalized model via a two-step procedure. We establish consistency and asymptotic normality of the resulting estimators. The proposed approach is illustrated through simulation studies and an application to data from the Global Retinoblastoma Outcome Study.

stat.ME

Anchoring-Based Causal Design (ABCD): Estimating the Effects of Beliefs

A central challenge in any study of the effects of beliefs on outcomes, such as decisions and behavior, is the risk of omitted variables bias. Omitted variables, frequently unmeasured or even unknown, can induce correlations between beliefs and decisions that are not genuinely causal, in which case the omitted variables are referred to as confounders. To address the challenge of causal inference, researchers frequently rely on information provision experiments to randomly manipulate beliefs. The information supplied in these experiments can serve as an instrumental variable (IV), enabling causal inference, so long as it influences decisions exclusively through its impact on beliefs. However, providing varying information to participants to shape their beliefs can raise both methodological and ethical concerns. Methodological concerns arise from potential violations of the exclusion restriction assumption. Such violations may stem from information source effects, when attitudes toward the source affect the outcome decision directly, thereby introducing a confounder. An ethical concern arises from manipulating the provided information, as it may involve deceiving participants. This paper proposes and empirically demonstrates a new method for treating beliefs and estimating their effects, the Anchoring-Based Causal Design (ABCD), which avoids deception and source influences. ABCD combines the cognitive mechanism known as anchoring with instrumental variable (IV) estimation. Instead of providing substantive information, the method employs a deliberately non-informative procedure in which participants compare their self-assessment of a concept to a randomly assigned anchor value. We present the method and the results of eight experiments demonstrating its application, strengths, and limitations. We conclude by discussing the potential of this design for advancing experimental social science.

econ.GN

Pseudo-Observations for Bivariate Survival Data

The pseudo-observations approach has been gaining popularity as a method to estimate covariate effects on censored survival data. It is used regularly to estimate covariate effects on quantities such as survival probabilities, restricted mean life, cumulative incidence, and others. In this work, we propose to generalize the pseudo-observations approach to situations where a bivariate failure-time variable is observed, subject to right censoring. The idea is to first estimate the joint survival function of both failure times and then use it to define the relevant pseudo-observations. Once the pseudo-observations are calculated, they are used as the response in a generalized linear model. We consider two common nonparametric estimators of the joint survival function: the estimator of Lin and Ying (1993) and the Dabrowska estimator (Dabrowska, 1988). For both estimators, we show that our bivariate pseudo-observations approach produces regression estimates that are consistent and asymptotically normal. Our proposed method enables estimation of covariate effects on quantities such as the joint survival probability at a fixed bivariate time point, or simultaneously at several time points, and consequentially can estimate covariate-adjusted conditional survival probabilities. We demonstrate the method using simulations and an analysis of two real-world datasets.

stat.ME

Estimating Mean Viral Load Trajectory from Intermittent Longitudinal Data and Unknown Time Origins

Viral load (VL) in the respiratory tract is the leading proxy for assessing infectiousness potential. Understanding the dynamics of disease-related VL within the host is very important and help to determine different policy and health recommendations. However, often only partial followup data are available with unknown infection date. In this paper we introduce a discrete time likelihood-based approach to modeling and estimating partial observed longitudinal samples. We model the VL trajectory by a multivariate normal distribution that accounts for possible correlation between measurements within individuals. We derive an expectation-maximization (EM) algorithm which treats the unknown time origins and the missing measurements as latent variables. Our main motivation is the reconstruction of the daily mean SARS-Cov-2 VL, given measurements performed on random patients, whose VL was measured multiple times on different days. The method is applied to SARS-Cov-2 cycle-threshold-value data collected in Israel.

stat.AP

Spatial modeling of randomly acquired characteristics on outsoles with application to forensic shoeprint analysis

Footwear comparison is used to link between a suspect's shoe and a footprint found at a crime scene. Investigators compare the two items using randomly acquired characteristics (RACs), such as scratches or holes. However, to date, the distribution of RAC characteristics has not been investigated thoroughly, and the evidential value of RACs is yet to be explored. An important question concerns the distribution of the location of RACs on shoe soles, which can serve as a benchmark for comparison. The location of RACs is modeled here as a point process over the shoe sole and a data set of 386 independent shoes is used to estimate its rate function. The analysis is somewhat complicated as the shoes are differentiated by shape, level of wear and tear and contact surface. This paper presents methods that take into account these challenges, either by using natural cubic splines on high resolution data, or by using a piecewise-constant model on larger regions defined by experts' knowledge. It is shown that RACs are likely to appear at certain locations, corresponding to the foot's morphology. The results can guide investigators in determining the evidential value of footprint comparison.

stat.AP

Testing Independence under Biased Sampling

Testing for association or dependence between pairs of random variables is a fundamental problem in statistics. In some applications, data are subject to selection bias that causes dependence between observations even when it is absent from the population. An important example is truncation models, in which observed pairs are restricted to a specific subset of the X-Y plane. Standard tests for independence are not suitable in such cases, and alternative tests that take the selection bias into account are required. To deal with this issue, we generalize the notion of quasi-independence with respect to the sampling mechanism, and study the problem of detecting any deviations from it. We develop two test statistics motivated by the classic Hoeffding's statistic, and use two approaches to compute their distribution under the null: (i) a bootstrap-based approach, and (ii) a permutation-test with non-uniform probability of permutations, sampled using either MCMC or importance sampling with various proposal distributions. We show that our tests can tackle cases where the biased sampling mechanism is estimated from the data, with an important application to the case of censoring with truncation. We prove the validity of the tests, and show, using simulations, that they perform well for important special cases of the problem and improve power compared to competing methods. The tests are applied to four datasets, two that are subject to truncation, with and without censoring, and two to positive bias mechanisms related to length bias.

stat.ME

The Scaled Uniform Model Revisited

Sufficiency, Conditionality and Invariance are basic principles of statistical inference. Current mathematical statistics courses do not devote much teaching time to these classical principles, and even ignore the latter two, in order to teach modern methods. However, being the philosophical cornerstones of statistical inference, a minimal understanding of these principles should be part of any curriculum in statistics. The scaled uniform model is used here to demonstrate the importance and usefulness of the principles. The main focus is on the conditionality principle that is probably the most basic and less familiar among the three. The appendix discusses the invariance principle and the conditionality principle in the case of sampling from a finite population.

math.ST

Variance function estimation in quantitative mass spectrometry with application to iTRAQ labeling

This paper describes and compares two methods for estimating the variance function associated with iTRAQ (isobaric tag for relative and absolute quantitation) isotopic labeling in quantitative mass spectrometry based proteomics. Measurements generated by the mass spectrometer are proportional to the concentration of peptides present in the biological sample. However, the iTRAQ reporter signals are subject to errors that depend on the peptide amounts. The variance function of the errors is therefore an essential parameter for evaluating the results, but estimating it is complicated, as the number of nuisance parameters increases with sample size while the number of replicates for each peptide remains small. Two experiments that were conducted with the sole goal of estimating the variance function and its stability over time are analyzed, and the resulting estimated variance function is used to analyze an experiment targeting aberrant signaling cascades in cells harboring distinct oncogenic mutations. Methods for constructing conservative $p$-values and confidence intervals are discussed.

stat.AP

Are adaptive allocation designs beneficial for improving power in binary response trials?

We consider the classical problem of selecting the best of two treatments in clinical trials with binary response. The target is to find the design that maximizes the power of the relevant test. Many papers use a normal approximation to the power function and claim that Neyman allocation that assigns subjects to treatment groups according to the ratio of the responses' standard deviations, should be used. As the standard deviations are unknown, an adaptive design is often recommended. The asymptotic justification of this approach is arguable, since it uses the normal approximation in tails where the error in the approximation is larger than the estimated quantity. We consider two different approaches for optimality of designs that are related to Pitman and Bahadur definitions of relative efficiency of tests. We prove that the optimal allocation according to the Pitman criterion is the balanced allocation and that the optimal allocation according to the Bahadur approach depends on the unknown parameters. Exact calculations reveal that the optimal allocation according to Bahadur is often close to the balanced design, and the powers of both are comparable to the Neyman allocation for small sample sizes and are generally better for large experiments. Our findings have important implications to the design of experiments, as the balanced design is proved to be optimal or close to optimal and the need for the complications involved in following an adaptive design for the purpose of increasing the power of tests is therefore questionable.

math.ST

Nonparametric estimation of a distribution function under biased sampling and censoring

This paper derives the nonparametric maximum likelihood estimator (NPMLE) of a distribution function from observations which are subject to both bias and censoring. The NPMLE is obtained by a simple EM algorithm which is an extension of the algorithm suggested by Vardi (Biometrika, 1989) for size biased data. Application of the algorithm to many models is discussed and a simulation study compares the estimator's performance to that of the product-limit estimator (PLE). An example demonstrates the utility of the NPMLE to data where the PLE is inappropriate.

math.ST