SearcharxivSearch

arXiv subjects

Xiao-Li Meng

Publications and source records attributed to Xiao-Li Meng.

At least 19 recordsLinked to original sources

The LIRA-Ising Model: Estimating the boundaries of irregularly shaped X-ray sources

Mapping the boundary of an extended source is a key step in the study of its morphology. The background contamination and statistical fluctuations of typical astronomical images make this a challenging statistical task, particularly for X-ray images with low surface brightness. We develop a three-step Bayesian procedure to identify the boundaries of irregularly shaped sources. We first apply a Bayesian multiscale reconstruction algorithm known as LIRA to obtain posterior pixelwise probability distributions of the source intensity that properly account for known structures, astrophysical background, and the effect of the telescope point spread function. Next, we adopt an Ising model to group pixels with similar intensities into cohesive regions corresponding to background and source. Finally, the boundary is derived on the basis of the most likely aggregation of pixels into the source region. Because the overall model combines LIRA and the Ising model, we call it LIRA-Ising. We verify the proposed method using a set of simulation studies. We then apply it to the Chandra X-ray Observatory images of two high redshift quasars, PKS J1421-0643 and 0730+257, to determine the extent and morphology of X-ray jets. Our method shows a uniform X-ray surface brightness of PKS J1421-0643 jet, and identifies knotty structure in the X-ray jet of 0730+257.

astro-ph.IM

Making high-order asymptotics practical: correcting goodness-of-fit test for astronomical count data

The C statistic is a widely used likelihood-ratio statistic for model fitting and goodness-of-fit assessments with Poisson data in high-energy physics and astrophysics. Although it enjoys convenient asymptotic properties, the statistic is routinely applied in cases where its nominal null distribution relies on unwarranted assumptions. Because researchers do not typically carry out robustness checks, their scientific findings are left vulnerable to misleading significance calculations. With an emphasis on low-count scenarios, we present a comprehensive study of the theoretical properties of C statistics and related goodness-of-fit algorithms. We focus on common ``plug-in'' algorithms where moments of C are obtained by assuming the true parameter equals its estimate. To correct such methods, we provide a suite of new principled user-friendly algorithms and well-calibrated p-values that are ready for immediate deployment in the (astro)physics data-analysis pipeline. Using both theoretical and numerical results, we show (a) standard $\chi^2$-based goodness-of-fit assessments are invalid in low-count settings, (b) naive methods (e.g., vanilla bootstrap) result in biased null distributions, and (c) the corrected Z-test based on conditioning and high-order asymptotics gives the best precision with low computational cost. We illustrate our methods via a suite of simulations and applied astrophysical analyses. An open-source Python package is provided in a GitHub repository.

stat.ME

Differential Privacy Meets Invariant Statistics: Some Conundrums in Quantifying Trade-Offs

This work was inspired by the question of whether data swapping, a popular form of statistical disclosure control used to protect many data products including three recent US Decennial Censuses, can satisfy differential privacy (DP). Given the existence of more than 200 formulations of DP (and counting), as a precondition to answering this question one must precisely specify what it actually means to be DP. Motivated by this observation, we first conduct a theoretical investigation into DP's fundamental essence, resulting in a five-building-block system explicating the who, where, what, how and how much aspects of DP. Instantiating this system in the context of the US Decennial Census, we then demonstrate the broad applicability and relevance of DP by comparing a swapping strategy like that used in 2010 with the TopDown Algorithm--the main DP method adopted in the 2020 Census. This chapter provides nontechnical summaries of these two pieces of work (developed elsewhere), as well as extended discussions on a number of issues they unearth that complicate the formulation and the navigation of the so-called privacy-utility trade-off: How can greater awareness of the five building blocks thwart privacy theatrics? How can invariants (statistics that are released as is, without any privacy protection) align with DP's philosophy of relative privacy? How do our results bridging traditional statistical disclosure control and DP allow a data custodian to reap the benefits of both these fields? And how can removing the implicit reliance on aleatoric uncertainty lead to new generalizations of DP? Our ultimate goal with these discussions is to deepen the theoretical basis, broaden the practical applicability, and reduce the misperception of DP--all without shaking its core foundations.

cs.CR

A Refreshment Stirred, Not Shaken: Invariant-Preserving Deployments of Differential Privacy for the U.S. Decennial Census

Protecting an individual's privacy when releasing their data is inherently an exercise in relativity, regardless of how privacy is qualified or quantified. This is because we can only limit the gain in information about an individual relative to what could be derived from other sources. This framing is the essence of differential privacy (DP), through which this article examines two statistical disclosure control (SDC) methods for the United States Decennial Census: the Permutation Swapping Algorithm (PSA), which resembles the 2010 Census's disclosure avoidance system (DAS), and the TopDown Algorithm (TDA), which was used in the 2020 DAS. To varying degrees, both methods leave unaltered certain statistics of the confidential data (their invariants) and hence neither can be readily reconciled with DP, at least as originally conceived. Nevertheless, we show how invariants can naturally be integrated into DP and use this to establish that the PSA satisfies pure DP subject to the invariants it necessarily induces, thereby proving that this traditional SDC method can, in fact, be understood from the perspective of DP. By a similar modification to zero-concentrated DP, we also provide a DP specification for the TDA. Finally, as a point of comparison, we consider a counterfactual scenario in which the PSA was adopted for the 2020 Census, resulting in a reduction in the nominal protection loss budget but at the cost of releasing many more invariants. This highlights the pervasive danger of comparing budgets without accounting for the other dimensions on which DP formulations vary (such as the invariants they permit). Therefore, while our results articulate the mathematical guarantees of SDC provided by the PSA, the TDA, and the 2020 DAS in general, care must be taken in translating these guarantees into actual privacy protection$\unicode{x2014}$just as is the case for any DP deployment.

cs.CR

A Heavily Right Strategy for Statistical Inference with Dependent Studies in Any Dimension

We leverage recent advances in heavy-tail approximations for global hypothesis testing with dependent studies to construct approximate confidence regions without modeling or estimating their dependence structures. A non-rejection region is a confidence region but it may not be convex. Convexity is appealing because it ensures any one-dimensional linear projection of the region is a confidence interval, easy to compute and interpret. We show why convexity fails for nearly all heavy-tail combination tests proposed in recent years, including the influential Cauchy combination test. These insights motivate a \textit{heavily right} strategy: truncating the left half of the Cauchy distribution to obtain the Half-Cauchy combination test. The harmonic mean test also corresponds to a heavily right distribution with a Cauchy-like tail, namely a Pareto distribution with unit power. We prove that both approaches guarantee convexity when individual studies are summarized by Hotelling $T^2$ or $\chi^{2}$ statistics (regardless of the validity of this summary) and provide efficient, \textit{exact} algorithms for implementation. Applying these methods, we develop a divide-and-combine strategy for mean estimation in any dimension and construct simultaneous confidence intervals in a network meta-analysis for treatment effect comparisons across multiple clinical trials. We also present many open problems and conclude with epistemic reflections.

stat.ME

A Heisenberg-esque Uncertainty Principle for Simultaneous (Machine) Learning and Error Assessment?

A highly cited and inspiring article by Bates et al (2024) demonstrates that the prediction errors estimated through cross-validation, Bootstrap or Mallow's $C_P$ can all be independent of the actual prediction errors. This essay hypothesizes that these occurrences signify a broader, Heisenberg-like uncertainty principle for learning: optimizing learning and assessing actual errors using the same data are fundamentally at odds. Only suboptimal learning preserves untapped information for actual error assessments, and vice versa, reinforcing the `no free lunch' principle. To substantiate this intuition, a Cramer-Rao-style lower bound is established under the squared loss, which shows that the relative regret in learning is bounded below by the square of the correlation between any unbiased error assessor and the actual learning error. Readers are invited to explore generalizations, develop variations, or even uncover genuine `free lunches.' The connection with the Heisenberg uncertainty principle is more than metaphorical, because both share an essence of the Cramer-Rao inequality: marginal variations cannot manifest individually to arbitrary degrees when their underlying co-variation is constrained, whether the co-variation is about individual states or their generating mechanisms, as in the quantum realm. A practical takeaway of such a learning principle is that it may be prudent to reserve some information specifically for error assessment rather than pursue full optimization in learning, particularly when intentional randomness is introduced to mitigate overfitting.

math.ST

Six Maxims of Statistical Acumen for Astronomical Data Analysis

The production of complex astronomical data is accelerating, especially with newer telescopes producing ever more large-scale surveys. The increased quantity, complexity, and variety of astronomical data demand a parallel increase in skill and sophistication in developing, deciding, and deploying statistical methods. Understanding limitations and appreciating nuances in statistical and machine learning methods and the reasoning behind them is essential for improving data-analytic proficiency and acumen. Aiming to facilitate such improvement in astronomy, we delineate cautionary tales in statistics via six maxims, with examples drawn from the astronomical literature. Inspired by the significant quality improvement in business and manufacturing processes by the routine adoption of Six Sigma, we hope the routine reflection on these Six Maxims will improve the quality of both data analysis and scientific findings in astronomy.

astro-ph.IM

Channelling Multimodality Through a Unimodalizing Transport: Warp-U Sampler and Stochastic Bridge Sampling Estimator

Monte Carlo integration is a powerful tool for scientific and statistical computation, but faces significant challenges when the integrand is a multi-modal distribution, even when the mode locations are known. This work introduces novel Monte Carlo sampling and integration estimation strategies for the multi-modal context by leveraging a generalized version of the stochastic Warp-U transformation Wang et al. [2022]. We propose two flexible classes of Warp-U transformations, one based on a general location-scale-skew mixture model and a second using neural ordinary differential equations. We develop an efficient sampling strategy called Warp-U sampling, which applies a Warp-U transformation to map a multi-modal density into a uni-modal one, then inverts the transformation with injected stochasticity. In high dimensions, our approach relies on information about the mode locations, but requires minimal tuning and demonstrates better mixing properties than conventional methods with identical mode information. To improve normalizing constant estimation once samples are obtained, we propose a stochastic Warp-U bridge sampling estimator, which we demonstrate has higher asymptotic precision per CPU second compared to the original approach proposed by Wang et al. [2022]. We also establish the ergodicity of our sampling algorithm. The effectiveness and current limitations of our methods are illustrated through simulation studies and an application to exoplanet detection.

stat.ME

BELIEF in Dependence: Leveraging Atomic Linearity in Data Bits for Rethinking Generalized Linear Models

Two linearly uncorrelated binary variables must be also independent because non-linear dependence cannot manifest with only two possible states. This inherent linearity is the atom of dependency constituting any complex form of relationship. Inspired by this observation, we develop a framework called binary expansion linear effect (BELIEF) for understanding arbitrary relationships with a binary outcome. Models from the BELIEF framework are easily interpretable because they describe the association of binary variables in the language of linear models, yielding convenient theoretical insight and striking Gaussian parallels. With BELIEF, one may study generalized linear models (GLM) through transparent linear models, providing insight into how the choice of link affects modeling. For example, setting a GLM interaction coefficient to zero does not necessarily lead to the kind of no-interaction model assumption as understood under their linear model counterparts. Furthermore, for a binary response, maximum likelihood estimation for GLMs paradoxically fails under complete separation, when the data are most discriminative, whereas BELIEF estimation automatically reveals the perfect predictor in the data that is responsible for complete separation. We explore these phenomena and provide related theoretical results. We also provide preliminary empirical demonstration of some theoretical results.

math.ST

Six Statistical Senses

This article proposes a set of categories, each one representing a particular distillation of important statistical ideas. Each category is labeled a "sense" because we think of these as essential in helping every statistical mind connect in constructive and insightful ways with statistical theory, methodologies, and computation, toward the ultimate goal of building statistical phronesis. The illustration of each sense with statistical principles and methods provides a sensical tour of the conceptual landscape of statistics, as a leading discipline in the data science ecosystem.

stat.OT

Scalable Spike-and-Slab

Spike-and-slab priors are commonly used for Bayesian variable selection, due to their interpretability and favorable statistical properties. However, existing samplers for spike-and-slab posteriors incur prohibitive computational costs when the number of variables is large. In this article, we propose Scalable Spike-and-Slab ($S^3$), a scalable Gibbs sampling implementation for high-dimensional Bayesian regression with the continuous spike-and-slab prior of George and McCulloch (1993). For a dataset with $n$ observations and $p$ covariates, $S^3$ has order $\max\{ n^2 p_t, np \}$ computational cost at iteration $t$ where $p_t$ never exceeds the number of covariates switching spike-and-slab states between iterations $t$ and $t-1$ of the Markov chain. This improves upon the order $n^2 p$ per-iteration cost of state-of-the-art implementations as, typically, $p_t$ is substantially smaller than $p$. We apply $S^3$ on synthetic and real-world datasets, demonstrating orders of magnitude speed-ups over existing exact samplers and significant gains in inferential quality over approximate samplers with comparable cost.

stat.CO

Concordance: In-flight Calibration of X-ray Telescopes without Absolute References

We describe a process for cross-calibrating the effective areas of X-ray telescopes that observe common targets. The targets are not assumed to be "standard candles" in the classic sense, in that we assume that the source fluxes have well-defined, but {\it a priori} unknown values. Using a technique developed by Chen et al. (2019, arXiv:1711.09429) that involves a statistical method called {\em shrinkage estimation}, we determine effective area correction factors for each instrument that brings estimated fluxes into the best agreement, consistent with prior knowledge of their effective areas. We expand the technique to allow unique priors on systematic uncertainties in effective areas for each X-ray astronomy instrument and to allow correlations between effective areas in different energy bands. We demonstrate the method with several data sets from various X-ray telescopes.

astro-ph.IM

Unrepresentative Big Surveys Significantly Overestimate US Vaccine Uptake

Surveys are a crucial tool for understanding public opinion and behavior, and their accuracy depends on maintaining statistical representativeness of their target populations by minimizing biases from all sources. Increasing data size shrinks confidence intervals but magnifies the impact of survey bias, an instance of the Big Data Paradox (Meng 2018). Here we demonstrate this paradox in estimates of first-dose COVID-19 vaccine uptake in US adults: Delphi-Facebook (about 250,000 responses per week) and Census Household Pulse (about 75,000 per week). By May 2021, Delphi-Facebook overestimated uptake by 17 percentage points and Census Household Pulse by 14, compared to a benchmark from the Centers for Disease Control and Prevention (CDC). Moreover, their large data sizes led to minuscule margins of error on the incorrect estimates. In contrast, an Axios-Ipsos online panel with about 1,000 responses following survey research best practices (AAPOR) provided reliable estimates and uncertainty. We decompose observed error using a recent analytic framework to explain the inaccuracy in the three surveys. We then analyze the implications for vaccine hesitancy and willingness. We show how a survey of 250,000 respondents can produce an estimate of the population mean that is no more accurate than an estimate from a simple random sample of size 10. Our central message is that data quality matters far more than data quantity, and compensating the former with the latter is a mathematically provable losing proposition.

stat.AP

A Multi-resolution Theory for Approximating Infinite-$p$-Zero-$n$: Transitional Inference, Individualized Predictions, and a World Without Bias-Variance Trade-off

Transitional inference is an empiricism concept, rooted and practiced in clinical medicine since ancient Greece. Knowledge and experiences gained from treating one entity are applied to treat a related but distinctively different one. This notion of "transition to the similar" renders individualized treatments an operational meaning, yet its theoretical foundation defies the familiar inductive inference framework. The uniqueness of entities is the result of potentially an infinite number of attributes (hence $p=\infty$), which entails zero direct training sample size (i.e., $n=0$) because genuine guinea pigs do not exist. However, the literature on wavelets and on sieve methods suggests a principled approximation theory for transitional inference via a multi-resolution (MR) perspective, where we use the resolution level to index the degree of approximation to ultimate individuality. MR inference seeks a primary resolution indexing an indirect training sample, which provides enough matched attributes to increase the relevance of the results to the target individuals and yet still accumulate sufficient indirect sample sizes for robust estimation. Theoretically, MR inference relies on an infinite-term ANOVA-type decomposition, providing an alternative way to model sparsity via the decay rate of the resolution bias as a function of the primary resolution level. Unexpectedly, this decomposition reveals a world without variance when the outcome is a deterministic function of potentially infinitely many predictors. In this deterministic world, the optimal resolution prefers over-fitting in the traditional sense when the resolution bias decays sufficiently rapidly. Furthermore, there can be many "descents" in the prediction error curve, when the contributions of predictors are inhomogeneous and the ordering of their importance does not align with the order of their inclusion in prediction.

math.ST

Double Happiness: Enhancing the Coupled Gains of L-lag Coupling via Control Variates

The recently proposed L-lag coupling for unbiased Markov chain Monte Carlo (MCMC) calls for a joint celebration by MCMC practitioners and theoreticians. For practitioners, it circumvents the thorny issue of deciding the burn-in period or when to terminate an MCMC sampling process, and opens the door for safe parallel implementation. For theoreticians, it provides a powerful tool to establish elegant and easily estimable bounds on the exact error of an MCMC approximation at any finite number of iterates. A serendipitous observation about the bias-correcting term leads us to introduce naturally available control variates into the L-lag coupling estimators. In turn, this extension enhances the coupled gains of L-lag coupling, because it results in more efficient unbiased estimators, as well as a better bound on the total variation error of MCMC iterations, albeit the gains diminish as L increases. Specifically, the new upper bound is theoretically guaranteed to never exceed the one given previously. We also argue that L-lag coupling represents a coupling for the future, breaking from the coupling-from-the-past type of perfect sampling, by reducing the generally unachievable requirement of being perfect to one of being unbiased, a worthwhile trade-off for ease of implementation in most practical situations. The theoretical analysis is supported by numerical experiments that show tighter bounds and a gain in efficiency when control variates are introduced.

stat.CO

Congenial Differential Privacy under Mandated Disclosure

Differentially private data releases are often required to satisfy a set of external constraints that reflect the legal, ethical, and logical mandates to which the data curator is obligated. The enforcement of constraints, when treated as post-processing, adds an extra phase in the production of privatized data. It is well understood in the theory of multi-phase processing that congeniality, a form of procedural compatibility between phases, is a prerequisite for the end users to straightforwardly obtain statistically valid results. Congenial differential privacy is theoretically principled, which facilitates transparency and intelligibility of the mechanism that would otherwise be undermined by ad-hoc post-processing procedures. We advocate for the systematic integration of mandated disclosure into the design of the privacy mechanism via standard probabilistic conditioning on the invariant margins. Conditioning automatically renders congeniality because any extra post-processing phase becomes unnecessary. We provide both initial theoretical guarantees and a Markov chain algorithm for our proposal. We also discuss intriguing theoretical issues that arise in comparing congenital differential privacy and optimization-based post-processing, as well as directions for further research.

cs.CR

Designing Test Information and Test Information in Design

DeGroot (1962) developed a general framework for constructing Bayesian measures of the expected information that an experiment will provide for estimation. We propose an analogous framework for measures of information for hypothesis testing. In contrast to estimation information measures that are typically used for surface estimation, test information measures are more useful in experimental design for hypothesis testing and model selection. In particular, we obtain a probability based measure, which has more appealing properties than variance based measures in design contexts where decision problems are of interest. The underlying intuition of our design proposals is straightforward: to distinguish between models we should collect data from regions of the covariate space for which the models differ most. Nicolae et al. (2008) gave an asymptotic equivalence between their test information measures and Fisher information. We extend this result to all test information measures under our framework. Simulation studies and an application in astronomy demonstrate the utility of our approach, and provide comparison to other methods including that of Box and Hill (1967).

math.ST