SearcharxivSearch

arXiv subjects

Radu V. Craiu

Publications and source records attributed to Radu V. Craiu.

At least 19 recordsLinked to original sources

Constraining the mass of the M31 ionized baryon Halo using CHIME/FRB Catalog 2

The circumgalactic medium (CGM) surrounding galaxies is believed to be a significant reservoir of baryons, yet its total mass remains poorly constrained. We present a novel approach to probe the CGM of the Andromeda galaxy (M31) using fast radio bursts (FRBs) from the CHIME/FRB Catalog 2. By comparing the dispersion measures (DMs) of 171 FRBs whose sightlines intersect M31's halo (within $r_{\rm vir}= 302\,\mathrm{kpc}$) to a control sample of 684 FRBs, we estimate the DM contribution from M31's CGM. We find evidence for an excess DM of $δ\mathrm{DM} = 5.9$--$59.6\,\mathrm{pc\,cm^{-3}}$ in the inner halo ($\approx 0$--$151\,\mathrm{kpc}$) and $δ\mathrm{DM} = 25.7$--$64.6\,\mathrm{pc\,cm^{-3}}$ in the outer halo ($\approx 151$--$302\,\mathrm{kpc}$). Using a generalized halo model parameterized by a closure radius $r_{\mathrm{close}}$, we constrain the baryon distribution and infer a best-fit value for $r_{\mathrm{close}}$ to be $9.2^{+9.9}_{-4.9}\,r_{\rm vir}$ ($1σ$), corresponding to a total CGM mass of $M_{b,\mathrm{halo}} = 18.6^{+7.9}_{-8.4} \times 10^{10}\,M_\odot$. Our results suggest that M31 may harbor a substantial fraction of its cosmic baryon budget in diffuse, ionized gas. This work demonstrates the potential of FRBs as a powerful tool for studying the CGM of nearby galaxies, with future larger samples expected to provide tighter constraints on the baryon content of galactic halos.

astro-ph.GA

Bayesian nonparametric mixtures of Archimedean copulas

Copula-based dependence modeling often relies on parametric formulations. This is mathematically convenient, but can be statistically inefficient when the parametric families are not suitable for the data and model in focus. A Bayesian nonparametric mixture of Archimedean copulas is introduced to increase the flexibility of copula-based dependence modeling. Specifically, the Poisson-Dirichlet process is used as a mixing distribution over the Archimedean copulas' parameter. Properties of the mixture model are studied for the main Archimedean families, and posterior distributions are sampled via their full conditional distributions. The performance of the model is illustrated via numerical experiments involving simulated and real data.

stat.ME

Rare Event Classification with Weighted Logistic Regression for Identifying Repeating Fast Radio Bursts

An important task in the study of fast radio bursts (FRBs) remains the automatic classification of repeating and non-repeating sources based on their morphological properties. We propose a statistical model that considers a modified logistic regression to classify FRB sources. The classical logistic regression model is modified to accommodate the small proportion of repeaters in the data, a feature that is likely due to the sampling procedure and duration and is not a characteristic of the population of FRB sources. The weighted logistic regression hinges on the choice of a tuning parameter that represents the true proportion $τ$ of repeating FRB sources in the entire population. The proposed method has a sound statistical foundation, direct interpretability, and operates with only 5 parameters, enabling quicker retraining with added data. Using the CHIME/FRB Collaboration sample of repeating and non-repeating FRBs and numerical experiments, we achieve a classification accuracy for repeaters of nearly 75\% or higher when $τ$ is set in the range of $50$ to $60$\%. This implies a tentative high proportion of repeaters, which is surprising, but is also in agreement with recent estimates of $τ$ that are obtained using other methods.

astro-ph.HE

Constraining the selection corrected luminosity function and total pulse count for radio transients

Studying transient phenomena, such as individual pulses from pulsars, has garnered considerable attention in the era of astronomical big data. Of specific interest to this study are Rotating Radio Transients (RRATs), nulling, and intermittent pulsars. This study introduces a new algorithm named LuNfit, tailored to correct the selection biases originating from the telescope and detection pipelines. Ultimately LuNfit estimates the intrinsic luminosity distribution and nulling fraction of the single pulses emitted by pulsars. LuNfit relies on Bayesian nested sampling so that the parameter space can be fully explored. Bayesian nested sampling also provides the additional benefit of simplifying model comparisons through the Bayes ratio. The robustness of LuNfit is shown through simulations and applying LuNfit onto pulsars with known nulling fractions. LuNfit is then applied to three RRATs, J0012+5431, J1538+1523, and J2355+1523, extracting their intrinsic luminosity distribution and burst rates. We find that their nulling fraction is 0.4(2), 0.749(5) and 0.995(2) respectively. We further find that a log-normal distribution likely describes the single pulse luminosity distribution of J0012+5431 and J1538+1523, while the Bayes ratio for J2355+1523 slightly favors an exponential distribution. We show the conventional method of correcting selection effects by "scaling up" the missed fraction of radio transients can be unreliable when the mean luminosity of the source is faint relative to the telescope sensitivity. Finally, we discuss the limitations of the current implementation of LuNfit while also delving into potential enhancements that would enable LuNfit to be applied to sources with complex pulse morphologies.

astro-ph.HE

Approximate Methods for Bayesian Computation

Rich data generating mechanisms are ubiquitous in this age of information and require complex statistical models to draw meaningful inference. While Bayesian analysis has seen enormous development in the last 30 years, benefitting from the impetus given by the successful application of Markov chain Monte Carlo (MCMC) sampling, the combination of big data and complex models conspire to produce significant challenges for the traditional MCMC algorithms. We review modern algorithmic developments addressing the latter and compare their performance using numerical experiments.

stat.CO

Six Statistical Senses

This article proposes a set of categories, each one representing a particular distillation of important statistical ideas. Each category is labeled a "sense" because we think of these as essential in helping every statistical mind connect in constructive and insightful ways with statistical theory, methodologies, and computation, toward the ultimate goal of building statistical phronesis. The illustration of each sense with statistical principles and methods provides a sensical tour of the conceptual landscape of statistics, as a leading discipline in the data science ecosystem.

stat.OT

Copula Modelling of Serially Correlated Multivariate Data with Hidden Structures

We propose a copula-based extension of the hidden Markov model (HMM) which applies when the observations recorded at each time in the sample are multivariate. The joint model produced by the copula extension allows decoding of the hidden states based on information from multiple observations. However, unlike the case of independent marginals, the copula dependence structure embedded into the likelihood poses additional computational challenges. We tackle the latter using a theoretically-justified variation of the EM algorithm developed within the framework of inference functions for margins. We illustrate the method using numerical experiments and an analysis of house occupancy.

stat.ME

Measuring the severity of multi-collinearity in high dimensions

Multi-collinearity is a wide-spread phenomenon in modern statistical applications and when ignored, can negatively impact model selection and statistical inference. Classic tools and measures that were developed for "$n>p$" data are not applicable nor interpretable in the high-dimensional regime. Here we propose 1) new individualized measures that can be used to visualize patterns of multi-collinearity, and subsequently 2) global measures to assess the overall burden of multi-collinearity without limiting the observed data dimensions. We applied these measures to genomic applications to investigate patterns of multi-collinearity in genetic variations across individuals with diverse ancestral backgrounds. The measures were able to visually distinguish genomic regions of excessive multi-collinearity and contrast the level of multi-collinearity between different continental populations.

stat.ME

Exploring dimension learning via a penalized probabilistic principal component analysis

Establishing a low-dimensional representation of the data leads to efficient data learning strategies. In many cases, the reduced dimension needs to be explicitly stated and estimated from the data. We explore the estimation of dimension in finite samples as a constrained optimization problem, where the estimated dimension is a maximizer of a penalized profile likelihood criterion within the framework of a probabilistic principal components analysis. Unlike other penalized maximization problems that require an "optimal" penalty tuning parameter, we propose a data-averaging procedure whereby the estimated dimension emerges as the most favourable choice over a range of plausible penalty parameters. The proposed heuristic is compared to a large number of alternative criteria in simulations and an application to gene expression data. Extensive simulation studies reveal that none of the methods uniformly dominate the other and highlight the importance of subject-specific knowledge in choosing statistical methods for dimension learning. Our application results also suggest that gene expression data have a higher intrinsic dimension than previously thought. Overall, our proposed heuristic strikes a good balance and is the method of choice when model assumptions deviated moderately.

stat.ME

Living on the Edge: An Unified Approach to Antithetic Sampling

We identify recurrent ingredients in the antithetic sampling literature leading to a unified sampling framework. We introduce a new class of antithetic schemes that includes the most used antithetic proposals. This perspective enables the derivation of new properties of the sampling schemes: i) optimality in the Kullback-Leibler sense; ii) closed-form multivariate Kendall's $τ$ and Spearman's $ρ$; iii)ranking in concordance order and iv) a central limit theorem that characterizes stochastic behavior of Monte Carlo estimators when the sample size tends to infinity. Finally, we provide applications to Monte Carlo integration and Markov Chain Monte Carlo Bayesian estimation.

stat.ME

The X Factor: A Robust and Powerful Approach to X-chromosome-Inclusive Whole-genome Association Studies

The X-chromosome is often excluded from genome-wide association studies because of analytical challenges. Some of the problems, such as the random, skewed or no X-inactivation model uncertainty, have been investigated. Other considerations have received little to no attention, such as the value in considering non-additive and gene-sex interaction effects, and the inferential consequence of choosing different baseline alleles (i.e.\ the reference vs.\ the alternative allele). Here we propose a unified and flexible regression-based association test for X-chromosomal variants. We provide theoretical justifications for its robustness in the presence of various model uncertainties, as well as for its improved power when compared with the existing approaches under certain scenarios. For completeness, we also revisit the autosomes and show that the proposed framework leads to a more robust approach than the standard method. Finally, we provide supporting evidence by revisiting several published association studies. Supplementary materials for this article are available online.

stat.AP

Double Happiness: Enhancing the Coupled Gains of L-lag Coupling via Control Variates

The recently proposed L-lag coupling for unbiased Markov chain Monte Carlo (MCMC) calls for a joint celebration by MCMC practitioners and theoreticians. For practitioners, it circumvents the thorny issue of deciding the burn-in period or when to terminate an MCMC sampling process, and opens the door for safe parallel implementation. For theoreticians, it provides a powerful tool to establish elegant and easily estimable bounds on the exact error of an MCMC approximation at any finite number of iterates. A serendipitous observation about the bias-correcting term leads us to introduce naturally available control variates into the L-lag coupling estimators. In turn, this extension enhances the coupled gains of L-lag coupling, because it results in more efficient unbiased estimators, as well as a better bound on the total variation error of MCMC iterations, albeit the gains diminish as L increases. Specifically, the new upper bound is theoretically guaranteed to never exceed the one given previously. We also argue that L-lag coupling represents a coupling for the future, breaking from the coupling-from-the-past type of perfect sampling, by reducing the generally unachievable requirement of being perfect to one of being unbiased, a worthwhile trade-off for ease of implementation in most practical situations. The theoretical analysis is supported by numerical experiments that show tighter bounds and a gain in efficiency when control variates are introduced.

stat.CO

Finding our Way in the Dark: Approximate MCMC for Approximate Bayesian Methods

With larger data at their disposal, scientists are emboldened to tackle complex questions that require sophisticated statistical models. It is not unusual for the latter to have likelihood functions that elude analytical formulations. Even under such adversity, when one can simulate from the sampling distribution, Bayesian analysis can be conducted using approximate methods such as Approximate Bayesian Computation (ABC) or Bayesian Synthetic Likelihood (BSL). A significant drawback of these methods is that the number of required simulations can be prohibitively large, thus severely limiting their scope. In this paper we design perturbed MCMC samplers that can be used within the ABC and BSL paradigms to significantly accelerate computation while maintaining control on computational efficiency. The proposed strategy relies on recycling samples from the chain's past. The algorithmic design is supported by a theoretical analysis while practical performance is examined via a series of simulation examples and data analyses.

stat.CO

Likelihood Inflating Sampling Algorithm

Markov Chain Monte Carlo (MCMC) sampling from a posterior distribution corresponding to a massive data set can be computationally prohibitive since producing one sample requires a number of operations that is linear in the data size. In this paper, we introduce a new communication-free parallel method, the Likelihood Inflating Sampling Algorithm (LISA), that significantly reduces computational costs by randomly splitting the dataset into smaller subsets and running MCMC methods independently in parallel on each subset using different processors. Each processor will be used to run an MCMC chain that samples sub-posterior distributions which are defined using an "inflated" likelihood function. We develop a strategy for combining the draws from different sub-posteriors to study the full posterior of the Bayesian Additive Regression Trees (BART) model. The performance of the method is tested using both simulated and real data.

stat.ML

Bayesian Model Averaging for the X-Chromosome Inactivation Dilemma in Genetic Association Study

X-chromosome is often excluded from the so called `whole-genome' association studies due to its intrinsic difference between males and females. One particular analytical challenge is the unknown status of X-inactivation, where one of the two X-chromosome variants in females may be randomly selected to be silenced. In the absence of biological evidence in favour of one specific model, we consider a Bayesian model averaging framework that offers a principled way to account for the inherent model uncertainty, providing model averaging-based posterior density intervals and Bayes factors. We examine the inferential properties of the proposed methods via extensive simulation studies, and we apply the methods to a genetic association study of an intestinal disease occurring in about twenty percent of Cystic Fibrosis patients. Compared with the results previously reported assuming the presence of inactivation, we show that the proposed Bayesian methods provide more feature-rich quantities that are useful in practice.

stat.AP

Gaussian Process Single Index Models for Conditional Copulas

Parametric conditional copula models allow the copula parameters to vary with a set of covariates according to an unknown calibration function. Flexible Bayesian inference for the calibration function of a bivariate conditional copula is proposed via a sparse Gaussian process (GP) prior distribution over the set of smooth calibration functions for the single index model (SIM). The estimation of parameters from the marginal distributions and the calibration function is done jointly via Markov Chain Monte Carlo sampling from the full posterior distribution. A new Conditional Cross Validated Pseudo-Marginal (CCVML) criterion is introduced in order to perform copula selection and is modified using a permutation-based procedure to assess data support for the simplifying assumption. The performance of the estimation method and model selection criteria is studied via a series of simulations using correct and misspecified models with Clayton, Frank and Gaussian copulas and a numerical application involving red wine features.

stat.ME

A scalable and efficient covariate selection criterion for mixed effects regression models with unknown random effects structure

We propose a new model selection criterion for mixed effects regression models that is computable when the model is fitted with a two-step method, even when the structure and the distribution of the random effects are unknown. The criterion is especially useful in the early stage of the model building process when one needs to decide which covariates should be included in a mixed effects regression model but has no knowledge of the random effect structure. This is particularly relevant in substantive fields where variable selection is guided by information criteria rather than regularization. The calculation of the criterion requires only the evaluation of cluster-level log-likelihoods and does not rely on heavy numerical integration. We provide theoretical and numerical arguments to justify the method and we illustrate its usefulness by analyzing data on a socio-economic study of young American Indians.

stat.ME