SearcharxivSearch

arXiv subjects

David E. Jones

Publications and source records attributed to David E. Jones.

11 recordsLinked to original sources

Variable selection in balance regression with applications to microbiome compositional data

Compositional data, where only relative abundances are available, are common in microbiome and other high-throughput sequencing studies. Log ratios between groups of variables serve as key biomarkers in these settings. However, selecting predictive log ratios is a combinatorial challenge, and existing greedy search-based methods are computationally expensive, limiting their applicability to high-dimensional data. We introduce the supervised log ratio (SLR) method, a novel and efficient approach for selecting predictive log ratios in high-dimensional settings. SLR first screens active variables using univariate regression on log ratio transformed data and then applies principal balance analysis to define balance biomarkers. Our approach leverages both the relationship between the response and predictors and the correlations among the predictors to improve accuracy in variable selection and prediction. Through simulations and two case studies -- one on inflammatory bowel disease (IBD) and another on colorectal cancer (CRC) -- we demonstrate that SLR outperforms existing methods, particularly in high-dimensional settings. SLR is implemented in an R package, publicly available on GitHub.

stat.ME

Channelling Multimodality Through a Unimodalizing Transport: Warp-U Sampler and Stochastic Bridge Sampling Estimator

Monte Carlo integration is a powerful tool for scientific and statistical computation, but faces significant challenges when the integrand is a multi-modal distribution, even when the mode locations are known. This work introduces novel Monte Carlo sampling and integration estimation strategies for the multi-modal context by leveraging a generalized version of the stochastic Warp-U transformation Wang et al. [2022]. We propose two flexible classes of Warp-U transformations, one based on a general location-scale-skew mixture model and a second using neural ordinary differential equations. We develop an efficient sampling strategy called Warp-U sampling, which applies a Warp-U transformation to map a multi-modal density into a uni-modal one, then inverts the transformation with injected stochasticity. In high dimensions, our approach relies on information about the mode locations, but requires minimal tuning and demonstrates better mixing properties than conventional methods with identical mode information. To improve normalizing constant estimation once samples are obtained, we propose a stochastic Warp-U bridge sampling estimator, which we demonstrate has higher asymptotic precision per CPU second compared to the original approach proposed by Wang et al. [2022]. We also establish the ergodicity of our sampling algorithm. The effectiveness and current limitations of our methods are illustrated through simulation studies and an application to exoplanet detection.

stat.ME

Functional PCA with Covariate Dependent Mean and Covariance Structure

Incorporating covariates into functional principal component analysis (PCA) can substantially improve the representation efficiency of the principal components and predictive performance. However, many existing functional PCA methods do not make use of covariates, and those that do often have high computational cost or make overly simplistic assumptions that are violated in practice. In this article, we propose a new framework, called Covariate Dependent Functional Principal Component Analysis (CD-FPCA), in which both the mean and covariance structure depend on covariates. We propose a corresponding estimation algorithm, which makes use of spline basis representations and roughness penalties, and is substantially more computationally efficient than competing approaches of adequate estimation and prediction accuracy. A key aspect of our work is our novel approach for modeling the covariance function and ensuring that it is symmetric positive semi-definite. We demonstrate the advantages of our methodology through a simulation study and an astronomical data analysis.

stat.ME

Practical Guidance for Bayesian Inference in Astronomy

In the last two decades, Bayesian inference has become commonplace in astronomy. At the same time, the choice of algorithms, terminology, notation, and interpretation of Bayesian inference varies from one sub-field of astronomy to the next, which can lead to confusion to both those learning and those familiar with Bayesian statistics. Moreover, the choice varies between the astronomy and statistics literature, too. In this paper, our goal is two-fold: (1) provide a reference that consolidates and clarifies terminology and notation across disciplines, and (2) outline practical guidance for Bayesian inference in astronomy. Highlighting both the astronomy and statistics literature, we cover topics such as notation, specification of the likelihood and prior distributions, inference using the posterior distribution, and posterior predictive checking. It is not our intention to introduce the entire field of Bayesian data analysis -- rather, we present a series of useful practices for astronomers who already have an understanding of the Bayesian "nuts and bolts" and wish to increase their expertise and extend their knowledge. Moreover, as the field of astrostatistics and astroinformatics continues to grow, we hope this paper will serve as both a helpful reference and as a jumping off point for deeper dives into the statistics and astrostatistics literature.

astro-ph.IM

eBASCS: Disentangling Overlapping Astronomical Sources II, using Spatial, Spectral, and Temporal Information

The analysis of individual X-ray sources that appear in a crowded field can easily be compromised by the misallocation of recorded events to their originating sources. Even with a small number of sources, that nonetheless have overlapping point spread functions, the allocation of events to sources is a complex task that is subject to uncertainty. We develop a Bayesian method designed to sift high-energy photon events from multiple sources with overlapping point spread functions, leveraging the differences in their spatial, spectral, and temporal signatures. The method probabilistically assigns each event to a given source. Such a disentanglement allows more detailed spectral or temporal analysis to focus on the individual component in isolation, free of contamination from other sources or the background. We are also able to compute source parameters of interest like their locations, relative brightness, and background contamination, while accounting for the uncertainty in event assignments. Simulation studies that include event arrival time information demonstrate that the temporal component improves event disambiguation beyond using only spatial and spectral information. The proposed methods correctly allocate up to 65% more events than the corresponding algorithms that ignore event arrival time information. We apply our methods to two stellar X-ray binaries, UV Cet and HBC515 A, observed with Chandra. We demonstrate that our methods are capable of removing the contamination due to a strong flare on UV Cet B in its companion approximately 40 times weaker during that event, and that evidence for spectral variability at timescales of a few ks can be determined in HBC515 Aa and HBC515 Ab.

astro-ph.IM

Toward Extremely Precise Radial Velocities: II. A Tool For Using Multivariate Gaussian Processes to Model Stellar Activity

The radial velocity method is one of the most successful techniques for the discovery and characterization of exoplanets. Modern spectrographs promise measurement precision of ~0.2-0.5 m/s for an ideal target star. However, the intrinsic variability of stellar spectra can mimic and obscure true planet signals at these levels. Rajpaul et al. (2015) and Jones et al. (2017) proposed applying a physically motivated, multivariate Gaussian process (GP) to jointly model the apparent Doppler shift and multiple indicators of stellar activity as a function of time, so as to separate the planetary signal from various forms of stellar variability. These methods are promising, but performing the necessary calculations can be computationally intensive and algebraically tedious. In this work, we present a flexible and computationally efficient software package, GPLinearOdeMaker.jl, for modeling multivariate time series using a linear combination of univariate GPs and their derivatives. The package allows users to easily and efficiently apply various multivariate GP models and different covariance kernel functions. We demonstrate GPLinearOdeMaker.jl by applying the Jones et al. (2017) model to fit measurements of the apparent Doppler shift and activity indicators derived from simulated active solar spectra time series affected by many evolving star spots. We show how GPLinearOdeMaker.jl makes it easy to explore the effect of different choices for the GP kernel. We find that local kernels could significantly increase the sensitivity and precision of Doppler planet searches relative to the widely used quasi-periodic kernel.

astro-ph.IM

Improving Exoplanet Detection Power: Multivariate Gaussian Process Models for Stellar Activity

The radial velocity method is one of the most successful techniques for detecting exoplanets. It works by detecting the velocity of a host star induced by the gravitational effect of an orbiting planet, specifically the velocity along our line of sight, which is called the radial velocity of the star. Low-mass planets typically cause their host star to move with radial velocities of 1 m/s or less. By analyzing a time series of stellar spectra from a host star, modern astronomical instruments can in theory detect such planets. However, in practice, intrinsic stellar variability (e.g., star spots, convective motion, pulsations) affects the spectra and often mimics a radial velocity signal. This signal contamination makes it difficult to reliably detect low-mass planets. A principled approach to recovering planet radial velocity signals in the presence of stellar activity was proposed by Rajpaul et al. (2015). It uses a multivariate Gaussian process model to jointly capture time series of the apparent radial velocity and multiple indicators of stellar activity. We build on this work in two ways: (i) we propose using dimension reduction techniques to construct new high-information stellar activity indicators; and (ii) we extend the Rajpaul et al. (2015) model to a larger class of models and use a power-based model comparison procedure to select the best model. Despite significant interest in exoplanets, previous efforts have not performed large-scale stellar activity model selection or attempted to evaluate models based on planet detection power. In the case of main sequence G2V stars, we find that our method substantially improves planet detection power compared to previous state-of-the-art approaches.

astro-ph.IM

Designing Test Information and Test Information in Design

DeGroot (1962) developed a general framework for constructing Bayesian measures of the expected information that an experiment will provide for estimation. We propose an analogous framework for measures of information for hypothesis testing. In contrast to estimation information measures that are typically used for surface estimation, test information measures are more useful in experimental design for hypothesis testing and model selection. In particular, we obtain a probability based measure, which has more appealing properties than variance based measures in design contexts where decision problems are of interest. The underlying intuition of our design proposals is straightforward: to distinguish between models we should collect data from regions of the covariate space for which the models differ most. Nicolae et al. (2008) gave an asymptotic equivalence between their test information measures and Fisher information. We extend this result to all test information measures under our framework. Simulation studies and an application in astronomy demonstrate the utility of our approach, and provide comparison to other methods including that of Box and Hill (1967).

math.ST

Warp Bridge Sampling: The Next Generation

Bridge sampling is an effective Monte Carlo method for estimating the ratio of normalizing constants of two probability densities, a routine computational problem in statistics, physics, chemistry, and other fields. The Monte Carlo error of the bridge sampling estimator is determined by the amount of overlap between the two densities. In the case of uni-modal densities, Warp-I, II, and III transformations (Meng and Schilling, 2002) are effective for increasing the initial overlap, but they are less so for multi-modal densities. This paper introduces Warp-U transformations that aim to transform multi-modal densities into Uni-modal ones without altering their normalizing constants. The construction of a Warp-U transformation starts with a Normal (or other convenient) mixture distribution $ϕ_{\text{mix}}$ that has reasonable overlap with the target density $p$, whose normalizing constant is unknown. The stochastic transformation that maps $ϕ_{\text{mix}}$ back to its generating distribution $N(0,1)$ is then applied to $p$ yielding its Warp-U version, which we denote $\tilde{p}$. Typically, $\tilde{p}$ is uni-modal and has substantially increased overlap with $N(0,1)$. Furthermore, we prove that the overlap between $\tilde{p}$ and $N(0,1)$ is guaranteed to be no less than the overlap between $p$ and $ϕ_{\text{mix}}$, in terms of any $f$-divergence. We propose a computationally efficient method to find an appropriate $ϕ_{\text{mix}}$, and a simple but effective approach to remove the bias which results from estimating the normalizing constants and fitting $ϕ_{\text{mix}}$ with the same data. We illustrate our findings using 10 and 50 dimensional highly irregular multi-modal densities, and demonstrate how Warp-U sampling can be used to improve the final estimation step of the Generalized Wang-Landau algorithm (Liang, 2005), a powerful sampling and estimation method.

stat.ME

The AMIGA sample of isolated galaxies XIII. The HI content of an almost "nurture free" sample

We present the largest catalogue of HI single dish observations of isolated galaxies to date and the corresponding HI scaling relations, as part of the multi-wavelength project AMIGA (Analysis of the interstellar Medium in Isolated GAlaxies). Despite numerous studies of the HI content of galaxies, no revision has been made for the most isolated L* galaxies since 1984. In total we have measurements or constraints on the HI masses of 844 galaxies from the Catalogue of Isolated Galaxies (CIG), obtained with our own observations at Arecibo, Effelsberg, Nancay and GBT, and spectra from the literature. Cuts are made to this sample to ensure isolation and a high level of completeness. We then fit HI scaling relations based on luminosity, optical diameter and morphology. Our regression model incorporates all the data, including upper limits, and accounts for uncertainties in both variables, as well as distance uncertainties. The scaling relation of HI mass with optical diameter is in good agreement with that of Haynes & Giovanelli 1984, but our relation with luminosity is considerably steeper. This is attributed to the large uncertainties in the luminosities, which introduce a bias when using OLS regression (used previously), and the different morphology distributions of the samples. We find that the main effect of morphology on the relations is to increase the intercept and flatten the slope towards later types. These trends were not evident in previous works due to the small number of detected early-type galaxies. The HI scaling relations of the AMIGA sample define an up-to-date metric of the HI content of almost "nurture free" galaxies. These relations allow the expected HI mass, in the absence of interactions, of a galaxy to be predicted to within 0.25 dex, and are thus suitable for use as statistical measures of the impact of interactions on the neutral gas content of galaxies. (Abridged)

astro-ph.GA

Disentangling Overlapping Astronomical Sources using Spatial and Spectral Information

We present a powerful new algorithm that combines both spatial information (event locations and the point spread function) and spectral information (photon energies) to separate photons from overlapping sources. We use Bayesian statistical methods to simultaneously infer the number of overlapping sources, to probabilistically separate the photons among the sources, and to fit the parameters describing the individual sources. Using the Bayesian joint posterior distribution, we are able to coherently quantify the uncertainties associated with all these parameters. The advantages of combining spatial and spectral information are demonstrated through a simulation study. The utility of the approach is then illustrated by analysis of observations of FK Aqr and FL Aqr with the XMM-Newton Observatory and the central region of the Orion Nebula Cluster with the Chandra X-ray Observatory.

astro-ph.IM