SearcharxivSearch

arXiv subjects

David A. Campbell

Publications and source records attributed to David A. Campbell.

9 recordsLinked to original sources

Testing Hypotheses of Covariate Effects on Topics of Discourse

We introduce an approach to topic modelling with document-level covariates that remains tractable in the face of large text corpora. This is achieved by de-emphasizing the role of parameter estimation in an underlying probabilistic model, assuming instead that the data come from a fixed but unknown distribution whose statistical functionals are of interest. We propose combining a convex formulation of non-negative matrix factorization with standard regression techniques as a fast-to-compute and useful estimate of such a functional. Uncertainty quantification can then be achieved by reposing non-parametric resampling methods on top of this scheme. This is in contrast to popular topic modelling paradigms, which posit a complex and often hard-to-fit generative model of the data. We argue that the simple, non-parametric approach advocated here is faster, more interpretable, and enjoys better inferential justification than said generative models. Finally, our methods are demonstrated with an application analysing covariate effects on discourse of flavours attributed to Canadian beers.

stat.ME

A survey for variable young stars with small telescopes: VIII -- Properties of 1687 Gaia selected members in 21 nearby clusters

The Hunting Outbursting Young Stars (HOYS) project performs long-term, optical, multi-filter, high cadence monitoring of 25 nearby young clusters and star forming regions. Utilising Gaia DR3 data we have identified about 17000 potential young stellar members in 45 coherent astrometric groups in these fields. Twenty one of them are clear young groups or clusters of stars within one kiloparsec and they contain 9143 Gaia selected potential members. The cluster distances, proper motions and membership numbers are determined. We analyse long term (about 7yr) V, R, and I-band light curves from HOYS for 1687 of the potential cluster members. One quarter of the stars are variable in all three optical filters, and two thirds of these have light curves that are symmetric around the mean. Light curves affected by obscuration from circumstellar materials are more common than those affected by accretion bursts, by a factor of 2-4. The variability fraction in the clusters ranges from 10 to almost 100 percent, and correlates positively with the fraction of stars with detectable inner disks, indicating that a lot of variability is driven by the disk. About one in six variables shows detectable periodicity, mostly caused by magnetic spots. Two thirds of the periodic variables with disk excess emission are slow rotators, and amongst the stars without disk excess two thirds are fast rotators - in agreement with rotation being slowed down by the presence of a disk.

astro-ph.SR

A survey for variable young stars with small telescopes: VI -- Analysis of the outbursting Be stars NSW284, Gaia19eyy, and VES263

This paper is one in a series reporting results from small telescope observations of variable young stars. Here, we study the repeating outbursts of three likely Be stars based on long-term optical, near-infrared, and mid-infrared photometry for all three objects, along with follow-up spectra for two of the three. The sources are characterised as rare, truly regularly outbursting Be stars. We interpret the photometric data within a framework for modelling light curve morphology, and find that the models correctly predict the burst shapes, including their larger amplitudes and later peaks towards longer wavelengths. We are thus able to infer the start and end times of mass loading into the circumstellar disks of these stars. The disk sizes are typically 3-6 times the areas of the central star. The disk temperatures are ~40%, and the disk luminosities are ~10% of those of the central Be star, respectively. The available spectroscopy is consistent with inside-out evolution of the disk. Higher excitation lines have larger velocity widths in their double-horned shaped emission profiles. Our observations and analysis support the decretion disk model for outbursting Be stars.

astro-ph.SR

Parallel Tempering via Simulated Tempering Without Normalizing Constants

In this paper we develop a new general Bayesian methodology that simultaneously estimates parameters of interest and the marginal likelihood of the model. The proposed methodology builds on Simulated Tempering, which is a powerful algorithm that enables sampling from multi-modal distributions. However, Simulated Tempering comes with the practical limitation of needing to specify a prior for the temperature along a chosen discretization schedule that will allow calculation of normalizing constants at each temperature. Our proposed model defines the prior for the temperature so as to remove the need for calculating normalizing constants at each temperature and thereby enables a continuous temperature schedule, while preserving the sampling efficiency of the Simulated Tempering algorithm. The resulting algorithm simultaneously estimates parameters while estimating marginal likelihoods through thermodynamic integration. We illustrate the applicability of the new algorithm to different examples involving mixture models of Gaussian distributions and ordinary differential equation models.

stat.CO

Incremental Mixture Importance Sampling with Shotgun optimization

This paper proposes a general optimization strategy, which combines results from different optimization or parameter estimation methods to overcome shortcomings of a single method. Shotgun optimization is developed as a framework which employs different optimization strategies, criteria, or conditional targets to enable wider likelihood exploration. The introduced Shotgun optimization approach is embedded into an incremental mixture importance sampling algorithm to produce improved posterior samples for multimodal densities and creates robustness in cases where the likelihood and prior are in disagreement. Despite using different optimization approaches, the samples are combined into samples from a single target posterior. The diversity of the framework is demonstrated on parameter estimation from differential equation models employing diverse strategies including numerical solutions and approximations thereof. Additionally the approach is demonstrated on mixtures of discrete and continuous parameters and is shown to ease estimation from synthetic likelihood models. R code of the implemented examples is stored in a zipped archive (codeSubmit.zip).

stat.CO

Bayesian Solution Uncertainty Quantification for Differential Equations

We explore probability modelling of discretization uncertainty for system states defined implicitly by ordinary or partial differential equations. Accounting for this uncertainty can avoid posterior under-coverage when likelihoods are constructed from a coarsely discretized approximation to system equations. A formalism is proposed for inferring a fixed but a priori unknown model trajectory through Bayesian updating of a prior process conditional on model information. A one-step-ahead sampling scheme for interrogating the model is described, its consistency and first order convergence properties are proved, and its computational complexity is shown to be proportional to that of numerical explicit one-step solvers. Examples illustrate the flexibility of this framework to deal with a wide variety of complex and large-scale systems. Within the calibration problem, discretization uncertainty defines a layer in the Bayesian hierarchy, and a Markov chain Monte Carlo algorithm that targets this posterior distribution is presented. This formalism is used for inference on the JAK-STAT delay differential equation model of protein dynamics from indirectly observed measurements. The discussion outlines implications for the new field of probabilistic numerics.

stat.ME

Sequentially Constrained Monte Carlo

Constraints can be interpreted in a broad sense as any kind of explicit restriction over the parameters. While some constraints are defined directly on the parameter space, when they are instead defined by known behaviour on the model, transformation of constraints into features on the parameter space may not be possible. Difficulties in sampling from the posterior distribution as a result of incorporation of constraints into the model is a common challenge leading to truncations in the parameter space and inefficient sampling algorithms. We propose a variant of sequential Monte Carlo algorithm for posterior sampling in presence of constraints by defining a sequence of densities through the imposition of the constraint. Particles generated from an unconstrained or mildly constrained distribution are filtered and moved through sampling and resampling steps to obtain a sample from the fully constrained target distribution. General and model specific forms of constraints enforcing strategies are defined. The Sequentially Constrained Monte Carlo algorithm is demonstrated on constraints defined by monotonicity of a function, densities constrained to low dimensional manifolds, adherence to a theoretically derived model, and model feature matching.

stat.ME

Transdimensional Approximate Bayesian Computation for Inference on Invasive Species Models with Latent Variables of Unknown Dimension

Accurate information on patterns of introduction and spread of non-native species is essential for making predictions and management decisions. In many cases, estimating unknown rates of introduction and spread from observed data requires evaluating intractable variable-dimensional integrals. In general, inference on the large class of models containing latent variables of large or variable dimension precludes exact sampling techniques. Approximate Bayesian computation (ABC) methods provide an alternative to exact sampling but rely on inefficient conditional simulation of the latent variables. To accomplish this task efficiently, a new transdimensional Monte Carlo sampler is developed for approximate Bayesian model inference and used to estimate rates of introduction and spread for the non-native earthworm species Dendrobaena octaedra (Savigny) along roads in the boreal forest of northern Alberta. Using low and high estimates of introduction and spread rates, the extent of earthworm invasions in northeastern Alberta was simulated to project the proportion of suitable habitat invaded in the year following data collection.

stat.ME

Monotone Function Estimation for Computer Experiments

In statistical modeling of computer experiments sometimes prior information is available about the underlying function. For example, the physical system simulated by the computer code may be known to be monotone with respect to some or all inputs. We develop a Bayesian approach to Gaussian process modelling capable of incorporating monotonicity information for computer model emulation. Markov chain Monte Carlo methods are used to sample from the posterior distribution of the process given the simulator output and monotonicity information. The performance of the proposed approach in terms of predictive accuracy and uncertainty quantification is demonstrated in a number of simulated examples as well as a real queueing system application.

stat.ME