SearcharxivSearch

arXiv subjects

Ewan Cameron

Publications and source records attributed to Ewan Cameron.

At least 19 recordsLinked to original sources

'The Order in the Horse's Heart': A Case Study in LLM-Assisted Stylometry for the Discovery of Biblical Allusion in Modern Literary Fiction

We present a dual-track pipeline for detecting biblical allusions in literary fiction and apply it to the novels of Cormac McCarthy. A bottom-up embedding track uses inverse document frequency to identify rare vocabulary shared with the King James Bible, embeds occurrences in their local context for sense disambiguation, and passes candidate passage pairs through cascaded LLM review. A top-down register track asks an LLM to read McCarthy's prose undirected to any specific biblical passage for comparison, catching allusions not distinguished by word or phrase rarity. Both tracks are cross-validated by a long-context model that holds entire novels alongside the KJV in a single pass, and every finding is checked against published scholarship. Restricting attention to allusions that carry a textual echo--shared phrasing, reworked vocabulary, or transplanted cadence--and distinguishing literary allusions proper from signposted biblical references (similes naming biblical figures, characters overtly citing scripture), the pipeline surfaces 349 allusions across the corpus. Among a target set of 115 previously documented allusions retrieved through human review of the academic literature, the pipeline independently recovers 62 (54% recall), with recall varying by connection type from 30% (transformed imagery) to 80% (register collisions). We contextualise these results with respect to the value-add from LLMs as assistants to mechanical stylometric analyses, and their potential to facilitate the statistical study of intertextuality in massive literary corpora.

cs.CL

Statistical modelling under differential privacy constraints: A case study in fine-scale geographical analysis with Australian Bureau of Statistics TableBuilder data

Guided by the principles of differential privacy protection the Australian Bureau of Statistics modifies the data summaries from the Australian Census provided through TableBuilder to researchers at approved institutions. This modification algorithm includes the injection of a small degree of artificial noise to every nonzero cell count followed by the suppression of very small cell counts to zero. Researchers working with small area TableBuilder outputs with a high suppression fraction have proposed various algorithmic solutions to reconciling these with less suppressed outputs from larger enclosing areas. Here we propose that a Bayesian, likelihood-based statistical approach in which the perturbation algorithm itself is explicitly represented is well suited to analyses with such randomly perturbed data. Using both real (TableBuilder) and mock datasets representing dwelling classifications in the Perth Greater Capital City Area we demonstrate the feasibility and utility of multi-scale Bayesian reconstruction of modified cell counts in a spatial setting.

stat.ME

nazgul: A statistical approach to gamma-ray burst localization. Triangulation via non-stationary time-series models

Context. Gamma-ray bursts can be located via arrival time signal triangulation using gamma-ray detectors in orbit throughout the solar system. The classical approach based on cross-correlations of binned light curves ignores the Poisson nature of the time-series data, and is unable to model the full complexity of the problem. Aims. To present a statistically proper and robust GRB timing/triangulation algorithm as a modern update to the original procedures used for the Interplanetary Network (IPN). Methods. A hierarchical Bayesian forward model for the unknown temporal signal evolution is learned via random Fourier features (RFF) and fitted to each detector's time-series data with time-differences that correspond to GRB's position on the sky via the appropriate Poisson likelihood. Results. Our novel method can robustly estimate the position of a GRB as verified via simulations. The uncertainties generated by the method are robust and in many cases more precise compared to the classical method. Thus, we have a method that can become a valuable tool for gravitational wave follow-up. All software and analysis scripts are made publicly available here (https://github.com/grburgess/nazgul) for the purpose of replication.

astro-ph.IM

Spatiotemporal mapping of malaria prevalence in Madagascar using routine surveillance and health survey data

Malaria transmission in Madagascar is highly heterogeneous, exhibiting spatial, seasonal and long-term trends. Previous efforts to map malaria risk in Madagascar used prevalence data from Malaria Indicator Surveys. These cross-sectional surveys, conducted during the high transmission season most recently in 2013 and 2016, provide nationally representative prevalence data but cover relatively short time frames. Conversely, monthly case data are collected at health facilities but suffer from biases, including incomplete reporting. We combined survey and case data to make monthly maps of prevalence between 2013 and 2016. Health facility catchments were estimated and incidence surfaces, environmental and socioeconomic covariates, and survey data informed a Bayesian prevalence model. Prevalence estimates were consistently high in the coastal regions and low in the highlands. Prevalence was lowest in 2014 and peaked in 2015, highlighting the importance of estimates between survey years. Seasonality was widely observed. Similar multi-metric approaches may be applicable across sub-Saharan Africa.

stat.AP

A simulation study of disaggregation regression for spatial disease mapping

Disaggregation regression has become an important tool in spatial disease mapping for making fine-scale predictions of disease risk from aggregated response data. By including high resolution covariate information and modelling the data generating process on a fine scale, it is hoped that these models can accurately learn the relationships between covariates and response at a fine spatial scale. However, validating these high resolution predictions can be a challenge, as often there is no data observed at this spatial scale. In this study, disaggregation regression was performed on simulated data in various settings and the resulting fine-scale predictions are compared to the simulated ground truth. Performance was investigated with varying numbers of data points, sizes of aggregated areas and levels of model misspecification. The effectiveness of cross validation on the aggregate level as a measure of fine-scale predictive performance was also investigated. Predictive performance improved as the number of observations increased and as the size of the aggregated areas decreased. When the model was well-specified, fine-scale predictions were accurate even with small numbers of observations and large aggregated areas. Under model misspecification predictive performance was significantly worse for large aggregated areas but remained high when response data was aggregated over smaller regions. Cross-validation correlation on the aggregate level was a moderately good predictor of fine-scale predictive performance. While the simulations are unlikely to capture the nuances of real-life response data, this study gives insight into the effectiveness of disaggregation regression in different contexts.

stat.AP

Nonparametric Causal Feature Selection for Spatiotemporal Risk Mapping of Malaria Incidence in Madagascar

Modern disease mapping draws upon a wealth of high resolution spatial data products reflecting environmental and/or socioeconomic factors as covariates, or `features', within a geostatistical framework to improve predictions of disease risk. Feature selection is an important step in building these models, helping to reduce overfitting and computational complexity, and to improve model interpretability. Selecting only features that have a causal relationship with the response variable could potentially improve predictions and generalisability, but identifying these causal features from non-interventional, spatiotemporal data is a challenging problem. Here we examine the performance of a causal feature selection procedure with regard to estimating malaria incidence in Madagascar. The studied procedure designed for this task combines the PC algorithm with spatiotemporal prewhitening and kernel-based independence tests extended to accommodate aggregated data. This case study reveals a clear advantage for causal feature selection in terms of the out-of-sample predictive accuracy in a forward temporal estimation task, but not in a spatiotemporal interpolation task, in comparison with thresholded spike-and-slab, for both linear and non-linear regression models. Compared to no feature selection, causal feature selection was most beneficial in settings wherein the volume of available data was low relative to the model complexity.

stat.AP

Black Hole Mass Scaling Relations for Spiral Galaxies. II. $M_{\rm BH}$-$M_{\rm *,tot}$ and $M_{\rm BH}$-$M_{\rm *,disk}$

Black hole mass ($M_{BH}$) scaling relations are typically derived using the properties of a galaxy's bulge and samples dominated by (high-mass) early-type galaxies. Studying late-type galaxies should provide greater insight into the mutual growth of black holes and galaxies in more gas-rich environments. We have used 40 spiral galaxies to establish how $M_{BH}$ scales with both the total stellar mass ($M_{*,tot}$) and the disk's stellar mass, having measured the spheroid (bulge) stellar mass ($M_{*,sph}$) and presented the $M_{BH}$-$M_{*,sph}$ relation in Paper I. The relation involving $M_{*,tot}$ may be beneficial for estimating $M_{BH}$ either from pipeline data or at higher redshift, conditions that are not ideal for the accurate isolation of the bulge. A symmetric Bayesian analysis finds $\log\left(M_{BH}/M_{\odot}\right)=\left(3.05_{-0.49}^{+0.57}\right)\log\left\{M_{*,tot}/[\upsilon(6.37\times10^{10}\,M_{\odot})]\right\}+(7.25_{-0.14}^{+0.13})$. The scatter from the regression of $M_{BH}$ on $M_{*,tot}$ is 0.66 dex; compare 0.56 dex for $M_{BH}$ on $M_{*,sph}$ and $0.57$ dex for $M_{BH}$ on $σ_*$. The slope is $>2$ times that obtained using core-Sérsic early-type galaxies, echoing a similar result involving $M_{*,sph}$, and supporting a varied growth mechanism among different morphological types. This steeper relation has consequences for galaxy/black hole formation theories, simulations, and predicting black hole masses. We caution that (i) an $M_{BH}$-$M_{*,tot}$ relation built from a mixture of early- and late-type galaxies will find an arbitrary slope of approximately 1-3, with no physical meaning beyond one's sample selection, and (ii) evolutionary studies of the $M_{BH}$-$M_{*,tot}$ relation need to be mindful of the galaxy types included at each epoch. We additionally update the $M_{*,tot}$-($\textit{face-on}$ spiral arm pitch angle) relation.

astro-ph.GA

Black Hole Mass Scaling Relations for Spiral Galaxies. I. $M_{\rm BH}$-$M_{\rm *,sph}$

The (supermassive black hole mass, $M_\text{BH}$)-(bulge stellar mass, $M_{\rm*,sph}$) relation is, obviously, derived using two quantities. We endeavor to provide accurate values for the latter via detailed multicomponent galaxy decompositions for the current full sample of 43 spiral galaxies having directly measured $M_\text{BH}$ values; 35 of these galaxies have been alleged to contain pseudobulges, 21 have water maser measurements, and three appear bulgeless. This more than doubles the previous sample size of spiral galaxies with a finessed image analysis. We have analyzed near-infrared images, accounting for not only the bulge, disk (exponential, truncated, or inclined), and bar but also for spiral arms and rings and additional central components (active galactic nuclei (AGNs), etc.). A symmetric Bayesian analysis finds $\log\left(M_\text{BH}/M_{\odot}\right)=\left(2.44_{-0.31}^{+0.35}\right)\log\left\{M_{\rm*,sph}/[\upsilon(1.15\times10^{10}\,M_{\odot})]\right\}+(7.24\pm0.12)$, with $\upsilon$ a stellar mass-to-light ratio term. The level of scatter equals that about the $M_{\rm BH}$-$σ_*$ relation. The nonlinear slope rules out the idea that many mergers, coupled with the central limit theorem, produced this scaling relation, and it corroborates previous observational studies and simulations, which have reported a near-quadratic slope at the low-mass end of the $M_\text{BH}$-$M_{\rm*,sph}$ diagram. Furthermore, bulges with AGNs follow this relation; they are not offset by an order of magnitude, and models that have invoked AGN feedback to establish a linear $M_{\rm BH}$-$M_{\rm*,sph}$ relation need revisiting. We additionally present an updated $M_\text{BH}$-(Sérsic index, $n_\text{sph}$) relation for spiral galaxy bulges with a comparable level of scatter and a new $M_{\rm*,sph}$-(spiral-arm pitch angle, $ϕ$) relation.

astro-ph.GA

Mapping malaria seasonality: a case study from Madagascar

Many malaria-endemic areas experience seasonal fluctuations in case incidence as Anopheles mosquito and Plasmodium parasite life cycles respond to changing environmental conditions. While most existing maps of malaria seasonality use fixed thresholds of rainfall, temperature, and/or vegetation indices to identify suitable transmission months, we develop a statistical modelling framework for characterising the seasonal patterns derived directly from case data. The procedure involves a spatiotemporal regression model for estimating the monthly proportions of total annual cases and an algorithm to identify operationally relevant characteristics such as the transmission start and peak months. A seasonality index combines the monthly proportion estimates and existing estimates of annual case incidence to provide a summary of "how seasonal" locations are relative to their surroundings. An advancement upon past seasonality mapping endeavours is the presentation of the uncertainty associated with each map, which will enable policymakers to make more statistically sound decisions. The methodology is illustrated using health facility data from Madagascar.

stat.AP

Variational Learning on Aggregate Outputs with Gaussian Processes

While a typical supervised learning framework assumes that the inputs and the outputs are measured at the same levels of granularity, many applications, including global mapping of disease, only have access to outputs at a much coarser level than that of the inputs. Aggregation of outputs makes generalization to new inputs much more difficult. We consider an approach to this problem based on variational learning with a model of output aggregation and Gaussian processes, where aggregation leads to intractability of the standard evidence lower bounds. We propose new bounds and tractable approximations, leading to improved prediction accuracy and scalability to large datasets, while explicitly taking uncertainty into account. We develop a framework which extends to several types of likelihoods, including the Poisson model for aggregated count data. We apply our framework to a challenging and important problem, the fine-scale spatial modelling of malaria incidence, with over 1 million observations.

stat.ML

Improved prediction accuracy for disease risk mapping using Gaussian Process stacked generalisation

Maps of infectious disease---charting spatial variations in the force of infection, degree of endemicity, and the burden on human health---provide an essential evidence base to support planning towards global health targets. Contemporary disease mapping efforts have embraced statistical modelling approaches to properly acknowledge uncertainties in both the available measurements and their spatial interpolation. The most common such approach is that of Gaussian process regression, a mathematical framework comprised of two components: a mean function harnessing the predictive power of multiple independent variables, and a covariance function yielding spatio-temporal shrinkage against residual variation from the mean. Though many techniques have been developed to improve the flexibility and fitting of the covariance function, models for the mean function have typically been restricted to simple linear terms. For infectious diseases, known to be driven by complex interactions between environmental and socio-economic factors, improved modelling of the mean function can greatly boost predictive power. Here we present an ensemble approach based on stacked generalisation that allows for multiple, non-linear algorithmic mean functions to be jointly embedded within the Gaussian process framework. We apply this method to mapping Plasmodium falciparum prevalence data in Sub-Saharan Africa and show that the generalised ensemble approach markedly out-performs any individual method.

stat.AP

The star cluster mass--galactocentric radius relation: Implications for cluster formation

Whether or not the initial star cluster mass function is established through a universal, galactocentric-distance-independent stochastic process, on the scales of individual galaxies, remains an unsolved problem. This debate has recently gained new impetus through the publication of a study that concluded that the maximum cluster mass in a given population is not solely determined by size-of-sample effects. Here, we revisit the evidence in favor and against stochastic cluster formation by examining the young ($\lesssim$ a few $\times 10^8$ yr-old) star cluster mass--galactocentric radius relation in M33, M51, M83, and the Large Magellanic Cloud. To eliminate size-of-sample effects, we first adopt radial bin sizes containing constant numbers of clusters, which we use to quantify the radial distribution of the first- to fifth-ranked most massive clusters using ordinary least-squares fitting. We supplement this analysis with an application of quantile regression, a binless approach to rank-based regression taking an absolute-value-distance penalty. Both methods yield, within the $1σ$ to $3σ$ uncertainties, near-zero slopes in the diagnostic plane, largely irrespective of the maximum age or minimum mass imposed on our sample selection, or of the radial bin size adopted. We conclude that, at least in our four well-studied sample galaxies, star cluster formation does not necessarily require an environment-dependent cluster formation scenario, which thus supports the notion of stochastic star cluster formation as the dominant star cluster-formation process within a given galaxy.

astro-ph.GA

Recursive Pathways to Marginal Likelihood Estimation with Prior-Sensitivity Analysis

We investigate the utility to computational Bayesian analyses of a particular family of recursive marginal likelihood estimators characterized by the (equivalent) algorithms known as "biased sampling" or "reverse logistic regression" in the statistics literature and "the density of states" in physics. Through a pair of numerical examples (including mixture modeling of the well-known galaxy data set) we highlight the remarkable diversity of sampling schemes amenable to such recursive normalization, as well as the notable efficiency of the resulting pseudo-mixture distributions for gauging prior sensitivity in the Bayesian model selection context. Our key theoretical contributions are to introduce a novel heuristic ("thermodynamic integration via importance sampling") for qualifying the role of the bridging sequence in this procedure and to reveal various connections between these recursive estimators and the nested sampling technique.

stat.ME

What we talk about when we talk about fields

In astronomical and cosmological studies one often wishes to infer some properties of an infinite-dimensional field indexed within a finite-dimensional metric space given only a finite collection of noisy observational data. Bayesian inference offers an increasingly-popular strategy to overcome the inherent ill-posedness of this signal reconstruction challenge. However, there remains a great deal of confusion within the astronomical community regarding the appropriate mathematical devices for framing such analyses and the diversity of available computational procedures for recovering posterior functionals. In this brief research note I will attempt to clarify both these issues from an "applied statistics" perpective, with insights garnered from my post-astronomy experiences as a computational Bayesian / epidemiological geostatistician.

astro-ph.IM

A Generalized Savage-Dickey Ratio

In this brief research note I present a generalized version of the Savage-Dickey Density Ratio for representation of the Bayes factor (or marginal likelihood ratio) of nested statistical models; the new version takes the form of a Radon-Nikodym derivative and is thus applicable to a wider family of probability spaces than the original (restricted to those admitting an ordinary Lebesgue density). A derivation is given following the measure-theoretic construction of Marin & Robert (2010), and the equivalent estimator is demonstrated in application to a distributional modeling problem.

stat.ME

On the Evidence for Cosmic Variation of the Fine Structure Constant (I): A Parametric Bayesian Model Selection Analysis of the Quasar Dataset

We review the evidence behind recent claims of spatial variation in the fine structure constant deriving from observations of ionic absorption lines in the light from distant quasars. To this end we expand upon previous non-Bayesian analyses limited by the assumptions of an unbiased and strictly Normal distribution for the "unexplained errors" of the benchmark quasar dataset. Through the technique of reverse logistic regression we estimate and compare marginal likelihoods for three competing hypotheses---(i) the null hypothesis (no cosmic variation), (ii) the monopole hypothesis (a constant Earth-to-quasar offset), and (iii) the monopole+dipole hypothesis (a cosmic variation manifest to the Earth-bound observer as a North-South divergence)---under a variety of candidate parametric forms for the unexplained error term. Our analysis reveals weak support for a skeptical interpretation in which the apparent dipole effect is driven solely by systematic errors of opposing sign inherent in measurements from the two telescopes employed to obtain these observations. Throughout we seek to exemplify a 'best practice' approach to Bayesian model selection with prior-sensitivity analysis; in a companion paper we extend this methodology to a semi-parametric framework using the infinite-dimensional Dirichlet process.

astro-ph.CO

On the Evidence for Cosmic Variation of the Fine Structure Constant (II): A Semi-Parametric Bayesian Model Selection Analysis of the Quasar Dataset

In the second paper of this series we extend our Bayesian reanalysis of the evidence for a cosmic variation of the fine structure constant to the semi-parametric modelling regime. By adopting a mixture of Dirichlet processes prior for the unexplained errors in each instrumental subgroup of the benchmark quasar dataset we go some way towards freeing our model selection procedure from the apparent subjectivity of a fixed distributional form. Despite the infinite-dimensional domain of the error hierarchy so constructed we are able to demonstrate a recursive scheme for marginal likelihood estimation with prior-sensitivity analysis directly analogous to that presented in Paper I, thereby allowing the robustness of our posterior Bayes factors to hyper-parameter choice and model specification to be readily verified. In the course of this work we elucidate various similarities between unexplained error problems in the seemingly disparate fields of astronomy and clinical meta-analysis, and we highlight a number of sophisticated techniques for handling such problems made available by past research in the latter. It is our hope that the novel approach to semi-parametric model selection demonstrated herein may serve as a useful reference for others exploring this potentially difficult class of error model.

astro-ph.IM

The near-IR $M_{bh}$ - L and $M_{bh}$ - n relations

We present near-IR surface photometry (2D-profiling) for a sample of 29 nearby galaxies for which super-massive black hole (SMBH) masses are constrained. The data is derived from the UKIDSS-LASS survey representing a significant improvement in image quality and depth over previous studies based on 2MASS data. We derive the spheroid luminosity and spheroid Sérsic index for each galaxy with GALFIT3 and use these data to construct SMBH mass -bulge luminosity ($M_{\rm bh}$--$L$) and SMBH - Sérsic index ($M_{\rm bh}$--$n$) relations. The best fit K-band relation for elliptical and disk galaxies is $\log(M_{\rm bh}/M_{\odot})= -0.36(\pm 0.03) (M_{\rm K} + 18) + 6.17(\pm 0.16)$ with an intrinsic scatter of 0.4$^{+0.09}_{-0.06}$dex whilst for elliptical galaxies we find $\log(M_{\rm bh}/M_{\odot})= -0.42(\pm 0.06) (M_{\rm K} + 22) + 7.5(\pm 0.15)$ with an intrinsic scatter of 0.31$^{+0.087}_{-0.047}$dex. Our revised $M_{\rm bh}$--$L$ relation agrees closely with the previous near-IR constraint by \citet{tex:G07}. The lack of improvement in the intrinsic scatter in moving to higher quality near-IR data suggests that the SMBH relations are not currently limited by the quality of the imaging data but is either intrinsic or a result of uncertainty in the precise number of required components required in the profiling process. Contrary to expectation (see \citealt{tex:GD07a}) a relation between SMBH mass and the Sérsic index was not found at near-IR wavelengths. This latter outcome is believed to be explained by the generic inconsistencies between 1D and 2D galaxy profiling which are currently under further investigation.

astro-ph.CO