Searcharxiv⌕ Search

arXiv subjects

Richard S. Savage

Publications and source records attributed to Richard S. Savage.

12 recordsLinked to original sources

Non-linear Multitask Learning with Deep Gaussian Processes

We present a multi-task learning formulation for Deep Gaussian processes (DGPs), through non-linear mixtures of latent processes. The latent space is composed of private processes that capture within-task information and shared processes that capture across-task dependencies. We propose two different methods for segmenting the latent space: through hard coding shared and task-specific processes or through soft sharing with Automatic Relevance Determination kernels. We show that our formulation is able to improve the learning performance and transfer information between the tasks, outperforming other probabilistic multi-task learning models across real-world and benchmarking settings.

stat.ML↗

Heterogeneous large datasets integration using Bayesian factor regression

Two key challenges in modern statistical applications are the large amount of information recorded per individual, and that such data are often not collected all at once but in batches. These batch effects can be complex, causing distortions in both mean and variance. We propose a novel sparse latent factor regression model to integrate such heterogeneous data. The model provides a tool for data exploration via dimensionality reduction while correcting for a range of batch effects. We study the use of several sparse priors (local and non-local) to learn the dimension of the latent factors. Our model is fitted in a deterministic fashion by means of an EM algorithm for which we derive closed-form updates, contributing a novel scalable algorithm for non-local priors of interest beyond the immediate scope of this paper. We present several examples, with a focus on bioinformatics applications. Our results show an increase in the accuracy of the dimensionality reduction, with non-local priors substantially improving the reconstruction of factor cardinality, as well as the need to account for batch effects to obtain reliable results. Our model provides a novel approach to latent factor regression that balances sparsity with sensitivity and is highly computationally efficient.

stat.AP↗

Identifying cancer subtypes in glioblastoma by combining genomic, transcriptomic and epigenomic data

We present a nonparametric Bayesian method for disease subtype discovery in multi-dimensional cancer data. Our method can simultaneously analyse a wide range of data types, allowing for both agreement and disagreement between their underlying clustering structure. It includes feature selection and infers the most likely number of disease subtypes, given the data. We apply the method to 277 glioblastoma samples from The Cancer Genome Atlas, for which there are gene expression, copy number variation, methylation and microRNA data. We identify 8 distinct consensus subtypes and study their prognostic value for death, new tumour events, progression and recurrence. The consensus subtypes are prognostic of tumour recurrence (log-rank p-value of $3.6 \times 10^{-4}$ after correction for multiple hypothesis tests). This is driven principally by the methylation data (log-rank p-value of $2.0 \times 10^{-3}$) but the effect is strengthened by the other 3 data types, demonstrating the value of integrating multiple data types. Of particular note is a subtype of 47 patients characterised by very low levels of methylation. This subtype has very low rates of tumour recurrence and no new events in 10 years of follow up. We also identify a small gene expression subtype of 6 patients that shows particularly poor survival outcomes. Additionally, we note a consensus subtype that showly a highly distinctive data signature and suggest that it is therefore a biologically distinct subtype of glioblastoma. The code is available from https://sites.google.com/site/multipledatafusion/

q-bio.GN↗

The Far-Infrared Surveyor (FIS) for AKARI

The Far-Infrared Surveyor (FIS) is one of two focal plane instruments on the AKARI satellite. FIS has four photometric bands at 65, 90, 140, and 160 um, and uses two kinds of array detectors. The FIS arrays and optics are designed to sweep the sky with high spatial resolution and redundancy. The actual scan width is more than eight arcmin, and the pixel pitch is matches the diffraction limit of the telescope. Derived point spread functions (PSFs) from observations of asteroids are similar to the optical model. Significant excesses, however, are clearly seen around tails of the PSFs, whose contributions are about 30% of the total power. All FIS functions are operating well in orbit, and its performance meets the laboratory characterizations, except for the two longer wavelength bands, which are not performing as well as characterized. Furthermore, the FIS has a spectroscopic capability using a Fourier transform spectrometer (FTS). Because the FTS takes advantage of the optics and detectors of the photometer, it can simultaneously make a spectral map. This paper summarizes the in-flight technical and operational performance of the FIS.

astro-ph↗

The Far-Infrared Properties of Spatially Resolved AKARI Observations

We present the spatially resolved observations of IRAS sources from the Japanese infrared astronomy satellite AKARI All-Sky Survey during the performance verification (PV) phase of the mission. We extracted reliable point sources matched with IRAS point source catalogue. By comparing IRAS and AKARI fluxes, we found that the flux measurements of some IRAS sources could have been over or underestimated and affected by the local background rather than the global background. We also found possible candidates for new AKARI sources and confirmed that AKARI observations resolved IRAS sources into multiple sources. All-Sky Survey observations are expected to verify the accuracies of IRAS flux measurements and to find new extragalactic point sources.

astro-ph↗

Bayesian methods of astronomical source extraction

We present two new source extraction methods, based on Bayesian model selection and using the Bayesian Information Criterion (BIC). The first is a source detection filter, able to simultaneously detect point sources and estimate the image background. The second is an advanced photometry technique, which measures the flux, position (to sub-pixel accuracy), local background and point spread function. We apply the source detection filter to simulated Herschel-SPIRE data and show the filter's ability to both detect point sources and also simultaneously estimate the image background. We use the photometry method to analyse a simple simulated image containing a source of unknown flux, position and point spread function; we not only accurately measure these parameters, but also determine their uncertainties (using Markov-Chain Monte Carlo sampling). The method also characterises the nature of the source (distinguishing between a point source and extended source). We demonstrate the effect of including additional prior knowledge. Prior knowledge of the point spread function increase the precision of the flux measurement, while prior knowledge of the background has onlya small impact. In the presence of higher noise levels, we show that prior positional knowledge (such as might arise from a strong detection in another waveband) allows us to accurately measure the source flux even when the source is too faint to be detected directly. These methods are incorporated in SUSSEXtractor, the source extraction pipeline for the forthcoming Akari FIS far-infrared all-sky survey. They are also implemented in a stand-alone, beta-version public tool that can be obtained at http://astronomy.sussex.ac.uk/$\sim$rss23/sourceMiner\_v0.1.2.0.tar.gz

astro-ph↗

Parametric modelling of the 3.6um to 8um colour distributions of galaxies in the SWIRE Survey

We fit a parametric model comprising a mixture of multi-dimensional Gaussian functions to the 3.6 to 8um colour and optical photo-z distribution of galaxy populations in the ELAIS-N1 and Lockman Fields of SWIRE. For 16,698 sources in ELAIS-N1 we find our data are best modelled (in the sense of the Bayesian Information Criterion) by the sum of four Gaussian distributions or modes (C_a, C_b, C_c and C_d). We compare the fit of our empirical model with predictions from existing semi-analytic and phenomological models. We infer that our empirical model provides a better description of the mid-infrared colour distribution of the SWIRE survey than these existing models. This colour distribution test is thus a powerful model discriminator and complementary to comparisons of number counts. We use our model to provide a galaxy classification scheme and explore the nature of the galaxies in the different modes of the model. C_a consists of dusty star-forming systems such as ULIRG's. Low redshift late-type spirals are found in C_b, where PAH emission dominates at 8um. C_c consists of dusty starburst systems at intermediate redshifts. Low redshift early-type spirals and ellipticals dominate C_d. We thus find a greater variety of galaxy types than one can with optical photometry alone. Finally we develop a new technique to identify unusual objects, and find a selection of outliers with very red IRAC colours. These objects are not detected in the optical, but have very strong detections in the mid-infrared. These sources are modelled as dust-enshrouded, strongly obscured AGN, where the high mid-infrared emission may either be attributed to dust heated by the AGN or substantial star-formation. These sources have z_ph ~ 2-4, making them incredibly infrared luminous, with a L_IR ~ 10^(12.6-14.1) L_sun.

astro-ph↗

Non-Gaussianity in the Very Small Array CMB maps with Smooth-Goodness-of-fit tests

(Abridged) We have used the Rayner & Best (1989) smooth tests of goodness-of-fit to study the Gaussianity of the Very Small Array (VSA) data. Out of the 41 published VSA individual pointings dedicated to cosmological observations, 37 are found to be consistent with Gaussianity, whereas four pointings show deviations from Gaussianity. In two of them, these deviations can be explained as residual systematic effects of a few visibility points which, when corrected, have a negligible impact on the angular power spectrum. The non-Gaussianity found in the other two (adjacent) pointings seems to be associated to a local deviation of the power spectrum of these fields with respect to the common power spectrum of the complete data set, at angular scales of the third acoustic peak (l = 700-900). No evidence of residual systematics is found in this case, and unsubstracted point sources are not a plausible explanation either. If those visibilities are removed, a cosmological analysis based on this new VSA power spectrum alone shows no differences in the parameter constraints with respect to our published results, except for the physical baryon density, which decreases by 10 percent. Finally, the method has been also used to analyse the VSA observations in the Corona Borealis supercluster region (Genova-Santos et al. 2005), which show a strong decrement which cannot be explained as primordial CMB. Our method finds a clear deviation (99.82%) with respect to Gaussianity in the second-order moment of the distribution, and which can not be explained as systematic effects. A detailed study shows that the non-Gaussianity is produced in scales of l~500, and that this deviation is intrinsic to the data (in the sense that can not be explained in terms of a Gaussian field with a different power spectrum).

astro-ph↗

Statistical constraints on the IR galaxy number counts and cosmic IR background from the Spitzer GOODS survey

We perform fluctuation analyses on the data from the Spitzer GOODS survey (epoch one) in the Hubble Deep Field North (HDF-N). We fit a parameterised power-law number count model of the form dN/dS = N_o S^{-δ} to data from each of the four Spitzer IRAC bands, using Markov Chain Monte Carlo (MCMC) sampling to explore the posterior probability distribution in each case. We obtain best-fit reduced chi-squared values of (3.43 0.86 1.14 1.13) in the four IRAC bands. From this analysis we determine the likely differential faint source counts down to $10^{-8} Jy$, over two orders of magnitude in flux fainter than has been previously determined. From these constrained number count models, we estimate a lower bound on the contribution to the Infra-Red (IR) background light arising from faint galaxies. We estimate the total integrated background IR light in the Spitzer GOODS HDF-N field due to faint sources. By adding the estimates of integrated light given by Fazio et al (2004), we calculate the total integrated background light in the four IRAC bands. We compare our 3.6 micron results with previous background estimates in similar bands and conclude that, subject to our assumptions about the noise characteristics, our analyses are able to account for the vast majority of the 3.6 micron background. Our analyses are sensitive to a number of potential systematic effects; we discuss our assumptions with regards to noise characteristics, flux calibration and flat-fielding artifacts.

astro-ph↗

High sensitivity measurements of the CMB power spectrum with the extended Very Small Array

We present deep Ka-band ($ν\approx 33$ GHz) observations of the CMB made with the extended Very Small Array (VSA). This configuration produces a naturally weighted synthesized FWHM beamwidth of $\sim 11$ arcmin which covers an $\ell$-range of 300 to 1500. On these scales, foreground extragalactic sources can be a major source of contamination to the CMB anisotropy. This problem has been alleviated by identifying sources at 15 GHz with the Ryle Telescope and then monitoring these sources at 33 GHz using a single baseline interferometer co-located with the VSA. Sources with flux densities $\gtsim 20$ mJy at 33 GHz are subtracted from the data. In addition, we calculate a statistical correction for the small residual contribution from weaker sources that are below the detection limit of the survey. The CMB power spectrum corrected for Galactic foregrounds and extragalactic point sources is presented. A total $\ell$-range of 150-1500 is achieved by combining the complete extended array data with earlier VSA data in a compact configuration. Our resolution of $Δ\ell \approx 60$ allows the first 3 acoustic peaks to be clearly delineated. The is achieved by using mosaiced observations in 7 regions covering a total area of 82 sq. degrees. There is good agreement with WMAP data up to $\ell=700$ where WMAP data run out of resolution. For higher $\ell$-values out to $\ell = 1500$, the agreement in power spectrum amplitudes with other experiments is also very good despite differences in frequency and observing technique.

astro-ph↗

Estimating the bispectrum of the Very Small Array data

We estimate the bispectrum of the Very Small Array data from the compact and extended configuration observations released in December 2002, and compare our results to those obtained from Gaussian simulations. There is a slight excess of large bispectrum values for two individual fields, but this does not appear when the fields are combined. Given our expected level of residual point sources, we do not expect these to be the source of the discrepancy. Using the compact configuration data, we put an upper limit of 5400 on the value of f_NL, the non-linear coupling parameter, at 95 per cent confidence. We test our bispectrum estimator using non-Gaussian simulations with a known bispectrum, and recover the input values.

astro-ph↗

Cosmological parameter estimation using Very Small Array data out to l=1500

We estimate cosmological parameters using data obtained by the Very Small Array (VSA) in its extended configuration, in conjunction with a variety of other CMB data and external priors. Within the flat $Λ$CDM model, we find that the inclusion of high resolution data from the VSA modifies the limits on the cosmological parameters as compared to those suggested by WMAP alone, while still remaining compatible with their estimates. We find that $Ω_{\rm b}h^2=0.0234^{+0.0012}_{-0.0014}$, $Ω_{\rm dm}h^2=0.111^{+0.014}_{-0.016}$, $h=0.73^{+0.09}_{-0.05}$, $n_{\rm S}=0.97^{+0.06}_{-0.03}$, $10^{10}A_{\rm S}=23^{+7}_{-3}$ and $τ=0.14^{+0.14}_{-0.07}$ for WMAP and VSA when no external prior is included.On extending the model to include a running spectral index of density fluctuations, we find that the inclusion of VSA data leads to a negative running at a level of more than 95% confidence ($n_{\rm run}=-0.069\pm 0.032$), something which is not significantly changed by the inclusion of a stringent prior on the Hubble constant. Inclusion of prior information from the 2dF galaxy redshift survey reduces the significance of the result by constraining the value of $Ω_{\rm m}$. We discuss the veracity of this result in the context of various systematic effects and also a broken spectral index model. We also constrain the fraction of neutrinos and find that $f_ν< 0.087$ at 95% confidence which corresponds to $m_ν<0.32{\rm eV}$ when all neutrino masses are the equal. Finally, we consider the global best fit within a general cosmological model with 12 parameters and find consistency with other analyses available in the literature. The evidence for $n_{\rm run}<0$ is only marginal within this model.

astro-ph↗