Searcharxiv⌕ Search

arXiv subjects

Ben Hoyle

Publications and source records attributed to Ben Hoyle.

At least 37 records · Page 2Linked to original sources

The XMM Cluster Survey: The Halo Occupation Number of BOSS galaxies in X-ray clusters

We present a direct measurement of the mean halo occupation distribution (HOD) of galaxies taken from the eleventh data release (DR11) of the Sloan Digital Sky Survey-III Baryon Oscillation Spectroscopic Survey (BOSS). The HOD of BOSS low-redshift (LOWZ: $0.2 < z < 0.4$) and Constant-Mass (CMASS: $0.43 <z <0.7$) galaxies is inferred via their association with the dark-matter halos of 174 X-ray-selected galaxy clusters drawn from the XMM Cluster Survey (XCS). Halo masses are determined for each galaxy cluster based on X-ray temperature measurements, and range between ${\rm log_{10}} (M_{180}/M_{\odot}) = 13-15$. Our directly measured HODs are consistent with the HOD-model fits inferred via the galaxy-clustering analyses of Parejko et al. for the BOSS LOWZ sample and White et al. for the BOSS CMASS sample. Under the simplifying assumption that the other parameters that describe the HOD hold the values measured by these authors, we have determined a best-fit alpha-index of 0.91$\pm$0.08 and $1.27^{+0.03}_{-0.04}$ for the CMASS and LOWZ HOD, respectively. These alpha-index values are consistent with those measured by White et al. and Parejko et al. In summary, our study provides independent support for the HOD models assumed during the development of the BOSS mock-galaxy catalogues that have subsequently been used to derive BOSS cosmological constraints.

astro-ph.CO↗

The XMM Cluster Survey: evolution of the velocity dispersion -- temperature relation over half a Hubble time

We measure the evolution of the velocity dispersion--temperature ($σ_{\rm v}$--$T_{\rm X}$) relation up to $z = 1$ using a sample of 38 galaxy clusters drawn from the \textit{XMM} Cluster Survey. This work improves upon previous studies by the use of a homogeneous cluster sample and in terms of the number of high redshift clusters included. We present here new redshift and velocity dispersion measurements for 12 $z > 0.5$ clusters observed with the GMOS instruments on the Gemini telescopes. Using an orthogonal regression method, we find that the slope of the relation is steeper than that expected if clusters were self-similar, and that the evolution of the normalisation is slightly negative, but not significantly different from zero ($σ_{\rm v} \propto T^{0.86 \pm 0.14} E(z)^{-0.37 \pm 0.33}$). We verify our results by applying our methods to cosmological hydrodynamical simulations. The lack of evolution seen in our data is consistent with simulations that include both feedback and radiative cooling.

astro-ph.CO↗

Stacking for machine learning redshifts applied to SDSS galaxies

We present an analysis of a general machine learning technique called 'stacking' for the estimation of photometric redshifts. Stacking techniques can feed the photometric redshift estimate, as output by a base algorithm, back into the same algorithm as an additional input feature in a subsequent learning round. We shown how all tested base algorithms benefit from at least one additional stacking round (or layer). To demonstrate the benefit of stacking, we apply the method to both unsupervised machine learning techniques based on self-organising maps (SOMs), and supervised machine learning methods based on decision trees. We explore a range of stacking architectures, such as the number of layers and the number of base learners per layer. Finally we explore the effectiveness of stacking even when using a successful algorithm such as AdaBoost. We observe a significant improvement of between 1.9% and 21% on all computed metrics when stacking is applied to weak learners (such as SOMs and decision trees). When applied to strong learning algorithms (such as AdaBoost) the ratio of improvement shrinks, but still remains positive and is between 0.4% and 2.5% for the explored metrics and comes at almost no additional computational cost.

astro-ph.IM↗

Measuring photometric redshifts using galaxy images and Deep Neural Networks

We propose a new method to estimate the photometric redshift of galaxies by using the full galaxy image in each measured band. This method draws from the latest techniques and advances in machine learning, in particular Deep Neural Networks. We pass the entire multi-band galaxy image into the machine learning architecture to obtain a redshift estimate that is competitive with the best existing standard machine learning techniques. The standard techniques estimate redshifts using post-processed features, such as magnitudes and colours, which are extracted from the galaxy images and are deemed to be salient by the user. This new method removes the user from the photometric redshift estimation pipeline. However we do note that Deep Neural Networks require many orders of magnitude more computing resources than standard machine learning architectures.

astro-ph.IM↗

Anomaly detection for machine learning redshifts applied to SDSS galaxies

We present an analysis of anomaly detection for machine learning redshift estimation. Anomaly detection allows the removal of poor training examples, which can adversely influence redshift estimates. Anomalous training examples may be photometric galaxies with incorrect spectroscopic redshifts, or galaxies with one or more poorly measured photometric quantity. We select 2.5 million 'clean' SDSS DR12 galaxies with reliable spectroscopic redshifts, and 6730 'anomalous' galaxies with spectroscopic redshift measurements which are flagged as unreliable. We contaminate the clean base galaxy sample with galaxies with unreliable redshifts and attempt to recover the contaminating galaxies using the Elliptical Envelope technique. We then train four machine learning architectures for redshift analysis on both the contaminated sample and on the preprocessed 'anomaly-removed' sample and measure redshift statistics on a clean validation sample generated without any preprocessing. We find an improvement on all measured statistics of up to 80% when training on the anomaly removed sample as compared with training on the contaminated sample for each of the machine learning routines explored. We further describe a method to estimate the contamination fraction of a base data sample.

astro-ph.CO↗

Tuning target selection algorithms to improve galaxy redshift estimates

We showcase machine learning (ML) inspired target selection algorithms to determine which of all potential targets should be selected first for spectroscopic follow up. Efficient target selection can improve the ML redshift uncertainties as calculated on an independent sample, while requiring less targets to be observed. We compare the ML targeting algorithms with the Sloan Digital Sky Survey (SDSS) target order, and with a random targeting algorithm. The ML inspired algorithms are constructed iteratively by estimating which of the remaining target galaxies will be most difficult for the machine learning methods to accurately estimate redshifts using the previously observed data. This is performed by predicting the expected redshift error and redshift offset (or bias) of all of the remaining target galaxies. We find that the predicted values of bias and error are accurate to better than 10-30% of the true values, even with only limited training sample sizes. We construct a hypothetical follow-up survey and find that some of the ML targeting algorithms are able to obtain the same redshift predictive power with 2-3 times less observing time, as compared to that of the SDSS, or random, target selection algorithms. The reduction in the required follow up resources could allow for a change to the follow-up strategy, for example by obtaining deeper spectroscopy, which could improve ML redshift estimates for deeper test data.

astro-ph.IM↗

Accurate photometric redshift probability density estimation - method comparison and application

We introduce an ordinal classification algorithm for photometric redshift estimation, which significantly improves the reconstruction of photometric redshift probability density functions (PDFs) for individual galaxies and galaxy samples. As a use case we apply our method to CFHTLS galaxies. The ordinal classification algorithm treats distinct redshift bins as ordered values, which improves the quality of photometric redshift PDFs, compared with non-ordinal classification architectures. We also propose a new single value point estimate of the galaxy redshift, that can be used to estimate the full redshift PDF of a galaxy sample. This method is competitive in terms of accuracy with contemporary algorithms, which stack the full redshift PDFs of all galaxies in the sample, but requires orders of magnitudes less storage space. The methods described in this paper greatly improve the log-likelihood of individual object redshift PDFs, when compared with a popular Neural Network code (ANNz). In our use case, this improvement reaches 50\% for high redshift objects ($z \geq 0.75$). We show that using these more accurate photometric redshift PDFs will lead to a reduction in the systematic biases by up to a factor of four, when compared with less accurate PDFs obtained from commonly used methods. The cosmological analyses we examine and find improvement upon are the following: gravitational lensing cluster mass estimates, modelling of angular correlation functions, and modelling of cosmic shear correlation functions.

astro-ph.CO↗

Data augmentation for machine learning redshifts applied to SDSS galaxies

We present analyses of data augmentation for machine learning redshift estimation. Data augmentation makes a training sample more closely resemble a test sample, if the two base samples differ, in order to improve measured statistics of the test sample. We perform two sets of analyses by selecting 800k (1.7M) SDSS DR8 (DR10) galaxies with spectroscopic redshifts. We construct a base training set by imposing an artificial r band apparent magnitude cut to select only bright galaxies and then augment this base training set by using simulations and by applying the K-correct package to artificially place training set galaxies at a higher redshift. We obtain redshift estimates for the remaining faint galaxy sample, which are not used during training. We find that data augmentation reduces the error on the recovered redshifts by 40% in both sets of analyses, when compared to the difference in error between the ideal case and the non augmented case. The outlier fraction is also reduced by at least 10% and up to 80% using data augmentation. We finally quantify how the recovered redshifts degrade as one probes to deeper magnitudes past the artificial magnitude limit of the bright training sample. We find that at all apparent magnitudes explored, the use of data augmentation with tree based methods provide a estimate of the galaxy redshift with a negligible bias, although the error on the recovered values increases as we probe to deeper magnitudes. These results have applications for surveys which have a spectroscopic training set which forms a biased sample of all photometric galaxies, for example if the spectroscopic detection magnitude limit is shallower than the photometric limit.

astro-ph.CO↗

Feature importance for machine learning redshifts applied to SDSS galaxies

We present an analysis of importance feature selection applied to photometric redshift estimation using the machine learning architecture Decision Trees with the ensemble learning routine Adaboost (hereafter RDF). We select a list of 85 easily measured (or derived) photometric quantities (or `features') and spectroscopic redshifts for almost two million galaxies from the Sloan Digital Sky Survey Data Release 10. After identifying which features have the most predictive power, we use standard artificial Neural Networks (aNN) to show that the addition of these features, in combination with the standard magnitudes and colours, improves the machine learning redshift estimate by 18% and decreases the catastrophic outlier rate by 32%. We further compare the redshift estimate using RDF with those from two different aNNs, and with photometric redshifts available from the SDSS. We find that the RDF requires orders of magnitude less computation time than the aNNs to obtain a machine learning redshift while reducing both the catastrophic outlier rate by up to 43%, and the redshift error by up to 25%. When compared to the SDSS photometric redshifts, the RDF machine learning redshifts both decreases the standard deviation of residuals scaled by 1/(1+z) by 36% from 0.066 to 0.041, and decreases the fraction of catastrophic outliers by 57% from 2.32% to 0.99%.

astro-ph.IM↗

Combining clustering and abundances of galaxy clusters to test cosmology and primordial non-Gaussianity

We present the clustering of galaxy clusters as a useful addition to the common set of cosmological observables. The clustering of clusters probes the large-scale structure of the Universe, extending galaxy clustering analysis to the high-peak, high-bias regime. Clustering of galaxy clusters complements the traditional cluster number counts and observable-mass relation analyses, significantly improving their constraining power by breaking existing calibration degeneracies. We use the maxBCG galaxy clusters catalogue to constrain cosmological parameters and cross-calibrate the mass-observable relation, using cluster abundances in richness bins and weak-lensing mass estimates. We then add the redshift-space power spectrum of the sample, including an effective modelling of the weakly non-linear contribution and allowing for an arbitrary photometric redshift smoothing. The inclusion of the power spectrum data allows for an improved self-calibration of the scaling relation. We find that the inclusion of the power spectrum typically brings a $\sim 50$ per cent improvement in the errors on the fluctuation amplitude $σ_8$ and the matter density $Ω_{\mathrm{m}}$. Finally, we apply this method to constrain models of the early universe through the amount of primordial non-Gaussianity of the local type, using both the variation in the halo mass function and the variation in the cluster bias. We find a constraint on the amount of skewness $f_{\mathrm{NL}} = 12 \pm 157 $ ($1σ$) from the cluster data alone.

astro-ph.CO↗

Testing Homogeneity with Galaxy Star Formation Histories

Observationally confirming spatial homogeneity on sufficiently large cosmological scales is of importance to test one of the underpinning assumptions of cosmology, and is also imperative for correctly interpreting dark energy. A challenging aspect of this is that homogeneity must be probed inside our past lightcone, while observations take place on the lightcone. The star formation history (SFH) in the galaxy fossil record provides a novel way to do this. We calculate the SFH of stacked Luminous Red Galaxy (LRG) spectra obtained from the Sloan Digital Sky Survey. We divide the LRG sample into 12 equal area contiguous sky patches and 10 redshift slices (0.2 < z < 0.5), which correspond to 120 blocks of volume 0.04Gpc3. Using the SFH in a time period which samples the history of the Universe between look-back times 11.5 to 13.4 Gyrs as a proxy for homogeneity, we calculate the posterior distribution for the excess large-scale variance due to inhomogeneity, and find that the most likely solution is no extra variance at all. At 95% credibility, there is no evidence of deviations larger than 5.8%.

astro-ph.CO↗

Probing the bias of radio sources at high redshift

The relationship between the clustering of dark matter and that of luminous matter is often described using the bias parameter. Here, we provide a new method to probe the bias of intermediate to high-redshift radio continuum sources for which no redshift information is available. We matched radio sources from the Faint Images of the Radio Sky at Twenty centimetres (FIRST) survey data to their optical counterparts in the Sloan Digital Sky Survey (SDSS) to obtain photometric redshifts for the matched radio sources. We then use the publicly available semi-empirical simulation of extragalactic radio continuum sources (S3) to infer the redshift distribution for all FIRST sources and estimate the redshift distribution of unmatched sources by subtracting the matched distribution from the distribution of all sources. We infer that the majority of unmatched sources are at higher redshifts than the optically matched sources and demonstrate how the angular scales of the angular two-point correlation function can be used to probe different redshift ranges. We compare the angular clustering of radio sources with that expected for dark matter and estimate the bias of different samples.

astro-ph.CO↗

The XMM Cluster Survey: The Stellar Mass Assembly of Fossil Galaxies

This paper presents both the result of a search for fossil systems (FSs) within the XMM Cluster Survey and the Sloan Digital Sky Survey and the results of a study of the stellar mass assembly and stellar populations of their fossil galaxies. In total, 17 groups and clusters are identified at z < 0.25 with large magnitude gaps between the first and fourth brightest galaxies. All the information necessary to classify these systems as fossils is provided. For both groups and clusters, the total and fractional luminosity of the brightest galaxy is positively correlated with the magnitude gap. The brightest galaxies in FSs (called fossil galaxies) have stellar populations and star formation histories which are similar to normal brightest cluster galaxies (BCGs). However, at fixed group/cluster mass, the stellar masses of the fossil galaxies are larger compared to normal BCGs, a fact that holds true over a wide range of group/cluster masses. Moreover, the fossil galaxies are found to contain a significant fraction of the total optical luminosity of the group/cluster within 0.5R200, as much as 85%, compared to the non-fossils, which can have as little as 10%. Our results suggest that FSs formed early and in the highest density regions of the universe and that fossil galaxies represent the end products of galaxy mergers in groups and clusters. The online FS catalog can be found at http://www.astro.ljmu.ac.uk/~xcs/Harrison2012/XCSFSCat.html.

astro-ph.CO↗

The similar stellar populations of quiescent spiral and elliptical galaxies

We compare the stellar population properties in the central regions of visually classified non-starforming spiral and elliptical galaxies from Galaxy Zoo and SDSS DR7. The galaxies lie in the redshift range $0.04<z<0.1$ and have stellar masses larger than $logM_*=10.4$. We select only face-on spiral galaxies in order to avoid contamination by light from the disk in the SDSS fiber and enabling the robust visual identification of spiral structure. Overall, we find that galaxies with larger central stellar velocity dispersions, regardless of morphological type, have older ages, higher metallicities, and an increased overabundance of alpha-elements. Age and alpha-enhancement, at fixed velocity dispersion, do not depend on morphological type. The only parameter that, at a given velocity dispersion, correlates with morphological type is metallicity, where the metallicity of the bulges of spiral galaxies is 0.07 dex higher than that of the ellipticals. However, for galaxies with a given total stellar mass, this dependence on morphology disappears. Under the assumption that, for our sample, the velocity dispersion traces the mass of the bulge alone, as opposed to the total mass (bulge+disk) of the galaxy, our results imply that the formation epoch of galaxy and the duration of its star-forming period are linked to the mass of the bulge. The extent to which metals are retained within the galaxy, and not removed as a result of outflows, is determined by the total mass of the galaxy.

astro-ph.CO↗

The XMM Cluster Survey: Evidence for energy injection at high redshift from evolution of the X-ray luminosity-temperature relation

We measure the evolution of the X-ray luminosity-temperature (L_X-T) relation since z~1.5 using a sample of 211 serendipitously detected galaxy clusters with spectroscopic redshifts drawn from the XMM Cluster Survey first data release (XCS-DR1). This is the first study spanning this redshift range using a single, large, homogeneous cluster sample. Using an orthogonal regression technique, we find no evidence for evolution in the slope or intrinsic scatter of the relation since z~1.5, finding both to be consistent with previous measurements at z~0.1. However, the normalisation is seen to evolve negatively with respect to the self-similar expectation: we find E(z)^{-1} L_X = 10^{44.67 +/- 0.09} (T/5)^{3.04 +/- 0.16} (1+z)^{-1.5 +/- 0.5}, which is within 2 sigma of the zero evolution case. We see milder, but still negative, evolution with respect to self-similar when using a bisector regression technique. We compare our results to numerical simulations, where we fit simulated cluster samples using the same methods used on the XCS data. Our data favour models in which the majority of the excess entropy required to explain the slope of the L_X-T relation is injected at high redshift. Simulations in which AGN feedback is implemented using prescriptions from current semi-analytic galaxy formation models predict positive evolution of the normalisation, and differ from our data at more than 5 sigma. This suggests that more efficient feedback at high redshift may be needed in these models.

astro-ph.CO↗

Galaxy Zoo: The Environmental Dependence of Bars and Bulges in Disc Galaxies

We present an analysis of the environmental dependence of bars and bulges in disc galaxies, using a volume-limited catalogue of 15810 galaxies at z<0.06 from the Sloan Digital Sky Survey with visual morphologies from the Galaxy Zoo 2 project. We find that the likelihood of having a bar, or bulge, in disc galaxies increases when the galaxies have redder (optical) colours and larger stellar masses, and observe a transition in the bar and bulge likelihoods, such that massive disc galaxies are more likely to host bars and bulges. We use galaxy clustering methods to demonstrate statistically significant environmental correlations of barred, and bulge-dominated, galaxies, from projected separations of 150 kpc/h to 3 Mpc/h. These environmental correlations appear to be independent of each other: i.e., bulge-dominated disc galaxies exhibit a significant bar-environment correlation, and barred disc galaxies show a bulge-environment correlation. We demonstrate that approximately half (50 +/- 10%) of the bar-environment correlation can be explained by the fact that more massive dark matter haloes host redder disc galaxies, which are then more likely to have bars. Likewise, we show that the environmental dependence of stellar mass can only explain a small fraction (25 +/- 10%) of the bar-environment correlation. Therefore, a significant fraction of our observed environmental dependence of barred galaxies is not due to colour or stellar mass dependences, and hence could be due to another galaxy property. Finally, by analyzing the projected clustering of barred and unbarred disc galaxies with halo occupation models, we argue that barred galaxies are in slightly higher-mass haloes than unbarred ones, and some of them (approximately 25%) are satellite galaxies in groups. We also discuss implications about the effects of minor mergers and interactions on bar formation.

astro-ph.CO↗

The fraction of early-type galaxies in low redshift groups and clusters of galaxies

We examine the fraction of early-type (and spiral) galaxies found in groups and clusters of galaxies as a function of dark matter halo mass. We use morphological classifications from the Galaxy Zoo project matched to halo masses from both the C4 cluster catalogue and the Yang et al (2007) group catalogue. We find that the fraction of early-type (or spiral) galaxies remains constant (changing by less than 10%) over three orders of magnitude in halo mass (13<log MH/Msol/h<15.8). This result is insensitive to our choice of halo mass measure, from velocity dispersions or summed optical luminosity. Furthermore, we consider the morphology-halo mass relations in bins of galaxy stellar mass M*, and find that while the trend of constant fraction remains unchanged, the early-type fraction amongst the most massive galaxies (11<log M*/Msol/h <12) is a factor of three greater than lower mass galaxies (10<logM*/Msol/h<10.7). We compare our observational results with those of simulations presented in De Lucia et al (2011), as well as previous observational analyses, and semi-analytic bulge (or disc) dominated galaxies from the Millennium Simulation. We find the simulations recover similar trends as observed, but may over-predict the abundances of the most massive bulge dominated (early-type) galaxies. Our results suggest that most morphological transformation is happening on the group scale before groups merge into massive clusters. However, we show that within each halo a morphology-density relation remains: it is summing the total fraction to a self-similar scaled radius which results in a flat morphology-halo mass relationship.

astro-ph.CO↗

The XMM Cluster Survey: The interplay between the brightest cluster galaxy and the intra-cluster medium via AGN feedback

Using a sample of 123 X-ray clusters and groups drawn from the XMM-Cluster Survey first data release, we investigate the interplay between the brightest cluster galaxy (BCG), its black hole, and the intra-cluster/group medium (ICM). It appears that for groups and clusters with a BCG likely to host significant AGN feedback, gas cooling dominates in those with Tx > 2 keV while AGN feedback dominates below. This may be understood through the sub-unity exponent found in the scaling relation we derive between the BCG mass and cluster mass over the halo mass range 10^13 < M500 < 10^15Msol and the lack of correlation between radio luminosity and cluster mass, such that BCG AGN in groups can have relatively more energetic influence on the ICM. The Lx - Tx relation for systems with the most massive BCGs, or those with BCGs co-located with the peak of the ICM emission, is steeper than that for those with the least massive and most offset, which instead follows self-similarity. This is evidence that a combination of central gas cooling and powerful, well fuelled AGN causes the departure of the ICM from pure gravitational heating, with the steepened relation crossing self-similarity at Tx = 2 keV. Importantly, regardless of their black hole mass, BCGs are more likely to host radio-loud AGN if they are in a massive cluster (Tx > 2 keV) and again co-located with an effective fuel supply of dense, cooling gas. This demonstrates that the most massive black holes appear to know more about their host cluster than they do about their host galaxy. The results lead us to propose a physically motivated, empirical definition of 'cluster' and 'group', delineated at 2 keV.

astro-ph.CO↗