SearcharxivSearch

arXiv subjects

Bryan Ostdiek

Publications and source records attributed to Bryan Ostdiek.

At least 19 recordsLinked to original sources

Creating Simple, Interpretable Anomaly Detectors for New Physics in Jet Substructure

Anomaly detection with convolutional autoencoders is a popular method to search for new physics in a model-agnostic manner. These techniques are powerful, but they are still a "black box," since we do not know what high-level physical observables determine how anomalous an event is. To address this, we adapt a recently proposed technique by Faucett et al., which maps out the physical observables learned by a neural network classifier, to the case of anomaly detection. We propose two different strategies that use a small number of high-level observables to mimic the decisions made by the autoencoder on background events, one designed to directly learn the output of the autoencoder, and the other designed to learn the difference between the autoencoder's outputs on a pair of events. Despite the underlying differences in their approach, we find that both strategies have similar ordering performance as the autoencoder and independently use the same six high-level observables. From there, we compare the performance of these networks as anomaly detectors. We find that both strategies perform similarly to the autoencoder across a variety of signals, giving a nontrivial demonstration that learning to order background events transfers to ordering a variety of signal events.

hep-ph

Neural Embedding: Learning the Embedding of the Manifold of Physics Data

In this paper, we present a method of embedding physics data manifolds with metric structure into lower dimensional spaces with simpler metrics, such as Euclidean and Hyperbolic spaces. We then demonstrate that it can be a powerful step in the data analysis pipeline for many applications. Using progressively more realistic simulated collisions at the Large Hadron Collider, we show that this embedding approach learns the underlying latent structure. With the notion of volume in Euclidean spaces, we provide for the first time a viable solution to quantifying the true search capability of model agnostic search algorithms in collider physics (i.e. anomaly detection). Finally, we discuss how the ideas presented in this paper can be employed to solve many practical challenges that require the extraction of physically meaningful representations from information in complex high dimensional datasets.

hep-ph

Substructure Detection Reanalyzed: Dark Perturber shown to be a Line-of-Sight Halo

Observations of structure at sub-galactic scales are crucial for probing the properties of dark matter, which is the dominant source of gravity in the universe. It will become increasingly important for future surveys to distinguish between line-of-sight halos and subhalos to avoid wrong inferences on the nature of dark matter. We reanalyze a sub-galactic structure (in lens JVAS B1938+666) that has been previously found using the gravitational imaging technique in galaxy-galaxy lensing systems. This structure has been assumed to be a satellite in the halo of the main lens galaxy. We fit the redshift of the perturber of the system as a free parameter, using the multi-plane thin-lens approximation, and find that the redshift of the perturber is $z_\mathrm{int} = 1.42\substack{+0.10 \\ -0.15}$ (with a main lens redshift of $z=0.881$). Our analysis indicates that this structure is more massive than the previous result by an order of magnitude. This constitutes the first dark perturber shown to be a line-of-sight halo with a gravitational lensing method.

astro-ph.CO

Revealing the Milky Way's Most Recent Major Merger with a Gaia EDR3 Catalog of Machine-Learned Line-of-Sight Velocities

Machine learning can play a powerful role in inferring missing line-of-sight velocities from astrometry in surveys such as Gaia. In this paper, we apply a neural network to Gaia Early Data Release 3 (EDR3) and obtain line-of-sight velocities and associated uncertainties for ~92 million stars. The network, which takes as input a star's parallax, angular coordinates, and proper motions, is trained and validated on ~6.4 million stars in Gaia with complete phase-space information. The network's uncertainty on its velocity prediction is a key aspect of its design; by properly convolving these uncertainties with the inferred velocities, we obtain accurate stellar kinematic distributions. As a first science application, we use the new network-completed catalog to identify candidate stars that belong to the Milky Way's most recent major merger, Gaia-Sausage-Enceladus (GSE). We present the kinematic, energy, angular momentum, and spatial distributions of the ~450,000 GSE candidates in this sample, and also study the chemical abundances of those with cross matches to GALAH and APOGEE. The network's predictive power will only continue to improve with future Gaia data releases as the training set of stars with complete phase-space information grows. This work provides a first demonstration of how to use machine learning to exploit high-dimensional correlations on data to infer line-of-sight velocities, and offers a template for how to train, validate and apply such a neural network when complete observational data is not available.

astro-ph.GA

Chasing Accreted Structures within Gaia DR2 using Deep Learning

In previous work, we developed a deep neural network classifier that only relies on phase-space information to obtain a catalog of accreted stars based on the second data release of Gaia (DR2). In this paper, we apply two clustering algorithms to identify velocity substructure within this catalog. We focus on the subset of stars with line-of-sight velocity measurements that fall in the range of Galactocentric radii $r \in [6.5, 9.5]$ kpc and vertical distances $|z| < 3$ kpc. Known structures such as Gaia Enceladus and the Helmi stream are identified. The largest previously-unknown structure, Nyx, is a vast stream consisting of at least 90 stars in the region of interest. This study displays the power of the machine learning approach by not only successfully identifying known features, but also discovering new kinematic structures that may shed light on the merger history of the Milky Way.

astro-ph.GA

Extracting the Subhalo Mass Function from Strong Lens Images with Image Segmentation

Detecting substructure within strongly lensed images is a promising route to shed light on the nature of dark matter. However, it is a challenging task, which traditionally requires detailed lens modeling and source reconstruction, taking weeks to analyze each system. We use machine-learning to circumvent the need for lens and source modeling and develop a neural network to both locate subhalos in an image as well as determine their mass using the technique of image segmentation. The network is trained on images with a single subhalo located near the Einstein ring across a wide range of apparent source magnitudes. The network is then able to resolve subhalos with masses $m\gtrsim 10^{8.5} M_{\odot}$. Training in this way allows the network to learn the gravitational lensing of light, and remarkably, it is then able to detect entire populations of substructure, even for locations further away from the Einstein ring than those used in training. Over a wide range of the apparent source magnitude, the false-positive rate is around three false subhalos per 100 images, coming mostly from the lightest detectable subhalo for that signal-to-noise ratio. With good accuracy and a low false-positive rate, counting the number of pixels assigned to each subhalo class over multiple images allows for a measurement of the subhalo mass function (SMF). When measured over three mass bins from $10^9M_{\odot}$--$10^{10} M_{\odot}$ the SMF slope is recovered with an error of 36% for 50 images, and this improves to 10% for 1000 images with Hubble Space Telescope-like noise.

astro-ph.CO

Evidence for a Vast Prograde Stellar Stream in the Solar Vicinity

Massive dwarf galaxies that merge with the Milky Way on prograde orbits can be dragged into the disk plane before being completely disrupted. Such mergers can contribute to an accreted stellar disk and a dark matter disk. We present evidence for Nyx, a vast new stellar stream in the vicinity of the Sun, that may provide the first indication that such an event occurred in the Milky Way. We identify about 500 stars that have coherent radial and prograde motion in this stream using a catalog of accreted stars built by applying deep learning methods to the second Gaia data release. Nyx is concentrated within $\pm 2$ kpc of the Galactic midplane and spans the full radial range studied (6.5-9.5 kpc). The kinematics of Nyx stars are distinct from those of both the thin and thick disk. In particular, its rotational speed lags the disk by $\sim 80$ km/s and its stars follow more eccentric orbits. A small number of Nyx stars have chemical abundances or inferred ages; from these, we deduce that Nyx stars have a peak metallicity of [Fe/H] $\sim -0.5$ and ages $\sim $10-13 Gyr. Taken together with the kinematic observations, these results strongly favor the interpretation that Nyx is the remnant of a disrupted dwarf galaxy. To further justify this interpretation, we explicitly demonstrate that metal-rich, prograde streams like Nyx can be found in the disk plane of Milky Way-like galaxies using the FIRE hydrodynamic simulations. Future spectroscopic studies will be able to validate whether Nyx stars originate from a single progenitor.

astro-ph.GA

Image segmentation for analyzing galaxy-galaxy strong lensing systems

The goal of this paper is to develop a machine learning model to analyze the main gravitational lens and detect dark substructure (subhalos) within simulated images of strongly lensed galaxies. Using the technique of image segmentation, we turn the task of identifying subhalos into a classification problem, where we label each pixel in an image as coming from the main lens, a subhalo within a binned mass range, or neither. Our network is only trained on images with a single smooth lens and either zero or one subhalo near the Einstein ring. On an independent test set with lenses with large ellipticities, quadrupole and octopole moments, and for source apparent magnitudes between 17-25, the area of the main lens is recovered accurately. On average, only 1.3% of the true area is missed and 1.2% of the true area is added to another part of the lens. In addition, subhalos as light as $10^{8.5}M_{\odot}$ can be detected if they lie in bright pixels along the Einstein ring. Furthermore, the model is able to generalize to new contexts it has not been trained on, such as locating multiple subhalos with varying masses or more than one large smooth lens.

astro-ph.CO

Deep Set Auto Encoders for Anomaly Detection in Particle Physics

There is an increased interest in model agnostic search strategies for physics beyond the standard model at the Large Hadron Collider. We introduce a Deep Set Variational Autoencoder and present results on the Dark Machines Anomaly Score Challenge. We find that the method attains the best anomaly detection ability when there is no decoding step for the network, and the anomaly score is based solely on the representation within the encoded latent space. This method was one of the top-performing models in the Dark Machines Challenge, both for the open data sets as well as the blinded data sets.

hep-ph

Challenges for Unsupervised Anomaly Detection in Particle Physics

Anomaly detection relies on designing a score to determine whether a particular event is uncharacteristic of a given background distribution. One way to define a score is to use autoencoders, which rely on the ability to reconstruct certain types of data (background) but not others (signals). In this paper, we study some challenges associated with variational autoencoders, such as the dependence on hyperparameters and the metric used, in the context of anomalous signal (top and $W$) jets in a QCD background. We find that the hyperparameter choices strongly affect the network performance and that the optimal parameters for one signal are non-optimal for another. In exploring the networks, we uncover a connection between the latent space of a variational autoencoder trained using mean-squared-error and the optimal transport distances within the dataset. We then show that optimal transport distances to representative events in the background dataset can be used directly for anomaly detection, with performance comparable to the autoencoders. Whether using autoencoders or optimal transport distances for anomaly detection, we find that the choices that best represent the background are not necessarily best for signal identification. These challenges with unsupervised anomaly detection bolster the case for additional exploration of semi-supervised or alternative approaches.

hep-ph

Parameter Inference from Event Ensembles and the Top-Quark Mass

One of the key tasks of any particle collider is measurement. In practice, this is often done by fitting data to a simulation, which depends on many parameters. Sometimes, when the effects of varying different parameters are highly correlated, a large ensemble of data may be needed to resolve parameter-space degeneracies. An important example is measuring the top-quark mass, where other physical and unphysical parameters in the simulation must be marginalized over when fitting the top-quark mass parameter. We compare three different methodologies for top-quark mass measurement: a classical histogram fitting procedure, similar to one commonly used in experiment optionally augmented with soft-drop jet grooming; a machine-learning method called DCTR; and a linear regression approach, either using a least-squares fit or with a dense linearly-activated neural network. Despite the fact that individual events are totally uncorrelated, we find that the linear regression methods work most effectively when we input an ensemble of events sorted by mass, rather than training them on individual events. Although all methods provide robust extraction of the top-quark mass parameter, the linear network does marginally best and is remarkably simple. For the top study, we conclude that the Monte-Carlo-based uncertainty on current extractions of the top-quark mass from LHC data can be reduced significantly (by perhaps a factor of 2) using networks trained on sorted event ensembles. More generally, machine learning from ensembles for parameter estimation has broad potential for collider physics measurements.

hep-ph

Machine Learning the 6th Dimension: Stellar Radial Velocities from 5D Phase-Space Correlations

The Gaia satellite will observe the positions and velocities of over a billion Milky Way stars. In the early data releases, the majority of observed stars do not have complete 6D phase-space information. In this Letter, we demonstrate the ability to infer the missing line-of-sight velocities until more spectroscopic observations become available. We utilize a novel neural network architecture that, after being trained on a subset of data with complete phase-space information, takes in a star's 5D astrometry (angular coordinates, proper motions, and parallax) and outputs a predicted line-of-sight velocity with an associated uncertainty. Working with a mock Gaia catalog, we show that the network can successfully recover the distributions and correlations of each velocity component for stars that fall within ~5 kpc of the Sun. We also demonstrate that the network can accurately reconstruct the velocity distribution of a kinematic substructure in the stellar halo that is spatially uniform, even when it comprises a small fraction of the total star count.

astro-ph.GA

On the ATLAS Top Mass Measurements and the Potential for Stealth Stop Contamination

The discovery of the stop - the Supersymmetric partner of the top quark - is a key goal of the physics program enabled by the Large Hadron Collider. Although much of the accessible parameter space has already been probed, all current searches assume the top mass is known. This is relevant for the "stealth stop" regime, which is characterized by decay kinematics that force the final state top quark off its mass shell; such decays would contaminate the top mass measurements. We investigate the resulting bias imparted to the template method based ATLAS approach. A careful recasting of these results shows that effect can be as large as 2.0 GeV, comparable to the current quoted uncertainty on the top mass. Thus, a robust exploration of the stealth stop splinter requires the simultaneous consideration of the impact on the top mass. Additionally, we explore the robustness of the template technique, and point out a simple strategy for improving the methodology implemented for the semi-leptonic channel.

hep-ph

Cataloging Accreted Stars within Gaia DR2 using Deep Learning

The goal of this study is to present the development of a machine learning based approach that utilizes phase space alone to separate the Gaia DR2 stars into two categories: those accreted onto the Milky Way from those that are in situ. Traditional selection methods that have been used to identify accreted stars typically rely on full 3D velocity, metallicity information, or both, which significantly reduces the number of classifiable stars. The approach advocated here is applicable to a much larger portion of Gaia DR2. A method known as "transfer learning" is shown to be effective through extensive testing on a set of mock Gaia catalogs that are based on the FIRE cosmological zoom-in hydrodynamic simulations of Milky Way-mass galaxies. The machine is first trained on simulated data using only 5D kinematics as inputs and is then further trained on a cross-matched Gaia/RAVE data set, which improves sensitivity to properties of the real Milky Way. The result is a catalog that identifies around 767,000 accreted stars within Gaia DR2. This catalog can yield empirical insights into the merger history of the Milky Way and could be used to infer properties of the dark matter distribution.

astro-ph.GA

Mass Agnostic Jet Taggers

Searching for new physics in large data sets needs a balance between two competing effects---signal identification vs background distortion. In this work, we perform a systematic study of both single variable and multivariate jet tagging methods that aim for this balance. The methods preserve the shape of the background distribution by either augmenting the training procedure or the data itself. Multiple quantitative metrics to compare the methods are considered, for tagging 2-, 3-, or 4-prong jets from the QCD background. This is the first study to show that the data augmentation techniques of Planing and PCA based scaling deliver similar performance as the augmented training techniques of Adversarial NN and uBoost, but are both easier to implement and computationally cheaper.

hep-ph

Dark Mesons at the LHC

A new, strongly-coupled dark sector could be accessible to LHC searches now. These dark sectors consist of composites formed from constituents that are charged under the electroweak group and interact with the Higgs, but are neutral under Standard Model color. In these scenarios, the most promising target is the dark meson sector, consisting of dark vector-mesons as well as dark pions. In this paper we study dark meson production and decay at the LHC in theories that preserve a global SU(2) dark flavor symmetry. Dark pions can be pair-produced through resonant dark vector meson production, $p p\toρ_D\toπ_Dπ_D$, and decay in one of two distinct ways: gaugephobic, when $π_D\to f\bar{f}'$ generally dominates; or gaugephilic, when $π_D\to W+h,Z+h$ dominates once kinematically open. Unlike QCD, the decay $π^0_D\toγγ$ is virtually absent due to the dark flavor symmetry. We recast a vast set of LHC searches to determine the current constraints on dark meson production and decay. When $m_{ρ_D}$ is slightly heavier than $2 m_{π_D}$ and $ρ_D^{\pm,0}$ kinetically mixes with the weak gauge bosons, the 8 TeV same-sign lepton search strategy sets the best bound, $m_{π_D}>500$ GeV. Yet, when only the $ρ^0_D$ kinetically mixes with hypercharge, we find the strongest LHC bound is $m_{π_D}>130$ GeV, that is only slightly better than what LEP II achieved. We find the relative insensitivity of LHC searches, especially at 13 TeV, can be blamed mainly on their penchant for high mass objects or large MET. Dedicated searches would undoubtedly yield substantially improved sensitivity. We provide a GitHub page to speed the implementation of these searches in future LHC analyses. Our findings provide a strong motivation for model-independent searches of the form $pp\to A\to B+C\to SM\, SM+SM\, SM$ where the theoretical prejudice is for SM to be a t,b,$τ$ or W,Z,h.

hep-ph

Magnifying the ATLAS Stealth Stop Splinter: Impact of Spin Correlations and Finite Widths

In this paper, we recast a "stealth stop" search in the notoriously difficult region of the stop-neutralino Simplified Model parameter space for which $m(\tilde{t}) - m(\tildeχ) \simeq m_t$. The properties of the final state are nearly identical for tops and stops, while the rate for stop pair production is $\mathcal{O}(10\%)$ of that for $t\bar{t}$. Stop searches away from this stealth region have left behind a "splinter" of open parameter space when $m(\tilde{t}) \simeq m_t$. Removing this splinter requires surgical precision: the ATLAS constraint on stop pair production reinterpreted here treats the signal as a contaminant to the measurement of the top pair production cross section using data from $\sqrt{s} = 7 \text{ TeV}$ and $8 \text{ TeV}$ in a correlated way to control for some systematic errors. ATLAS fixed $m(\tilde{t}) \simeq m_t$ and $m(\tildeχ)= 1 \text{ GeV}$, implying that a careful recasting of these results into the full $m(\tilde{t}) - m(\tildeχ)$ plane is warranted. We find that the parameter space with $m(\tildeχ)\lesssim 55 \text{ GeV}$ is excluded for $m(\tilde{t}) \simeq m_t$ --- although this search does cover new parameter space, it is unable to fully pull the splinter. Along the way, we review a variety of interesting physical issues in detail: (i) when the two-body width is a good approximation; (ii) what the impact on the total rate from taking the narrow width is a good approximation; (iii) how the production rate is affected when the wrong widths are used; (iv) what role the spin correlations play in the limits. In addition, we provide a guide to using MadGraph for implementing the full production including finite width and spin correlation effects, and we survey a variety of pitfalls one might encounter.

hep-ph

(Machine) Learning to Do More with Less

Determining the best method for training a machine learning algorithm is critical to maximizing its ability to classify data. In this paper, we compare the standard "fully supervised" approach (that relies on knowledge of event-by-event truth-level labels) with a recent proposal that instead utilizes class ratios as the only discriminating information provided during training. This so-called "weakly supervised" technique has access to less information than the fully supervised method and yet is still able to yield impressive discriminating power. In addition, weak supervision seems particularly well suited to particle physics since quantum mechanics is incompatible with the notion of mapping an individual event onto any single Feynman diagram. We examine the technique in detail -- both analytically and numerically -- with a focus on the robustness to issues of mischaracterizing the training samples. Weakly supervised networks turn out to be remarkably insensitive to systematic mismodeling. Furthermore, we demonstrate that the event level outputs for weakly versus fully supervised networks are probing different kinematics, even though the numerical quality metrics are essentially identical. This implies that it should be possible to improve the overall classification ability by combining the output from the two types of networks. For concreteness, we apply this technology to a signature of beyond the Standard Model physics to demonstrate that all these impressive features continue to hold in a scenario of relevance to the LHC.

hep-ph