SearcharxivSearch

arXiv subjects

Sung Hak Lim

Publications and source records attributed to Sung Hak Lim.

At least 19 recordsLinked to original sources

Anomaly detection for multijet scenarios

Signals of physics beyond the Standard Model continue to resist discovery at the LHC. Recent years have seen the proliferation of new anomaly detection techniques, promising discovery with significantly fewer model assumptions than traditional approaches. Strategies based on weak supervision have been especially successful, but so far were largely limited in scope by their reliance on a resonance manifesting a decay into a pair of jets. In this work, we demonstrate that a well-established idea from jet substructure physics - recursive soft drop - in combination with the CATHODE technique for anomaly detection can be used to simultaneously perform anomaly detection for signals with an arbitrary number of jets in the final state, greatly increasing the scope of such searches.

hep-ph

ClearPotential: Revealing Local Dark Matter in Three Dimensions

We present ClearPotential, a data-driven, three-dimensional measurement of the gravitational potential of the local Milky Way using unsupervised machine learning, without the symmetry assumptions, specific functional forms, and binning required in previous work. The potential is modeled as a neural network, optimized to solve the equilibrium collisionless Boltzmann equation for the observed phase space density of Gaia DR3 Red Clump stars within 4 kpc of the Sun. This density is obtained from data using normalizing flows, and our unsupervised solution to the Boltzmann equation automatically corrects for selection effects from crowding and the dust-driven extinction of starlight. Our fully-differentiable model of the gravitational potential allows us to map the acceleration and mass density of the Galaxy in the volume around the Sun, including in the dust-obscured disk towards the Galactic Center. We determine the dark matter density at the Solar radius to be $(0.84 \pm 0.08)\times 10^{-2}\,{M}_\odot/{\rm pc}^3$, and analyze the structure of the dark matter halo. We find strong evidence for a tilted oblate halo, weak preference for a cored inner profile, and the strongest constraints to date on a possible dark matter disk. We place a bound on the timescale of disequilibrium in the local Milky Way, and find mild evidence for disequilibrium using independent acceleration measurements from timings of binary pulsar systems. This work provides the clearest map of the local Galactic potential to date and marks an important step in the era of data-driven astrometry.

astro-ph.GA

Sweeping the Dust Away -- Correcting the Phase Space Density of the Milky Way with Unsupervised Machine Learning

The Boltzmann equation relates the equilibrium phase space distribution of stars in the Milky Way to the Galaxy's gravitational potential. However, observations of stellar populations are biased by extinction from foreground dust, which complicates measurements of the potential in the disk and towards the Galactic center. Using the kinematics of Red Clump and Red Branch stars in Gaia DR3, we use machine learning to simultaneously estimate both the unbiased stellar phase space density and the gravitational potential. The unbiased phase space density is obtained through a learned "dust efficiency factor" -- an observational selection function that accounts for dust extinction. The potential and the dust efficiency are parameterized by fully connected neural networks and are completely data driven. We validate the dust efficiency using a recent three-dimensional dust map in this work, and examine the potential in a companion paper.

astro-ph.GA

Mapping Dark Matter in the Milky Way using Normalizing Flows and Gaia DR3

We present a novel, data-driven analysis of Galactic dynamics, using unsupervised machine learning -- in the form of density estimation with normalizing flows -- to learn the underlying phase space distribution of 6 million nearby stars from the Gaia DR3 catalog. Solving the equilibrium collisionless Boltzmann equation, we calculate -- for the first time ever -- a model-free, unbinned estimate of the local acceleration and mass density fields within a 3 kpc sphere around the Sun. As our approach makes no assumptions about symmetries, we can test for signs of disequilibrium in our results. We find our results are consistent with equilibrium at the 10% level, limited by the current precision of the normalizing flows. After subtracting the known contribution of stars and gas from the calculated mass density, we find clear evidence for dark matter throughout the analyzed volume. Assuming spherical symmetry and averaging mass density measurements, we find a local dark matter density of $0.47\pm 0.05$ GeV/cm$^3$. We compute the dark matter density at four radii in the stellar halo and fit to a generalized NFW profile. Although the uncertainties are large, we find a profile broadly consistent with recent analyses.

astro-ph.GA

Reweighting and Analysing Event Generator Systematics by Neural Networks on High-Level Features

The state-of-the-art deep learning (DL) models for jet classification use jet constituent information directly, improving performance tremendously. This draws attention to interpretability, namely, the decision-making process, correlations contributing to the classification, and high-level features (HLFs) representing the difference between signal and background. We address the interpretability issue using a modular architecture called the analysis model (AM), which combines several motivated HLFs as the input. We focus on the generator systematics of the top vs. QCD classification by one of the best classifiers, Particle Transformer (ParT). Taking commonly used event generators Pythia (PY) and Herwig (HW) as examples, we demonstrate that the event weights estimated by the AM generator classifier align the HW classification score distribution to PY ones for QCD jets, with small training uncertainty. This suggests the AM is sufficient to describe simulated QCD jet features with relatively few observables, and generator systematics would also be reduced by reweighting the simulation by data. On the other hand, large event weights are required for QCD-like top jets, which leads to imperfect reweighting for both AM and ParT generator classifiers. Moreover, the AM HLFs are insufficient for describing PY and HW differences, causing lower reweighting accuracy compared with ParT. The missing features are the correlation among the collimated high-energy jet constituents, which are strongly correlated to the energy flow polynomials (EFPs) selected for top vs. QCD classification, showing the complementarity between AM HLFs and the selected EFPs.

hep-ph

JFlow: Model-Independent Spherical Jeans Analysis using Equivariant Continuous Normalizing Flows

The kinematics of stars in dwarf spheroidal galaxies have been studied to understand the structure of dark matter halos. However, the kinematic information of these stars is often limited to celestial positions and line-of-sight velocities, making full phase space analysis challenging. Conventional methods rely on projected analytic phase space density models with several parameters and infer dark matter halo structures by solving the spherical Jeans equation. In this paper, we introduce an unsupervised machine learning method for solving the spherical Jeans equation in a model-independent way as a first step toward model-independent analysis of dwarf spheroidal galaxies. Using equivariant continuous normalizing flows, we demonstrate that spherically symmetric stellar phase space densities and velocity dispersions can be estimated without model assumptions. As a proof of concept, we apply our method to Gaia challenge datasets for spherical models and measure dark matter mass densities for given velocity anisotropy profiles. Our method can identify halo structures accurately, even with a small number of tracer stars.

astro-ph.GA

GalaxyFlow: Upsampling Hydrodynamical Simulations for Realistic Mock Stellar Catalogs

Cosmological N-body simulations of galaxies operate at the level of "star particles" with a mass resolution on the scale of thousands of solar masses. Turning these simulations into stellar mock catalogs requires "upsampling" the star particles into individual stars following the same phase-space density. In this paper, we introduce two new upsampling methods. First, we describe GalaxyFlow, a sophisticated upsampling method that utilizes normalizing flows to both estimate the stellar phase space density and sample from it. Second, we improve on existing upsamplers based on adaptive kernel density estimation, using maximum likelihood estimation to fine-tune the bandwidth for such algorithms in a way that improves both the density estimation accuracy and upsampling results. We demonstrate our upsampling techniques on a neighborhood of the Solar location in two simulated galaxies: Auriga 6 and h277. Both yield smooth stellar distributions that closely resemble the stellar densities seen in the Gaia DR3 catalog. Furthermore, we introduce a novel multi-model classifier test to compare the accuracy of different upsampling methods quantitatively. This test confirms that GalaxyFlow estimates the density of the underlying star particles more accurately than methods based on kernel density estimation, at the cost of being more computationally intensive.

astro-ph.GA

Jet Classification Using High-Level Features from Anatomy of Top Jets

Recent advancements in deep learning models have significantly enhanced jet classification performance by analyzing low-level features (LLFs). However, this approach often leads to less interpretable models, emphasizing the need to understand the decision-making process and to identify the high-level features (HLFs) crucial for explaining jet classification. To address this, we consider the top jet tagging problems and introduce an analysis model (AM) that analyzes selected HLFs designed to capture important features of top jets. Our AM mainly consists of the following three modules: a relation network analyzing two-point energy correlations, mathematical morphology and Minkowski functionals for generalizing jet constituent multiplicities, and a recursive neural network analyzing subjet constituent multiplicity to enhance sensitivity to subjet color charges. We demonstrate that our AM achieves performance comparable to the Particle Transformer (ParT) while requiring fewer computational resources in a comparison of top jet tagging using jets simulated at the hadronic calorimeter angular resolution scale. Furthermore, as a more constrained architecture than ParT, the AM exhibits smaller training uncertainties because of the bias-variance tradeoff. We also compare the information content of AM and ParT by decorrelating the features already learned by AM. Lastly, we briefly comment on the results of AM with finer angular resolution inputs.

hep-ph

Measuring Galactic Dark Matter through Unsupervised Machine Learning

Measuring the density profile of dark matter in the Solar neighborhood has important implications for both dark matter theory and experiment. In this work, we apply autoregressive flows to stars from a realistic simulation of a Milky Way-type galaxy to learn -- in an unsupervised way -- the stellar phase space density and its derivatives. With these as inputs, and under the assumption of dynamic equilibrium, the gravitational acceleration field and mass density can be calculated directly from the Boltzmann Equation without the need to assume either cylindrical symmetry or specific functional forms for the galaxy's mass density. We demonstrate our approach can accurately reconstruct the mass density and acceleration profiles of the simulated galaxy, even in the presence of Gaia-like errors in the kinematic measurements.

astro-ph.GA

Morphology for Jet Classification

We introduce a jet tagger based on a neural network analyzing the Minkowski Functionals (MFs) of pixellated jet images. The MFs are geometric measures of binary images, and they can be regarded as a generalization of the particle multiplicity, which is an important quantity in jet tagging. Their changes by dilation encode the jet constituents' geometric structures that appear at various angular scales. We explicitly show that this analysis using the MFs together with mathematical morphology can be considered a constrained convolutional neural network (CNN). Conversely, CNN could model the MFs in a certain limit, and we show their correlation in the example of tagging semi-visible jets emerging from the strong interaction of a hidden valley scenario. The MFs are independent of the IRC-safe observables commonly used in jet physics. We combine this morphological analysis with an IRC-safe relation network which models two-point energy correlations. While the resulting network uses constrained input parameters, it shows comparable dark jet and top jet tagging performances to the CNN. The architecture has significant computational advantages when the available data is limited. We show that its tagging performance is much better than that of the CNN with a small number of training samples. We also qualitatively discuss their parton-shower model dependency. The results suggest that the MFs can be an efficient parameterization of the IRC-unsafe feature space of jets.

hep-ph

Neural Network-based Top Tagger with Two-Point Energy Correlations and Geometry of Soft Emissions

Deep neural networks trained on jet images have been successful in classifying different kinds of jets. In this paper, we identify the crucial physics features that could reproduce the classification performance of the convolutional neural network in the top jet vs. QCD jet classification. We design a neural network that considers two types of substructural features: two-point energy correlations, and the IRC unsafe counting variables of a morphological analysis of jet images. The new set of IRC unsafe variables can be described by Minkowski functionals from integral geometry. To integrate these features into a single framework, we reintroduce two-point energy correlations in terms of a graph neural network and provide the other features to the network afterward. The network shows a comparable classification performance to the convolutional neural network. Since both networks are using IRC unsafe features at some level, the results based on simulations are often dependent on the event generator choice. We compare the classification results of Pythia 8 and Herwig 7, and a simple reweighting on the distribution of IRC unsafe features reduces the difference between the results from the two simulations.

hep-ph

Interpretable Deep Learning for Two-Prong Jet Classification with Jet Spectra

Classification of jets with deep learning has gained significant attention in recent times. However, the performance of deep neural networks is often achieved at the cost of interpretability. Here we propose an interpretable network trained on the jet spectrum $S_{2}(R)$ which is a two-point correlation function of the jet constituents. The spectrum can be derived from a functional Taylor series of an arbitrary jet classifier function of energy flows. An interpretable network can be obtained by truncating the series. The intermediate feature of the network is an infrared and collinear safe C-correlator which allows us to estimate the importance of a $S_{2}(R)$ deposit at an angular scale R in the classification. The performance of the architecture is comparable to that of a convolutional neural network (CNN) trained on jet images, although the number of inputs and complexity of architecture is significantly simpler than the CNN classifier. We consider two examples: one is the classification of two-prong jets which differ in color charge of the mother particle, and the other is a comparison between Pythia 8 and Herwig 7 generated jets.

hep-ph

Spectral Analysis of Jet Substructure with Neural Networks: Boosted Higgs Case

Jets from boosted heavy particles have a typical angular scale which can be used to distinguish them from QCD jets. We introduce a machine learning strategy for jet substructure analysis using a spectral function on the angular scale. The angular spectrum allows us to scan energy deposits over the angle between a pair of particles in a highly visual way. We set up an artificial neural network (ANN) to find out characteristic shapes of the spectra of the jets from heavy particle decays. By taking the Higgs jets and QCD jets as examples, we show that the ANN of the angular spectrum input has similar performance to existing taggers. In addition, some improvement is seen when additional extra radiations occur. Notably, the new algorithm automatically combines the information of the multi-point correlations in the jet.

hep-ph

Monojet Signatures from Heavy Colored Particles: Future Collider Sensitivities and Theoretical Uncertainties

In models with colored particle $\mathcal{Q}$ that can decay into a dark matter candidate $X$, the relevant collider process $pp\to \mathcal{Q}\bar{\mathcal{Q}}\rightarrow X\bar{X}+$jets gives rise to events with significant transverse momentum imbalance. When the masses of $\mathcal{Q}$ and $X$ are very close, the relevant signature becomes monojet-like, and Large Hadron Collider (LHC) search limits become much less constraining. In this paper, we study the current and anticipated experimental sensitivity to such particles at the High-Luminosity LHC at $\sqrt{s}=14\,\mathrm{TeV}$ with $\mathcal{L}=3\,\mathrm{ab}^{-1}$ of data and the proposed High-Energy LHC at $\sqrt{s}=27\,\mathrm{TeV}$ with $\mathcal{L}=15\,\mathrm{ab}^{-1}$ of data. We estimate the reach for various Lorentz and QCD color representations of $\mathcal{Q}$. Identifying the nature of $\mathcal{Q}$ is very important to understanding the physics behind the monojet signature. Therefore, we also study the dependence of the observables built from the $pp\to\mathcal{Q}\bar{\mathcal{Q}} + j $ process on $\mathcal{Q}$ itself. Using the state-of-the-art Monte Carlo suites MadGraph5_aMC@NLO+Pythia8 and Sherpa, we find that when these observables are calculated at NLO in QCD with parton shower matching and multijet merging, the residual theoretical uncertainties are comparable to differences observed when varying the quantum numbers of $\mathcal{Q}$ itself. We find, however, that the precision achievable with NNLO calculations, where available, can resolve this dilemma.

hep-ph

Identifying a new particle with jet substructures

We investigate a potential of measuring properties of a heavy resonance X, exploiting jet substructure techniques. Motivated by heavy higgs boson searches, we focus on the decays of X into a pair of (massive) electroweak gauge bosons. More specifically, we consider a hadronic Z boson, which makes it possible to determine properties of X at an earlier stage. For $m_X$ of O(1) TeV, two quarks from a Z boson would be captured as a "merged jet" in a significant fraction of events. The use of the merged jet enables us to consider a Z-induced jet as a reconstructed object without any combinatorial ambiguity. We apply a conventional jet substructure method to extract four-momenta of subjets from a merged jet. We find that jet substructure procedures may enhance features in some kinematic observables formed with subjets. Subjet momenta are fed into the matrix element associated with a given hypothesis on the nature of X, which is further processed to construct a matrix element method (MEM)-based observable. For both moderately and highly boosted Z bosons, we demonstrate that the MEM with current jet substructure techniques can be a very powerful discriminator in identifying the physics nature of X. We also discuss effects from choosing different jet sizes for merged jets and jet-grooming parameters upon the MEM analyses.

hep-ph

Identifying the production process of new physics at colliders; symmetric or asymmetric?

We propose a class of kinematic variables, which is a smooth generalization of min-max type mass variables such as the Cambridge-$M_{T2}$ and $M_2$, for measuring a mass spectrum of intermediate resonances in a semi-invisibly decaying pair production. While kinematic endpoints of min-max type mass variables are only sensitive to a heavier resonance mass, kinematic endpoints of new variables are sensitive to all masses. These new mass variables can be used to resolve a mass spectrum, so that if the true mass spectrum is asymmetric, then the kinematic endpoints are separate while the endpoints are the same for the symmetric true mass spectrum. We demonstrate the behavior of kinematic endpoint of these new variables in pair production of two-body and three-body decays with one invisible particle.

hep-ph

The 750 GeV Diphoton Excess May Not Imply a 750 GeV Resonance

We discuss non-standard interpretations of the 750 GeV diphoton excess recently reported by the ATLAS and CMS Collaborations which do not involve a new, relatively broad, resonance with a mass near 750 GeV. Instead, we consider the sequential cascade decay of a much heavier, possibly quite narrow, resonance into two photons along with one or more invisible particles. The resulting diphoton invariant mass signal is generically rather broad, as suggested by the data. We examine three specific event topologies - the antler, the sandwich, and the 2-step cascade decay, and show that they all can provide a good fit to the observed published data. In each case, we delineate the preferred mass parameter space selected by the best fit. In spite of the presence of invisible particles in the final state, the measured missing transverse energy is moderate, due to its anti- correlation with the diphoton invariant mass. We comment on the future prospects of discriminating with higher statistics between our scenarios, as well as from more conventional interpretations.

hep-ph

OPTIMASS: A Package for the Minimization of Kinematic Mass Functions with Constraints

Reconstructed mass variables, such as $M_2$, $M_{2C}$, $M_T^\star$, and $M_{T2}^W$, play an essential role in searches for new physics at hadron colliders. The calculation of these variables generally involves constrained minimization in a large parameter space, which is numerically challenging. We provide a C++ code, OPTIMASS, which interfaces with the MINUIT library to perform this constrained minimization using the Augmented Lagrangian Method. The code can be applied to arbitrarily general event topologies and thus allows the user to significantly extend the existing set of kinematic variables. We describe this code and its physics motivation, and demonstrate its use in the analysis of the fully leptonic decay of pair-produced top quarks using the $M_2$ variables.

hep-ph