SearcharxivSearch

arXiv subjects

Avik Roy

Publications and source records attributed to Avik Roy.

At least 19 recordsLinked to original sources

Evidential Deep Learning for Uncertainty Quantification and Out-of-Distribution Detection in Jet Identification using Deep Neural Networks

Current methods commonly used for uncertainty quantification (UQ) in deep learning (DL) models utilize Bayesian methods which are computationally expensive and time-consuming. In this paper, we provide a detailed study of UQ based on evidential deep learning (EDL) for deep neural network models designed to identify jets in high energy proton-proton collisions at the Large Hadron Collider and explore its utility in anomaly detection. EDL is a DL approach that treats learning as an evidence acquisition process designed to provide confidence (or epistemic uncertainty) about test data. Using publicly available datasets for jet classification benchmarking, we explore hyperparameter optimizations for EDL applied to the challenge of UQ for jet identification. We also investigate how the uncertainty is distributed for each jet class, how this method can be implemented for the detection of anomalies, how the uncertainty compares with Bayesian ensemble methods, and how the uncertainty maps onto latent spaces for the models. Our studies uncover some pitfalls of EDL applied to anomaly detection and a more effective way to quantify uncertainty from EDL as compared with the foundational EDL setup. These studies illustrate a methodological approach to interpreting EDL in jet classification models, providing new insights on how EDL quantifies uncertainty and detects out-of-distribution data which may lead to improved EDL methods for DL models applied to classification tasks.

hep-ex

FAIR AI Models in High Energy Physics

The findable, accessible, interoperable, and reusable (FAIR) data principles provide a framework for examining, evaluating, and improving how data is shared to facilitate scientific discovery. Generalizing these principles to research software and other digital products is an active area of research. Machine learning (ML) models -- algorithms that have been trained on data without being explicitly programmed -- and more generally, artificial intelligence (AI) models, are an important target for this because of the ever-increasing pace with which AI is transforming scientific domains, such as experimental high energy physics (HEP). In this paper, we propose a practical definition of FAIR principles for AI models in HEP and describe a template for the application of these principles. We demonstrate the template's use with an example AI model applied to HEP, in which a graph neural network is used to identify Higgs bosons decaying to two bottom quarks. We report on the robustness of this FAIR AI model, its portability across hardware architectures and software frameworks, and its interpretability.

hep-ex

FAIR Principles for data and AI models in high energy physics research and education

In recent years, digital object management practices to support findability, accessibility, interoperability, and reusability (FAIR) have begun to be adopted across a number of data-intensive scientific disciplines. These digital objects include datasets, AI models, software, notebooks, workflows, documentation, etc. With the collective dataset at the Large Hadron Collider scheduled to reach the zettabyte scale by the end of 2032, the experimental particle physics community is looking at unprecedented data management challenges. It is expected that these grand challenges may be addressed by creating end-to-end AI frameworks that combine FAIR and AI-ready datasets, advances in AI, modern computing environments, and scientific data infrastructure. In this work, the FAIR4HEP collaboration explores the interpretation of FAIR principles in the context of data and AI models for experimental high energy physics research. We investigate metrics to quantify the FAIRness of experimental datasets and AI models, and provide open source notebooks to guide new users on the use of FAIR principles in practice.

hep-ex

Interpretability of an Interaction Network for identifying $H \rightarrow b\bar{b}$ jets

Multivariate techniques and machine learning models have found numerous applications in High Energy Physics (HEP) research over many years. In recent times, AI models based on deep neural networks are becoming increasingly popular for many of these applications. However, neural networks are regarded as black boxes -- because of their high degree of complexity it is often quite difficult to quantitatively explain the output of a neural network by establishing a tractable input-output relationship and information propagation through the deep network layers. As explainable AI (xAI) methods are becoming more popular in recent years, we explore interpretability of AI models by examining an Interaction Network (IN) model designed to identify boosted $H\to b\bar{b}$ jets amid QCD background. We explore different quantitative methods to demonstrate how the classifier network makes its decision based on the inputs and how this information can be harnessed to reoptimize the model-making it simpler yet equally effective. We additionally illustrate the activity of hidden layers within the IN model as Neural Activation Pattern (NAP) diagrams. Our experiments suggest NAP diagrams reveal important information about how information is conveyed across the hidden layers of deep model. These insights can be useful to effective model reoptimization and hyperparameter tuning.

hep-ex

Deep Learning for the Matrix Element Method

Extracting scientific results from high-energy collider data involves the comparison of data collected from the experiments with synthetic data produced from computationally-intensive simulations. Comparisons of experimental data and predictions from simulations increasingly utilize machine learning (ML) methods to try to overcome these computational challenges and enhance the data analysis. There is increasing awareness about challenges surrounding interpretability of ML models applied to data to explain these models and validate scientific conclusions based upon them. The matrix element (ME) method is a powerful technique for analysis of particle collider data that utilizes an \textit{ab initio} calculation of the approximate probability density function for a collision event to be due to a physics process of interest. The ME method has several unique and desirable features, including (1) not requiring training data since it is an \textit{ab initio} calculation of event probabilities, (2) incorporating all available kinematic information of a hypothesized process, including correlations, without the need for feature engineering and (3) a clear physical interpretation in terms of transition probabilities within the framework of quantum field theory. These proceedings briefly describe an application of deep learning that dramatically speeds-up ME method calculations and novel cyberinfrastructure developed to execute ME-based analyses on heterogeneous computing platforms.

hep-ex

A Detailed Study of Interpretability of Deep Neural Network based Top Taggers

Recent developments in the methods of explainable AI (XAI) allow researchers to explore the inner workings of deep neural networks (DNNs), revealing crucial information about input-output relationships and realizing how data connects with machine learning models. In this paper we explore interpretability of DNN models designed to identify jets coming from top quark decay in high energy proton-proton collisions at the Large Hadron Collider (LHC). We review a subset of existing top tagger models and explore different quantitative methods to identify which features play the most important roles in identifying the top jets. We also investigate how and why feature importance varies across different XAI metrics, how correlations among features impact their explainability, and how latent space representations encode information as well as correlate with physically meaningful quantities. Our studies uncover some major pitfalls of existing XAI methods and illustrate how they can be overcome to obtain consistent and meaningful interpretation of these models. We additionally illustrate the activity of hidden layers as Neural Activation Pattern (NAP) diagrams and demonstrate how they can be used to understand how DNNs relay information across the layers and how this understanding can help to make such models significantly simpler by allowing effective model reoptimization and hyperparameter tuning. These studies not only facilitate a methodological approach to interpreting models but also unveil new insights about what these models learn. Incorporating these observations into augmented model design, we propose the Particle Flow Interaction Network (PFIN) model and demonstrate how interpretability-inspired model augmentation can improve top tagging performance.

hep-ex

FAIR for AI: An interdisciplinary and international community building perspective

A foundational set of findable, accessible, interoperable, and reusable (FAIR) principles were proposed in 2016 as prerequisites for proper data management and stewardship, with the goal of enabling the reusability of scholarly data. The principles were also meant to apply to other digital assets, at a high level, and over time, the FAIR guiding principles have been re-interpreted or extended to include the software, tools, algorithms, and workflows that produce data. FAIR principles are now being adapted in the context of AI models and datasets. Here, we present the perspectives, vision, and experiences of researchers from different countries, disciplines, and backgrounds who are leading the definition and adoption of FAIR principles in their communities of practice, and discuss outcomes that may result from pursuing and incentivizing FAIR AI research. The material for this report builds on the FAIR for AI Workshop held at Argonne National Laboratory on June 7, 2022.

cs.CY

Making Digital Objects FAIR in High Energy Physics: An Implementation for Universal FeynRules Output (UFO) Models

Research in the data-intensive discipline of high energy physics (HEP) often relies on domain-specific digital contents. Reproducibility of research relies on proper preservation of these digital objects. This paper reflects on the interpretation of principles of Findability, Accessibility, Interoperability, and Reusability (FAIR) in such context and demonstrates its implementation by describing the development of an end-to-end support infrastructure for preserving and accessing Universal FeynRules Output (UFO) models guided by the FAIR principles. UFO models are custom-made python libraries used by the HEP community for Monte Carlo simulation of collider physics events. Our framework provides simple but robust tools to preserve and access the UFO models and corresponding metadata in accordance with the FAIR principles.

hep-ph

Data Science and Machine Learning in Education

The growing role of data science (DS) and machine learning (ML) in high-energy physics (HEP) is well established and pertinent given the complex detectors, large data, sets and sophisticated analyses at the heart of HEP research. Moreover, exploiting symmetries inherent in physics data have inspired physics-informed ML as a vibrant sub-field of computer science research. HEP researchers benefit greatly from materials widely available materials for use in education, training and workforce development. They are also contributing to these materials and providing software to DS/ML-related fields. Increasingly, physics departments are offering courses at the intersection of DS, ML and physics, often using curricula developed by HEP researchers and involving open software and data used in HEP. In this white paper, we explore synergies between HEP research and DS/ML education, discuss opportunities and challenges at this intersection, and propose community activities that will be mutually beneficial.

physics.ed-ph

Explainable AI for High Energy Physics

Neural Networks are ubiquitous in high energy physics research. However, these highly nonlinear parameterized functions are treated as \textit{black boxes}- whose inner workings to convey information and build the desired input-output relationship are often intractable. Explainable AI (xAI) methods can be useful in determining a neural model's relationship with data toward making it \textit{interpretable} by establishing a quantitative and tractable relationship between the input and the model's output. In this letter of interest, we explore the potential of using xAI methods in the context of problems in high energy physics.

hep-ex

Non-resonant Diagrams for Single Production of Top and Bottom Partners

The search for Top and Bottom Partners is a major focus of analyses at both the ATLAS and CMS experiments. When singly produced, these vector-like partners of the Standard Model third generation quarks retain a sizeable cross-section that makes them attractive search candidates in their respective topologies. While most efforts have concentrated on the resonant mode for single production of these hypothetical particles, the most dominant mode at narrow widths, recent studies have revealed a wide and rich phenomenology involving the non-resonant diagrams. In this letter we categorically investigate the impact of the non-resonant diagrams on different Top and Bottom partner single production topologies and their impact on the inclusive cross-section estimation. We also propose a parameterization for calculating the the correction factor to the single vector-like quark production cross-section when such diagrams are included.

hep-ph

Recipes for when Physics Fails: Recovering Robust Learning of Physics Informed Neural Networks

Physics-informed Neural Networks (PINNs) have been shown to be effective in solving partial differential equations by capturing the physics induced constraints as a part of the training loss function. This paper shows that a PINN can be sensitive to errors in training data and overfit itself in dynamically propagating these errors over the domain of the solution of the PDE. It also shows how physical regularizations based on continuity criteria and conservation laws fail to address this issue and rather introduce problems of their own causing the deep network to converge to a physics-obeying local minimum instead of the global minimum. We introduce Gaussian Process (GP) based smoothing that recovers the performance of a PINN and promises a robust architecture against noise/errors in measurements. Additionally, we illustrate an inexpensive method of quantifying the evolution of uncertainty based on the variance estimation of GPs on boundary data. Robust PINN performance is also shown to be achievable by choice of sparse sets of inducing points based on sparsely induced GPs. We demonstrate the performance of our proposed methods and compare the results from existing benchmark models in literature for time-dependent Schr\"odinger and Burgers' equations.

cs.LG

Invariance-based Multi-Clustering of Latent Space Embeddings for Equivariant Learning

Variational Autoencoders (VAEs) have been shown to be remarkably effective in recovering model latent spaces for several computer vision tasks. However, currently trained VAEs, for a number of reasons, seem to fall short in learning invariant and equivariant clusters in latent space. Our work focuses on providing solutions to this problem and presents an approach to disentangle equivariance feature maps in a Lie group manifold by enforcing deep, group-invariant learning. Simultaneously implementing a novel separation of semantic and equivariant variables of the latent space representation, we formulate a modified Evidence Lower BOund (ELBO) by using a mixture model pdf like Gaussian mixtures for invariant cluster embeddings that allows superior unsupervised variational clustering. Our experiments show that this model effectively learns to disentangle the invariant and equivariant representations with significant improvements in the learning rate and an observably superior image recognition and canonical state reconstruction compared to the currently best deep learning models.

cs.LG

Novel Interpretation Strategy for Searches of Singly Produced Vector-like Quarks at the LHC

Vector-like Quarks (VLQs) are potential signatures of physics beyond the Standard Model at the TeV energy scale and major efforts have been put forward at both ATLAS and CMS experiments in search of these particles. In order to make these search results more relatable in the context of most plausible theories of VLQs, it is deemed important to present the analysis results in a general fashion. We investigate the challenges associated with such interpretations of singly produced VLQ searches and propose a generalized, semi-analytical framework that allows a model-independent casting of the results in terms of unconstrained, free parameters of the VLQ Lagrangian. We also propose a simple parameterization of the correction factor to the single VLQ production cross-section at large decay widths. We illustrate how the proposed framework can be used to conveniently represent statistical limits by numerically reinterpreting results from benchmark ATLAS and CMS analyses.

hep-ph

Modifications of the Page Curve from correlations within Hawking radiation

We investigate quantum correlations between successive steps of black hole evaporation and investigate whether they might resolve the black hole information paradox. 'Small' corrections in various models were shown to be unable to restore unitarity. We study a toy qubit model of evaporation that allows small quantum correlations between successive steps and reaffirm previous results. Then, we relax the 'smallness' condition and find a nontrivial upper and lower bound on the entanglement entropy change during the evaporation process. This gives a quantitative measure of the size of the correction needed to restore unitarity. We find that these entanglement entropy bounds lead to a significant deviation from the expected Page curve.

hep-th

Does Considering Quantum Correlations Resolve the Information Paradox?

In this paper, we analyze whether quantum correlations between successive steps of evaporation can open any way to resolve the black hole information paradox. Recently a celebrated result in literature shows that `small' correction to leading order Hawking analysis fails to restore unitarity in black hole evaporation. We study a toy qubit model of evaporation allowing small quantum correlations between successive steps and verify the previous result. Then we generalize the concept of correction to Hawking state by relaxing the `smallness' condition. Our result generates a nontrivial upper and lower bound on change in entanglement entropy in the evaporation process. This gives us a quantitative measure of correction that would mathematically facilitate restoration of unitarity in black hole evaporation. We then investigate whether this result is compatible to the established physical constraints of unitary evolution of a state in a subsystem. We find that the generalized bound on entanglement entropy leads to significant deviation from Page curve. This leads us to agree with the recent claim in literature that no amount of correction in the form of Bell pair states would lead to any resolution to the information paradox.

hep-th

Tunneling across the quantum horizon does not resolve the information paradox

Parikh and Wilczek formulated Hawking radiation as quantum tunneling across the event horizon proving the spectrum to be nonthermal. These nonthermality factors emerging due to back reaction effects have been claimed to be responsible for correlations among the emitted quanta. It has been proposed by several authors in literature that these correlations actually carry out information locked in a black hole and hence provide a resolution to the long debated black hole information paradox. This paper demonstrates that this is a fallacious proposition. Finally, it formulates the implications of the no-hair theorem in the context of Parikh-Wilczek spectrum.

hep-th

Comments on Information Erasure in Black Hole

We analyze the Kim, Lee & Lee model of information erasure by black holes and find contradictions with standard physical laws. We demonstrate that the erasure model leads to arbitrarily fast information erasure; the proposed physical interpretation of information freezing at the event horizon as observed by an asymptotic observer is problematic; and information erasure, whatever the process may be, near the black hole horizon leads to contradictions with quantum mechanics if Landauer's principle is assumed. The later part of the work demonstrates the significance of the "erasure entropy." We show that the erasure entropy is the mutual information between two subsystems.

hep-th