SearcharxivSearch

arXiv subjects

Joosep Pata

Publications and source records attributed to Joosep Pata.

At least 19 recordsLinked to original sources

ParticleTransformer is all you need for reconstructing hadronic tau leptons

The large number of $Z \rightarrow \tau\tau$ events expected during the TeraZ program at FCC-ee will allow for precision measurements and searches for physics beyond the Standard Model, requiring accurate reconstruction of hadronically decaying tau leptons. This reconstruction is particularly challenging due to the presence of undetected neutrinos and the diverse topology of hadronic tau decays, making the design of robust heuristic reconstruction algorithms challenging. In this work, we present the first fully machine learned hadronic tau reconstruction approach tuned for FCC-ee studies. The reconstruction is formulated as a set of complementary tasks, including tau identification, decay mode classification, charge reconstruction, and full four-momentum regression. The algorithms are evaluated on fully simulated electron-positron collision samples with realistic detector effects using the CLD detector setup. We compare dedicated task-specific models with a unified multi-task model and quantify their performance in a granular manner across all reconstruction tasks. Both approaches achieve per-mille-level tau mis-identification rates at high signal efficiency, decay mode classification F1 scores of up to 0.95 for the dominant channels, and sub-per-mille charge mis-identification rates, outperforming a conventional jet-charge estimator by up to two orders of magnitude. For the full kinematic reconstruction, the models achieve per-mille-level angular resolution and percent-level visible transverse momentum resolution, exceeding the performance of reconstruction-level jet observables. The resulting models provide a realistic high-performance solution for hadronic tau reconstruction at FCC-ee, offering identification, charge discrimination, decay mode analysis and full kinematic reconstruction.

hep-ex

Machine-learned particle flow as a foundation model for collider physics

The workflow from particle collision to physics analysis passes through a series of reconstruction steps that are traditionally modular and disconnected, with no shared representation linking low-level detector data to high-level analysis tasks. We show that casting event reconstruction as a machine learning problem naturally produces such a shared representation. We repurpose a machine learning model trained for particle-flow reconstruction (MLPF) to perform three distinct analysis tasks: jet flavor identification, jet energy regression, and missing momentum regression. By appending the per-particle latent representations learned during reconstruction as additional input features, we substantially improve over baselines that use kinematic features alone. We further demonstrate that a single linear layer trained using only the latent representations achieves competitive performance against state-of-the-art baseline architectures, and outperforms the baseline for missing momentum regression with approximately 35 times fewer parameters. These results demonstrate that the latent representations learned during reconstruction encode essential physics information needed for downstream analysis, establishing MLPF as a foundation model and offering a concrete step toward an end-to-end pipeline from detector data to physics analysis.

hep-ex

Fine-tuning machine-learned particle-flow reconstruction for new detector geometries in future colliders

We demonstrate transfer learning capabilities in a machine-learned algorithm trained for particle-flow reconstruction in high energy particle colliders. This paper presents a cross-detector fine-tuning study, where we initially pretrain the model on a large full simulation dataset from one detector design, and subsequently fine-tune the model on a sample with a different collider and detector design. Specifically, we use the Compact Linear Collider detector (CLICdet) model for the initial training set and demonstrate successful knowledge transfer to the CLIC-like detector (CLD) proposed for the Future Circular Collider in electron-positron mode. We show that with an order of magnitude less samples from the second dataset, we can achieve the same performance as a costly training from scratch, across particle-level and event-level performance metrics, including jet and missing transverse momentum resolution. Furthermore, we find that the fine-tuned model achieves comparable performance to the traditional rule-based particle-flow approach on event-level metrics after training on 100,000 CLD events, whereas a model trained from scratch requires at least 1 million CLD events to achieve similar reconstruction performance. To our knowledge, this represents the first full-simulation cross-detector transfer learning study for particle-flow reconstruction. These findings offer valuable insights towards building large foundation models that can be fine-tuned across different detector designs and geometries, helping to accelerate the development cycle for new detectors and opening the door to rapid detector design and optimization using machine learning.

hep-ex

Reconstructing hadronically decaying tau leptons with a jet foundation model

The limited availability and accuracy of simulated data has motivated the use of foundation models in high energy physics, with the idea to first train a task-agnostic model on large and potentially unlabeled datasets. This enables the subsequent fine-tuning of the learned representation for specific downstream tasks, potentially requiring much smaller dataset sizes to reach the performance of models trained from scratch. We study how OmniJet-$α$, one of the proposed foundation models for particle jets, can be used on a new set of tasks, and in a new dataset, in order to reconstruct hadronically decaying $τ$ leptons. We show that the pretraining can successfully be utilized for this multi-task problem, improving the resolution of momentum reconstruction by about 50\% when the pretrained weights are fine-tuned, compared to training the model from scratch. While much work remains ahead to develop generic foundation models for high-energy physics, this early result of generalizing an existing model to a new dataset and to previously unconsidered tasks highlights the importance of testing the approaches on a diverse set of datasets and tasks.

hep-ex

On the detection of stellar wakes in the Milky Way: a deep learning approach

Due to poor observational constraints on the low-mass end of the subhalo mass function, the detection of dark matter (DM) subhalos on sub-galactic scales would provide valuable information about the nature of DM. Stellar wakes, induced by passing DM subhalos, encode information about the mass of the inducing perturber and thus serve as an indirect probe for the DM substructure within the Milky Way (MW). Our aim is to assess the viability and performance of deep learning searches for stellar wakes in the Galactic stellar halo caused by DM subhalos of varying mass. We simulate massive objects (subhalos) moving through a homogeneous medium of DM and star particles, with phase-space parameters tailored to replicate the conditions of the Galaxy at a specific distance from the Galactic center. The simulation data is used to train deep neural networks with the purpose of inferring both the presence and mass of the moving perturber, and assess subhalo detectability in varying conditions of the Galactic stellar and DM halos. We find that our binary classifier is able to infer the presence of subhalos, showing non-trivial performance down to a subhalo mass of $5 \times 10^7 \rm \, M_\odot$. We also find that our binary classifier is generalisable to datasets describing subhalo orbits at different Galactocentric distances. In a multiple-hypothesis case, we are able to discern between samples containing subhalos of different masses. Out of the phase-space observables available to us, we conclude that overdensity and velocity divergence are the most important features for subhalo detection performance.

astro-ph.GA

A unified machine learning approach for reconstructing hadronically decaying tau leptons

Tau leptons serve as an important tool for studying the production of Higgs and electroweak bosons, both within and beyond the Standard Model of particle physics. Accurate reconstruction and identification of hadronically decaying tau leptons is a crucial task for current and future high energy physics experiments. Given the advances in jet tagging, we demonstrate how tau lepton reconstruction can be decomposed into tau identification, kinematic reconstruction, and decay mode classification in a multi-task machine learning setup. Based on an electron-positron collision dataset with full detector simulation and reconstruction, we show that common jet tagging architectures can be effectively used for these sub-tasks. We achieve comparable momentum resolutions of 2-3% with all the tested models, while the precision of reconstructing individual decay modes is between 80-95%. We find ParticleTransformer to be the best-performing approach, significantly outperforming the heuristic baseline. This paper also serves as an introduction to a new publicly available $\mathtt{Fu}τ\mathtt{ure}$ dataset for the development of tau reconstruction algorithms. This allows to further study the resilience of ML models to domain shifts and the efficient use of foundation models for such tasks.

hep-ex

Improved particle-flow event reconstruction with scalable neural networks for current and future particle detectors

Efficient and accurate algorithms are necessary to reconstruct particles in the highly granular detectors anticipated at the High-Luminosity Large Hadron Collider and the Future Circular Collider. We study scalable machine learning models for event reconstruction in electron-positron collisions based on a full detector simulation. Particle-flow reconstruction can be formulated as a supervised learning task using tracks and calorimeter clusters. We compare a graph neural network and kernel-based transformer and demonstrate that we can avoid quadratic operations while achieving realistic reconstruction. We show that hyperparameter tuning significantly improves the performance of the models. The best graph neural network model shows improvement in the jet transverse momentum resolution by up to 50% compared to the rule-based algorithm. The resulting model is portable across Nvidia, AMD and Habana hardware. Accurate and fast machine-learning based reconstruction can significantly improve future measurements at colliders.

physics.data-an

Tau lepton identification and reconstruction: a new frontier for jet-tagging ML algorithms

Identifying and reconstructing hadronic $τ$ decays ($τ_{\textrm{h}}$) is an important task at current and future high-energy physics experiments, as $τ_{\textrm{h}}$ represent an important tool to analyze the production of Higgs and electroweak bosons as well as to search for physics beyond the Standard Model. The identification of $τ_{\textrm{h}}$ can be viewed as a generalization and extension of jet-flavour tagging, which has in the recent years undergone significant progress due to the use of deep learning. Based on a granular simulation with realistic detector effects and a particle flow-based event reconstruction, we show in this paper that deep learning-based jet-flavour-tagging algorithms are powerful $τ_{\textrm{h}}$ identifiers. Specifically, we show that jet-flavour-tagging algorithms such as LorentzNet and ParticleTransformer can be adapted in an end-to-end fashion for discriminating $τ_{\textrm{h}}$ from quark and gluon jets. We find that the end-to-end transformer-based approach significantly outperforms contemporary state-of-the-art $τ_{\textrm{h}}$ reconstruction and identification algorithms currently in use at the Large Hadron Collider.

hep-ex

A Bayesian estimation of the Milky Way's circular velocity curve using Gaia DR3

Our goal is to calculate the circular velocity curve of the Milky Way, along with corresponding uncertainties that quantify various sources of systematic uncertainty in a self-consistent manner. The observed rotational velocities are described as circular velocities minus the asymmetric drift. The latter is described by the radial axisymmetric Jeans equation. We thus reconstruct the circular velocity curve between Galactocentric distances from 5 kpc to 14 kpc using a Bayesian inference approach. The estimated error bars quantify uncertainties in the Sun's Galactocentric distance and the spatial-kinematic morphology of the tracer stars. As tracers, we used a sample of roughly 0.6 million stars on the red giant branch stars with six-dimensional phase-space coordinates from Gaia data release 3 (DR3). More than 99% of the sample is confined to a quarter of the stellar disc with mean radial, rotational, and vertical velocity dispersions of $(35\pm 18)\,\rm km/s$, $(25\pm 13)\,\rm km/s$, and $(19\pm 9)\,\rm km/s$, respectively. We find a circular velocity curve with a slope of $0.4\pm 0.6\,\rm km/s/kpc$, which is consistent with a flat curve within the uncertainties. We further estimate a circular velocity at the Sun's position of $v_c(R_0)=233\pm7\, \rm km/s$ and that a region in the Sun's vicinity, characterised by a physical length scale of $\sim 1\,\rm kpc$, moves with a bulk motion of $V_{LSR} =7\pm 7\,\rm km/s$. Finally, we estimate that the dark matter (DM) mass within 14 kpc is $\log_{10}M_{\rm DM}(R<14\, {\rm kpc})/{\rm M_{\odot}}= \left(11.2^{+2.0}_{-2.3}\right)$ and the local spherically averaged DM density is $ρ_{\rm DM}(R_0)=\left(0.41^{+0.10}_{-0.09}\right)\,{\rm GeV/cm^3}=\left(0.011^{+0.003}_{-0.002}\right)\,{\rm M_\odot/pc^3}$. In addition, the effect of biased distance estimates on our results is assessed.

astro-ph.GA

Dynamics of false vacuum bubbles with trapped particles

We study the impact of the ambient fluid on the evolution of collapsing false vacuum bubbles by simulating the dynamics of a coupled bubble-particle system. A significant increase in the mass of the particles across the bubble wall leads to a buildup of those particles inside the false vacuum bubble. We show that the backreaction of the particles on the bubble slows or even reverses the collapse. Consequently, if the particles in the true vacuum become heavier than in the false vacuum, the particle-wall interactions always decrease the compactness that the false vacuum bubbles can reach making their collapse to black holes less likely.

hep-ph

Progress towards an improved particle flow algorithm at CMS with machine learning

The particle-flow (PF) algorithm, which infers particles based on tracks and calorimeter clusters, is of central importance to event reconstruction in the CMS experiment at the CERN LHC, and has been a focus of development in light of planned Phase-2 running conditions with an increased pileup and detector granularity. In recent years, the machine learned particle-flow (MLPF) algorithm, a graph neural network that performs PF reconstruction, has been explored in CMS, with the possible advantages of directly optimizing for the physical quantities of interest, being highly reconfigurable to new conditions, and being a natural fit for deployment to heterogeneous accelerators. We discuss progress in CMS towards an improved implementation of the MLPF reconstruction, now optimized using generator/simulation-level particle information as the target for the first time. This paves the way to potentially improving the detector response in terms of physical quantities of interest. We describe the simulation-based training target, progress and studies on event-based loss terms, details on the model hyperparameter tuning, as well as physics validation with respect to the current PF algorithm in terms of high-level physical quantities such as the jet and missing transverse momentum resolutions. We find that the MLPF algorithm, trained on a generator/simulator level particle information for the first time, results in broadly compatible particle and jet reconstruction performance with the baseline PF, setting the stage for improving the physics performance by additional training statistics and model tuning.

physics.data-an

Sensitivity Estimation for Dark Matter Subhalos in Synthetic Gaia DR2 using Deep Learning

The abundance of dark matter (DM) subhalos orbiting a host galaxy is a generic prediction of the cosmological framework, and is a promising way to constrain the nature of DM. In this paper, we investigate the use of machine learning-based tools to quantify the magnitude of phase-space perturbations caused by the passage of DM subhalos. A simple binary classifier and an anomaly detection model are proposed to estimate if stars or star particles close to DM subhalos are statistically detectable in simulations. The simulated datasets are three Milky Way-like galaxies and nine synthetic Gaia DR2 surveys derived from these. Firstly, we find that the anomaly detection algorithm, trained on a simulated galaxy with full 6D kinematic observables and applied on another galaxy, is nontrivially sensitive to the DM subhalo population. On the other hand, the classification-based approach is not sufficiently sensitive due to the extremely low statistics of signal stars for supervised training. Finally, the sensitivity of both algorithms in the Gaia-like surveys is negligible. The enormous size of the Gaia dataset motivates the further development of scalable and accurate data analysis methods that could be used to select potential regions of interest for DM searches to ultimately constrain the Milky Way's subhalo mass function, as well as simulations where to study the sensitivity of such methods under different signal hypotheses.

astro-ph.GA

Hyperparameter optimization of data-driven AI models on HPC systems

In the European Center of Excellence in Exascale computing "Research on AI- and Simulation-Based Engineering at Exascale" (CoE RAISE), researchers develop novel, scalable AI technologies towards Exascale. This work exercises High Performance Computing resources to perform large-scale hyperparameter optimization using distributed training on multiple compute nodes. This is part of RAISE's work on data-driven use cases which leverages AI- and HPC cross-methods developed within the project. In response to the demand for parallelizable and resource efficient hyperparameter optimization methods, advanced hyperparameter search algorithms are benchmarked and compared. The evaluated algorithms, including Random Search, Hyperband and ASHA, are tested and compared in terms of both accuracy and accuracy per compute resources spent. As an example use case, a graph neural network model known as MLPF, developed for the task of Machine-Learned Particle-Flow reconstruction in High Energy Physics, acts as the base model for optimization. Results show that hyperparameter optimization significantly increased the performance of MLPF and that this would not have been possible without access to large-scale High Performance Computing resources. It is also shown that, in the case of MLPF, the ASHA algorithm in combination with Bayesian optimization gives the largest performance increase per compute resources spent out of the investigated algorithms.

physics.data-an

Machine Learning for Particle Flow Reconstruction at CMS

We provide details on the implementation of a machine-learning based particle flow algorithm for CMS. The standard particle flow algorithm reconstructs stable particles based on calorimeter clusters and tracks to provide a global event reconstruction that exploits the combined information of multiple detector subsystems, leading to strong improvements for quantities such as jets and missing transverse energy. We have studied a possible evolution of particle flow towards heterogeneous computing platforms such as GPUs using a graph neural network. The machine-learned PF model reconstructs particle candidates based on the full list of tracks and calorimeter clusters in the event. For validation, we determine the physics performance directly in the CMS software framework when the proposed algorithm is interfaced with the offline reconstruction of jets and missing transverse energy. We also report the computational performance of the algorithm, which scales approximately linearly in runtime and memory usage with the input size.

physics.data-an

Explaining machine-learned particle-flow reconstruction

The particle-flow (PF) algorithm is used in general-purpose particle detectors to reconstruct a comprehensive particle-level view of the collision by combining information from different subdetectors. A graph neural network (GNN) model, known as the machine-learned particle-flow (MLPF) algorithm, has been developed to substitute the rule-based PF algorithm. However, understanding the model's decision making is not straightforward, especially given the complexity of the set-to-set prediction task, dynamic graph building, and message-passing steps. In this paper, we adapt the layerwise-relevance propagation technique for GNNs and apply it to the MLPF algorithm to gauge the relevant nodes and features for its predictions. Through this process, we gain insight into the model's decision-making.

physics.data-an

MLPF: Efficient machine-learned particle-flow reconstruction using graph neural networks

In general-purpose particle detectors, the particle-flow algorithm may be used to reconstruct a comprehensive particle-level view of the event by combining information from the calorimeters and the trackers, significantly improving the detector resolution for jets and the missing transverse momentum. In view of the planned high-luminosity upgrade of the CERN Large Hadron Collider (LHC), it is necessary to revisit existing reconstruction algorithms and ensure that both the physics and computational performance are sufficient in an environment with many simultaneous proton-proton interactions (pileup). Machine learning may offer a prospect for computationally efficient event reconstruction that is well-suited to heterogeneous computing platforms, while significantly improving the reconstruction quality over rule-based algorithms for granular detectors. We introduce MLPF, a novel, end-to-end trainable, machine-learned particle-flow algorithm based on parallelizable, computationally efficient, and scalable graph neural networks optimized using a multi-task objective on simulated events. We report the physics and computational performance of the MLPF algorithm on a Monte Carlo dataset of top quark-antiquark pairs produced in proton-proton collisions in conditions similar to those expected for the high-luminosity LHC. The MLPF algorithm improves the physics response with respect to a rule-based benchmark algorithm and demonstrates computationally scalable particle-flow reconstruction in a high-pileup environment.

physics.data-an

Graph Neural Networks for Particle Reconstruction in High Energy Physics detectors

Pattern recognition problems in high energy physics are notably different from traditional machine learning applications in computer vision. Reconstruction algorithms identify and measure the kinematic properties of particles produced in high energy collisions and recorded with complex detector systems. Two critical applications are the reconstruction of charged particle trajectories in tracking detectors and the reconstruction of particle showers in calorimeters. These two problems have unique challenges and characteristics, but both have high dimensionality, high degree of sparsity, and complex geometric layouts. Graph Neural Networks (GNNs) are a relatively new class of deep learning architectures which can deal with such data effectively, allowing scientists to incorporate domain knowledge in a graph structure and learn powerful representations leveraging that structure to identify patterns of interest. In this work we demonstrate the applicability of GNNs to these two diverse particle reconstruction problems.

physics.ins-det

Processing Columnar Collider Data with GPU-Accelerated Kernels

At high energy physics experiments, processing billions of records of structured numerical data from collider events to a few statistical summaries is a common task. The data processing is typically more complex than standard query languages allow, such that custom numerical codes are used. At present, these codes mostly operate on individual event records and are parallelized in multi-step data reduction workflows using batch jobs across CPU farms. Based on a simplified top quark pair analysis with CMS Open Data, we demonstrate that it is possible to carry out significant parts of a collider analysis at a rate of around a million events per second on a single multicore server with optional GPU acceleration. This is achieved by representing HEP event data as memory-mappable sparse arrays of columns, and by expressing common analysis operations as kernels that can be used to process the event data in parallel. We find that only a small number of relatively simple functional kernels are needed for a generic HEP analysis. The approach based on columnar processing of data could speed up and simplify the cycle for delivering physics results at HEP experiments. We release the \texttt{hepaccelerate} prototype library as a demonstrator of such methods.

physics.data-an