SearcharxivSearch

arXiv subjects

Inês Ochoa

Publications and source records attributed to Inês Ochoa.

10 recordsLinked to original sources

Searching for HWW Anomalous Couplings with Simulation-Based Inference

Understanding the source of the universe's asymmetry between matter and antimatter is one of the major open questions in particle physics. In this work, the sensitivity of novel machine-learning-based inference techniques to CP-odd and CP-even $HWW$ anomalous couplings is studied in the $\\WH \rightarrow \ell νb\bar{b}$ channel ($\ell = e, μ$), within the Standard Model Effective Field Theory (SMEFT) framework. Two machine-learning simulation-based inference (SBI) methods are explored: a per-event likelihood-ratio estimator, which directly approximates the ratio of probability densities between competing hypotheses, is benchmarked against a per-event optimal-observable estimator optimized for sensitivity to the parameters of interest. Both approaches are also compared to traditional summary statistics, in this case histograms of kinematic and angular observables, as commonly used in experimental analyses. SBI methods provide tighter constraints than one-dimensional summary statistics, though their performance is comparable to two-dimensional histogram analysis. The optimal-observable approach remains promising for its ability to probe multiple couplings simultaneously. Restricting the analysis to a region of high $S/B$ also enhances sensitivity to CP-odd operators while preserving sensitivity to CP-even operators, which histogram analyses often lose. Although the likelihood-ratio estimator sometimes struggles with likelihood minima and shapes, optimisations that target its robustness could make it more sensitive than both the optimal-observable estimator and the histogram method. These results underscore the potential of advanced simulation-based inference techniques, encouraging further exploration with LHC Run 3 data to surpass current ATLAS and CMS sensitivities.

hep-ph

Design optimization of hadronic calorimeters for future colliders

Calorimeters are a crucial component in modern particle detectors. They are responsible for providing accurate energy measurements of particles produced in high-energy collisions. The demanding requirements set for next-generation collider experiments impose new challenges on the design of new detectors, and a systematic approach to their optimization is increasingly necessary. The performance of calorimeters is primarily characterized by their energy resolution, parameterized by a stochastic and a constant term, related to sampling fluctuations and non-uniformities respectively. To improve the reconstruction quality of physics objects in the calorimeter, both terms need to be taken into account. Changes in a longitudinally constrained design usually result in a trade-off between these terms, making optimization a non-trivial task. This work focuses on the optimization of a hadronic sampling calorimeter, based on the FCC-ee ALLEGRO detector concept. By controlling the absorber layer thickness in a Geant4 simulation, the impact of the passive to active material proportion on the deposited energy distribution and resolution can be analyzed. Our methodology aims at exploring the design space with practical considerations, paving the way for the development of a closed optimization framework that can evaluate multiple designs against physics performance targets.

physics.ins-det

The Critical Importance of Software for HEP

Particle physics has an ambitious and broad global experimental programme for the coming decades. Large investments in building new facilities are already underway or under consideration. Scaling the present processing power and data storage needs by the foreseen increase in data rates in the next decade for HL-LHC is not sustainable within the current budgets. As a result, a more efficient usage of computing resources is required in order to realise the physics potential of future experiments. Software and computing are an integral part of experimental design, trigger and data acquisition, simulation, reconstruction, and analysis, as well as related theoretical predictions. A significant investment in computing and software is therefore critical. Advances in software and computing, including artificial intelligence (AI) and machine learning (ML), will be key for solving these challenges. Making better use of new processing hardware such as graphical processing units (GPUs) or ARM chips is a growing trend. This forms part of a computing solution that makes efficient use of facilities and contributes to the reduction of the environmental footprint of HEP computing. The HEP community already provided a roadmap for software and computing for the last EPPSU, and this paper updates that, with a focus on the most resource critical parts of our data processing chain.

hep-ex

Is Tokenization Needed for Masked Particle Modelling?

In this work, we significantly enhance masked particle modeling (MPM), a self-supervised learning scheme for constructing highly expressive representations of unordered sets relevant to developing foundation models for high-energy physics. In MPM, a model is trained to recover the missing elements of a set, a learning objective that requires no labels and can be applied directly to experimental data. We achieve significant performance improvements over previous work on MPM by addressing inefficiencies in the implementation and incorporating a more powerful decoder. We compare several pre-training tasks and introduce new reconstruction methods that utilize conditional generative models without data tokenization or discretization. We show that these new methods outperform the tokenized learning objective from the original MPM on a new test bed for foundation models for jets, which includes using a wide variety of downstream tasks relevant to jet physics, such as classification, secondary vertex finding, and track identification.

hep-ph

Fitting a Collider in a Quantum Computer: Tackling the Challenges of Quantum Machine Learning for Big Datasets

Current quantum systems have significant limitations affecting the processing of large datasets with high dimensionality, typical of high energy physics. In the present paper, feature and data prototype selection techniques were studied to tackle this challenge. A grid search was performed and quantum machine learning models were trained and benchmarked against classical shallow machine learning methods, trained both in the reduced and the complete datasets. The performance of the quantum algorithms was found to be comparable to the classical ones, even when using large datasets. Sequential Backward Selection and Principal Component Analysis techniques were used for feature's selection and while the former can produce the better quantum machine learning models in specific cases, it is more unstable. Additionally, we show that such variability in the results is caused by the use of discrete variables, highlighting the suitability of Principal Component analysis transformed data for quantum machine learning applications in the high energy physics context.

hep-ph

Differentiable Vertex Fitting for Jet Flavour Tagging

We propose a differentiable vertex fitting algorithm that can be used for secondary vertex fitting, and that can be seamlessly integrated into neural networks for jet flavour tagging. Vertex fitting is formulated as an optimization problem where gradients of the optimized solution vertex are defined through implicit differentiation and can be passed to upstream or downstream neural network components for network training. More broadly, this is an application of differentiable programming to integrate physics knowledge into neural network models in high energy physics. We demonstrate how differentiable secondary vertex fitting can be integrated into larger transformer-based models for flavour tagging and improve heavy flavour jet classification.

hep-ex

High-dimensional Anomaly Detection with Radiative Return in $e^{+}e^{-}$ Collisions

Experiments at a future $e^{+}e^{-}$ collider will be able to search for new particles with masses below the nominal centre-of-mass energy by analyzing collisions with initial-state radiation (radiative return). We show that machine learning methods that use imperfect or missing training labels can achieve sensitivity to generic new particle production in radiative return events. In addition to presenting an application of the classification without labels (CWoLa) search method in $e^{+}e^{-}$ collisions, our study combines weak supervision with variable-dimensional information by deploying a deep sets neural network architecture. We have also investigated some of the experimental aspects of anomaly detection in radiative return events and discuss these in the context of future detector design.

hep-ph

Anomalous Jet Identification via Sequence Modeling

This paper presents a novel method of searching for boosted hadronically decaying objects by treating them as anomalous elements of a contaminated dataset. A Variational Recurrent Neural Network (VRNN) is used to model jets as sequences of constituent four-vectors. After applying a pre-processing method which boosts each jet to the same reference mass and energy, the VRNN provides each jet an Anomaly Score that distinguishes between the structure of signal and background jets. The model is trained in an entirely unsupervised setting and without high level variables, making the score more robust against mass and $p_{T}$ correlations when compared to methods based primarily on jet substructure. Performance is evaluated on the jet level, as well as in an analysis context by searching for a heavy resonance with a final state of two boosted jets. The Anomaly Score shows consistent performance along a wide range of signal contamination amounts, for both two and three-pronged jet substructure hypotheses. Analysis results demonstrate that the use of Anomaly Score as a classifier enhances signal sensitivity while retaining a smoothly falling background jet mass distribution. The model's discriminatory performance resulting from an unsupervised training scenario opens up the possibility to train directly on data without a pre-defined signal hypothesis.

hep-ph

The LHC Olympics 2020: A Community Challenge for Anomaly Detection in High Energy Physics

A new paradigm for data-driven, model-agnostic new physics searches at colliders is emerging, and aims to leverage recent breakthroughs in anomaly detection and machine learning. In order to develop and benchmark new anomaly detection methods within this framework, it is essential to have standard datasets. To this end, we have created the LHC Olympics 2020, a community challenge accompanied by a set of simulated collider events. Participants in these Olympics have developed their methods using an R&D dataset and then tested them on black boxes: datasets with an unknown anomaly (or not). This paper will review the LHC Olympics 2020 challenge, including an overview of the competition, a description of methods deployed in the competition, lessons learned from the experience, and implications for data analyses with future datasets as well as future colliders.

hep-ph

Boosted Higgs $\rightarrow b\bar{b}$ in vector-boson associated production at 14 TeV

The production of the Standard Model Higgs boson in association with a vector boson, followed by the dominant decay to $H \rightarrow b\bar{b}$, is a strong prospect for confirming and measuring the coupling to $b$-quarks in $pp$ collisions at $\sqrt{s}=14$ TeV. We present an updated study of the prospects for this analysis, focussing on the most sensitive highly Lorentz-boosted region. The evolution of the efficiency and composition of the signal and main background processes as a function of the transverse momentum of the vector boson are studied covering the region $200-1000$ GeV, comparing both a conventional dijet and jet substructure selection. The lower transverse momentum region ($200-400$ GeV) is identified as the most sensitive region for the Standard Model search, with higher transverse momentum regions not improving the statistical sensitivity. For much of the studied region ($200-600$ GeV), a conventional dijet selection performs as well as the substructure approach, while for the highest transverse momentum regions ($> 600$ GeV), which are particularly interesting for Beyond the Standard Model and high luminosity measurements, the jet substructure techniques are essential.

hep-ph