SearcharxivSearch

arXiv subjects

Marco Zanetti

Publications and source records attributed to Marco Zanetti.

17 recordsLinked to original sources

Learning to Validate Generative Models: a Goodness-of-Fit Approach

Generative models are increasingly central to scientific workflows, yet their systematic use and interpretation require a proper understanding of their limitations through rigorous validation. Classic approaches struggle with scalability, statistical power, or interpretability when applied to high-dimensional data, making it difficult to certify the reliability of these models in realistic, high-dimensional scientific settings. Here, we propose the use of the New Physics Learning Machine (NPLM), a learning-based approach to goodness-of-fit testing inspired by the Neyman--Pearson construction, to test generative networks trained on high-dimensional scientific data. We demonstrate the performance of NPLM for validation in two benchmark cases: generative models trained on mixtures of Gaussian models with increasing dimensionality, and a public end-to-end model, known as FlowSim, developed to generate high-energy physics collision events. We demonstrate that the NPLM can serve as a powerful validation method while also providing a means to diagnose sub-optimally modeled regions of the data.

stat.ML

Ultra-low latency quantum-inspired machine learning predictors implemented on FPGA

Tensor Networks (TNs) are a computational paradigm used for representing quantum many-body systems. Recent works have shown how TNs can also be applied to perform Machine Learning (ML) tasks, yielding comparable results to standard supervised learning techniques. In this work, we study the use of Tree Tensor Networks (TTNs) in high-frequency real-time applications by exploiting the low-latency hardware of the Field-Programmable Gate Array (FPGA) technology. We present different implementations of TTN classifiers, capable of performing inference on classical ML datasets as well as on complex physics data. A preparatory analysis of bond dimensions and weight quantization is realized in the training phase, together with entanglement entropy and correlation measurements, that help setting the choice of the TTN architecture. The generated TTNs are then deployed on a hardware accelerator; using an FPGA integrated into a server, the inference of the TTN is completely offloaded. Eventually, a classifier for High Energy Physics (HEP) applications is implemented and executed fully pipelined with sub-microsecond latency.

hep-ex

Triggerless data acquisition pipeline for Machine Learning based statistical anomaly detection

This work describes an online processing pipeline designed to identify anomalies in a continuous stream of data collected without external triggers from a particle detector. The processing pipeline begins with a local reconstruction algorithm, employing neural networks on an FPGA as its first stage. Subsequent data preparation and anomaly detection stages are accelerated using GPGPUs. As a practical demonstration of anomaly detection, we have developed a data quality monitoring application using a cosmic muon detector. Its primary objective is to detect deviations from the expected operational conditions of the detector. This serves as a proof-of-concept for a system that can be adapted for use in large particle physics experiments, enabling anomaly detection on datasets with reduced bias.

hep-ex

A fast and flexible machine learning approach to data quality monitoring

We present a machine learning based approach for real-time monitoring of particle detectors. The proposed strategy evaluates the compatibility between incoming batches of experimental data and a reference sample representing the data behavior in normal conditions by implementing a likelihood-ratio hypothesis test. The core model is powered by recent large-scale implementations of kernel methods, nonparametric learning algorithms that can approximate any continuous function given enough data. The resulting algorithm is fast, efficient and agnostic about the type of potential anomaly in the data. We show the performance of the model on multivariate data from a drift tube chambers muon detector.

hep-ex

Fast kernel methods for Data Quality Monitoring as a goodness-of-fit test

We here propose a machine learning approach for monitoring particle detectors in real-time. The goal is to assess the compatibility of incoming experimental data with a reference dataset, characterising the data behaviour under normal circumstances, via a likelihood-ratio hypothesis test. The model is based on a modern implementation of kernel methods, nonparametric algorithms that can learn any continuous function given enough data. The resulting approach is efficient and agnostic to the type of anomaly that may be present in the data. Our study demonstrates the effectiveness of this strategy on multivariate data from drift tube chamber muon detectors.

hep-ex

Learning new physics efficiently with nonparametric methods

We present a machine learning approach for model-independent new physics searches. The corresponding algorithm is powered by recent large-scale implementations of kernel methods, nonparametric learning algorithms that can approximate any continuous function given enough data. Based on the original proposal by D'Agnolo and Wulzer (arXiv:1806.02350), the model evaluates the compatibility between experimental data and a reference model, by implementing a hypothesis testing procedure based on the likelihood ratio. Model-independence is enforced by avoiding any prior assumption about the presence or shape of new physics components in the measurements. We show that our approach has dramatic advantages compared to neural network implementations in terms of training times and computational resources, while maintaining comparable performances. In particular, we conduct our tests on higher dimensional datasets, a step forward with respect to previous studies.

hep-ph

A horizontally scalable online processing system for trigger-less data acquisition

The majority of high energy physics experiments rely on data acquisition and hardware-based trigger systems performing a number of stringent selections before storing data for offline analysis. The online reconstruction and selection performed at the trigger level are bound to the synchronous nature of the data acquisition system, resulting in a trade-off between the amount of data collected and the complexity of the online reconstruction performed. Exotic physics processes, such as long-lived and slow-moving particles, are rarely targeted by online triggers as they require complex and nonstandard online reconstruction, usually incompatible with the time constraints of most data acquisition systems. The online trigger selection can thus impact as one of the main limiting factors to the experimental reach for exotic signatures. Alternative data acquisition solutions based on the continuous and asynchronous processing of the stream of data from the detectors are therefore foreseeable. Trigger-less data readout systems, paired with efficient streaming data processing solutions, can provide a viable alternative. In this document, an end-to-end implementation of a fully trigger-less data acquisition and online processing system is discussed. An easily scalable and deployable implementation of such an architecture is proposed, based on open-source distributed computing frameworks capable of performing asynchronous online processing of streaming data. The proposed schema can be suitable for deployment as a fully integrated system for small-scale experimental apparatus, or to complement the trigger-based data acquisition systems of larger experiments. A muon telescope setup consisting of a set of gaseous detectors is used as the experimental development testbed in this work, and a fully integrated online processing pipeline deployed on cloud computing resources is implemented and described.

physics.ins-det

Learning New Physics from an Imperfect Machine

We show how to deal with uncertainties on the Standard Model predictions in an agnostic new physics search strategy that exploits artificial neural networks. Our approach builds directly on the specific Maximum Likelihood ratio treatment of uncertainties as nuisance parameters for hypothesis testing that is routinely employed in high-energy physics. After presenting the conceptual foundations of our method, we first illustrate all aspects of its implementation and extensively study its performances on a toy one-dimensional problem. We then show how to implement it in a multivariate setup by studying the impact of two typical sources of experimental uncertainties in two-body final states at the LHC.

hep-ph

Learning Multivariate New Physics

We discuss a method that employs a multilayer perceptron to detect deviations from a reference model in large multivariate datasets. Our data analysis strategy does not rely on any prior assumption on the nature of the deviation. It is designed to be sensitive to small discrepancies that arise in datasets dominated by the reference model. The main conceptual building blocks were introduced in Ref. [1]. Here we make decisive progress in the algorithm implementation and we demonstrate its applicability to problems in high energy physics. We show that the method is sensitive to putative new physics signals in di-muon final states at the LHC. We also compare our performances on toy problems with the ones of alternative methods proposed in the literature.

hep-ph

Conceptual Design Report for the LUXE Experiment

This Conceptual Design Report describes LUXE (Laser Und XFEL Experiment), an experimental campaign that aims to combine the high-quality and high-energy electron beam of the European XFEL with a powerful laser to explore the uncharted terrain of quantum electrodynamics characterised by both high energy and high intensity. We will reach this hitherto inaccessible regime of quantum physics by analysing high-energy electron-photon and photon-photon interactions in the extreme environment provided by an intense laser focus. The physics background and its relevance are presented in the science case which in turn leads to, and justifies, the ensuing plan for all aspects of the experiment: Our choice of experimental parameters allows (i) effective field strengths to be probed at and beyond the Schwinger limit and (ii) a precision to be achieved that permits a detailed comparison of the measured data with calculations. In addition, the high photon flux predicted will enable a sensitive search for new physics beyond the Standard Model. The initial phase of the experiment will employ an existing 40 TW laser, whereas the second phase will utilise an upgraded laser power of 350 TW. All expectations regarding the performance of the experimental set-up as well as the expected physics results are based on detailed numerical simulations throughout.

hep-ex

Muon trigger with fast Neural Networks on FPGA, a demonstrator

The online reconstruction of muon tracks in High Energy Physics experiments is a highly demanding task, typically performed with programmable logic boards, such as FPGAs. Complex analytical algorithms are executed in a quasi-real-time environment to identify, select and reconstruct local tracks in often noise-rich environments. A novel approach to the generation of local triggers based on an hybrid combination of Artificial Neural Networks and analytical methods is proposed, targeting the muon reconstruction for drift tube detectors. The proposed algorithm exploits Neural Networks to solve otherwise computationally expensive analytical tasks for the unique identification of coherent signals and the removal of the geometrical ambiguities. The proposed approach is deployed on state-of-the-art FPGA and its performances are evaluated on simulation and on data collected from cosmic rays.

hep-ex

Machine Learning Pipelines with Modern Big Data Tools for High Energy Physics

The effective utilization at scale of complex machine learning (ML) techniques for HEP use cases poses several technological challenges, most importantly on the actual implementation of dedicated end-to-end data pipelines. A solution to these challenges is presented, which allows training neural network classifiers using solutions from the Big Data and data science ecosystems, integrated with tools, software, and platforms common in the HEP environment. In particular, Apache Spark is exploited for data preparation and feature engineering, running the corresponding (Python) code interactively on Jupyter notebooks. Key integrations and libraries that make Spark capable of ingesting data stored using ROOT format and accessed via the XRootD protocol, are described and discussed. Training of the neural network models, defined using the Keras API, is performed in a distributed fashion on Spark clusters by using BigDL with Analytics Zoo and also by using TensorFlow, notably for distributed training on CPU and GPU resourcess. The implementation and the results of the distributed training are described in detail in this work.

cs.DC

Using Big Data Technologies for HEP Analysis

The HEP community is approaching an era were the excellent performances of the particle accelerators in delivering collision at high rate will force the experiments to record a large amount of information. The growing size of the datasets could potentially become a limiting factor in the capability to produce scientific results timely and efficiently. Recently, new technologies and new approaches have been developed in industry to answer to the necessity to retrieve information as quickly as possible to analyze PB and EB datasets. Providing the scientists with these modern computing tools will lead to rethinking the principles of data analysis in HEP, making the overall scientific process faster and smoother. In this paper, we are presenting the latest developments and the most recent results on the usage of Apache Spark for HEP analysis. The study aims at evaluating the efficiency of the application of the new tools both quantitatively, by measuring the performances, and qualitatively, focusing on the user experience. The first goal is achieved by developing a data reduction facility: working together with CERN Openlab and Intel, CMS replicates a real physics search using Spark-based technologies, with the ambition of reducing 1 PB of public data in 5 hours, collected by the CMS experiment, to 1 TB of data in a format suitable for physics analysis. The second goal is achieved by implementing multiple physics use-cases in Apache Spark using as input preprocessed datasets derived from official CMS data and simulation. By performing different end-analyses up to the publication plots on different hardware, feasibility, usability and portability are compared to the ones of a traditional ROOT-based workflow.

cs.DC

Fitting the Higgs to Natural SUSY

We present a fit to the 2012 LHC Higgs data in different supersymmetric frameworks using naturalness as a guiding principle. We consider the MSSM and its D-term and F-term extensions that can raise the tree-level Higgs mass. When adding an extra chiral superfield to the MSSM, three parameters are needed determine the tree-level couplings of the lightest Higgs. Two more parameters cover the most relevant loop corrections, that affect the hγγand hgg vertexes. Motivated by this consideration, we present the results of a five parameters fit encompassing a vast class of complete supersymmetric theories. We find meaningful bounds on singlet mixing and on the mass of the pseudoscalar Higgs mA as a function of tanβ in the MSSM. We show that in the (mA, tanβ) plane, Higgs couplings measurements are probing areas of parameter space currently inaccessible to direct searches. We also consider separately the two cases in which only loop effects or only tree-level effects are sizeable. In the former case we study in detail stops' and charginos' contributions to Higgs couplings, while in the latter we show that the data point to the decoupling limit of the Higgs sector. In a particular realization of the decoupling limit, with an approximate PQ symmetry, we obtain constraints on the heavy scalar Higgs mass in a general type-II Two Higgs Doublet Model.

hep-ph

Prospective Studies for LEP3 with the CMS Detector

On July 4, 2012, the discovery of a new boson, with mass around 125 GeV/c2 and with properties compatible with those of a standard-model Higgs boson, was announced at CERN. In this context, a high-luminosity electron-positron collider ring, operating in the LHC tunnel at a centre-of-mass energy of 240 GeV and called LEP3, becomes an attractive opportunity both from financial and scientific point of views. The performance and the suitability of the CMS detector are evaluated, with emphasis on an accurate measurement of the Higgs boson properties. The precision expected for the Higgs boson couplings is found to be significantly better than that predicted by Linear Collider studies.

hep-ex

Commissioning of the CMS High Level Trigger

The CMS experiment will collect data from the proton-proton collisions delivered by the Large Hadron Collider (LHC) at a centre-of-mass energy up to 14 TeV. The CMS trigger system is designed to cope with unprecedented luminosities and LHC bunch-crossing rates up to 40 MHz. The unique CMS trigger architecture only employs two trigger levels. The Level-1 trigger is implemented using custom electronics, while the High Level Trigger (HLT) is based on software algorithms running on a large cluster of commercial processors, the Event Filter Farm. We present the major functionalities of the CMS High Level Trigger system as of the starting of LHC beams operations in September 2008. The validation of the HLT system in the online environment with Monte Carlo simulated data and its commissioning during cosmic rays data taking campaigns are discussed in detail. We conclude with the description of the HLT operations with the first circulating LHC beams before the incident occurred the 19th September 2008.

physics.ins-det

The CMS High Level Trigger: Commissioning and First Operation with LHC Beams

The CMS experiment will collect data from the proton-proton collisions delivered by the Large Hadron Collider (LHC) at a centre-of-mass energy up to 14 TeV. The CMS trigger system is designed to cope with unprecedented luminosities and LHC bunch-crossing rates up to 40 MHz. The unique CMS trigger architecture only employs two trigger levels. The Level-1 trigger is implemented using custom electronics. The High Level Trigger is implemented on a large cluster of commercial processors, the Filter Farm. Trigger menus have been developed for detector calibration and for fulfilment of the CMS physics program, at start-up of LHC operations, as well as for operations with higher luminosities. A complete multipurpose trigger menu developed for an early instantaneous luminosity of 10^{32}cm{-2}s{-1} has been tested in the HLT system under realistic online running conditions. The required computing power needed to process with no dead time a maximum HLT input rate of 50 kHz, as expected at startup, has been measured, using the most recent commercially available processors. The Filter Farm has been equipped with 720 such processors, providing a computing power at least a factor two larger than expected to be needed at startup. Results for the commissioning of the full-scale trigger and data acquisition system with cosmic muon runs are reported. The trigger performance during operations with LHC circulating proton beams, delivered in September 2008, is outlined and first results are shown.

physics.ins-det