SearcharxivSearch

arXiv subjects

Vinicius Mikuni

Publications and source records attributed to Vinicius Mikuni.

At least 19 recordsLinked to original sources

Future of Artificial Intelligence for Science in Japan 2024 Community Report

This white paper summarizes scientific challenges and AI/ML research opportunities identified through the FAIRS Japan 2024 unconference process. The discussion focuses on three major physics domains: accelerator physics, cosmology and astrophysics, and neutrino physics. Although each domain has distinct scientific goals and experimental constraints, several common technical themes emerge: high-dimensional reconstruction, fast and accurate simulation, uncertainty propagation, simulation-to-data mismatch, anomaly detection, real-time decision-making, and shared infrastructure.

hep-ph

OmniCosmos: Transferring Particle Physics Knowledge Across the Cosmos

Foundation models build an effective representations of data that can be deployed on diverse downstream tasks. Previous research developed the OmniLearned foundation model for collider physics and showed that it could significantly advance discovery potential across collider experiments. In this paper we go beyond collider physics and show that Foundation Models trained on collider data can help improve the prediction of cosmological parameters and to predict halo and galaxy velocities in different datasets from CosmoBench. This is the first time a collider physics model is shown to generalize across scientific fields.

astro-ph.CO

Generation of Imaging Air Cherenkov Telescope images using Diffusion Models

Substantial amounts of air-shower simulations are needed to derive the instrument response for analyzing Imaging Air Cherenkov Telescope (IACT) data. This process is both computationally intensive and requires repetition under varying observation conditions, due to detector aging, changes in the atmosphere, or the instrument hardware. Generative models offer an efficient alternative, significantly accelerating simulations while compactly storing extensive simulation libraries, and providing a differentiable surrogate model of the instrument. However, their applicability has so far been limited in gamma-ray astronomy, particularly for modeling hadronic showers that dominate the background and exhibit significant intrinsic fluctuations that are challenging to model. In this study, we present the first application of score-based diffusion models to generate monoscopic $γ$-ray and proton shower images with nearly 2,000 pixels and benchmark the performance against Wasserstein GANs using H.E.S.S. simulations. We examine quality using both low-level parameters and well-established shower-shape observables, and take the first step towards analysis readiness by investigating $γ$-hadron separation. While GAN-based approaches can reproduce $γ$-ray showers with high fidelity, they fail to generate proton events of comparable quality, leading to a measurable degradation in analysis performance. In contrast, score-based diffusion modles achieve significantly superior quality for $γ$-ray and proton showers, accurately reproducing high-level correlations and generating events that are statistically indistinguishable from simulations at the analysis level. These results establish diffusion-based models as the first analysis-ready surrogate model of a single IACT, opening new prospects for fast instrument response generation, detector optimization, and connected downstream tasks.

astro-ph.IM

Cross-Domain Transfer with Particle Physics Foundation Models: From Jets to Neutrino Interactions

Future AI-based studies in particle physics will likely start from a foundation model to accelerate training and enhance sensitivity. As a step toward a general-purpose foundation model for particle physics, we investigate whether the OmniLearned and ParticleViT foundation models pretrained on diverse high-$Q^2$ simulated and real $pp$ and $ep$ collisions retain useful knowledge to a few-GeV fixed-target neutrino experiment. We process MINERvA neutrino--nucleus scattering events and evaluate pretrained models on two types of tasks: regression of available energy and binary classification of charged-current pion final states ($\mathrm{CC1π^{\pm}}$, $\mathrm{CCNπ^{\pm}}$, and $\mathrm{CC1π^{0}}$). Pretrained OmniLearned and ParticleViT models outperform similarly sized models trained from scratch at the same compute budget, with the largest gains for OmniLearned on regression and for ParticleViT on classification. When the same transformer architecture is instead initialized from unrelated text pretraining (BERT), this advantage appears only marginally for classification in terms of compute efficiency and not in any way for regression. These results suggest that particle-level foundation models acquire inductive biases that generalize across large differences in energy scale, detector technology, and underlying physics processes, pointing toward detector-agnostic inference in particle physics.

hep-ex

Pre-Training for Simulation-Based Science: A Study on Jet Foundation Model Training Objectives

Foundation models (FMs) trained on large datasets and fine-tuned on downstream tasks have emerged as a powerful paradigm in AI for science. Industrial FMs are typically trained using self-supervision with masking due to the lack of labels. In many scientific domains, accurate simulations are plentiful and facilitate large, labeled datasets. This opens up new possibilities for pre-training. We present a systematic comparison of pre-training methods using the OmniLearned High Energy Physics FM framework. We test supervised classification, flow-matching generation, and self-supervised masked particle modeling. All models are pre-trained on the JetClass dataset and fine-tuned on two representative downstream tasks, top jet classification and JetNet conditional generation. Among other observations, for classification tasks, we find that pure classifier pre-training is optimal when downstream labels and model capacity are plentiful, but combining it with self-supervised masked particle modeling (MPM) is uniquely powerful in the low-finetuning label regime. Flow matching-based generative pre-training seems to provide little benefit for downstream classification, and interestingly, for downstream generation, we find that flow matching must be in the pre-training objective to see a significant finetuning advantage, hinting at the orthogonality of classification and generation tasks. That is, for a model to transfer to both generative and classification downstream tasks, it must be pre-trained on both. This study provides a template for controlled scaling analysis of pre-training objectives for foundation models in simulation-based sciences.

hep-ph

Explicit or Implicit? Encoding Physics at the Precision Frontier

High-performance machine learning tools in particle physics rest on two complementary directions: encoding symmetries explicitly in the architecture, and implicitly learning the structure of the data through large-scale (pre-) training. We compare the performance of the representative L-GATr and OmniLearn models on three especially challenging tasks: reweighting-based unfolding, likelihood-ratio estimation, and weakly supervised anomaly detection. Across all benchmarks, both methods achieve comparable performance given the statistical precision of the finetuning datasets, suggesting that the significant efficiency gains from encoding known particle physics structures are largely method-independent.

hep-ph

Are We Ready for AI-Driven Discovery? AI Verification Before the Next Fundamental Physics Breakthrough

Machine learning (ML) has become integral to fundamental physics, accelerating statistical workflows from data acquisition through inference and hypothesis testing. As ML systems grow increasingly autonomous, ensuring their reliability for discovery claims becomes critical. This review synthesizes the VERaiPHY (Validation & Evaluation for Robust AI in PHYsics) initiative's frameworks for rigorous ML assessment across particle physics, astrophysics, and cosmology. We establish when verification is essential by contextualizing ML within the statistical discovery workflow. We emphasize fundamental limitations: inductive bias is unavoidable, sample complexity bounds learning, and experimental constraints limit discovery. We reflect on physicists' evolving role as both experimental designers and evaluators whose judgments encode scientific rigor into AI systems. Responsible integration requires understanding ML's transformative potential alongside its intrinsic boundaries.

physics.data-an

Machine Learning-based Unfolding for Cross Section Measurements in the Presence of Nuisance Parameters

Statistically correcting measured cross sections for detector effects is an important step across many applications. In particle physics, this inverse problem is known as unfolding. In cases with complex instruments, the distortions they introduce are often known only implicitly through simulations of the detector. Modern machine learning has enabled efficient simulation-based approaches for unfolding high-dimensional data. Among these, one of the first methods successfully deployed on experimental data is the OmniFold algorithm, a classifier-based Expectation-Maximization procedure. In practice, however, the forward model is only approximately specified, and the corresponding uncertainty is encoded through nuisance parameters. Building on the well-studied OmniFold algorithm, we show how to extend machine learning-based unfolding to incorporate nuisance parameters. Our new algorithm, called Profile OmniFold, is demonstrated using a Gaussian example as well as a particle physics case study using simulated data from the CMS Experiment at the Large Hadron Collider.

stat.AP

Searching for Anomalies with Foundation Models

Foundation models have the potential to extend the discovery reach for anomaly detection searches. When studying the large OmniLearned foundation model on data from the CMS experiment, unexpected behavior was observed in a mass sideband. The purpose of this paper is to perform a full analysis, including a complete background estimate, on the phase space picked out by the large model. We find that the background estimation describes the data well in validation regions, but is unable to accurately model the signal region. We invite further scrutiny of these events and our methods.

hep-ex

OmniMol: Transferring Particle Physics Knowledge to Molecular Dynamics with Point-Edge Transformers

We present OmniMol, a state-of-the-art all-to-all transformer-based small molecule machine-learned interatomic potential (MLIP). OmniMol is built by adapting Omnilearned, a foundation model for particle jets found in high-energy physics (HEP) experiments such as at the Large Hadron Collider (LHC). Omnilearned is built with a Point-Edge-Transformer (PET) and pre-trained using a diverse set of one billion particle jets. It includes an interaction-matrix attention bias that injects pairwise sub-nuclear (HEP) or atomic (molecular-dynamics) physics directly into the transformer's attention logits, steering the network toward physically meaningful neighborhoods without sacrificing expressivity. We demonstrate OmniMol using the oMol dataset and find excellent performance even with relatively few examples for fine-tuning. Further, due to architectural transfer from Omnilearned, we demonstrate uniquely fast inference. This study lays the foundation for building interdisciplinary connections given datasets represented as collections of point clouds.

physics.chem-ph

Diffusion-Based Point-Cloud Generation of Heavy-Ion Events

Heavy-ion collisions produce final states with thousands to tens of thousands of particles, making their simulation among the most computationally intensive tasks in high-energy nuclear physics. We present a fast, high-fidelity generative model for heavy-ion events based on a score-driven diffusion process and the Point-Edge Transformer architecture within the OmniLearn framework. A two-stage training strategy is performed: Stage-1 training on lower-multiplicity O-O collisions allowing the model to learn a stable event and particles representation, followed by fine-tuning on challenging high-multiplicity Pb-Pb collisions. We benchmark the generator with a broad set of closure checks, including agreement of event- and particle-level observables in one and two dimensions, flow consistency reconstructed from the generated particles, end-to-end jet finding with FastJet including key jet and substructure observables, and a classifier-based application to quantify the sample fidelity. The results are promising, showing that a compact generative model can produce realistic, high-multiplicity heavy-ion events, at a level that makes local-scale generation for heavy-ion collisions at high energies a practical goal.

hep-ph

OmniLearned: A Foundation Model Framework for All Tasks Involving Jet Physics

Foundation models use large datasets to build an effective representation of data that can be deployed on diverse downstream tasks. Previous research developed the OmniLearn foundation model for jet physics, using unique properties of particle physics, and showed that it could significantly advance discovery potential across collider experiments. This paper introduces a major upgrade, resulting in the OmniLearned framework. This framework has three new elements: (1) updates to the model architecture and training, (2) using over one billion jets used for training, and (3) providing well-documented software for accessing all datasets and models. We demonstrate OmniLearned with three representative tasks: top-quark jet tagging with the community Delphes-based benchmark dataset, b-tagging with ATLAS full simulation, and anomaly detection with CMS experimental data. In each case, OmniLearned is the state of the art, further expanding the discovery potential of past, current, and future collider experiments.

hep-ph

EveNet: A Foundation Model for Particle Collision Data Analysis

While deep learning is transforming data analysis in high-energy physics, computational challenges limit its potential. We address these challenges in the context of collider physics by introducing EveNet, an event-level foundation model pretrained on 500 million simulated collision events using a hybrid objective of self-supervised learning and physics-informed supervision. By leveraging a shared particle-cloud representation, EveNet outperforms state-of-the-art baselines across diverse tasks, including searches for heavy resonances and exotic Higgs decays, and demonstrates exceptional data efficiency in low-statistics regimes. Crucially, we validate the transferability of the model to experimental data by rediscovering the $Υ$ meson in CMS Open Data and show its capacity for precision physics through the robust extraction of quantum correlation observables stable against systematic uncertainties. These results indicate that EveNet can successfully encode the fundamental physical structure of particle interactions, which offers a unified and resource-efficient framework to accelerate discovery at current and future colliders.

hep-ex

CaloChallenge 2022: A Community Challenge for Fast Calorimeter Simulation

We present the results of the "Fast Calorimeter Simulation Challenge 2022" - the CaloChallenge. We study state-of-the-art generative models on four calorimeter shower datasets of increasing dimensionality, ranging from a few hundred voxels to a few tens of thousand voxels. The 31 individual submissions span a wide range of current popular generative architectures, including Variational AutoEncoders (VAEs), Generative Adversarial Networks (GANs), Normalizing Flows, Diffusion models, and models based on Conditional Flow Matching. We compare all submissions in terms of quality of generated calorimeter showers, as well as shower generation time and model size. To assess the quality we use a broad range of different metrics including differences in 1-dimensional histograms of observables, KPD/FPD scores, AUCs of binary classifiers, and the log-posterior of a multiclass classifier. The results of the CaloChallenge provide the most complete and comprehensive survey of cutting-edge approaches to calorimeter fast simulation to date. In addition, our work provides a uniquely detailed perspective on the important problem of how to evaluate generative models. As such, the results presented here should be applicable for other domains that use generative AI and require fast and faithful generation of samples in a large phase space.

physics.ins-det

SEAL - A Symmetry EncourAging Loss for High Energy Physics

Physical symmetries provide a strong inductive bias for constructing functions to analyze data. In particular, this bias may improve robustness, data efficiency, and interpretability of machine learning models. However, building machine learning models that explicitly respect symmetries can be difficult due to the dedicated components required. Moreover, real-world experiments may not exactly respect fundamental symmetries at the level of finite granularities and energy thresholds. In this work, we explore an alternative approach to create symmetry-aware machine learning models. We introduce soft constraints that allow the model to decide the importance of added symmetries during the learning process instead of enforcing exact symmetries. We investigate two complementary approaches, one that penalizes the model based on specific transformations of the inputs and one inspired by group theory and infinitesimal transformations of the inputs. Using top quark jet tagging and Lorentz equivariance as examples, we observe that the addition of the soft constraints leads to more robust performance while requiring negligible changes to current state-of-the-art models.

hep-ph

Unbinned measurement of thrust in $e^+e^-$ collisions at $\sqrt{s}$ = 91.2 GeV with ALEPH archived data

The strong coupling constant ($α_{S}$) is a fundamental parameter of quantum chromodynamics (QCD), the theory of the strong force. Some of the earliest precise constraints on $α_{S}$ came from measurements of event shape observables, such as thrust ($T$), using hadronic $Z$ boson decays produced in $e^+e^-$ collisions. However, recent work has revealed discrepancies between event-shape-based extractions of $α_{S}$ and values determined using other experimental methods. This work reexamines archived $e^+e^-$ data collected at a collision energy of $\sqrt{s}=91.2$ GeV by the ALEPH detector at the Large Electron-Positron Collider. Modern machine learning techniques are used to correct for detector effects in an unbinned manner, allowing the $T$ distribution to be measured with higher granularity than previous ALEPH measurements. The new measurement reveals a small but systematic shift towards larger values of $τ=1-T$, and the potential implications of this shift for $α_{S}$ extractions are illustrated by comparing to state-of-the-art theoretical calculations. In addition, the region of $-6<\logτ<-2$, where poorly-understood non-perturbative effects are large, is compared to modern parton shower Monte Carlo simulations. This measurement provides unique new inputs for $α_{S}$ extractions and also improves constraints on phenomenological models of QCD dynamics such as parton fragmentation and hadronization.

hep-ex

Neural Posterior Unfolding

Differential cross section measurements are the currency of scientific exchange in particle and nuclear physics. A key challenge for these analyses is the correction for detector distortions, known as deconvolution or unfolding. Binned unfolding of cross section measurements traditionally rely on the regularized inversion of the response matrix that represents the detector response, mapping pre-detector (`particle level') observables to post-detector (`detector level') observables. In this paper we introduce Neural Posterior Unfolding, a modern, Bayesian approach that leverages normalizing flows for unfolding. By using normalizing flows for neural posterior estimation, NPU offers several key advantages including implicit regularization through the neural network architecture, fast amortized inference that eliminates the need for repeated retraining, and direct access to the full uncertainty in the unfolded result. In addition to introducing NPU, we implement a classical Bayesian unfolding method called Fully Bayesian Unfolding (FBU) in modern Python so it can also be studied. These tools are validated on simple Gaussian examples and then tested on simulated jet substructure examples from the Large Hadron Collider (LHC). We find that the Bayesian methods are effective and worth additional development to be analysis ready for cross section measurements at the LHC and beyond.

hep-ph

Analysis-ready Generative Unfolding

Machine Learning (ML)-based unfolding methods have enabled high-dimensional and unbinned differential cross section measurements. While a suite of such methods has been proposed, most focus exclusively on the challenge of statistically removing resolution effects. In practice, unfolding methods must also account for impurities and finite acceptance and efficiency effects. In this paper, we extend a class of unfolding methods based on generative ML to include the full suite of effects relevant for cross section measurements. Our new methods include fully generative solutions as well as generative-discriminative hybrid approaches (GenFoldG and GenFoldC). We demonstrate these new techniques in both Gaussian and simulated LHC examples. Overall, we find that both methods are able to accommodate all effects, thus adding a complementary and analysis-ready method to the unfolding toolkit.

hep-ph