SearcharxivSearch

arXiv subjects

Tilman Plehn

Publications and source records attributed to Tilman Plehn.

At least 19 recordsLinked to original sources

MadAgents

We uncover an effective and communicative set of agents working with MadGraph. Agentic installation, learning-by-doing training, and user support provide easy access to state-of-the-art simulations and accelerate LHC research. We show in detail how MadAgents interact with inexperienced and advanced users, support a range of simulation tasks, and analyze the results. In a second step, we illustrate how MadAgents automatize event generation and run an autonomous simulation campaign, starting from a pdf file of a paper. We also present successive MadAgents updates, including a Claude Code implementation with a self-improvement loop.

hep-ph

Unfolding without Iterations, Adversaries, or Surrogates

Correcting measurements for detector effects and constructing appropriate public data representations is a pressing problem in LHC physics. Current methods solve this inverse problem by relying on iterations, minimax optimization, or a surrogate forward mapping. We introduce Adversary-free Unfolding SanS Iteration or Emulation (AUSSIE), which dispenses with these mechanisms while remaining asymptotically correct. AUSSIE replaces the second OmniFold step with a new loss function that directly yields solutions with minimal dependence on the reference simulation. We showcase AUSSIE on various unfolding tasks, including full-phase-space jet substructure.

hep-ph

Know What You Don't Flow

Calibrated learned uncertainties are a key requirement also for generative neural networks in LHC physics. For a toy model with an explicit likelihood we show how a heteroscedastic and a Bayesian normalizing flow learn the systematic and statistical uncertainties on the underlying phase space density. Without an explicit likelihood we train the heteroscedastic loss on a classifier-reweighted approximate generative network. We illustrate our comprehensive approach for top pair events and show how a conditional heteroscedastic flow propagates calibrated uncertainties to all phase space directions.

hep-ph

NAE, Statistically

Searches for new physics using neural anomaly scores have transformative potential, but suffer from a lack of statistical interpretability. The normalized autoencoder (NAE) provides a probabilistic interpretation of the standard bottleneck architecture, tying the anomaly score to a learned likelihood. We validate this relation for a toy model, test it for jets using a dual-NAE setup, and show how a Bayesian NAE learns this likelihood with an uncertainty.

hep-ph

VERaiPHY -- Validation & Evaluation for Robust AI in PHYsics

Modern machine learning is leading to substantial gains in precision, flexibility, and computational efficiency in fundamental physics. Statistical validation, uncertainty quantification, and robustness assessment are less systematically addressed. The VERaiPHY initiative (Validation & Evaluation for Robust AI in PHYsics) is a series of articles developed within the PHYSTAT programme, aimed at establishing statistical standards for the development, evaluation, and deployment of ML techniques. Each article focuses on a specific methodological domain from a statistics perspective and clarifies statistical questions, tests, and the interpretation of results. This opening article establishes the probabilistic, statistical, and machine learning foundations that the later contributions assume, together with the notation used throughout.

hep-ph

Unbinning global LHC analyses

Neural simulation-based inference has been shown to outperform traditional, histogram-based inference in numerous phenomenological and experimental studies at the LHC. So far, these analyses have focused on individual processes. We study the combination of four different di-boson processes in terms of the Standard Model Effective Field Theory. Our results demonstrate how neural simulation-based inference also wins over traditional methods for more global LHC analyses.

hep-ph

Neural Control Variates at LO and NLO

We employ neural control variates to minimize the range of event weights and avoid negative weights for phase-space integration and event generation. A signed control variate, built from two normalizing flows, fulfills both tasks. Combined with neural importance sampling, it significantly reduces the computational cost of LO and NLO predictions. For the NLO case, our conditional neural control variate can be viewed as a trainable subtraction term, complementing the established physics subtraction schemes for enhanced sampling performance.

hep-ph

Machine Learning is Good for Physics - and Vice Versa

Scientific AI is rapidly transforming fundamental physics research and challenging defining aspects of the fundamental physics methodology. We discuss opportunities and dangers of this transformation and find exciting benefits from a close interaction between AI and fundamental physics, provided that we remain aware of the scientific methodologies of the respective fields. For fundamental physics, we discuss two such aspects: statistical validation and a generalizing theory description, both with the goal of discovering new physics in vast datasets.

hep-ph

Generative Amplification with Surrogate Monte Carlo

Amplitude surrogates for LHC simulations build on generative amplification, the fact that a surrogate trained on an expensive and small training dataset describes the smooth amplitude more precisely than the training data does. Applying techniques developed for generative networks, we quantify this amplification for gluon-associated $Z$ production. Significant amplification appears in sparsely populated kinematic tails, where it matters most. Our results show how generative amplification from surrogate Monte Carlo far outperforms the density estimation in current generative networks.

hep-ph

Virtues and Vices of Equivariant Transformers

We study for the first time the benefit of Lorentz-equivariant transformers for large-size jet tagging and flavor tagging. To control their computing demands, we optimize all implementations for inference cost metrics. In our scaling studies, we find that Lorentz-equivariant networks outperform standard transformers, provided geometric features are relevant. This holds true in an idealized world as well as for limited resources. The conditional gain from Lorentz equivariance provides interesting input to the development of foundation models for LHC data.

hep-ph

Explicit or Implicit? Encoding Physics at the Precision Frontier

High-performance machine learning tools in particle physics rest on two complementary directions: encoding symmetries explicitly in the architecture, and implicitly learning the structure of the data through large-scale (pre-) training. We compare the performance of the representative L-GATr and OmniLearn models on three especially challenging tasks: reweighting-based unfolding, likelihood-ratio estimation, and weakly supervised anomaly detection. Across all benchmarks, both methods achieve comparable performance given the statistical precision of the finetuning datasets, suggesting that the significant efficiency gains from encoding known particle physics structures are largely method-independent.

hep-ph

Agentic Re-Casting using Agentic Re-Simulations

Analysis re-casting at the LHC is highly standardized and nevertheless requires resources, time, and physics input. Building on the new MadAgents.v3, we show how a global SFitter analysis can be updated by an agentic system with a physicist in the loop. The agentic interface allows us to make the advanced SFitter methodology available to a wider audience. All physical and technical aspects of this agentic re-casting study can be trivially generalized beyond SFitter.

hep-ph

FASTColor -- Full-color Amplitude Surrogate Toolkit for QCD

High-multiplicity events remain a bottleneck for LHC simulations due to their computational cost. We present a ML-surrogate approach to accelerate matrix element reweighting from leading-color (LC) to full-color (FC) accuracy, building on recent advancements in LC event generation. Comparing a variety of modern network architectures for representative QCD processes, we achieve speed-up of around a factor two over the current LC-to-FC baseline. We also show how transformers learn and exploit underlying symmetries, to improve generalization. Given the gained trust in trained networks and developments in learned uncertainties, the LC-to-FC approach will eventually benefit further from not needing a final classic unweighting step.

hep-ph

Forecasting Generative Amplification

Generative networks are perfect tools to enhance the speed and precision of LHC simulations. Especially when generating events beyond the size of the training dataset, it is important to understand their statistical precision. We present two complementary methods to estimate the amplification factor without large holdout datasets. Averaging amplification uses Bayesian networks or ensembling to estimate amplification from the precision of integrals over given phase-space volumes. Differential amplification uses hypothesis testing to quantify amplification without any resolution loss. Applied to state-of-the-art event generators, both methods indicate that amplification is already possible in specific regions of phase space.

hep-ph

Local Conformal Predictions for Calibrated Surrogates

Neural network surrogates for LHC scattering amplitudes require trustworthy uncertainty estimates, a challenging task given the non-Gaussian systematics. We target it using conformal prediction, a distribution-free post-processing to complement trained surrogates with calibrated uncertainties. We find that standard conformal predictions struggle to provide locally calibrated uncertainties. This leads us to introduce FALCON, a novel conformal prediction method that learns locally calibrated confidence intervals. Our simple examples illustrate the power of distribution-free uncertainty quantification for ultra-fast event generation at the LHC.

hep-ph

One Generator, Any Process: LLM-Conditioning for the LHC

Neural network training for LHC event generation should, ideally, benefit from common high-level patterns in different processes. We propose novel conditioning schemes for continuous parameters, process labels, and Feynman diagrams. We employ pre-trained LLMs as multi-modal foundation models to provide descriptive embeddings for an autoregressive transformer. With such high-level physics-inductive bias the generative networks converge faster, provide better result, and generalize to unseen processes.

hep-ph

How to Trust Learned Loop Amplitudes

Higher-order theory predictions are crucial for the precision LHC program, but the time-consuming amplitude evaluation challenges the corresponding Monte-Carlo simulations. Machine-learned amplitude surrogates can resolve this problem, if we can guarantee their precision over the entire phase space. First, we show that our surrogates provide a calibrated learned uncertainty, even for non-Gaussian systematics; second, we describe how less accurate phase space regions can be identified; third, we demonstrate how the precision in these regions can be improved reliably.

hep-ph

The Latent Information Geometry of Jet Classification

Latent representations are an important theme in modern machine learning. Any network training with the notion of locality introduces a latent geometry which we can analyze with the help of differential geometry, specifically information geometry. We introduce the main concepts needed to analyze learned latent geometries, specifically curvature and nonmetricities, and show how they can be used for decoder and classifier geometries. We then apply our new methods to understand the physics behind binary quark-gluon classification and three-fold fat jet tagging.

hep-ph