SearcharxivSearch

arXiv subjects

Manuel Szewc

Publications and source records attributed to Manuel Szewc.

At least 19 recordsLinked to original sources

Defining a Minimum Resolution for Unbinned Analyses

Collider analyses combine rigorous statistical techniques with state-of-the-art Machine Learning models. However, when the latter are used directly to estimate the likelihood function of the background, hard to quantify systematic effects may bias the estimation of the relevant signal parameters. To address this problem, we present the Minimum Resolution Likelihood (MRL) method, which defines a Fiducial Signal Region that effectively turns the systematic effects into statistical uncertainties. We show with examples that the resulting signal strength estimation is either unbiased or consistent with zero. We consider both toy examples and a realistic application based on the HI-SIGMA technique applied to di-Higgs searches.

hep-ph

Vistas: A Visualization Interface for Particle Collision Simulations

We introduce Vistas, a tool for visualizing high-energy particle physics collisions simulated by the Pythia Monte-Carlo event generator. Vistas utilizes the browser-based event display framework Phoenix to show distinct computational stages of a high-energy collision event simulation: the hard process, parton shower, hadronization, and particle decays. Particles produced from each of these stages are represented as lines in an interactive three-dimensional graph structure, where each line is along the direction of its particle's three-momentum vector. The event can be rotated, translated and zoomed, and details for each particle can be accessed by selecting the relevant particle line. Additionally, particle lines from all stages of the simulation can be toggled on and off and can be filtered by particle-level kinematic selection requirements. This interactive environment provides an intuitive interpretation of Pythia simulation output, including detailed features such as color flow, beam remnants, and multiple parton interactions, making it a useful tool in physics education settings, from outreach activities to graduate particle-physics courses.

physics.ed-ph

HDSense: An efficient method for ranking observable sensitivity

Identifying which observables most effectively constrain model parameters can be computationally prohibitive when considering full likelihoods of many correlated observables. This is especially important for, e.g., hadronization models, where high precision is required to interpret the results of collider experiments. We introduce the High-Dimensional Sensitivity (HDSense) score, a computationally efficient metric for ranking observable sets using only one-dimensional histograms. Derived by profiling over unknown correlations in the Fisher information framework, the score balances total information content against redundancy between observables. We apply HDSense to rank a set observables in terms of their constraining power with respect to five parameters of the Lund string model of hadronization implemented in Pythia using simulated leptonic collider events at the $Z$ pole. Validation against machine-learning--based full-likelihood approximations demonstrates that HDSense successfully identifies near-optimal observable subsets. The framework naturally handles data from multiple experiments with different acceptances and incorporates detector effects. While demonstrated on hadronization models, the methodology applies broadly to generic parameter estimation problems where correlations are unknown or difficult to model.

hep-ph

Many Wrongs Make a Right: Leveraging Biased Simulations Towards Unbiased Parameter Inference

In particle physics, as in many areas of science, parameter inference relies on simulations to bridge the gap between theory and experiment. Recent developments in simulation-based inference have boosted the sensitivity of analyses; however, biases induced by simulation-data mismodeling can be difficult to control within standard inference pipelines. In this work, we propose a Template-Adapted Mixture Model to confront this problem in the context of signal fraction estimation: inferring the population proportion of signal in a mixed sample of signal and background, both of which follow arbitrarily complex distributions. We harness many biased simulations to perform data-driven estimates of each process distribution in the signal region, substantially reducing the bias on the signal fraction due to the domain shift between simulation and reality. We explore different methodological choices, including model selection, feature representation, and statistical method, and apply them to a Gaussian toy example and to a semi-realistic di-Higgs measurement. We find that the presented methods successfully leverage the biased simulations to provide estimates with well-calibrated uncertainties.

hep-ph

Iterative HOMER with uncertainties

We present iHOMER, an iterative version of the HOMER method to extract Lund fragmentation functions from experimental data. Through iterations, we address the information gap between latent and observable phase spaces and systematically remove bias. To quantify uncertainties on the inferred weights, we use a combination of Bayesian neural networks and uncertainty-aware regression. We find that the combination of iterations and uncertainty quantification produces well-calibrated weights that accurately reproduce the data distribution. A parametric closure test shows that the iteratively learned fragmentation function is compatible with the true fragmentation function.

hep-ph

Di-Higgs to 4b with Bayesian inference: improving simulation estimates

Measuring di-Higgs production in the four-bottom channel is challenged by overwhelming QCD backgrounds and imperfect simulations. We develop a Bayesian mixture model that simultaneously infers signal and background fractions and their individual shapes directly in the signal region. The likelihood is a nuanced combination of a one-dimensional kinematic discriminator and per-jet flavour scores; with their correlations incorporated via kinematic bins. Monte Carlo informs weak Dirichlet priors, while the posterior adjusts to the interplay of the model, priors and observed data. Using pseudo-data simulated with standard tools and with controlled mismatches, we show that the method corrects biased priors, delivers calibrated 68-95% credible intervals for the signal count, and improves dataset-level ROC/AUC relative to simple cut-and-count baselines. This study highlights how Bayesian inference can harvest information present in the signal region and self-calibrate model parameters, providing a robust route to increased sensitivity in di-Higgs searches.

hep-ph

Data-Driven High-Dimensional Statistical Inference with Generative Models

Crucial to many measurements at the LHC is the use of correlated multi-dimensional information to distinguish rare processes from large backgrounds, which is complicated by the poor modeling of many of the crucial backgrounds in Monte Carlo simulations. In this work, we introduce HI-SIGMA, a method to perform unbinned high-dimensional statistical inference with data-driven background distributions. In contradistinction to many applications of Simulation Based Inference in High Energy Physics, HI-SIGMA relies on generative ML models, rather than classifiers, to learn the signal and background distributions in the high-dimensional space. These ML models allow for interpretable inference while also incorporating model errors and other sources of systematic uncertainties. We showcase this methodology on a simplified version of a di-Higgs measurement in the $bbγγ$ final state, where the di-photon resonance allows for background interpolation from sidebands into the signal region. We demonstrate that HI-SIGMA provides improved sensitivity as compared to standard classifier-based methods, and that systematic uncertainties can be straightforwardly incorporated by extending methods which have been used for histogram based analyses.

hep-ph

Characterizing the hadronization of parton showers using the HOMER method

We update the HOMER method, a technique to solve a restricted version of the inverse problem of hadronization -- extracting the Lund string fragmentation function $f(z)$ from data using only observable information. Here, we demonstrate its utility by extracting $f(z)$ from synthetic Pythia simulations using high-level observables constructed on an event-by-event basis, such as multiplicities and shape variables. Four cases of increasing complexity are considered, corresponding to $e^+e^-$ collisions at a center-of-mass energy of $90$ GeV producing either a string stretched between a $q$ and $\bar{q}$ containing no gluons; the same string containing one gluon $g$ with fixed kinematics; the same but the gluon has varying kinematics; and the most realistic case, strings with an unrestricted number of gluons that is the end-result of a parton shower. We demonstrate the extraction of $f(z)$ in each case, with the result of only a relatively modest degradation in performance of the HOMER method with the increased complexity of the string system.

hep-ph

Inferring correlated distributions: boosted top jets

Improving the understanding of signal and background distributions in signal-region is a valuable key to enhance any analysis in collider physics. This is usually a difficult task because -- among others -- signal and backgrounds are hard to discriminate in signal-region, simulations may reach a limit of reliability if they need to model non-perturbative QCD, and distributions are multi-dimensional and many times may be correlated within each class. Bayesian density estimation is a technique that leverages prior knowledge and data correlations to effectively extract information from data in signal-region. In this work we extend previous works on data-driven mixture models for meaningful unsupervised signal extraction in collider physics to incorporate correlations between features. Using a standard dataset of top and QCD jets, we show how simulators, despite having an expected bias, can be used to inject sufficient inductive nuance into an inference model in terms of priors to then be corrected by data and estimate the true correlated distributions between features within each class. We compare the model with and without correlations to show how the signal extraction is sensitive to their inclusion and we quantify the improvement due to the inclusion of correlations using both supervised and unsupervised metrics.

hep-ph

Post-hoc reweighting of hadron production in the Lund string model

We present a method for reweighting flavor selection in the Lund string fragmentation model. This is the process of calculating and applying event weights enabling fast and exact variation of hadronization parameters on pre-generated event samples. The procedure is post hoc, requiring only a small amount of additional information stored per event, and allowing for efficient estimation of hadronization uncertainties without repeated simulation. Weight expressions are derived from the hadronization algorithm itself, and validated against direct simulation for a wide range of observables and parameter shifts. The hadronization algorithm can be viewed as a hierarchical Markov process with stochastic rejections, a structure common to many complex simulations outside of high-energy physics. This perspective makes the method modular, extensible, and potentially transferable to other domains. We demonstrate the approach in Pythia, including both numerical stability and timing benefits.

hep-ph

Describing Hadronization via Histories and Observables for Monte-Carlo Event Reweighting

We introduce a novel method for extracting a fragmentation model directly from experimental data without requiring an explicit parametric form, called Histories and Observables for Monte-Carlo Event Reweighting (HOMER), consisting of three steps: the training of a classifier between simulation and data, the inference of single fragmentation weights, and the calculation of the weight for the full hadronization chain. We illustrate the use of HOMER on a simplified hadronization problem, a $q\bar{q}$ string fragmenting into pions, and extract a modified Lund string fragmentation function $f(z)$. We then demonstrate the use of HOMER on three types of experimental data: (i) binned distributions of high level observables, (ii) unbinned event-by-event distributions of these observables, and (iii) full particle cloud information. After demonstrating that $f(z)$ can be extracted from data (the inverse of hadronization), we also show that, at least in this limited setup, the fidelity of the extracted $f(z)$ suffers only limited loss when moving from (i) to (ii) to (iii). Public code is available at https://gitlab.com/uchep/mlhad.

hep-ph

Rejection Sampling with Autodifferentiation - Case study: Fitting a Hadronization Model

We present an autodifferentiable rejection sampling algorithm termed Rejection Sampling with Autodifferentiation (RSA). In conjunction with reweighting, we show that RSA can be used for efficient parameter estimation and model exploration. Additionally, this approach facilitates the use of unbinned machine-learning-based observables, allowing for more precise, data-driven fits. To showcase these capabilities, we apply an RSA-based parameter fit to a simplified hadronization model.

hep-ph

Direct CKM determination from W decays at future lepton colliders

We project the reach of future lepton colliders for measuring CKM elements from direct observations of $W$ decays. We focus our attention to $|V_{cs}|$ and $|V_{cb}|$ determinations, using FCC-ee as case study. We employ state-of-the-art jet flavor taggers to obtain the projected sensitivity, and scan over tagger performances to show their effect. We conclude that future lepton collider can sizeably improve the sensitivity on $|V_{cs}|$ and $|V_{cb}|$, albeit the achievable reach will strongly depend on the level of systematic uncertainties on tagger parameters.

hep-ph

Improvement and generalization of ABCD method with Bayesian inference

To find New Physics or to refine our knowledge of the Standard Model at the LHC is an enterprise that involves many factors. We focus on taking advantage of available information and pour our effort in re-thinking the usual data-driven ABCD method to improve it and to generalize it using Bayesian Machine Learning tools. We propose that a dataset consisting of a signal and many backgrounds is well described through a mixture model. Signal, backgrounds and their relative fractions in the sample can be well extracted by exploiting the prior knowledge and the dependence between the different observables at the event-by-event level with Bayesian tools. We show how, in contrast to the ABCD method, one can take advantage of understanding some properties of the different backgrounds and of having more than two independent observables to measure in each event. In addition, instead of regions defined through hard cuts, the Bayesian framework uses the information of continuous distribution to obtain soft-assignments of the events which are statistically more robust. To compare both methods we use a toy problem inspired by $pp\to hh\to b\bar b b \bar b$, selecting a reduced and simplified number of processes and analysing the flavor of the four jets and the invariant mass of the jet-pairs, modeled with simplified distributions. Taking advantage of all this information, and starting from a combination of biased and agnostic priors, leads us to a very good posterior once we use the Bayesian framework to exploit the data and the mutual information of the observables at the event-by-event level. We show how, in this simplified model, the Bayesian framework outperforms the ABCD method sensitivity in obtaining the signal fraction in scenarios with $1\%$ and $0.5\%$ true signal fractions in the dataset. We also show that the method is robust against the absence of signal.

hep-ph

Flavor violating Higgs and $Z$ decays at FCC-ee

Recent advances in $b$, $c$, and $s$ quark tagging coupled with novel statistical analysis techniques will allow future high energy and high statistics electron-positron colliders, such as the FCC-ee, to place phenomenologically relevant bounds on flavor violating Higgs and $Z$ decays to quarks. We assess the FCC-ee reach for $Z/h\to bs, cu$ decays as a function of jet tagging performance. We also update the SM predictions for the corresponding branching ratios, as well as the indirect constraints on the flavor violating Higgs and $Z$ couplings to quarks. Using type III two Higgs doublet model as an example of beyond the standard model physics, we show that the searches for $h\to bs, cu$ decays at FCC-ee can probe new parameter space not excluded by indirect searches. We also reinterpret the FCC-ee reach for $Z\to bs , cu$ in terms of the constraints on models with vectorlike quarks.

hep-ph

Towards a data-driven model of hadronization using normalizing flows

We introduce a model of hadronization based on invertible neural networks that faithfully reproduces a simplified version of the Lund string model for meson hadronization. Additionally, we introduce a new training method for normalizing flows, termed MAGIC, that improves the agreement between simulated and experimental distributions of high-level (macroscopic) observables by adjusting single-emission (microscopic) dynamics. Our results constitute an important step toward realizing a machine-learning based model of hadronization that utilizes experimental data during training. Finally, we demonstrate how a Bayesian extension to this normalizing-flow architecture can be used to provide analysis of statistical and modeling uncertainties on the generated observable distributions.

hep-ph

Reweighting Monte Carlo Predictions and Automated Fragmentation Variations in Pythia 8

This work reports on a method for uncertainty estimation in simulated collider-event predictions. The method is based on a Monte Carlo-veto algorithm, and extends previous work on uncertainty estimates in parton showers by including uncertainty estimates for the Lund string-fragmentation model. This method is advantageous from the perspective of simulation costs: a single ensemble of generated events can be reinterpreted as though it was obtained using a different set of input parameters, where each event now is accompanied with a corresponding weight. This allows for a robust exploration of the uncertainties arising from the choice of input model parameters, without the need to rerun full simulation pipelines for each input parameter choice. Such explorations are important when determining the sensitivities of precision physics measurements. Accompanying code is available at https://gitlab.com/uchep/mlhad-weights-validation.

hep-ph

Accessing CKM suppressed top decays at the LHC

We present an strategy for measuring the off-diagonal elements of the third row of CKM matrix $|V_{tq}|$ through the branching fractions of top quark decays $t\to q W$, where $q$ is a light quark jet. This strategy is an extension of existing measurements, with the improvement rooted in the use of orthogonal $b$- and $q$-taggers that add a new observable, the number of light-quark-tagged jets, to the already commonly used observable, the fraction of $b$-tagged jets in an event. Careful inclusion of the additional complementary observable significantly increases the expected statistical power of the analysis, with the possibility of excluding a null $|V_{td}|^2+|V_{ts}|^2$ at $95\%$ C.L. at the HL-LHC.

hep-ph