SearcharxivSearch

arXiv subjects

Ezequiel Alvarez

Publications and source records attributed to Ezequiel Alvarez.

At least 19 recordsLinked to original sources

Many Wrongs Make a Right: Leveraging Biased Simulations Towards Unbiased Parameter Inference

In particle physics, as in many areas of science, parameter inference relies on simulations to bridge the gap between theory and experiment. Recent developments in simulation-based inference have boosted the sensitivity of analyses; however, biases induced by simulation-data mismodeling can be difficult to control within standard inference pipelines. In this work, we propose a Template-Adapted Mixture Model to confront this problem in the context of signal fraction estimation: inferring the population proportion of signal in a mixed sample of signal and background, both of which follow arbitrarily complex distributions. We harness many biased simulations to perform data-driven estimates of each process distribution in the signal region, substantially reducing the bias on the signal fraction due to the domain shift between simulation and reality. We explore different methodological choices, including model selection, feature representation, and statistical method, and apply them to a Gaussian toy example and to a semi-realistic di-Higgs measurement. We find that the presented methods successfully leverage the biased simulations to provide estimates with well-calibrated uncertainties.

hep-ph

Dark Matter Clumps as Sources of Gravitational-Wave Glitches in LIGO/Virgo/KAGRA data

We consider the hypothetical possibility that non-stationary glitch features in the noise of ground-based gravitational-wave detectors could be produced by small dark matter clumps that pass through the earth in the vicinity of gravitational-wave detectors. We first derive the gravitational-wave strain that would be generated by the passage of such a dark matter clump. We find that the strain is primarily sourced by the Newtonian gravitational acceleration of the mirrors toward the clump and by the Shapiro time delay of the photons in the laser beams as they pass through the gravitational potential created by the dark matter clump. We also find that the Newtonian acceleration effect dominates the gravitational-wave strain for both ground and space-based interferometers. We then compare our dark matter clump, gravitational-wave strain model to 84 Koi-Fish glitches detected during the second observing run of the LIGO/Virgo/KAGRA collaboration through a Markov Chain Monte Carlo Bayesian analysis. We find that all glitches but 9 can be confidently rejected as having originated from dark matter clumps. For the remaining glitches, the dark matter hypothesis cannot be excluded, and the maximum \textit{a posteriori} parameters yield minimum densities of about $10^{-7} {\rm{g}}/{\rm{cm}}^3$, within the model. These results allow us to place the first direct upper limits with gravitational-wave detectors on the local over-density of dark matter in the form of clumps in the local neighborhood of Earth, namely $ρ_{{\rm DM} \, {\rm clumps}} \lesssim 10^{-15} {\rm{g}}/{\rm{cm}}^{-3}$.

gr-qc

Di-Higgs to 4b with Bayesian inference: improving simulation estimates

Measuring di-Higgs production in the four-bottom channel is challenged by overwhelming QCD backgrounds and imperfect simulations. We develop a Bayesian mixture model that simultaneously infers signal and background fractions and their individual shapes directly in the signal region. The likelihood is a nuanced combination of a one-dimensional kinematic discriminator and per-jet flavour scores; with their correlations incorporated via kinematic bins. Monte Carlo informs weak Dirichlet priors, while the posterior adjusts to the interplay of the model, priors and observed data. Using pseudo-data simulated with standard tools and with controlled mismatches, we show that the method corrects biased priors, delivers calibrated 68-95% credible intervals for the signal count, and improves dataset-level ROC/AUC relative to simple cut-and-count baselines. This study highlights how Bayesian inference can harvest information present in the signal region and self-calibrate model parameters, providing a robust route to increased sensitivity in di-Higgs searches.

hep-ph

Inferring correlated distributions: boosted top jets

Improving the understanding of signal and background distributions in signal-region is a valuable key to enhance any analysis in collider physics. This is usually a difficult task because -- among others -- signal and backgrounds are hard to discriminate in signal-region, simulations may reach a limit of reliability if they need to model non-perturbative QCD, and distributions are multi-dimensional and many times may be correlated within each class. Bayesian density estimation is a technique that leverages prior knowledge and data correlations to effectively extract information from data in signal-region. In this work we extend previous works on data-driven mixture models for meaningful unsupervised signal extraction in collider physics to incorporate correlations between features. Using a standard dataset of top and QCD jets, we show how simulators, despite having an expected bias, can be used to inject sufficient inductive nuance into an inference model in terms of priors to then be corrected by data and estimate the true correlated distributions between features within each class. We compare the model with and without correlations to show how the signal extraction is sensitive to their inclusion and we quantify the improvement due to the inclusion of correlations using both supervised and unsupervised metrics.

hep-ph

Improvement and generalization of ABCD method with Bayesian inference

To find New Physics or to refine our knowledge of the Standard Model at the LHC is an enterprise that involves many factors. We focus on taking advantage of available information and pour our effort in re-thinking the usual data-driven ABCD method to improve it and to generalize it using Bayesian Machine Learning tools. We propose that a dataset consisting of a signal and many backgrounds is well described through a mixture model. Signal, backgrounds and their relative fractions in the sample can be well extracted by exploiting the prior knowledge and the dependence between the different observables at the event-by-event level with Bayesian tools. We show how, in contrast to the ABCD method, one can take advantage of understanding some properties of the different backgrounds and of having more than two independent observables to measure in each event. In addition, instead of regions defined through hard cuts, the Bayesian framework uses the information of continuous distribution to obtain soft-assignments of the events which are statistically more robust. To compare both methods we use a toy problem inspired by $pp\to hh\to b\bar b b \bar b$, selecting a reduced and simplified number of processes and analysing the flavor of the four jets and the invariant mass of the jet-pairs, modeled with simplified distributions. Taking advantage of all this information, and starting from a combination of biased and agnostic priors, leads us to a very good posterior once we use the Bayesian framework to exploit the data and the mutual information of the observables at the event-by-event level. We show how, in this simplified model, the Bayesian framework outperforms the ABCD method sensitivity in obtaining the signal fraction in scenarios with $1\%$ and $0.5\%$ true signal fractions in the dataset. We also show that the method is robust against the absence of signal.

hep-ph

Inferring flavor mixtures in multijet events

Multijet events with heavy-flavors are of central importance at the LHC since many relevant processes -- such as $t\bar t$, $hh$, $t\bar t h$ and others -- have a preferred branching ratio for this final state. Current techniques for tackling these processes use hard-assignment selections through $b$-tagging working points, and suffer from systematic uncertainties because of the difficulties in Monte Carlo simulations. We develop a flexible Bayesian mixture model approach to simultaneously infer $b$-tagging score distributions and the flavor mixture composition in the dataset. We model multidimensional jet events, and to enhance estimation efficiency, we design structured priors that leverages the continuity and unimodality of the $b$-tagging score distributions. Remarkably, our method eliminates the need for a parametric assumption and is robust against model misspecification -- It works for arbitrarily flexible continuous curves and is better if they are unimodal. We have run a toy inferential process with signal $bbbb$ and backgrounds $bbcc$ and $cccc$, and we find that with a few hundred events we can recover the true mixture fractions of the signal and backgrounds, as well as the true $b$-tagging score distribution curves, despite their arbitrariness and nonparametric shapes. We discuss prospects for taking these findings into a realistic scenario in a physics analysis. The presented results could be a starting point for a different and novel kind of analysis in multijet events, with a scope competitive with current state-of-the-art analyses. We also discuss the possibility of using these results in general cases of signals and backgrounds with approximately known continuous distributions and/or expected unimodality.

hep-ph

Exploring unsupervised top tagging using Bayesian inference

Recognizing hadronically decaying top-quark jets in a sample of jets, or even its total fraction in the sample, is an important step in many LHC searches for Standard Model and Beyond Standard Model physics as well. Although there exists outstanding top-tagger algorithms, their construction and their expected performance rely on Montecarlo simulations, which may induce potential biases. For these reasons we develop two simple unsupervised top-tagger algorithms based on performing Bayesian inference on a mixture model. In one of them we use as the observed variable a new geometrically-based observable $\tilde{A}_{3}$, and in the other we consider the more traditional $τ_{3}/τ_{2}$ $N$-subjettiness ratio, which yields a better performance. As expected, we find that the unsupervised tagger performance is below existing supervised taggers, reaching expected Area Under Curve AUC $\sim 0.80-0.81$ and accuracies of about 69% $-$ 75% in a full range of sample purity. However, these performances are more robust to possible biases in the Montecarlo that their supervised counterparts. Our findings are a step towards exploring and considering simpler and unbiased taggers.

hep-ph

Bayesian inference to study a signal with two or more decaying particles in a non-resonant background

We study the application of a Bayesian method to extract relevant information from data for the case of a signal consisting of two or more decaying particles and its background. The method takes advantage of the dependence that exists in the distributions of the decaying products at the event-by-event level and processes the information for the whole sample to infer the mixture fraction and the relevant parameters for signal and background distributions. The algorithm usually needs a numerical computation of the posterior, which we work out explicitly in a benchmark scenario of a simplified $pp\to hh \to b\bar b γγ$ search. We perform a posterior predictive check on the results and we show how the signal fraction is correctly extracted from the sample, as well as many parameters in the signal and background distributions. The presented framework could be used for other searches such as $pp\to ZZ,\ WW,\ ZW$ and pair of Leptoquarks, among many others.

hep-ph

Bayesian Probabilistic Modelling for Four-Tops at the LHC

Monte Carlo (MC) generators are crucial for analyzing data in particle collider experiments. However, often even a small mismatch between the MC simulations and the measurements can undermine the interpretation of the results. This is particularly important in the context of LHC searches for rare physics processes within and beyond the standard model (SM). One of the ultimate rare processes in the SM currently being explored at the LHC, $pp\to t\bar tt \bar t$ with its large multi-dimensional phase-space is an ideal testing ground to explore new ways to reduce the impact of potential MC mismodelling on experimental results. We propose a novel statistical method capable of disentangling the 4-top signal from the dominant backgrounds in the same-sign dilepton channel, while simultaneously correcting for possible MC imperfections in modelling of the most relevant discriminating observables -- the jet multiplicity distributions. A Bayesian mixture of multinomials is used to model the light-jet and $b$-jet multiplicities under the assumption of their conditional independence. The signal and background distributions generated from a deliberately mistuned MC simulator are used as model priors. The posterior distributions, as well as the signal and background fractions, are then learned from the data using Bayesian inference. We demonstrate that our method can mitigate the effects of large MC mismodellings in the context of a realistic $t\bar tt\bar t$ search, leading to corrected posterior distributions that better approximate the underlying truth-level spectra.

hep-ph

Unsupervised quark/gluon jet tagging with Poissonian Mixture Models

The classification of jets induced by quarks or gluons is important for New Physics searches at high-energy colliders. However, available taggers usually rely on modelling the data through Monte Carlo simulations, which could veil intractable theoretical and systematical uncertainties. To significantly reduce biases, we propose an unsupervised learning algorithm that, given a sample of jets, can learn the SoftDrop Poissonian rates for quark- and gluon-initiated jets and their fractions. We extract the Maximum Likelihood Estimates for the mixture parameters and the posterior probability over them. We then construct a quark-gluon tagger and estimate its accuracy in actual data to be in the $0.65-0.7$ range, below supervised algorithms but nevertheless competitive. We also show how relevant unsupervised metrics perform well, allowing for an unsupervised hyperparameter selection. Further, we find that this result is not affected by an angular smearing introduced to simulate detector effects for central jets. The presented unsupervised learning algorithm is simple; its result is interpretable and depends on very few assumptions.

hep-ph

Topping-up multilepton plus b-jets anomalies at the LHC with a $Z'$ boson

During the last years ATLAS and CMS have reported a number of slight to mild discrepancies in signatures of multileptons plus $b$-jets in analyses such as $t\bar t H$, $t\bar t W^\pm$, $t\bar t Z$ and $t\bar t t\bar t$. Among them, a recent ATLAS result on $t\bar t H$ production has also reported an excess in the charge asymmetry in the same-sign dilepton channel with two or more $b$-tagged jets. Motivated by these tantalizing discrepancies, we study a phenomenological New Physics model consisting of a $Z'$ boson that couples to up-type quarks via right-handed currents: $t_Rγ^μ\bar t_R$, $t_Rγ^μ\bar c_R$, and $t_R γ^μ\bar u_R$. The latter vertex allows to translate the charge asymmetry at the LHC initial state protons to a final state with top quarks which, decaying to a positive lepton and a $b$-jet, provides a crucial contribution to some of the observed discrepancies. Through an analysis at a detector level, we select the region in parameter space of our model that best reproduces the data in the aforementioned $t\bar t H$ study, and in a recent ATLAS $t\bar t t \bar t$ search. We find that our model provides a better fit to the experimental data than the Standard Model for a New Physics scale of approximately $\sim$500 GeV, and with a hierarchical coupling of the $Z'$ boson that favours the top quark and the presence of FCNC currents. In order to estimate the LHC sensitivity to this signal, we design a broadband search featuring many kinematic regions with different signal-to-background ratio, and perform a global analysis. We also define signal-enhanced regions and study observables that could further distinguish signal from background. We find that the region in parameter space of our model that best fits the analysed data could be probed with a significance exceeding 3 standard deviations with just the full Run-2 dataset.

hep-ph

Estimating COVID-19 cases and outbreaks on-stream through phone-calls

One of the main problems in controlling COVID-19 epidemic spread is the delay in confirming cases. Having information on changes in the epidemic evolution or outbreaks rise before lab-confirmation is crucial in decision making for Public Health policies. We present an algorithm to estimate on-stream the number of COVID-19 cases using the data from telephone calls to a COVID-line. By modeling the calls as background (proportional to population) plus signal (proportional to infected), we fit the calls in Province of Buenos Aires (Argentina) with coefficient of determination $R^2 > 0.85$. This result allows us to estimate the number of cases given the number of calls from a specific district, days before the lab results are available. We validate the algorithm with real data. We show how to use the algorithm to track on-stream the epidemic, and present the Early Outbreak Alarm to detect outbreaks in advance to lab results. One key point in the developed algorithm is a detailed track of the uncertainties in the estimations, since the alarm uses the significance of the observables as a main indicator to detect an anomaly. We present the details of the explicit example in Villa Azul (Quilmes) where this tool resulted crucial to control an outbreak on time. The presented tools have been designed in urgency with the available data at the time of the development, and therefore have their limitations which we describe and discuss. We consider possible improvements on the tools, many of which are currently under development.

q-bio.PE

Z'-explorer: a simple tool to probe Z' models against LHC data

New Physics model building requires a vast number of cross-checks against available experimental results. In particular, new neutral, colorless, spin-1 bosons $Z'$, can be found in many models. We introduce in this work a new easy-to-use software Z'-explorer which probes $Z'$ models to all available decay channels at LHC. This program scrutinizes the parameter space of the model to determine which part is still allowed, which is to be shortly explored, and which channel is the most sensitive in each region of parameter space. User does not need to implement the model nor run any Monte Carlo simulation, but instead just needs to use the $Z'$ mass and its couplings to Standard Model particles. We describe Z'-explorer backend and provide instructions to use it from its frontend, while applying it to a variety of $Z'$ models. In particular we show Z'-explorer application and utility in a sequential Standard Model, a B-L $Z'$ and a simplified two-sector or Warped/Composite model. The output of the program condenses the phenomenology of the model features, the experimental techniques and the search strategies in each channel in an enriching outcome. We find that compelling add-ons to the software would be to include correlation between decay channels, low-energy physics results, and Dark Matter searches. The software is open-source ready to use, and available for modifications, improvements and updates by the community.

hep-ph

Containing COVID-19 outbreaks using a Firewall

COVID-19 outbreaks have proven to be very difficult to isolate and extinguish before they spread out. An important reason behind this might be that epidemiological barriers consisting in stopping symptomatic people are likely to fail because of the contagion time before onset, mild cases and/or asymptomatics carriers. Motivated by these special COVID-19 features, we study a scheme for containing an outbreak in a city that consists in adding an extra firewall block between the outbreak and the rest of the city. We implement a coupled compartment model with stochastic noise to simulate a localized outbreak that is partially isolated and analyze its evolution with and without firewall for different plausible model parameters. We explore how further improvements could be achieved if the epidemic evolution would trigger policy changes for the flux and/or lock-down in the different blocks. Our results show that a substantial improvement is obtained by merely adding an extra block between the outbreak and the bulk of the city.

physics.soc-ph

COVID-19 mild cases determination from correlating COVID-line calls to reported cases

Background: One of the most challenging keys to understand COVID-19 evolution is to have a measure on those mild cases which are never tested because their few symptoms are soft and/or fade away soon. The problem is not only that they are difficult to identify and test, but also that it is believed that they may constitute the bulk of the cases and could be crucial in the pandemic equation. Methods: We present a novel algorithm to extract the number of these mild cases by correlating a COVID-line calls to reported cases in given districts. The key assumption is to realize that, being a highly contagious disease, the number of calls by mild cases should be proportional to the number of reported cases. Whereas a background of calls not related to infected people should be proportional to the district population. Results: We find that for Buenos Aires Province, in addition to the background, there are in signal 6.6 +/- 0.4 calls per each reported COVID-19 case. Using this we estimate in Buenos Aires Province 20 +/- 2 COVID-19 symptomatic cases for each one reported. Conclusions: A very simple algorithm that models the COVID-line calls as sum of signal plus background allows to estimate the crucial number of the rate of symptomatic to reported COVID-19 cases in a given district. The result from this method is an early and inexpensive estimate and should be contrasted to other methods such as serology and/or massive testing.

q-bio.PE

A Machine Learning alternative to placebo-controlled clinical trials upon new diseases: A primer

The appearance of a new dangerous and contagious disease requires the development of a drug therapy faster than what is foreseen by usual mechanisms. Many drug therapy developments consist in investigating through different clinical trials the effects of different specific drug combinations by delivering it into a test group of ill patients, meanwhile a placebo treatment is delivered to the remaining ill patients, known as the control group. We compare the above technique to a new technique in which all patients receive a different and reasonable combination of drugs and use this outcome to feed a Neural Network. By averaging out fluctuations and recognizing different patient features, the Neural Network learns the pattern that connects the patients initial state to the outcome of the treatments and therefore can predict the best drug therapy better than the above method. In contrast to many available works, we do not study any detail of drugs composition nor interaction, but instead pose and solve the problem from a phenomenological point of view, which allows us to compare both methods. Although the conclusion is reached through mathematical modeling and is stable upon any reasonable model, this is a proof-of-concept that should be studied within other expertises before confronting a real scenario. All calculations, tools and scripts have been made open source for the community to test, modify or expand it. Finally it should be mentioned that, although the results presented here are in the context of a new disease in medical sciences, these are useful for any field that requires a experimental technique with a control group.

q-bio.QM

Topic Model for four-top at the LHC

We study the implementation of a Topic Model algorithm in four-top searches at the LHC as a test-probe of a not ideal system for applying this technique. We study this Topic Model behavior as its different hypotheses such as mutual reducibility and equal distribution in all samples shift from true. The four-top final state at the LHC is not only relevant because it does not fulfill these conditions, but also because it is a difficult and inefficient system to reconstruct and current Monte Carlo modeling of signal and backgrounds suffers from non-negligible uncertainties. We implement this Topic Model algorithm in the Same-Sign lepton channel where S/B is of order one and all backgrounds cannot have more than two b-jets at parton level. We define different mixtures according to the number of b-jets and we use the total number of jets to demix. Since only the background has an anchor bin, we find that we can reconstruct the background in the signal region independently of Monte Carlo. We propose to use this information to tune the Monte Carlo in the signal region and then compare signal prediction with data. We also explore Machine Learning techniques applied to this Topic Model algorithm and find slight improvements as well as potential roads to investigate. Although our findings indicate that still with the full LHC run 3 data the implementation would be challenging, we pursue through this work to find ways to reduce the impact of Monte Carlo simulations in four-top searches at the LHC.

hep-ph

Intelligent Arxiv: Sort daily papers by learning users topics preference

Current daily paper releases are becoming increasingly large and areas of research are growing in diversity. This makes it harder for scientists to keep up to date with current state of the art and identify relevant work within their lines of interest. The goal of this article is to address this problem using Machine Learning techniques. We model a scientific paper to be built as a combination of different scientific knowledge from diverse topics into a new problem. In light of this, we implement the unsupervised Machine Learning technique of Latent Dirichlet Allocation (LDA) on the corpus of papers in a given field to: i) define and extract underlying topics in the corpus; ii) get the topics weight vector for each paper in the corpus; and iii) get the topics weight vector for new papers. By registering papers preferred by a user, we build a user vector of weights using the information of the vectors of the selected papers. Hence, by performing an inner product between the user vector and each paper in the daily Arxiv release, we can sort the papers according to the user preference on the underlying topics. We have created the website IArxiv.org where users can read sorted daily Arxiv releases (and more) while the algorithm learns each users preference, yielding a more accurate sorting every day. Current IArxiv.org version runs on Arxiv categories astro-ph, gr-qc, hep-ph and hep-th and we plan to extend to others. We propose several new useful and relevant implementations to be additionally developed as well as new Machine Learning techniques beyond LDA to further improve the accuracy of this new tool.

cs.LG