SearcharxivSearch

arXiv subjects

Tanvi Wamorkar

Publications and source records attributed to Tanvi Wamorkar.

5 recordsLinked to original sources

Explicit or Implicit? Encoding Physics at the Precision Frontier

High-performance machine learning tools in particle physics rest on two complementary directions: encoding symmetries explicitly in the architecture, and implicitly learning the structure of the data through large-scale (pre-) training. We compare the performance of the representative L-GATr and OmniLearn models on three especially challenging tasks: reweighting-based unfolding, likelihood-ratio estimation, and weakly supervised anomaly detection. Across all benchmarks, both methods achieve comparable performance given the statistical precision of the finetuning datasets, suggesting that the significant efficiency gains from encoding known particle physics structures are largely method-independent.

hep-ph

A Scientific Human-Agent Reproduction Pipeline

Reproducing scientific analyses is essential for preserving knowledge, building extensible codebases, and deepening researcher understanding - yet the effort often outweighs its academic recognition. We argue that the reproduction of scientific data analyses is fundamentally a translation task: converting human-readable knowledge (papers, documentation) into machine-readable analysis code. This makes it uniquely well-suited for AI agents. We present SHARP (Scientific Human-Agent Reproduction Pipeline), a structured framework for reproducing scientific analyses through human-agent collaboration. SHARP decomposes a reproduction task into discrete steps, which an AI agent executes autonomously using specialized subagents for code generation, testing, and quality assurance. At defined checkpoints, the researcher reviews progress, provides feedback, and steers the analysis - keeping the human firmly in control of scientific judgment while the agent handles implementation. We demonstrate SHARP by reproducing a jet classification task in particle physics from a published paper. We evaluate the reproduction along three axes: analysis performance against the original results, code quality and faithfulness, and the nature of the human-agent conversation. The latter is evaluated with a novel framework for characterizing human-agent interactions. Our work highlights a practical model for AI-assisted scientific reproduction where the researcher's role shifts from writing code to understanding, evaluating, and directing - elevating human understanding rather than replacing it.

hep-ph

Stabilizing Neural Likelihood Ratio Estimation

Likelihood ratios are used for a variety of applications in particle physics data analysis, including parameter estimation, unfolding, and anomaly detection. When the data are high-dimensional, neural networks provide an effective tools for approximating these ratios. However, neural network training has an inherent stochasticity that limits their precision. A widely-used approach to reduce these fluctuations is to train many times and average the output (ensembling). We explore different approaches to ensembling and pretraining neural networks for stabilizing likelihood ratio estimation. For numerical studies focus on unbinned unfolding with OmniFold, as it requires many likelihood ratio estimations. We find that ensembling approaches that aggregate the models at step 1, before pushing the weights to step 2 improve both bias and variance of final results. Variance can be further improved by pre-training, however at the cost increasing bias.

hep-ph

Tools for Unbinned Unfolding

Machine learning has enabled differential cross section measurements that are not discretized. Going beyond the traditional histogram-based paradigm, these unbinned unfolding methods are rapidly being integrated into experimental workflows. In order to enable widespread adaptation and standardization, we develop methods, benchmarks, and software for unbinned unfolding. For methodology, we demonstrate the utility of boosted decision trees for unfolding with a relatively small number of high-level features. This complements state-of-the-art deep learning models capable of unfolding the full phase space. To benchmark unbinned unfolding methods, we develop an extension of existing dataset to include acceptance effects, a necessary challenge for real measurements. Additionally, we directly compare binned and unbinned methods using discretized inputs for the latter in order to control for the binning itself. Lastly, we have assembled two software packages for the OmniFold unbinned unfolding method that should serve as the starting point for any future analyses using this technique. One package is based on the widely-used RooUnfold framework and the other is a standalone package available through the Python Package Index (PyPI).

hep-ph

Combined analysis of HPK 3.1 LGADs using a proton beam, beta source, and probe station towards establishing high volume quality control

The upgrades of the CMS and ATLAS experiments for the high luminosity phase of the Large Hadron Collider will employ precision timing detectors based on Low Gain Avalanche Detectors (LGADs). We present a suite of results combining measurements from the Fermilab Test Beam Facility, a beta source telescope, and a probe station, allowing full characterization of the HPK type 3.1 production of LGAD prototypes developed for these detectors. We demonstrate that the LGAD response to high energy test beam particles is accurately reproduced with a beta source. We further establish that probe station measurements of the gain implant accurately predict the particle response and operating parameters of each sensor, and conclude that the uniformity of the gain implant in this production is sufficient to produce full-sized sensors for the ATLAS and CMS timing detectors.

physics.ins-det