SearcharxivSearch

arXiv subjects

Lukas Heinrich

Publications and source records attributed to Lukas Heinrich.

At least 19 recordsLinked to original sources

The Analysis, not the Aperture: End-to-End Transformer Reconstruction for Imaging Atmospheric Cherenkov Telescopes

Imaging Atmospheric Cherenkov Telescopes (IACTs) detect very-high-energy gamma rays by imaging the nanosecond Cherenkov flash of the air shower they initiate in the Earth's atmosphere. For four decades the first steps of IACT event reconstruction have been essentially unchanged, relying on a heavy parameterisation and dimensionality reduction of the recorded images. This is reasonable when the image is bright, but discards important information when only a few tens of Cherenkov photons are recorded, which is a primary reason why small telescopes perform poorly at sub-TeV energies. We show that this limitation is a property of the analysis rather than of the hardware. We simulate a deliberately simple and idealised compact telescope and treat each event as a short movie that is passed directly to a video vision transformer with a factorised spatio-temporal encoder. A single composite network with a gradient-normalised multi-task loss performs gamma/hadron classification, energy regression and arrival-direction regression at once. This is the first application of a video vision transformer to IACT data. We compare it against an optimised standard analysis on the same dataset. The transformer lowers the energy threshold by a factor of three, from 0.22 to 0.07 TeV, and reconstructs arrival directions down to 0.05 TeV. At 0.2 TeV it increases the effective collection area by a factor of three, and raises the gamma/hadron separation power from an area under the receiver operating characteristic curve of 0.80 to 0.91. At 0.05 TeV, where the standard analysis retains almost nothing, that area grows by nearly two orders of magnitude. These results show promising new opportunities for compact and affordable telescopes operating at sub-TeV energies, paving the way for a broader exploration of time-domain astrophysics.

astro-ph.IM

pylhe: A Lightweight Python interface to Les Houches Event files

Les Houches Event files are a standard format for Monte Carlo event generators in high-energy physics. pylhe is a lightweight pure-Python library for reading and writing LHE event data. It supports .lhe and compressed .lhe.gz files, implements the widely used LHE 3.0 features including multiple event weights and generator metadata, and also supports the recent HDF5-based LHEH5 format. By exposing events through a pythonic iterator and shared data structures, pylhe enables memory-efficient processing of large event samples without loading entire files into memory. The library also supports format conversion and integration with modern analysis workflows, including columnar analysis with Awkward Array and downstream machine-learning applications.

hep-ph

Are We Ready for AI-Driven Discovery? AI Verification Before the Next Fundamental Physics Breakthrough

Machine learning (ML) has become integral to fundamental physics, accelerating statistical workflows from data acquisition through inference and hypothesis testing. As ML systems grow increasingly autonomous, ensuring their reliability for discovery claims becomes critical. This review synthesizes the VERaiPHY (Validation & Evaluation for Robust AI in PHYsics) initiative's frameworks for rigorous ML assessment across particle physics, astrophysics, and cosmology. We establish when verification is essential by contextualizing ML within the statistical discovery workflow. We emphasize fundamental limitations: inductive bias is unavoidable, sample complexity bounds learning, and experimental constraints limit discovery. We reflect on physicists' evolving role as both experimental designers and evaluators whose judgments encode scientific rigor into AI systems. Responsible integration requires understanding ML's transformative potential alongside its intrinsic boundaries.

physics.data-an

Gradient estimators for parameter inference in discrete stochastic kinetic models

Stochastic kinetic models are ubiquitous in physics, yet inferring their parameters from experimental data remains challenging. For deterministic models, parameter inference often relies on gradients, which can be obtained efficiently through automatic differentiation (AD). However, AD cannot be applied directly to the Gillespie stochastic simulation algorithm (SSA), since sampling from a discrete set of reactions introduces non-differentiable operations. In this work, we adopt three gradient estimators from machine learning for the Gillespie SSA: the Gumbel-Softmax Straight-Through (GS-ST) estimator, the Score Function estimator, and the Alternative Path estimator. We use the estimators to evaluate gradients of steady-state and time-dependent observables, and compare their performance in representative biophysical systems with relaxation dynamics (bimolecular association) and oscillatory dynamics (repressilator). We find that the GS-ST estimator generally yields well-behaved gradient estimates, but exhibits diverging variance in challenging parameter regimes, which can cause parameter inference to fail. In these cases, other estimators provide more robust, lower variance gradients. Our results demonstrate that gradient-based parameter inference can be effectively combined with the Gillespie SSA, with different estimators offering complementary advantages.

physics.comp-ph

Exploring the Boundaries of Differentiable Radiation Transport and Detector Simulation

We present an application of automatic differentiation for particle transport through matter using a Geant4-like radiation transport simulation with a full electromagnetic physics model. When differentiating this step-based transport, we observe exploding gradients driven by rare but extreme sensitivities at material boundaries, which propagate through subsequent transport and shower development. To obtain usable derivatives for optimization, we introduce a targeted mitigation strategy that stops gradient propagation through boundary-crossing operations under identifiable unstable conditions while leaving the forward (primal) simulation unchanged. We demonstrate that this enables stable, optimization-ready gradients in a detector-design problem.

physics.ins-det

It Just Takes Two: Scaling Amortized Inference to Large Sets

Neural posterior estimation has emerged as a powerful tool for amortized inference, with growing adoption across scientific and applied domains. In many of these applications, the conditioning variable is a set of observations whose elements depend not only on the target but also on unknown factors shared across the set. Optimal inference therefore requires treating the set jointly, which in turn requires training the estimator at the deployment set size -- a regime where memory and compute quickly become prohibitive. We introduce a simple, theoretically grounded strategy that decouples representation learning from posterior modeling. Our method trains a mean-pool Deep Set on sets of size at most two, producing an encoder that generalizes to arbitrary set sizes. The inference head is then finetuned on pre-aggregated embeddings, making training cost essentially independent of the deployment set size N. Across scalar, image, multi-view 3D, molecular, and high-dimensional conditional generation benchmarks with N in the thousands, our approach matches or outperforms standard baselines at a fraction of the compute.

cs.LG

BRICKS: Compositional Neural Markov Kernels for Zero-Shot Radiation-Matter Simulation

We introduce a new strategy for compositional neural surrogates for radiation-matter interactions, a key task spanning domains from particle physics through nuclear and space engineering to medical physics. Exploiting the locality and the Markov nature of particle interactions, we create a \emph{next-particle prediction} kernel using hybrid discrete-continuous transformer models based on Riemannian Flow Matching on product manifolds. The model generates variable-sized typed sets of particles and radiation side effects that are the result of the interaction of an incident particle with a material volume. The resulting kernel can be composed to simulate unseen large-scale material distributions in a zero-shot manner. Unlike mechanistic simulators, our model is designed to be differentiable, provides tractable likelihoods for future downstream applications. A significant computational speed-up on GPU compared to CPU-bound mechanistic simulation is observed for single-kernel execution. We evaluate the model at the kernel level and demonstrate predictive stability over multi-round autoregressive rollouts. We additionally release a novel 20M-event radiation-matter interaction dataset for further research.

cs.LG

maria goes NIFTy: Gaussian Process-Based Reconstruction and Denoising of Simulated (Sub-)Millimetre Single-Dish Telescope Data

(Sub-)millimetre single-dish telescopes feature faster mapping speeds and access larger spatial scales than their interferometric counterparts. However, atmospheric fluctuations tend to dominate their signals and complicate recovery of the astronomical sky. Here we develop a framework for Gaussian process-based sky reconstruction and separation of the atmospheric emission from the astronomical signal based on Numerical Information Field Theory (NIFTy). To validate this novel approach, we use the maria software to generate synthetic time-ordered observational data mimicking the MUSTANG-2 bolometric array. This approach leads to significantly improved sky reconstructions versus traditional methods.

astro-ph.IM

On the Codesign of Scientific Experiments and Industrial Systems

The optimization of large experiments in fundamental science, such as detectors for subnuclear physics at particle colliders, shares with the optimization of complex systems for industrial or societal applications the common issue of addressing the inter-relation between parameters describing the hardware used in data production and parameters used to analyse those data. While in many cases this coupling can be ignored -- when the problem can be successfully factored into simpler sub-tasks and the latter addressed serially -- there are situations in which that approach fails to converge to the absolute maximum of expected performance, as it results in a mis-alignment of the optimized hardware and software solutions. In this work we consider a few use cases of interest in fundamental science collected primarily from particle physics and related areas, and a pot-pourri of industrial and societal applications where the matter is similarly of relevance. We discuss the emergence of strong hardware-software coupling in some of those systems, as well as co-design procedures that may be deployed to identify the global maximum of their relevant utility functions. We observe how numerous opportunities exist to advance methods and tools for hardware-software co-design optimization, bridging fundamental science and industry through application- and challenge-driven projects, and shaping the future of scientific experiments and industrial systems.

physics.ins-det

Neural Scaling Laws for Boosted Jet Tagging

The success of Large Language Models (LLMs) has established that scaling compute, through joint increases in model capacity and dataset size, is the primary driver of performance in modern machine learning. While machine learning has long been an integral component of High Energy Physics (HEP) data analysis workflows, the compute used to train state-of-the-art HEP models remains orders of magnitude below that of industry foundation models. With scaling laws only beginning to be studied in the field, we investigate neural scaling laws for boosted jet classification using the public JetClass dataset. We derive compute optimal scaling laws and identify an effective performance limit that can be consistently approached through increased compute. We study how data repetition, common in HEP where simulation is expensive, modifies the scaling yielding a quantifiable effective dataset size gain. We then study how the scaling coefficients and asymptotic performance limits vary with the choice of input features and particle multiplicity, demonstrating that increased compute reliably drives performance toward an asymptotic limit, and that more expressive, lower-level features can raise the performance limit and improve results at fixed dataset size.

hep-ex

Differentiable quantum-trajectory simulation of Lindblad dynamics for QGP transport-coefficient inference

We study parameter estimation for the transport coefficients of the quark-gluon plasma by differentiating open-quantum-system-based Monte Carlo simulations of quarkonium suppression. The underlying simulator requires solving a Lindblad equation in a large Hilbert space, which makes parameter estimation computationally expensive. We approach the problem using gradient-based optimization. Specifically, we apply the score-function gradient estimator to differentiate through discrete jump sampling in the Monte Carlo wave-function algorithm used to solve the Lindblad equation. The resulting stochastic gradient estimator exhibits sufficiently low variance and can still be estimated in an embarrassingly parallel manner, enabling efficient scaling of the simulations. We implement this gradient estimator in the existing open-source quarkonium suppression code QTraj. To demonstrate its utility for parameter estimation, we infer the two transport coefficients $\hatκ$ and $\hatγ$ using gradient-based optimization on synthetic nuclear modification factor data.

physics.comp-ph

aim-resolve: Automatic Identification and Modeling for Bayesian Radio Interferometric Imaging

Modern radio interferometers deliver large volumes of data containing high-sensitivity sky maps over wide fields-of-view. These large area observations can contain various and superposed structures such as point sources, extended objects, and large-scale diffuse emission. To fully realize the potential of these observations, it is crucial to build appropriate sky emission models which separate and reconstruct the underlying astrophysical components. We introduce aim-resolve, an automatic and iterative method that combines the Bayesian imaging algorithm resolve with deep learning and clustering algorithms in order to jointly solve the reconstruction and source extraction problem. The method identifies and models different astrophysical components in radio observations while providing uncertainty quantification of the results. By using different model descriptions for point sources, extended objects, and diffuse background emission, the method efficiently separates the individual components and improves the overall reconstruction. We demonstrate the effectiveness of this method on synthetic image data containing multiple different sources. We further show the application of aim-resolve to an L-band (856 - 1712 MHz) MeerKAT observation of the radio galaxy ESO 137-006 and other radio galaxies in that environment. We observe a reasonable object identification for both applications, yielding a clean separation of the individual components and precise reconstructions of point sources and extended objects along with detailed uncertainty quantification. In particular, the method enables the creation of catalogs containing source positions and brightnesses and the corresponding uncertainties. The full decoupling of sky emission model and instrument response makes the method applicable to a wide variety of instruments or wavelength bands.

astro-ph.IM

Double Descent and Overparameterization in Particle Physics Data

Recently, the benefit of heavily overparameterized models has been observed in machine learning tasks: models with enough capacity to easily cross the \emph{interpolation threshold} improve in generalization error compared to the classical bias-variance tradeoff regime. We demonstrate this behavior for the first time in particle physics data and explore when and where `double descent' appears and under which circumstances overparameterization results in a performance gain.

hep-ex

Machine Learning for the Cluster Reconstruction in the CALIFA Calorimeter at R3B

The R3B experiment at FAIR studies nuclear reactions using high-energy radioactive beams. One key detector in R3B is the CALIFA calorimeter consisting of 2544 CsI(Tl) scintillator crystals designed to detect light charged particles and gamma rays with an energy resolution in the per cent range after Doppler correction. Precise cluster reconstruction from sparse hit patterns is a crucial requirement. Standard algorithms typically use fixed cluster sizes or geometric thresholds. To enhance performance, advanced machine learning techniques such as agglomerative clustering were implemented to use the full multi-dimensional parameter space including geometry, energy and time of individual interactions. An Edge Detection Neural Network exhibited significant differences. This study, based on Geant4 simulations, demonstrates improvements in cluster reconstruction efficiency of more than 30%, showcasing the potential of machine learning in nuclear physics experiments.

physics.ins-det

Flow Annealed Importance Sampling Bootstrap meets Differentiable Particle Physics

High-energy physics requires the generation of large numbers of simulated data samples from complex but analytically tractable distributions called matrix elements. Surrogate models, such as normalizing flows, are gaining popularity for this task due to their computational efficiency. We adopt an approach based on Flow Annealed importance sampling Bootstrap (FAB) that evaluates the differentiable target density during training and helps avoid the costly generation of training data in advance. We show that FAB reaches higher sampling efficiency with fewer target evaluations in high dimensions in comparison to other methods.

hep-ph

HGPflow: Extending Hypergraph Particle Flow to Collider Event Reconstruction

In high energy physics, the ability to reconstruct particles based on their detector signatures is essential for downstream data analyses. A particle reconstruction algorithm based on learning hypergraphs (HGPflow) has previously been explored in the context of single jets. In this paper, we expand the scope to full proton-proton and electron-positron collision events and study reconstruction quality using metrics at the particle, jet, and event levels. Instead of passing entire events through HGPflow, we train it on smaller partitions for scalability and to avoid potential bias from long-range correlations related to the physics process. We demonstrate that this approach is feasible and that on most metrics, HGPflow outperforms both traditional particle flow algorithms and a machine learning-based benchmark model.

hep-ex

Profile Likelihoods in Cosmology: When, Why and How illustrated with $Λ$CDM, Massive Neutrinos and Dark Energy

Frequentist parameter inference using profile likelihoods has received increased attention in the cosmology literature recently since it can give important complementary information to Bayesian credible intervals. Here, we give a pedagogical review of frequentist parameter inference in cosmology and focus on when the graphical profile likelihood construction gives meaningful constraints, i.e. confidence intervals with correct coverage. This construction rests on the assumption of the asymptotic limit of a large data set such as in Wilks' theorem. We assess the validity of this assumption in the context of three cosmological models with Planck 2018 Plik_lite data: While our tests for the $Λ$CDM model indicate that the profile likelihood method gives correct coverage, $Λ$CDM with the sum of neutrino masses as a free parameter appears consistent with a Gaussian near a boundary motivating the use of the boundary-corrected or Feldman-Cousins graphical method; for $w_0$CDM with the equation of state of dark energy, $w_0$, as a free parameter, we find indication of a violation of the assumptions. Finally, we compare frequentist and Bayesian constraints of these models. Our results motivate care when using the graphical profile likelihood method in cosmology. Along with this paper, we publish our profile-likelihood code "pinc".

astro-ph.CO

Reinterpretation and preservation of data and analyses in HEP

Data from particle physics experiments are unique and are often the result of a very large investment of resources. Given the potential scientific impact of these data, which goes far beyond the immediate priorities of the experimental collaborations that obtain them, it is imperative that the collaborations and the wider particle physics community publish and preserve sufficient information to ensure that this impact can be realised, now and into the future. The information to be published and preserved includes the algorithms, statistical information, simulations and the recorded data. This publication and preservation requires significant resources, and should be a strategic priority with commensurate planning and resource allocation from the earliest stages of future facilities and experiments.

hep-ph