Searcharxiv⌕ Search

arXiv subjects

C. Vico Villalba

Publications and source records attributed to C. Vico Villalba.

2 recordsLinked to original sources

FPGA Acceleration of Matrix-Element Calculations for Monte Carlo Event Generation

We present an FPGA-based study of matrix-element acceleration for Monte Carlo event generation, using MadGraph5_aMC@NLO as a benchmark framework. Two complementary scenarios are considered. First, we implement the full matrix-element workflow on an AMD Alveo U250 accelerator for the benchmark process $e^+e^- \to μ^+μ^-$, enabling an end-to-end evaluation of FPGA acceleration for a simple process. Second, for the more complex $gg \to t\bar{t}+X$ processes with increasing jet multiplicity, we investigate FPGA acceleration of the color-algebra kernels as a structured and scalable entry point for selective acceleration. In this second case, the reported speedups correspond to the isolated color-reduction kernel operating on precomputed amplitudes, rather than to the full matrix-element evaluation or the complete event-generation workflow. The proposed implementations are developed using High-Level Synthesis and are evaluated in terms of numerical accuracy, performance, energy efficiency, resource utilization, and scalability. Compared with CPU and GPU implementations available within the MG5aMC framework, the FPGA solutions achieve substantial speedups and significantly improved energy efficiency. For the considered benchmarks, the numerical results remain in close agreement with the corresponding CPU reference calculations, while the resource analysis highlights the importance of numerical representation in determining scalability on FPGA devices. These results support the use of FPGAs as a competitive architecture for selected Monte Carlo event-generation workloads in high-energy physics.

hep-ex↗

Cascade Pipeline for Leading-Order Matrix Element Evaluation on AMD Versal AI Engine Arrays

A major computational bottleneck in modern High Energy Physics event generators arises from the integration of the matrix element, which requires repeated evaluations at different phase-space points to cover all possible initial- and final-state configurations. As the Large Hadron Collider enters its High-Luminosity phase, the demand for energy-efficient acceleration is expected to exceed the limits of conventional CPU scaling, motivating the use of highly parallel computing platforms such as graphics processing units (GPUs). In this work, we present an alternative approach based on a cascade pipeline architecture for evaluating leading-order matrix elements of the \ggttg process on AMD Versal AI Engine (\aie) arrays. Due to the 16\,kB per-tile program memory constraint, the computation is decomposed into a five-stage pipeline, with stages communicating via a wavefunction-token protocol over the on-chip cascade interface. Mapping 80 independent pipelines onto the 400 \aie tiles of the VCK190 platform yields a projected throughput of $1.0\times10^6$ matrix element evaluations per second at 54.8\,W, corresponding to a $34\times$ speedup over a single CPU core and a $7.7\times$ improvement in energy efficiency. Numerical agreement with the \amcnlo double-precision reference is validated at the parts-per-million level in mean relative error.

hep-ex↗