SearcharxivSearch

arXiv subjects

Enzo Meneses

Publications and source records attributed to Enzo Meneses.

7 recordsLinked to original sources

Convex Hull 3D Filtering with GPU Ray Tracing and Tensor Cores

In recent years, applications such as real-time simulations, autonomous systems, and video games increasingly demand the processing of complex geometric models under stringent time constraints. Traditional geometric algorithms, including the convex hull, are subject to these challenges. A common approach to improve performance is scaling computational resources, which often results in higher energy consumption. Given the growing global concern regarding sustainable use of energy, this becomes a critical limitation. This work presents a 3D preprocessing filter for the convex hull algorithm using ray tracing and tensor core technologies. The filter builds a delimiter polyhedron based on Manhattan distances that discards points from the original set. The filter is evaluated on two point distributions: uniform and sphere. Experimental results show that the proposed filter, combined with convex hull construction, accelerates the computation of the 3D convex hull by up to 200x with respect to a CPU parallel implementation. This research demonstrates that geometric algorithms can be accelerated through massive parallelism while maintaining efficient energy utilization. Beyond execution time and speedup evaluation, we also analyze GPU energy consumption, showing that the proposed preprocessing filter not only reduces the computational workload but also achieves performance gains with controlled energy usage. These results highlight the dual benefit of the method in terms of both speed and energy efficiency, reinforcing its applicability in modern high-performance scenarios.

cs.CG

Ray Tracing Cores for General-Purpose Computing: A Literature Review

Recent research on ray tracing cores has explored repurposing these cores to solve non-graphical problems by reformulating them as geometric queries, leveraging the inherent parallelism of ray tracing. Although successful in specific cases, these applications lack a clear pattern, and the conditions under which RT cores can provide computational benefits are still not clearly understood. The objective of this literature review is to examine diverse applications of ray tracing cores in general-purpose computation, identifying common features, performance gains, and limitations. By categorizing these efforts, the review aims to provide guidance on the types of problems that can effectively exploit ray tracing hardware beyond traditional rendering tasks. This is achieved with a blibliometric review based on 59 research articles indexed in Scopus, and a systematic literature review on 35 of them which propose new RT solutions and compare them with state-of-the-art methods to solve 32 distinct problems, in some works achieving up to $200\times$ speedup. Most of the problems analyzed in this work have applications in physics simulations and in solving some geometric queries, but problems with potential applications in databases and AI can also be found. Analyzing the characteristics of the problems, it was found that nearest neighbor search, including its variants, benefit the most from ray tracing cores as well as problems that rely on heuristic to diminish the necessary work. This is aligned with the biggest strength of RT cores; discarding tree branches when traversing a tree to avoid unnecessary work. Also, it was found that many short-length rays should be preferred over a few large rays. The results found in this work can serve as a guide for knowing beforehand which applications are better potential candidates to benefit from RT Core computation.

cs.DC

Advancing RT Core-Accelerated Fixed-Radius Nearest Neighbor Search

In this work we introduce three ideas that can further improve particle FRNN physics simulations running on RT Cores; i) a real-time update/rebuild ratio optimizer for the bounding volume hierarchy (BVH) structure, ii) a new RT core use, with two variants, that eliminates the need of a neighbor list and iii) a technique that enables RT cores for FRNN with periodic boundary conditions (BC). Experimental evaluation using the Lennard-Jones FRNN interaction model as a case study shows that the proposed update/rebuild ratio optimizer is capable of adapting to the different dynamics that emerge during a simulation, leading to a RT core pipeline up to $\sim 3.4\times$ faster than with other known approaches to manage the BVH. In terms of simulation step performance, the proposed variants can significantly improve the speedup and energy efficiency (EE) of the base RT core idea; from $\sim1.3\times$ at small radius to $\sim2.0\times$ for log normal radius distributions. Furthermore, the proposed variants manage to simulate cases that would otherwise not fit in memory because of the use of neighbor lists, such as clusters of particles with log normal radius distribution. The proposed RT Core technique to support periodic BC is indeed effective as it does not introduce any significant penalty in performance. In terms of scaling, the proposed methods scale both their performance and EE across GPU generations. Throughout the experimental evaluation, we also identify the simulation cases were regular GPU computation should still be preferred, contributing to the understanding of the strengths and limitations of RT cores.

cs.DC

CAT: Cellular Automata on Tensor cores

Cellular automata (CA) are simulation models that can produce complex emergent behaviors from simple local rules. Although state-of-the-art GPU solutions are already fast due to their data-parallel nature, their performance can rapidly degrade in CA with a large neighborhood radius. With the inclusion of tensor cores across the entire GPU ecosystem, interest has grown in finding ways to leverage these fast units outside the field of artificial intelligence, which was their original purpose. In this work, we present CAT, a GPU tensor core approach that can accelerate CA in which the cell transition function acts on a weighted summation of its neighborhood. CAT is evaluated theoretically, using an extended PRAM cost model, as well as empirically using the Larger Than Life (LTL) family of CA as case studies. The results confirm that the cost model is accurate, showing that CAT exhibits constant time throughout the entire radius range $1 \le r \le 16$, and its theoretical speedups agree with the empirical results. At low radius $r=1,2$, CAT is competitive and is only surpassed by the fastest state-of-the-art GPU solution. Starting from $r=3$, CAT progressively outperforms all other approaches, reaching speedups of up to $101\times$ over a GPU baseline and up to $\sim 14\times$ over the fastest state-of-the-art GPU approach. In terms of energy efficiency, CAT is competitive in the range $1 \le r \le 4$ and from $r \ge 5$ it is the most energy efficient approach. As for performance scaling across GPU architectures, CAT shows a promising trend that if continues for future generations, it would increase its performance at a higher rate than classical GPU solutions. The results obtained in this work put CAT as an attractive GPU approach for scientists that need to study emerging phenomena on CA with large neighborhood radius.

cs.DC

Accelerating Range Minimum Queries with Ray Tracing Cores

During the last decade GPU technology has shifted from pure general purpose computation to the inclusion of application specific integrated circuits (ASICs), such as Tensor Cores and Ray Tracing (RT) cores. Although these special purpose GPU cores were designed to further accelerate specific fields such as AI and real-time rendering, recent research has managed to exploit them to further accelerate other tasks that typically used regular GPU computing. In this work we present RTXRMQ, a new approach that can compute range minimum queries (RMQs) with RT cores. The main contribution is the proposal of a geometric solution for RMQ, where elements become triangles that are placed and shaped according to the element's value and position in the array, respectively, such that the closest hit of a ray launched from a point given by the query parameters corresponds to the result of that query. Experimental results show that RTXRMQ is currently best suited for small query ranges relative to the problem size, achieving up to $5\times$ and $2.3\times$ of speedup over state of the art CPU (HRMQ) and GPU (LCA) approaches, respectively. Although for medium and large query ranges RTXRMQ is currently surpassed by LCA, it is still competitive by being $2.5\times$ and $4\times$ faster than HRMQ which is a highly parallel CPU approach. Furthermore, performance scaling experiments across the latest RTX GPU architectures show that if the current RT scaling trend continues, then RTXRMQ's performance would scale at a higher rate than HRMQ and LCA, making the approach even more relevant for future high performance applications that employ batches of RMQs.

cs.DC

Infinite conformal symmetry and emergent chiral (super)fields of topologically non-trivial configurations: From Yang-Mills-Higgs to the Skyrme model

The present manuscript discusses a remarkable phenomenon concerning non-linear and non-integrable field theories in $(3+1)$-dimensions, living at finite density and possessing non-trivial topological charges and non-Abelian internal symmetries (both local and global). With suitable types of ansätze, one can construct infinite-dimensional families of analytic solutions with non-vanishing topological charges (representing the Baryonic number) labelled by both two integers numbers and by free scalar fields in $(1+1)$-dimensions. These exact configurations represent $(3+1)$-dimensional topological solitons hosting $(1+1)$-dimensional chiral modes localized at the energy density peaks. First, we analyze the Yang-Mills-Higgs model, in which the fields depend on all the space-time coordinates (to keep alive the topological Chern-Simons charge), but in such a way to reduce the equations system to the field equations of two-dimensional free massless chiral scalar fields. Then, we move to the non-linear sigma model, showing that a suitable ansatz reduces the field equations to the one of a two-dimensional free massless scalar field. Then, we discuss the Skyrme model concluding that the inclusion of the Skyrme term gives rise to a chiral two-dimensional free massless scalar field (instead of a free massless field in two dimensions as in the non-linear sigma model) describing analytically spatially modulated Hadronic layers and tubes. The comparison of the present approach both with the instantons-dyons liquid approach and with Lattice QCD is shortly outlined.

hep-th

GGArray: A Dynamically Growable GPU Array

We present a dynamically Growable GPU array (GGArray) fully implemented in GPU that does not require synchronization with the host. The idea is to improve the programming of GPU applications that require dynamic memory, by offering a structure that does not require pre-allocating GPU VRAM for the worst case scenario. The GGArray is based on the LFVector, by utilizing an array of them in order to take advantage of the GPU architecture and the synchronization offered by thread blocks. This structure is compared to other state of the art ones such as a pre-allocated static array and a semi-static array that needs to be resized through communication with the host. Experimental evaluation shows that the GGArray has a competitive insertion and resize performance, but it is slower for regular parallel memory accesses. Given the results, the GGArray is a potentially useful structure for applications with high uncertainty on the memory usage as well as applications that have phases, such as an insertion phase followed by a regular GPU phase. In such cases, the GGArray can be used for the first phase and then data can be flattened for the second phase in order to allow the classical GPU memory accesses which are faster. These results constitute a step towards achieving a parallel efficient C++ like vector for modern GPU architectures.

cs.DC