Searcharxiv⌕ Search

arXiv subjects

E. Wes Bethel

Publications and source records attributed to E. Wes Bethel.

16 recordsLinked to original sources

Prefill/Decode-Aware Evaluation of LLM Inference on Emerging AI Accelerators

As large language models (LLMs) are increasingly deployed in latency- and cost-sensitive settings, inference efficiency has become a central systems challenge. While GPUs dominate current deployments, a growing number of AI accelerators claim advantages for LLM inference, yet it remains unclear under which conditions such accelerators outperform GPUs in practice. Recent inference systems decompose execution into Prefill and Decode phases, which exhibit distinct computational characteristics and latency metrics, commonly captured by time to first token (TTFT) and time per output token (TPOT). This paper presents a phase-aware evaluation of LLM inference performance across GPUs and emerging AI accelerators using a common model, Llama2-7B. By separately measuring Prefill and Decode performance, we reveal that accelerator advantages differ by phase and metric. Our results show that GPUs consistently excel in the compute-intensive Prefill phase, while GroqRack achieves significantly lower TPOT during Decode (batching not currently supported). However, GPUs regain an advantage in Decode throughput as batch size increases. These findings demonstrate that each platform exhibits distinct phase-dependent strengths. We further analyze heterogeneous Prefill/Decode disaggregation across different accelerator platforms, identifying performance gains and the workload and network conditions under which such gains are realized.

cs.AR↗

Sequence and Image Transformations with Monarq: Quantum Implementations for NISQ Devices

We introduce Monarq, a unified quantum data processing framework that combines QCrank encoding with the EHands protocol for polynomial transformations, and demonstrate its implementation on noisy intermediate-scale quantum (NISQ) hardware. This framework provides fundamental quantum building blocks for signal and image processing tasks, including convolution, discrete-time Fourier transform (DFT), squared gradient computation, and edge detection, serving as a reference for a broad class of data processing applications on near-term quantum devices.

quant-ph↗

Quantum Computing and Visualization Research Challenges and Opportunities

Quantum computing (QC) has experienced rapid growth in recent years with the advent of robust programming environments, readily accessible software simulators and cloud-based QC hardware platforms, and growing interest in learning how to design useful methods that leverage this emerging technology for practical applications. From the perspective of the field of visualization, this article examines research challenges and opportunities along the path from initial feasibility to practical use of QC platforms applied to meaningful problems.

quant-ph↗

EHands: Quantum Protocol for Polynomial Computation on Real-Valued Encoded States

We present EHands, a quantum-native protocol for implementing multivariable polynomial transformations on quantum processors. The protocol introduces four fundamental, reversible operators: multiplication, addition, negation, and parity flip, and employs the Expectation Value ENcoding (EVEN) scheme to represent real numbers as quantum states. Unlike discretization or binary encoding methods, EHands operates directly on vectorized real-valued inputs prepared in the initial state and applies a shallow quantum circuit that depends only on the polynomial coefficients. The result is obtained from the expectation value measured on a single qubit, enabling efficient parallel evaluation of a polynomial across multiple data points using a single circuit. We introduce both a reversible implementation for degree-$d$ polynomials, requiring $3d$ qubits, and a non-reversible variant that uses qubit resets to reduce the requirements to $d+1$ qubits. Both implementations exhibit linear depth scaling in $d$ and are explicitly decomposed into one- and two-qubit gates for direct execution on current quantum processing units. The protocol's effectiveness is demonstrated through experimental validation on IBM's Heron-class quantum processors, showing reliable polynomial approximations of functions like ReLU and arctan.

quant-ph↗

Intelligent IoT Attack Detection Design via ODLLM with Feature Ranking-based Knowledge Base

The widespread adoption of Internet of Things (IoT) devices has introduced significant cybersecurity challenges, particularly with the increasing frequency and sophistication of Distributed Denial of Service (DDoS) attacks. Traditional machine learning (ML) techniques often fall short in detecting such attacks due to the complexity of blended and evolving patterns. To address this, we propose a novel framework leveraging On-Device Large Language Models (ODLLMs) augmented with fine-tuning and knowledge base (KB) integration for intelligent IoT network attack detection. By implementing feature ranking techniques and constructing both long and short KBs tailored to model capacities, the proposed framework ensures efficient and accurate detection of DDoS attacks while overcoming computational and privacy limitations. Simulation results demonstrate that the optimized framework achieves superior accuracy across diverse attack types, especially when using compact models in edge computing environments. This work provides a scalable and secure solution for real-time IoT security, advancing the applicability of edge intelligence in cybersecurity.

cs.CR↗

Circuit Partitioning and Full Circuit Execution: A Comparative Study of GPU-Based Quantum Circuit Simulation

Executing large quantum circuits is not feasible using the currently available NISQ (noisy intermediate-scale quantum) devices. The high costs of using real quantum devices make it further challenging to research and develop quantum algorithms. As a result, performing classical simulations is usually the preferred method for researching and validating large-scale quantum algorithms. However, these simulations require a huge amount of resources, as each additional qubit exponentially increases the computational space required. Distributed Quantum Computing (DQC) is a promising alternative to reduce the resources required for simulating large quantum algorithms at the cost of increased runtime. This study presents a comparative analysis of two simulation methods: circuit-splitting and full-circuit execution using distributed memory, each having a different type of overhead. The first method, using CutQC, cuts the circuit into smaller subcircuits and allows us to simulate a large quantum circuit on smaller machines. The second method, using Qiskit-Aer-GPU, distributes the computational space across a distributed memory system to simulate the entire quantum circuit. Results indicate that full-circuit executions are faster than circuit-splitting for simulations performed on a single node. However, circuit-splitting simulations show promising results in specific scenarios as the number of qubits is scaled.

quant-ph↗

From Bits to Qubits: Challenges in Classical-Quantum Integration

While quantum computing holds immense potential for tackling previously intractable problems, its current practicality remains limited. A critical aspect of realizing quantum utility is the ability to efficiently interface with data from the classical world. This research focuses on the crucial phase of quantum encoding, which enables the transformation of classical information into quantum states for processing within quantum systems. We focus on three prominent encoding models: Phase Encoding, Qubit Lattice, and Flexible Representation of Quantum Images (FRQI) for cost and efficiency analysis. The aim of quantifying their different characteristics is to analyze their impact on quantum processing workflows. This comparative analysis offers valuable insights into their limitations and potential to accelerate the development of practical quantum computing solutions.

cs.ET↗

Towards a Scalable In Situ Fast Fourier Transform

The Fast Fourier Transform (FFT) is a numerical operation that transforms a function into a form comprised of its constituent frequencies and is an integral part of scientific computation and data analysis. The objective of our work is to enable use of the FFT as part of a scientific in situ processing chain to facilitate the analysis of data in the spectral regime. We describe the implementation of an FFT endpoint for the transformation of multi-dimensional data within the SENSEI infrastructure. Our results show its use on a sample problem in the context of a multi-stage in situ processing workflow.

cs.DC↗

Quantum Computing and Visualization: A Disruptive Technological Change Ahead

The focus of this Visualization Viewpoints article is to provide some background on Quantum Computing (QC), to explore ideas related to how visualization helps in understanding QC, and examine how QC might be useful for visualization with the growth and maturation of both technologies in the future. In a quickly evolving technology landscape, QC is emerging as a promising pathway to overcome the growth limits in classical computing. In some cases, QC platforms offer the potential to vastly outperform the familiar classical computer by solving problems more quickly or that may be intractable on any known classical platform. As further performance gains for classical computing platforms are limited by diminishing Moore's Law scaling, QC platforms might be viewed as a potential successor to the current field of exascale-class platforms. While present-day QC hardware platforms are still limited in scale, the field of quantum computing is robust and rapidly advancing in terms of hardware capabilities, software environments for developing quantum algorithms, and educational programs for training the next generation of scientists and engineers. After a brief introduction to QC concepts, the focus of this article is to explore the interplay between the fields of visualization and QC. First, visualization has played a role in QC by providing the means to show representations of the quantum state of single-qubits in superposition states and multiple-qubits in entangled states. Second, there are a number of ways in which the field of visual data exploration and analysis may potentially benefit from this disruptive new technology though there are challenges going forward.

quant-ph↗

Extensions to the SENSEI In situ Framework for Heterogeneous Architectures

The proliferation of GPUs and accelerators in recent supercomputing systems, so called heterogeneous architectures, has led to increased complexity in execution environments and programming models as well as to deeper memory hierarchies on these systems. In this work, we discuss challenges that arise in in situ code coupling on these heterogeneous architectures. In particular, we present data and execution model extensions to the SENSEI in situ framework that are targeted at the effective use of systems with heterogeneous architectures. We then use these new data and execution model extensions to investigate several in situ placement and execution configurations and to analyze the impact these choices have on overall performance.

cs.DC↗

Quantum-parallel vectorized data encodings and computations on trapped-ions and transmons QPUs

Compact quantum data representations are essential to the emerging field of quantum algorithms for data analysis. We introduce two new data encoding schemes, QCrank and QBArt, which have a high degree of quantum parallelism through uniformly controlled rotation gates. QCrank encodes a sequence of real-valued data as rotations of the data qubits, allowing for high storage density. QBArt directly embeds a binary representation of the data in the computational basis, requiring fewer quantum measurements and lending itself to well-understood arithmetic operations on binary data. We present several applications of the proposed encodings for different types of data. We demonstrate quantum algorithms for DNA pattern matching, Hamming weight calculation, complex value conjugation, and retrieving an O(400) bits image, all executed on the Quantinuum QPU. Finally, we use various cloud-accessible QPUs, including IBMQ and IonQ, to perform additional benchmarking experiments.

quant-ph↗

Quantum pixel representations and compression for $N$-dimensional images

We introduce a novel and uniform framework for quantum pixel representations that overarches many of the most popular representations proposed in the recent literature, such as (I)FRQI, (I)NEQR, MCRQI, and (I)NCQI. The proposed QPIXL framework results in more efficient circuit implementations and significantly reduces the gate complexity for all considered quantum pixel representations. Our method only requires a linear number of gates in terms of the number of pixels and does not use ancilla qubits. Furthermore, the circuits only consist of Ry gates and CNOT gates making them practical in the NISQ era. Additionally, we propose a circuit and image compression algorithm that is shown to be highly effective, being able to reduce the necessary gates to prepare an FRQI state for example scientific images by up to 90% without sacrificing image quality. Our algorithms are made publicly available as part of QPIXL++, a Quantum Image Pixel Library.

quant-ph↗

Performance Analysis of Traditional and Data-Parallel Primitive Implementations of Visualization and Analysis Kernels

Measurements of absolute runtime are useful as a summary of performance when studying parallel visualization and analysis methods on computational platforms of increasing concurrency and complexity. We can obtain even more insights by measuring and examining more detailed measures from hardware performance counters, such as the number of instructions executed by an algorithm implemented in a particular way, the amount of data moved to/from memory, memory hierarchy utilization levels via cache hit/miss ratios, and so forth. This work focuses on performance analysis on modern multi-core platforms of three different visualization and analysis kernels that are implemented in different ways: one is "traditional", using combinations of C++ and VTK, and the other uses a data-parallel approach using VTK-m. Our performance study consists of measurement and reporting of several different hardware performance counters on two different multi-core CPU platforms. The results reveal interesting performance differences between these two different approaches for implementing these kernels, results that would not be apparent using runtime as the only metric.

cs.DC↗

DPP-PMRF: Rethinking Optimization for a Probabilistic Graphical Model Using Data-Parallel Primitives

We present a new parallel algorithm for probabilistic graphical model optimization. The algorithm relies on data-parallel primitives (DPPs), which provide portable performance over hardware architecture. We evaluate results on CPUs and GPUs for an image segmentation problem. Compared to a serial baseline, we observe runtime speedups of up to 13X (CPU) and 44X (GPU). We also compare our performance to a reference, OpenMP-based algorithm, and find speedups of up to 7X (CPU).

cs.DC↗

Using High-Speed WANs and Network Data Caches to Enable Remote and Distributed Visualization

Visapult is a prototype application and framework for remote visualization of large scientific datasets. We approach the technical challenges of tera-scale visualization with a unique architecture that employs high speed WANs and network data caches for data staging and transmission. This architecture allows for the use of available cache and compute resources at arbitrary locations on the network. High data throughput rates and network utilization are achieved by parallelizing I/O at each stage in the application, and by pipelining the visualization process. On the desktop, the graphics interactivity is effectively decoupled from the latency inherent in network applications. We present a detailed performance analysis of the application, and improvements resulting from field-test analysis conducted as part of the DOE Combustion Corridor project.

cs.DC↗