Searcharxiv⌕ Search

arXiv subjects

Hiroshi Horii

Publications and source records attributed to Hiroshi Horii.

15 recordsLinked to original sources

Diffusion stabilises time-periodic solutions in conservation laws coupled to a relaxation oscillator

We study a viscous one-dimensional conservation law, coupled to a fast ordinary differential equation. For a vanishing viscosity, the system has an infinite-dimensional family of time-periodic solutions, corresponding to relaxation oscillations. We show that for symmetric initial conditions, a small positive viscosity selects a unique periodic solution, which we prove to be linearly stable. The proof exploits the slow-fast structure through an averaging strategy, as well as spectral-theoretic methods. The results are illustrated by numerical simulations.

math.AP↗

Crossing the 12,000-atom barrier with heterogeneous quantum-classical supercomputing: quantum chemistry of protein-ligand complexes

We develop a workflow decomposing a molecule into fragments via quantum embedding and simulating them with a heterogeneous quantum-classical (HQC) method. We sample fragment electronic configurations on two 156-qubit quantum processors (ibm$\_$cleveland, ibm$\_$kobe), using up to 94 qubits, running 21,006 circuits for over 239 hours, collecting $3.0 \cdot 10^9$ measurement outcomes -- the most resource-intensive HQC computation for quantum chemistry to date. We compute fragment wavefunctions via optimized subspace diagonalization on supercomputers Fugaku, Miyabi-G, and ROQUO, achieving 72.5$\%$ parallel efficiency with scalable distributed linear algebra kernels. We simulate two protein-ligand complexes spanning dispersion- and electrostatics-dominated regimes (11,608, 12,635 atoms), demonstrate $>40\times$ increase in system size and up to $210\times$ improvement in accuracy over previous state-of-the-art, with HQC matching coupled-cluster (CCSD) accuracy in fragment energies. We present the first HQC protein-ligand binding prediction using a mixed-basis set and an automated end-to-end workflow enabling practical HQC calculations of large protein systems.

quant-ph↗

Examining QRMI as a Unified Interface for Quantum-HPC Integration

The efficient and scalable integration of quantum resources into high-performance computing (HPC) environments requires standardized mechanisms for resource management, scheduling, and workflow orchestration across diverse and heterogeneous infrastructures. The Quantum Resource Management Interface (QRMI) addresses this challenge through a thin, vendor-agnostic middleware layer that provides standardized APIs for scheduling, executing, and monitoring quantum workloads while exposing quantum resources as first-class schedulable resources alongside CPUs and GPUs. Although previous work demonstrated QRMI integration with the Slurm workload manager, its applicability across other workload managers remained unexamined. This paper extends the validation of QRMI to a broad range of workload managers, including PBS, LSF, Grid Engine, Kubernetes, and the Flux Framework, encompassing traditional batch schedulers, a cloud-native orchestration platform, and a graph-based scheduler. We examine the integration patterns, implementation requirements, and scheduler-specific considerations associated with each environment and compare QRMI with alternative approaches to quantum resource integration. We demonstrate that QRMI provides a portable and flexible abstraction layer that minimizes scheduler-specific modifications while enabling consistent access to heterogeneous quantum resources across both on-premises and cloud environments.

cs.ET↗

Reference Architecture of a Quantum-Centric Supercomputer

Quantum computers have demonstrated utility in simulating quantum systems beyond brute-force classical approaches. As the community builds on these demonstrations to explore using quantum computing for applied research, algorithms and workflows have emerged that require leveraging both quantum computers and classical high-performance computing (HPC) systems to scale applications, especially in chemistry and materials, beyond what either system can simulate alone. Today, these disparate systems operate in isolation, forcing users to manually orchestrate workloads, coordinate job scheduling, and transfer data between systems -- a cumbersome process that hinders productivity and severely limits rapid algorithmic exploration. These challenges motivate the need for flexible and high-performance Quantum-Centric Supercomputing (QCSC) systems that integrate Quantum Processing Units (QPUs), Graphics Processing Units (GPUs), and Central Processing Units (CPUs) to accelerate discovery of such algorithms across applications. These systems will be co-designed across quantum and classical HPC infrastructure, middleware, and application layers to accelerate the adoption of quantum computing for solving critical computational problems. We envision QCSC evolution through three distinct phases: (1) quantum systems as specialized compute offload engines within existing HPC complexes; (2) heterogeneous quantum and classical HPC systems coupled through advanced middleware, enabling seamless execution of hybrid quantum-classical algorithms; and (3) fully co-designed heterogeneous quantum-HPC systems for hybrid computational workflows. This article presents a reference architecture and roadmap for these QCSC systems.

cs.ET↗

GPU-Accelerated Selected Basis Diagonalization with Thrust for SQD-based Algorithms

Selected Basis Diagonalization (SBD) plays a central role in Sample-based Quantum Diagonalization (SQD), where iterative diagonalization of the Hamiltonian in selected configuration subspaces forms the dominant classical workload. We present a GPU-accelerated implementation of SBD using the Thrust library. By restructuring key components -- including configuration processing, excitation generation, and matrix-vector operations -- around fine-grained data-parallel primitives and flattened GPU-friendly data layouts, the proposed approach efficiently exploits modern GPU architectures. In our experiments, the Thrust-based SBD achieves up to $\sim$40$\times$ speedup over CPU execution and substantially reduces the total runtime of SQD iterations. These results demonstrate that GPU-native parallel primitives provide a simple, portable, and high-performance foundation for accelerating SQD-based quantum-classical workflows.

cs.DC↗

Observability Architecture for Quantum-Centric Supercomputing Workflows

Quantum-centric supercomputing (QCSC) workflows often involve hybrid classical-quantum algorithms that are inherently probabilistic and executed on remote quantum hardware, making them difficult to interpret and limiting the ability to monitor runtime performance and behavior. The high cost of quantum circuit execution and large-scale high-performance computing (HPC) infrastructure further restricts the number of feasible trials, making comprehensive evaluation of execution results essential for iterative development. We propose an observability architecture tailored for QCSC workflows that decouples telemetry collection from workload execution, enabling persistent monitoring across system and algorithmic layers and retaining detailed execution data for reproducible and retrospective analysis, eliminating redundant runs. Applied to a representative workflow involving sample-based quantum diagonalization, our system reveals solver behavior across multiple iterations. This approach enhances transparency and reproducibility in QCSC environments, supporting infrastructure-aware algorithm design and systematic experimentation.

quant-ph↗

Closed-loop calculations of electronic structure on a quantum processor and a classical supercomputer at full scale

Quantum computers must operate in concert with classical computers to deliver on the promise of quantum advantage for practical problems. To achieve that, it is important to understand how quantum and classical computing can interact together, and how one can characterize the scalability and efficiency of hybrid quantum-classical workflows. So far, early experiments with quantum-centric supercomputing workflows have been limited in scale and complexity. Here, we use a Heron quantum processor deployed on premises with the entire supercomputer Fugaku to perform the largest computation of electronic structure involving quantum and classical high-performance computing. We design a closed-loop workflow between the quantum processors and 152,064 classical nodes of Fugaku, to approximate the electronic structure of chemistry models beyond the reach of exact diagonalization, with accuracy comparable to some all-classical approximation methods. Our work pushes the limits of the integration of quantum and classical high-performance computing, showcasing computational resource orchestration at the largest scale possible for current classical supercomputers.

quant-ph↗

Optimal Condition for Initialization Variance in Deep Neural Networks: An SGD Dynamics Perspective

Stochastic gradient descent (SGD), one of the most fundamental optimization algorithms in machine learning (ML), can be recast through a continuous-time approximation as a Fokker-Planck equation for Langevin dynamics, a viewpoint that has motivated many theoretical studies. Within this framework, we study the relationship between the quasi-stationary distribution derived from this equation and the initial distribution through the Kullback-Leibler (KL) divergence. As the quasi-steady-state distribution depends on the expected cost function, the KL divergence eventually reveals the connection between the expected cost function and the initialization distribution. By applying this to deep neural network models (DNNs), we can express the bounds of the expected loss function explicitly in terms of the initialization parameters. Then, by minimizing this bound, we obtain an optimal condition of the initialization variance in the Gaussian case. This result provides a concrete mathematical criterion, rather than a heuristic approach, to select the scale of weight initialization in DNNs. In addition, we experimentally confirm our theoretical results by using the classical SGD to train fully connected neural networks on the MNIST and Fashion-MNIST datasets. The result shows that if the variance of the initialization distribution satisfies our theoretical optimal condition, then the corresponding DNN model always achieves lower final training loss and higher test accuracy than the conventional He-normal initialization. Our work thus supplies a mathematically grounded indicator that guides the choice of initialization variance and clarifies its physical meaning of the dynamics of parameters in DNNs.

stat.ML↗

Design and architecture of the IBM Quantum Engine Compiler

In this work, we describe the design and architecture of the open-source Quantum Engine Compiler (qe-compiler) currently used in production for IBM Quantum systems. The qe-compiler is built using LLVM's Multi-Level Intermediate Representation (MLIR) framework and includes definitions for several dialects to represent parameterized quantum computation at multiple levels of abstraction. The compiler also provides Python bindings and a diagnostic system. An open-source LALR lexer and parser built using Bison and Flex generates an Abstract Syntax Tree that is translated to a high-level MLIR dialect. An extensible hierarchical target system for modeling the heterogeneous nature of control systems at compilation time is included. Target-based and generic compilation passes are added using a pipeline interface to translate the input down to low-level intermediate representations (including LLVM IR) and can take advantage of LLVM backends and tooling to generate machine executable binaries. The qe-compiler is built to be extensible, maintainable, performant, and scalable to support the future of quantum computing.

quant-ph↗

Chiral symmetry restoration at high matter density observed in pionic atoms

Modern theories of physics tell that the vacuum is not an empty space. Hidden in the vacuum is a structure of anti-quarks $\bar{q}$ and quarks $q$. The $\bar{q}$ and $q$ pair has the same quantum number as the vacuum and condensates in it since the strong interaction of the quantum chromodynamics (QCD) is too strong to leave it empty. The $\bar{q}q$ condensation breaks the chiral symmetry of the vacuum. The expectation value $<\bar{q}q>$ is an order parameter. For higher temperature or higher matter-density, $|<\bar{q}q>|$ decreases reflecting the restoration of the symmetry. In contrast to these clear-cut arguments, experimental evidence is so far limited. First of all, the $\bar{q}q$ is nothing but the vacuum itself. It is neither visible nor perceptible. In this article, we unravel this invisible existence by high precision measurement of pionic atoms, $π^-$-meson-nucleus bound systems. Using the $π^-$ as a probe, we demonstrate that $|<\bar{q}q>|$ is reduced in the nucleus at 58% of the normal nuclear density by a factor of 77 $\pm$ 2% compared with that in the vacuum. This reduction indicates that the chiral symmetry is partially restored due to the extremely high density of the nucleus. The present experimental result clearly exhibits the existence of the hidden structure, the chiral condensate, in the vacuum.

nucl-ex↗

Efficient techniques to GPU Accelerations of Multi-Shot Quantum Computing Simulations

Quantum computers are becoming practical for computing numerous applications. However, simulating quantum computing on classical computers is still demanding yet useful because current quantum computers are limited because of computer resources, hardware limits, instability, and noises. Improving quantum computing simulation performance in classical computers will contribute to the development of quantum computers and their algorithms. Quantum computing simulations on classical computers require long performance times, especially for quantum circuits with a large number of qubits or when simulating a large number of shots for noise simulations or circuits with intermediate measures. Graphical processing units (GPU) are suitable to accelerate quantum computer simulations by exploiting their computational power and high bandwidth memory and they have a large advantage in simulating relatively larger qubits circuits. However, GPUs are inefficient at simulating multi-shots runs with noises because the randomness prevents highly parallelization. In addition, GPUs have a disadvantage in simulating circuits with a small number of qubits because of the large overheads in GPU kernel execution. In this paper, we introduce optimization techniques for multi-shot simulations on GPUs. We gather multiple shots of simulations into a single GPU kernel execution to reduce overheads by scheduling randomness caused by noises. In addition, we introduce shot-branching that reduces calculations and memory usage for multi-shot simulations. By using these techniques, we speed up x10 from previous implementations.

quant-ph↗

Anomalous fluctuations of renewal-reward processes with heavy-tailed distributions

For renewal-reward processes with a power-law decaying waiting time distribution, anomalously large probabilities are assigned to atypical values of the asymptotic processes. Previous works have reveals that this anomalous scaling causes a singularity in the corresponding large deviation function. In order to further understand this problem, we study in this article the scaling of variance in several renewal-reward processes: counting processes with two different power-law decaying waiting time distributions and a Knudsen gas (a heat conduction model). Through analytical and numerical analyses of these models, we find that the variances show an anomalous scaling when the exponent of the power law is -3. For a counting process with the power-law exponent smaller than -3, this anomalous scaling does not take place: this indicates that the processes only fluctuate around the expectation with an error that is compatible with a standard large deviation scaling. In this case, we argue that anomalous scaling appears in higher order cumulants. Finally, many-body particles interacting through soft-core interactions with the boundary conditions employed in the Knudsen gas are studied using numerical simulations. We observe that the variance scaling becomes normal even though the power-law exponent in the boundary conditions is -3.

cond-mat.stat-mech↗

Large-time asymptotic of heavy tailed renewal processes

We study the large-time asymptotic of renewal-reward processes with a heavy-tailed waiting time distribution. It is known that the heavy tail of the distribution produces an extremely slow dynamics, resulting in a singular large deviation function. When the singularity takes place, the bottom of the large deviation function is flattened, manifesting anomalous fluctuations of the renewal-reward processes. In this article, we aim to study how these singularities emerge as the time increases. Using a classical result on the sum of random variables with regularly varying tail, we develop an expansion approach to prove an upper bound of the finite-time moment generating function for the Pareto waiting time distribution (power law) with an integer exponent. We perform numerical simulations using Pareto (with a real value exponent), inverse Rayleigh and log-normal waiting time distributions, and demonstrate similar results are anticipated in these waiting time distributions.

math-ph↗

Cache Blocking Technique to Large Scale Quantum Computing Simulation on Supercomputers

Classical computers require large memory resources and computational power to simulate quantum circuits with a large number of qubits. Even supercomputers that can store huge amounts of data face a scalability issue in regard to parallel quantum computing simulations because of the latency of data movements between distributed memory spaces. Here, we apply a cache blocking technique by inserting swap gates in quantum circuits to decrease data movements. We implemented this technique in the open source simulation framework Qiskit Aer. We evaluated our simulator on GPU clusters and observed good scalability.

quant-ph↗

Load Balancing for Skewed Streams on Heterogeneous Cluster

Streaming applications frequently encounter skewed workloads and execute on heterogeneous clusters. Optimal resource utilization in such adverse conditions becomes a challenge, as it requires inferring the resource capacities and input distribution at run time. In this paper, we tackle the aforementioned challenges by modeling them as a load balancing problem. We propose a novel partitioning strategy called Consistent Grouping (CG), which enables each processing element instance (PEI) to process the workload according to its capacity. The main idea behind CG is the notion of small, equal-sized virtual workers at the sources, which are assigned to workers based on their capacities. We provide a theoretical analysis of the proposed algorithm and show via extensive empirical evaluation that our proposed scheme outperforms the state-of-the-art approaches, like key grouping. In particular, CG achieves 3.44x better performance in terms of latency compared to key grouping.

cs.DC↗