SearcharxivSearch

arXiv subjects

Abhinav Anand

Publications and source records attributed to Abhinav Anand.

At least 19 recordsLinked to original sources

A Heterogeneous Distributed Architecture for Quantum Simulation

Architectural specialization and distribution can help scale fault-tolerant quantum computers, but may also introduce substantial overheads from communication, routing, and resource duplication. We introduce a heterogeneous distributed architecture in which a magic core is connected to an extensible storage system composed of one-dimensional lanes of specialized cold-storage nodes. The storage system supports parallel random access to Pauli string parities. This organization is particularly well suited to fermionic quantum simulation, enabling parallel execution of the highly non-local Pauli strings arising from these systems. We evaluate the architecture on fault-tolerant simulations of the dynamics of the Fermi-Hubbard and sparse Sachdev-Ye-Kitaev (SYK) models on systems of up to 450 logical qubits. These workloads exhibit complementary communication structures: Fermi-Hubbard produces a spectrum of interactions from local to non-local shaped by lattice geometry, whereas sparse SYK produces highly non-local and overlapping Pauli operators. For a Trotter step of a 450-logical-qubit Fermi-Hubbard workload, a six-lane system with 30 T-state factories is within approximately $1.4\times$ the wall-clock time of a homogeneous distributed architecture with 4 times as many T-state factories and substantially greater connectivity and sites for injecting magic. For matched T-factory counts, our architecture is $\sim 2\times$ faster.

quant-ph

Performance, Efficiency and Collapse -- Advantages and Challenges in Offline Post-training of Code LLMs

Post-training with reinforcement learning (RL) is a critical phase in the development of code-generating large language models (LLMs), as it ensures adherence to instructions and the production of functionally correct code. This process typically requires computationally intensive code sample generation from Transformer-based LLMs and substantial GPU-CPU communication for sequence verification. To address these computational challenges, this work examines whether RL-based post-training can be performed entirely offline by leveraging existing datasets rather than generating new samples. The findings indicate that, with only a few hours of training, zero-shot code generation performance of LLMs can be substantially improved without online sampling. Additionally, offline RL produces performance gains across models ranging from 0.5B to 7B parameters, although the extent of improvement varies among model families.

cs.LG

Towards Understanding What State Space Models Learn About Code

State Space Models (SSMs) have emerged as an efficient alternative to the Transformer architecture. Prior work shows that, when trained under comparable conditions, SSMs can match or surpass Transformers on code understanding tasks. However, their internal mechanisms remain a black box. We present the first systematic analysis of what SSM-based code models learn along with the direct comparison between SSM and Transformer models in this domain. Our analysis shows that SSMs capture syntactic and semantic structure more effectively than Transformers during pretraining but forgets certain relations during fine-tuning on some tasks. To investigate this behavior, we introduce SSM-Interpret, a frequency-domain framework that exposes a spectral shift toward short-range dependencies during fine-tuning. Guided by these findings, we propose architectural modifications that significantly improve the performance of SSM-based code model by upto +6 MRR on NLCodeSearch. This demonstrates that our analysis not only explains model behavior but also leads directly to better designs.

cs.AI

Efficient Post-training of LLMs for Code Generation With Offline Reinforcement Learning

Post-training using online reinforcement learning (RL) is an important training step for LLMs, including code-generating models. However, online RL for code generation involves LLM inference and verification of the generated output, which can take considerable time and resources. In this paper, we explore the application of offline RL to code-generating models by leveraging existing code datasets. Our experiments demonstrate that offline RL is an effective training strategy for improving LLM performance. We show that offline RL can be especially beneficial for small LLMs and challenging coding problems.

cs.AI

Domain-Adaptable Reinforcement Learning for Code Generation with Dense Rewards

Large language models show strong potential for automated code generation, but lack guarantees for correctness, quality, safety, and domain-specific constraints. For instance in robotics, where code generation is increasingly being used for planning and executing actions, awareness of the environment and physical constraints is critical. To facilitate the adaption of code-generating LLMs to diverse requirements, including domain-specific ones, we present a reinforcement learning framework that fine-tunes pre-trained LLMs using proximal policy optimization. Our customizable execution-aware reward formula captures and optimizes syntax, functional correctness, code style, security, and simulator executability. A token-level reward mapping mechanism enables effective credit assignment from execution outcomes to generated tokens. The framework is evaluated on general-purpose code generation (MBPP/MBPP+) and robotic program synthesis (RoboEval). The results show substantial improvements in functional correctness and simulator executability, including an absolute pass@1 increase of 19% on MBPP and a reduction in execution failures by 51% on RoboEval. These findings demonstrate that structured reinforcement learning can effectively align language models to correct program generation and domain-specific requirements.

cs.LG

INJEQT: Improved Magic-State Injection Protocol for Fault-Tolerant Quantum Extractor Architectures

Near-term FTQC system designs are constrained by limited error budgets and largely sequential execution of non-Clifford gates. As a result, reducing the number of the most-error prone instructions becomes critical for successful program execution. In this work, we study the extractor architecture, a recently proposed FTQC design that enables universal quantum computation on spatially-efficient QEC codes such as the BB code family. In these architectures, over $90\%$ of the total program error arises from the synthillation process, which involves $\lvert T\rangle$-state preparation and injection to implement non-Clifford gates. We observe that standard Rz synthillation requires multiple sequential $\lvert T\rangle$-state injections, each incurring an inter-module measurements, the most expensive instruction in the architecture, which cumulatively dominate the overall error budget. To address this bottleneck, we propose INJEQT, a $2$-factory design that uses an auxiliary code capable of synthesizing $Rz(θ)$ states with lower error rates. These states are then injected into the extractor modules using only a constant number of inter-module measurements. This approach reduces overall error rates by up to $22\times$. We further reduce the time overhead by a pre-fetching strategy that prepares the Rz states and their correction states in parallel. This approach improves the wall-clock time by up to $13\times$ and reduces the space-time cost by up to $7.2\times$, for an optimal choice of the number of INJEQT factories for each metric. We evaluate INJEQT for multiple state preparation techniques such as distillation, cultivation and STAR, and model the execution times for both lattice surgery-based and transversal CNOT based injections. Our results demonstrate that INJEQT is robust across factory choices and device technologies, enabling more efficient architectural designs for FTQC.

quant-ph

Architecting Early Fault Tolerant Neutral Atoms Systems with Quantum Advantage

Recent advancements in neutral atom platforms have enabled exploration of early fault-tolerant (FT) architectures for applications with quantum advantage, such as quantum dynamics simulations. An efficient fault-tolerant architecture has both spatially efficient quantum error correction codes (low qubit overhead), and efficient methodologies (transversal based gates, extractor based gates, etc.) for logical computation, to minimize overall execution time. Achieving the right balance between space and time can be critical for enabling early FT demonstrations of quantum advantage. In this work, we identify bottlenecks in existing spatially efficient schemes, which tend to be very serial, and do not take advantage of unutilized space. We introduce a teleportation-based scheme that leverages the reconfigurable connectivity of neutral atoms to parallelize logical operations. Our approach achieves up to \textbf{$\mathbf{\sim 3 \times}$ speedup} over extractor architectures at no extra space cost and achieves the best spacetime performance among other viable architectures before accounting for external \textit{resource-states}. To rigorously evaluate performance, we construct explicit quantum advantage benchmarks and \textit{simulate} compilation to a fault-tolerant instruction set, including low-level gate scheduling and shuttling patterns, and resource-state nondeterminism. We find that our speedups still apply and report exact space-time cost along with success probabilities, identifying architectures capable of achieving quantum advantage \textbf{with as little as $\mathbf{11,495}$ atoms and a runtime of $\mathbf{\sim 15}$ hours}.

quant-ph

Optimizing Logical Mappings for Quantum Low-Density Parity Check Codes

Early demonstrations of fault tolerant quantum systems have paved the way for logical-level compilation. For fault-tolerant applications to succeed, execution must finish with a low total program error rate (i.e., a low program failure rate). In this work, we study a promising candidate for future fault-tolerant architectures with low spatial overhead: the Gross code. Compilation for the Gross code entails compiling to Pauli Based Computation and then reducing the rotations and measurements to the Bicycle ISA. Depending on the configuration of modules and the placement of code modules on hardware, one can reduce the amount of resulting Bicycle instructions to produce a lower overall error rate. We find that NISQ-based, and existing FTQC mappers are insufficient for mapping logical qubits on Gross code architectures because 1. they do not account for the two-level nature of the logical qubit mapping problem, which separates into code modules with distinct measurements, and 2. they naively account only for length two interactions, whereas Pauli-Products are up to length $n$, where $n$ is the number of logical qubits in the circuit. For these reasons, we introduce a two-stage pipeline that first uses hypergraph partitioning to create in-module clusters, and then executes a priority-based algorithm to efficiently assign clusters onto hardware. We find that our mapping policy reduces the error contribution from inter-module measurements, the largest source of error in the Gross Code, by up to $\sim36\%$ in the best case, with an average reduction of $\sim13\%$. On average, we reduce the failure rates from inter-module measurements by $\sim22\%$ with localized factory availability, and by $\sim17\%$ on grid architectures, allowing hardware developers to be less constrained in developing scalable fault tolerant systems due to software driven reductions in program failure rates.

quant-ph

Analysis of Long Range Dependency Understanding in State Space Models

Although state-space models (SSMs) have demonstrated strong performance on long-sequence benchmarks, most research has emphasized predictive accuracy rather than interpretability. In this work, we present the first systematic kernel interpretability study of the diagonalized state-space model (S4D) trained on a real-world task (vulnerability detection in source code). Through time and frequency domain analysis of the S4D kernel, we show that the long-range modeling capability of S4D varies significantly under different model architectures, affecting model performance. For instance, we show that the depending on the architecture, S4D kernel can behave as low-pass, band-pass or high-pass filter. The insights from our analysis can guide future work in designing better S4D-based models.

cs.LG

Unbiased observable estimation with approximate channels in fault-tolerant quantum computation

Unitary errors, such as those arising from fault-tolerant compilation of quantum algorithms, systematically bias observable estimates. Correcting this bias typically requires additional resources, such as an increased number of non-Clifford gates. In this work, we present an alternative method for correcting bias in the expectation values of observables. The method leverages a decomposition of the ideal quantum channel into a probabilistic mixture of noisy quantum channels. Using this decomposition, we construct unbiased estimators as weighted sums of expectation values obtained from the noisy channels. We provide a detailed analysis of the method, identify the conditions under which it is effective, and validate its performance through numerical simulations. In particular, we demonstrate unbiased observable estimation in the presence of unitary errors by simulating the time dynamics of the Ising Hamiltonian. Our strategy offers a resource-efficient way to reduce the impact of unitary errors, improving methods for estimating observables in noisy near-term quantum devices and fault-tolerant implementation of quantum algorithms.

quant-ph

Dynamic local single-shot checks for toric codes

Quantum error correction typically requires repeated syndrome extraction due to measurement noise, which results in substantial time overhead in fault-tolerant computation. Single-shot error correction aims to suppress errors using only one round of syndrome extraction. However, for most codes, it requires high-weight checks, which significantly degrade, and often eliminate, single-shot performance at the circuit level. In this work, we introduce local single-shot checks, where we impose constraints on check weights. Using a dynamic measurement scheme, we show that the number of required measurement rounds can be reduced by a factor determined by this constraint. As an example, we show through numerical simulation that our scheme can improve decoding performance compared to conventional checks when using sliding-window decoding with a reduced window size under circuit-level noise models for toric codes. Our work provides a new direction for constructing checks that can reduce time overhead in large-scale fault-tolerant quantum computation.

quant-ph

Cyclone: Designing Efficient and Highly Parallel QCCD Architectural Codesigns for Fault Tolerant Quantum Memory

Modular trapped-ion quantum computing hardware, known as QCCDs require shuttling operations in order to maintain effective all-to-all connectivity. Each module or trap can perform only one operation at a time, resulting in low intra-trap parallelism, but there is no restriction on operations happening on independent traps, enabling high inter-trap parallelism. Unlike their superconducting counterparts, the design space for QCCDs is relatively flexible and can be explored beyond current grid designs. In particular, current grid-based architectures significantly limit the performance of many promising, high-rate codes such as HGP codes and BB codes, suffering from numerous trap to trap ``roadblocks", forcing serialization and destroying the inherent parallelism of these codes.. Many of these codes are highly parallelizable, meaning that with appropriate hardware layouts and matching software schedules, execution latency can be reduced. Faster execution, in turn, reduces error accumulation from decoherence and heating, ultimately improving code performance when mapped to realistic hardware. To address this, we propose Cyclone, a circular software-hardware codesign that departs from traditional 2D grids in favor of a flexible ring topology, where ancilla qubits move in lockstep. Cyclone eliminates roadblocks, bounds total movement, and enables high levels of parallelism, resulting in up to ~4$\times$ speedup in execution times. With HGP codes, Cyclone achieves up to a 2$\times$ order of magnitude improvement in logical error rate, and with BB codes, this improvement reaches up to a 3$\times$ in order of magnitude.Spatially, Cyclone reduces the number of required traps and ancilla qubits by $2\times$.The overall spacetime improvement over a standard grid is up to $\sim 20 \times$, demonstrating Cyclone as a scalable and efficient alternative to conventional 2D QCCD architectures.

quant-ph

CodeSSM: Towards State Space Models for Code Understanding

Although transformers dominate many code-specific tasks, they have significant limitations. This paper explores State Space Models (SSMs) as a promising alternative for code understanding tasks such as retrieval, classification, and clone detection. We introduce CodeSSM, the first SSM-based model trained on code corpora to assess its effectiveness. Our results demonstrate that SSMs are more sample-efficient and can extrapolate to longer contexts beyond the pretraining length. Extensive experiments show that SSMs offer a viable alternative to transformers, addressing several their limitations. Additionally, CodeSSM reduces memory usage by up to 64\% compared to transformers at a context length of 2048, with greater savings as context length grows.

cs.SE

Stabilizer configuration interaction: Finding molecular subspaces with error detection properties

In this work, we explore a new approach to designing both algorithms and error detection codes for preparing approximate ground states of molecules. We propose a classical algorithm to find the optimal stabilizer state by using excitations of the Hartree-Fock state, followed by constructing quantum error-detection codes based on this stabilizer state using codeword-stabilized codes. Through various numerical experiments, we confirm that our method finds the best stabilizer approximations to the true ground states of molecules up to 36 qubits in size. Additionally, we construct generalized stabilizer states that offer a better approximation to the true ground states. Furthermore, for a simple noise model, we demonstrate that both the stabilizer and (some) generalized stabilizer states can be prepared with higher fidelity using the error-detection codes we construct. Our work represents a promising step toward designing algorithms for early fault-tolerant quantum computation.

quant-ph

Leveraging commuting groups for an efficient variational Hamiltonian ansatz

Efficiently calculating the low-lying eigenvalues of Hamiltonians, written as sums of Pauli operators, is a fundamental challenge in quantum computing. While various methods have been proposed to reduce the complexity of quantum circuits for this task, there remains room for further improvement. In this article, we introduce a new circuit design using commuting groups within the Hamiltonian to further reduce the circuit complexity of Hamiltonian-based quantum circuits. Our approach involves partitioning the Pauli operators into mutually commuting clusters and finding Clifford unitaries that diagonalize each cluster. We then design an ansatz that uses these Clifford unitaries for efficient switching between the clusters, complemented by a layer of parameterized single qubit rotations for each individual cluster. By conducting numerical simulations, we demonstrate the effectiveness of our method in accurately determining the ground state energy of different quantum chemistry Hamiltonians. Our results highlight the applicability and potential of our approach for designing problem-inspired ansatz for various quantum computing applications.

quant-ph

Integrating Symbolic Execution into the Fine-Tuning of Code-Generating LLMs

Code-generating Large Language Models (LLMs) have become essential tools in modern software development, enhancing productivity and accelerating development. This paper aims to investigate the fine-tuning of code-generating LLMs using Reinforcement Learning and Direct Preference Optimization, further improving their performance. To achieve this, we enhance the training data for the reward model with the help of symbolic execution techniques, ensuring more comprehensive and objective data. With symbolic execution, we create a custom dataset that better captures the nuances in code evaluation. Our reward models, fine-tuned on this dataset, demonstrate significant improvements over the baseline, CodeRL, in estimating the quality of generated code. Our code-generating LLMs, trained with the help of reward model feedback, achieve similar results compared to the CodeRL benchmark.

cs.SE

Information flow in parameterized quantum circuits

In this work, we introduce a new way to quantify information flow in quantum systems, especially for parameterized quantum circuits. We use a graph representation of the circuits and propose a new distance metric using the mutual information between gate nodes. We then present an optimization procedure for variational algorithms using paths based on the distance measure. We explore the features of the algorithm by means of the variational quantum eigensolver, in which we compute the ground state energies of the Heisenberg model. In addition, we employ the method to solve a binary classification problem using variational quantum classification. From numerical simulations, we show that our method can be successfully used for optimizing the parameterized quantum circuits primarily used in near-term algorithms. We further note that information-flow based paths can be used to improve convergence of existing stochastic gradient based methods.

quant-ph

Hamiltonian-based graph-state ansatz for variational quantum algorithms

One promising application of near-term quantum devices is to prepare trial wavefunctions using short circuits for solving different problems via variational algorithms. For this purpose, we introduce a new circuit design that combines graph-based diagonalization circuits with arbitrary single-qubit rotation gates to get Hamiltonian-based graph states ansätze (H-GSA). We test the accuracy of the proposed ansatz in estimating ground state energies of various molecules of size up to 12-qubits. Additionally, we compare the gate count and parameter number complexity of the proposed ansatz against previously proposed schemes and find an order magnitude reduction in gate count complexity with slight increase in the number of parameters. Our work represents a significant step towards constructing compact quantum circuits with good trainability and convergence properties and applications in solving chemistry and physics problems.

quant-ph