Searcharxiv⌕ Search

arXiv subjects

Ulya R. Karpuzcu

Publications and source records attributed to Ulya R. Karpuzcu.

At least 19 recordsLinked to original sources

On the Limits of Ising Machines for Circuit-Derived SAT

Emerging Ising machines are promising for solving computationally hard optimization problems, yet their limits on structured, circuit-derived satisfiability (SAT) problems remain poorly understood. Using semiprime factorization as a representative benchmark, we show that tight constraints, when mapped into optimization form, fundamentally distort the energy landscape, and that these distortions are amplified when problems are decomposed to fit limited Ising machine capacity. To address this, we propose a hybrid flow that offloads Ising-harmful structure to lightweight preprocessing while reserving the genuinely hard search for the Ising machine. We further show that generic, circuit-structure-unaware decomposition is insufficient for circuit-derived instances, and that structure-aware partitioning is essential. These findings identify constraint handling as a central obstacle, highlighting hybrid hardware-software approaches as the path forward for scaling Ising machines to real-world SAT workloads. Evaluated on fabricated 45-spin all-to-all Ising chips, our flow extends solvable problem sizes from 8-bit (94 variables) to 11-bit (190 variables) without any hardware changes.

cs.ET↗

An FPGA-ASIC Co-Design Framework for Capacity-Constrained Physics-Based Ising Chips

When a problem exceeds an analog Ising machine's spin capacity it cannot be solved in one shot: a digital orchestration layer must iteratively decompose the graph, clamp boundary spins, and deliver hardware-sized subproblems to the solver. As per-subproblem solve times reach the us-scale regime, this digital layer (not the analog core) becomes the primary scalability bottleneck. On a 28 nm CMOS coupled-oscillator Ising chip (T_core ~ 77.5 us, 24 mW), a CPU orchestrator leaves the solver idle for over 98% of every iteration at N=750, and the gap survives OpenMP and AVX2 optimization because it is dominated by data-dependent memory access rather than arithmetic throughput. We argue that the orchestration layer is a first-class design object and propose a sizing methodology for hybrid analog-digital Ising systems: given a solver's core time, clock frequency, and target problem class, it gives an analytical baseline for hardware parallelism and memory-bandwidth dimensioning. Guided by the sizing laws, an FPGA orchestration layer co-located with the chip satisfies the pipeline condition with substantial margin. Across graph coloring, MaxCut, and SAT benchmarks it delivers 9.8x-46.5x raw end-to-end time-to-solution (TTS) improvements (14.8x-65.9x on per-repeat runtime) over the optimized CPU orchestrator, with device-level power reductions of 85x-116x.

cs.ET↗

HETRI: Heterogeneous Ising Multiprocessing

Ising machines are effective solvers for complex combinatorial optimization problems. The idea is mapping the optimal solution(s) to a combinatorial optimization problem to the minimum energy state(s) of a physical system, which naturally converges to a minimum energy state upon perturbance. The underlying mathematical abstraction, the Ising model, can capture the dynamic behavior of different physical systems by mapping each problem variable to a spin which can interact with other spins. Ising model as a mathematical abstraction can be mapped to hardware using traditional devices. In this paper we instead focus on Ising machines which represent a network of physical spins directly implemented in hardware using, e.g., quantum bits or electronic oscillators. To eliminate the scalability bottleneck due to the mismatch in problem vs. Ising machine size and connectivity, in this paper we make the case for HETRI: Heterogeneous Ising Multiprocessing. HETRI organizes the maximum number of physical spins that the underlying technology supports in Ising cores; and multiple independent Ising cores, in Ising chips. Ising cores in a chip feature different inter-spin connectivity or spin counts to match the problem characteristics. We provide a detailed design space exploration and quantify the performance in terms of time or energy to solution and solution accuracy with respect to homogeneous alternatives under the very same hardware budget and considering the very same spin technology.

cs.ET↗

Ising Acceleration for Multi-Robot Multi-Target Planning

Ising machines are emerging as promising hardware for combinatorial optimization. With recent advances in CMOS Ising technology, they are becoming attractive as low-power accelerator systems for robotics, where energy is limited and combinatorial optimization arises in multiple forms. However, a hardware-aware analysis of where such chips fit within a robotics planning stack is still missing. This paper studies the capabilities and limitations of CMOS Ising machines for low-power acceleration in multi-robot multi-target planning. We analyze three planning layers---target sharing, tour construction, and pathfinding---using real 45-spin all-to-all connected CMOS Ising chips as representative devices. We propose new Ising-based planning methods and a multi-mapping pipeline that uses spin merging, coefficient quantization, and spin-budget branching to adapt subproblems to spin- and coefficient-limited hardware. Our results show that the proposed recursive target-sharing method naturally matches the Ising hardware, achieving up to 8,000x lower energy than a classical baseline. End to end, the Ising pipeline produces routes within 9% of a strong classical baseline at 130x lower energy, showing that compact CMOS Ising machines can be effective in selected parts of the planning stack.

cs.ET↗

Breaking Local-Minimum Traps in Spiking Neural Network-Based Solvers for CSPs via Parallel Tempering

Spiking neural networks (SNNs) with stochastic neurons can solve constraint satisfaction problems (CSPs) by encoding constraints via connectivity and performing probabilistic search via spike dynamics. However, fixed-temperature stochastic dynamics often get trapped in local minima - near-satisfying configurations - a vulnerability that escalates with problem difficulty. To overcome this, we integrate parallel tempering (PT) into the neural sampling solver, running multiple parallel replicas at varying inverse temperatures. Replicas periodically exchange temperatures rather than network states, managing the trade-off between exploration and concentration around low-energy configurations while preserving asynchronous, spike-based computation. We evaluate this architecture against a parallel baseline of four independent, fixed-temperature solvers using equal computational resources across 1000 instances from the SATLIB uf20-91 benchmark. Parallel tempering improves success probability on 332 instances while worsening only 5. Crucially, these gains are concentrated on hard instances where independent solvers fail. Violation trajectory analysis confirms the underlying mechanism: temperature exchanges allow replicas to traverse energy barriers unreachable by fixed-temperature dynamics, successfully escaping the narrow basins that constrain the baseline. To our knowledge, this represents the first integration of parallel tempering into an SNN-based CSP solver.

cs.ET↗

Computing In Spintronic Memory: A Thermal Perspective

Computing-in-Memory (CiM) is a promising paradigm to address the memory bottleneck constraining traditional systems. Most power-efficient CiM variants can directly perform Boolean operations in non-volatile memory arrays. Higher microarchitectural activity due to CiM, however, can significantly increase power density (power per area) and result in thermal hotspots. In this paper, we provide a quantitative thermal characterization for CiM. We demonstrate that (i) the temperature remains mostly uniform due to lateral thermal conduction; (ii) the temperature increases linearly with the number of memory cells participating in computation; (iii) the temperature decreases linearly with the memory array size; (iv) the memory technology dictates the power density, hence the thermal characteristics.

cs.ET↗

Extractive summarization on a CMOS Ising machine

Extractive summarization (ES) aims to generate a concise summary by selecting a subset of sentences from a document while maximizing relevance and minimizing redundancy. Although modern ES systems achieve high accuracy using powerful neural models, their deployment typically relies on CPU or GPU infrastructures that are energy-intensive and poorly suited for real-time inference in resource-constrained environments. In this work, we explore the feasibility of implementing McDonald-style extractive summarization on a low-power CMOS coupled oscillator-based Ising machine (COBI) that supports integer-valued, all-to-all spin couplings. We first propose a hardware-aware Ising formulation that reduces the scale imbalance between local fields and coupling terms, thereby improving robustness to coefficient quantization: this method can be applied to any problem formulation that requires k of n variables to be chosen. We then develop a complete ES pipeline including (i) stochastic rounding and iterative refinement to compensate for precision loss, and (ii) a decomposition strategy that partitions a large ES problem into smaller Ising subproblems that can be efficiently solved on COBI and later combined. Experimental results on the CNN/DailyMail dataset show that our pipeline can produce high-quality summaries using only integer-coupled Ising hardware with limited precision. COBI achieves 3-4.5x runtime speedups compared to a brute-force method, which is comparable to software Tabu search, and two to three orders of magnitude reductions in energy, while maintaining competitive summary quality. These results highlight the potential of deploying CMOS Ising solvers for real-time, low-energy text summarization on edge devices.

cs.LG↗

Supporting Higher-Order Interactions in Practical Ising Machines

Ising machines as hardware solvers of combinatorial optimization problems (COPs) can efficiently explore large solution spaces due to their inherent parallelism and physics-based dynamics. Many important COP classes such as satisfiability (SAT) assume arbitrary interactions between problem variables, while most Ising machines only support pairwise (second-order) interactions. This necessitates translation of higher-order interactions to pair-wise, which typically results in extra variables not corresponding to problem variables, and a larger problem for the Ising machine to solve than the original problem. This in turn can significantly increase time-to-solution and/or degrade solution accuracy. In this paper, considering a representative CMOS-compatible class of Ising machines, we propose a practical design to enable direct hardware support for higher order interactions. By minimizing the overhead of problem translation and mapping, our design leads to up to 4x lower time-to-solution without compromising solution accuracy.

physics.comp-ph↗

DROID: Discrete-Time Simulation for Ring-Oscillator-Based Ising Design

Many combinatorial problems can be mapped to Ising machines, i.e., networks of coupled oscillators that settle to a minimum-energy ground state, from which the problem solution is inferred. This work proposes DROID, a novel event-driven method for simulating the evolution of a CMOS Ising machine to its ground state. The approach is accurate under general delay-phase relations that include the effects of the transistor nonlinearities and is computationally efficient. On a realistic-size all-to-all coupled ring oscillator array, DROID is nearly four orders of magnitude faster than a traditional HSPICE simulation in predicting the evolution of a coupled oscillator system and is demonstrated to attain a similar distribution of solutions as the hardware.

cs.ET↗

On Error Correction for Nonvolatile Processing-In-Memory

Processing in memory (PiM) represents a promising computing paradigm to enhance performance of numerous data-intensive applications. Variants performing computing directly in emerging nonvolatile memories can deliver very high energy efficiency. PiM architectures directly inherit the vulnerabilities of the underlying memory substrates, but they also are subject to errors due to the computation in place. Numerous well-established error correcting codes (ECC) for memory exist, and are also considered in the PiM context, however, they typically ignore errors that occur throughout computation. In this paper we revisit the error correction design space for nonvolatile PiM, considering both storage/memory and computation-induced errors, surveying several self-checking and homomorphic approaches. We propose several solutions and analyze their complex performance-area-coverage trade-off, using three representative nonvolatile PiM technologies. All of these solutions guarantee single error correction for both, bulk bitwise computations and ordinary memory/storage errors.

cs.ET↗

3SAT on an All-to-All-Connected CMOS Ising Solver Chip

This work solves 3SAT, a classical NP-complete problem, on a CMOS-based Ising hardware chip with all-to-all connectivity. The paper addresses practical issues in going from algorithms to hardware. It considers several degrees of freedom in mapping the 3SAT problem to the chip - using multiple Ising formulations for 3SAT; exploring multiple strategies for decomposing large problems into subproblems that can be accommodated on the Ising chip; and executing a sequence of these subproblems on CMOS hardware to obtain the solution to the larger problem. These are evaluated within a software framework, and the results are used to identify the most promising formulations and decomposition techniques. These best approaches are then mapped to the all-to-all hardware, and the performance of 3SAT is evaluated on the chip. Experimental data shows that the deployed decomposition and mapping strategies impact SAT solution quality: without our methods, the CMOS hardware cannot achieve 3SAT solutions on SATLIB benchmarks.

cs.ET↗

Effectiveness of Variable Distance Quantum Error Correcting Codes

Quantum error correction is capable of digitizing quantum noise and increasing the robustness of qubits. Typically, error correction is designed with the target of eliminating all errors - making an error so unlikely it can be assumed that none occur. In this work, we use statistical quantum fault injection on the quantum phase estimation algorithm to test the sensitivity to quantum noise events. Our work suggests that quantum programs can tolerate non-trivial errors and still produce usable output. We show that it may be possible to reduce error correction overhead by relaxing tolerable error rate requirements. In addition, we propose using variable strength (distance) error correction, where overhead can be reduced by only protecting more sensitive parts of the quantum program with high distance codes.

quant-ph↗

Towards Homomorphic Inference Beyond the Edge

Beyond edge devices can function off the power grid and without batteries, enabling them to operate in difficult to access regions. However, energy costly long-distance communication required for reporting results or offloading computation becomes a limitation. Here, we reduce this overhead by developing a beyond edge device which can effectively act as a nearby server to offload computation. For security reasons, this device must operate on encrypted data, which incurs a high overhead. We use energy-efficient and intermittent-safe in-memory computation to enable this encrypted computation, allowing it to provide a speedup for beyond edge applications within a power budget of a few milliWatts.

cs.CR↗

Exploring the Feasibility of Using 3D XPoint as an In-Memory Computing Accelerator

This paper describes how 3D XPoint memory arrays can be used as in-memory computing accelerators. We first show that thresholded matrix-vector multiplication (TMVM), the fundamental computational kernel in many applications including machine learning, can be implemented within a 3D XPoint array without requiring data to leave the array for processing. Using the implementation of TMVM, we then discuss the implementation of a binary neural inference engine. We discuss the application of the core concept to address issues such as system scalability, where we connect multiple 3D XPoint arrays, and power integrity, where we analyze the parasitic effects of metal lines on noise margins. To assure power integrity within the 3D XPoint array during this implementation, we carefully analyze the parasitic effects of metal lines on the accuracy of the implementations. We quantify the impact of parasitics on limiting the size and configuration of a 3D XPoint array, and estimate the maximum acceptable size of a 3D XPoint subarray.

cs.AR↗

Benchmarking Quantum Computers and the Impact of Quantum Noise

Benchmarking is how the performance of a computing system is determined. Surprisingly, even for classical computers this is not a straightforward process. One must choose the appropriate benchmark and metrics to extract meaningful results. Different benchmarks test the system in different ways and each individual metric may or may not be of interest. Choosing the appropriate approach is tricky. The situation is even more open ended for quantum computers, where there is a wider range of hardware, fewer established guidelines, and additional complicating factors. Notably, quantum noise significantly impacts performance and is difficult to model accurately. Here, we discuss benchmarking of quantum computers from a computer architecture perspective and provide numerical simulations highlighting challenges which suggest caution.

quant-ph↗

On Value Recomputation to Accelerate Invisible Speculation

Recent architectural approaches that address speculative side-channel attacks aim to prevent software from exposing the microarchitectural state changes of transient execution. The Delay-on-Miss technique is one such approach, which simply delays loads that miss in the L1 cache until they become non-speculative, resulting in no transient changes in the memory hierarchy. However, this costs performance, prompting the use of value prediction (VP) to regain some of the delay. However, the problem cannot be solved by simply introducing a new kind of speculation (value prediction). Value-predicted loads have to be validated, which cannot be commenced until the load becomes non-speculative. Thus, value-predicted loads occupy the same amount of precious core resources (e.g., reorder buffer entries) as Delay-on-Miss. The end result is that VP only yields marginal benefits over Delay-on-Miss. In this paper, our insight is that we can achieve the same goal as VP (increasing performance by providing the value of loads that miss) without incurring its negative side-effect (delaying the release of precious resources), if we can safely, non-speculatively, recompute a value in isolation (without being seen from the outside), so that we do not expose any information by transferring such a value via the memory hierarchy. Value Recomputation, which trades computation for data transfer was previously proposed in an entirely different context: to reduce energy-expensive data transfers in the memory hierarchy. In this paper, we demonstrate the potential of value recomputation in relation to the Delay-on-Miss approach of hiding speculation, discuss the trade-offs, and show that we can achieve the same level of security, reaching 93% of the unsecured baseline performance (5% higher than Delay-on-miss), and exceeding (by 3%) what even an oracular (100% accuracy and coverage) value predictor could do.

cs.AR↗

Read Mapping Near Non-Volatile Memory

DNA sequencing is the physical/biochemical process of identifying the location of the four bases (Adenine, Guanine, Cytosine, Thymine) in a DNA strand. As semiconductor technology revolutionized computing, modern DNA sequencing technology (termed Next Generation Sequencing, NGS)revolutionized genomic research. As a result, modern NGS platforms can sequence hundreds of millions of short DNA fragments in parallel. The sequenced DNA fragments, representing the output of NGS platforms, are termed reads. Besides genomic variations, NGS imperfections induce noise in reads. Mapping each read to (the most similar portion of) a reference genome of the same species, i.e., read mapping, is a common critical first step in a diverse set of emerging bioinformatics applications. Mapping represents a search-heavy memory-intensive similarity matching problem, therefore, can greatly benefit from near-memory processing. Intuition suggests using fast associative search enabled by Ternary Content Addressable Memory (TCAM) by construction. However, the excessive energy consumption and lack of support for similarity matching (under NGS and genomic variation induced noise) renders direct application of TCAM infeasible, irrespective of volatility, where only non-volatile TCAM can accommodate the large memory footprint in an area-efficient way. This paper introduces GeNVoM, a scalable, energy-efficient and high-throughput solution. Instead of optimizing an algorithm developed for general-purpose computers or GPUs, GeNVoM rethinks the algorithm and non-volatile TCAM-based accelerator design together from the ground up. Thereby GeNVoM can improve the throughput by up to 113.5 times (3.6); the energy consumption, by up to 210.9 times (1.36), when compared to a GPU (accelerator) baseline, which represents one of the highest-throughput implementations known.

cs.DC↗

Quantum Computing: An Overview Across the System Stack

Quantum computers, if fully realized, promise to be a revolutionary technology. As a result, quantum computing has become one of the hottest areas of research in the last few years. Much effort is being applied at all levels of the system stack, from the creation of quantum algorithms to the development of hardware devices. The quantum age appears to be arriving sooner rather than later as commercially useful small-to-medium sized machines have already been built. However, full-scale quantum computers, and the full-scale algorithms they would perform, remain out of reach for now. It is currently uncertain how the first such computer will be built. Many different technologies are competing to be the first scalable quantum computer.

quant-ph↗