SearcharxivSearch

arXiv subjects

Sven Groppe

Publications and source records attributed to Sven Groppe.

18 recordsLinked to original sources

Hardware-in-the-Loop Syndrome-to-Decoder Validation for Repetition, Surface, CSS-LDPC, and Digitized-GKP Codes

Quantum error-correction experiments increasingly require a verified interface between measured syndrome bits and decoder-native correction requests. We report a four-branch syndrome-to-decoder study spanning three IBM gate-model hardware circuits and one PennyLane-backed digitized-GKP model. The hardware branches implement a five-data-qubit repetition code, a distance-five rotated-surface-code Z-check extraction layer, and the Z-check half of the Steane CSS code as a compact CSS-LDPC benchmark. The GKP branch samples finite-squeezed Gaussian-CV q-readout and injected q-shifts, then bins wrapped quadrature coordinates into the same outer surface-code Z-check interface. All cases use 4096 shots per stream, clean and injected streams, LiDMaS+ request construction, and MWPM/minimum-weight correction as the plotted baseline, with union-find and hard-decision belief-propagation/min-sum policies replayed for interface validation. The correction-volume panels additionally report mean minimum-weight correction weight for each decoded stream. Repetition and CSS-LDPC hardware preserve the dominant expected syndrome and correction for every injected target. The routed 56-qubit surface circuit exhibits broad hardware-induced syndrome activation: exact localization drops to $0.003$--$0.108$, but target-containing localization remains $0.279$--$0.642$. The digitized-GKP study gives exact q-shift localization of $0.350$--$0.495$ and target-containing localization of $0.417$--$0.608$. The results support an auditable syndrome-to-decoder interface rather than a threshold claim.

quant-ph

Improved Join Order Optimization for Database Queries using Hybrid Quantum-Classical Approaches for QUBO Problems

Efficient query optimization is crucial for relational database systems, especially for optimizing join orders in complex queries. This work introduces a hybrid approach that integrates Eliminating Cartesian Products (ECP) with splitting the QUBO search space (SQSS) to reduce the size of the QUBO problem, minimizing binary variables and constraints. This improves the performance of the quantum algorithm while lowering hardware requirements. We evaluate our method using real-world SQL queries from the ErgastF1 dataset on quantum and classical algorithms, including Quantum Annealing (QA), Simulated Annealing (SA), QAOA, and VQE, implemented on D-Wave's Quantum Annealer and universal gate-based simulators. Additionally, we analyze the impact of selectivity and SQSS on QUBO weight distribution and algorithmic performance, highlighting optimization efficiency for QA and SA. Experimental results show consistent optimal join orders and enhanced query optimization for various selectivity conditions, and they also highlight the limitations of current quantum hardware for complex queries. This study further confirms the potential of hybrid quantum-classical methods for scalable quantum-enhanced database optimization.

cs.DB

Efficient Batch Search Algorithm for B+ Tree Index Structures with Level-Wise Traversal on FPGAs

This paper introduces a search algorithm for index structures based on a B+ tree, specifically optimized for execution on a field-programmable gate array (FPGA). Our implementation efficiently traverses and reuses tree nodes by processing a batch of search keys level by level. This approach reduces costly global memory accesses, improves reuse of loaded B+ tree nodes, and enables parallel search key comparisons directly on the FPGA. Using a high-level synthesis (HLS) approach, we developed a highly flexible and configurable search kernel design supporting variable batch sizes, customizable node sizes, and arbitrary tree depths. The final design was implemented on an AMD Alveo U250 Data Center Accelerator Card, and was evaluated against the B+ tree search algorithm from the TLX library running on an AMD EPYC 7542 processor (2.9 GHz). With a batch size of 1000 search keys, a B+ tree containing one million entries, and a tree order of 16, we measured a 4.9x speedup for the single-kernel FPGA design compared to a single-threaded CPU implementation. Running four kernel instances in parallel on the FPGA resulted in a 2.1$\times$ performance improvement over a CPU implementation using 16 threads.

cs.AR

A Unified Hardware-to-Decoder Architecture for Hybrid Continuous-Variable and Discrete-Variable Quantum Error Correction in LiDMaS+

We present an architecture-level hardware-to-logical-to-decoder execution stack for hybrid continuous-variable and discrete-variable quantum error correction in LiDMaS+. Provider-native records are normalized into a single decoder IO contract and replayed under fixed controls across MWPM, UF, BP, and neural-MWPM. In a Xanadu case study using fixture inputs and sampled public datasets, replay integrity was complete: 108/108 fixture and 4000/4000 real-slice request-response lines, with zero request-parse errors, zero response-parse errors, and zero decoder-name mismatches. Under matched inputs, decoder behavior is clearly regime-dependent. For weighted fixture summaries, average flip count was 1.296 (MWPM), 1.296 (UF), 0.667 (BP), and 1.296 (neural-MWPM). For weighted real-data summaries, average flip count was 0.641 (MWPM), 0.741 (UF), 0.318 (BP), and 0.641 (neural-MWPM); corresponding nonempty-flip rates were 0.490, 0.490, 0.318, and 0.490. Across fixture data, BP reduced weighted correction volume by 48.6\% versus MWPM; across real slices, BP reduced weighted correction volume by 50.4\% versus MWPM and 57.1\% versus UF. Quality controls show the central interpretability tradeoff: BP is intervention-conservative but leaves higher residual burden, while MWPM-family decoders intervene more aggressively and clear more syndrome. Warning-no-syndrome rates remained decoder-invariant and dataset-driven (fixture weighted 0.259; real weighted 0.510), confirming preserved sparsity semantics from hardware input to logical correction. Re-running analysis stages reproduced identical SHA-256 artifacts, enabling deterministic study iteration. These results establish a practical benchmarking foundation for photonic GKP-oriented hardware programs where decoder policy must be selected as a function of operating regime.

quant-ph

Decoder Dependence in Surface-Code Threshold Estimation with Native Gottesman-Kitaev-Preskill Digitization and Parallelized Sampling

We quantify decoder dependence in surface-code threshold studies under two matched regimes: Pauli noise and native GKP-style Gaussian displacement digitization. Using LiDMaS+ v1.1.0, we benchmark MWPM, Union-Find (UF), Belief Propagation (BP), and neural-guided MWPM with fixed seeds, identical sweep grids, and unified reporting across runs 06--14. At $d=5$ and $\sigma=0.20$, MWPM and UF define the Pareto frontier, with (runtime, LER) = (1.341 s, 0.2273) and (1.332 s, 0.2303); neural-guided MWPM is slower and less accurate (1.396 s, 0.3730), and BP is dominated (7.640 s, 0.6107). Crossing-bootstrap diagnostics are stable only for MWPM, with median $\sigma^\star_{3,5}=0.10$ (1911/2000 valid) and $\sigma^\star_{5,7}=0.1375$ (1941/2000 valid), while other decoders show no valid crossing samples. Dense-window scanning over $\sigma \in [0.08,0.24]$ returns NaN crossings for all decoders, confirming estimator- and window-sensitive threshold localization. Rank-stability and effect-size bootstrap analyses reinforce ordering robustness: BP remains rank 4, neural-guided MWPM rank 3, and MWPM-UF differences are small ($\Delta_{\mathrm{MWPM-UF}}=-0.00383$, 95\% interval $[-0.0104,0.00329]$) across $\sigma \in [0.05,0.35]$. Threaded execution preserves statistical fidelity while improving throughput: $1.34\times$ speedup in Pauli mode and $1.94\times$ in native GKP mode, with mean $|\Delta\mathrm{LER}|$ $6.07\times10^{-3}$ and $5.20\times10^{-3}$, respectively. We therefore recommend estimator-conditional threshold reporting coupled to runtime-fidelity checks for reproducible hardware-facing practical future decoder benchmarking workflows.

quant-ph

OMNIA: Closing the Loop by Leveraging LLMs for Knowledge Graph Completion

Knowledge Graphs (KGs) are widely used to represent structured knowledge, yet their automatic construction, especially with Large Language Models (LLMs), often results in incomplete or noisy outputs. Knowledge Graph Completion (KGC) aims to infer and add missing triples, but most existing methods either rely on structural embeddings that overlook semantics or language models that ignore the graph's structure and depend on external sources. In this work, we present OMNIA, a two-stage approach that bridges structural and semantic reasoning for KGC. It first generates candidate triples by clustering semantically related entities and relations within the KG, then validates them through lightweight embedding filtering followed by LLM-based semantic validation. OMNIA performs on the internal KG, without external sources, and specifically targets implicit semantics that are most frequent in LLM-generated graphs. Extensive experiments on multiple datasets demonstrate that OMNIA significantly improves F1-score compared to traditional embedding-based models. These results highlight OMNIA's effectiveness and efficiency, as its clustering and filtering stages reduce both search space and validation cost while maintaining high-quality completion.

cs.DB

Decoder Dependence in Surface-Code Threshold Estimation under Digitized Hybrid Continuous-Variable and Discrete Noise

Surface-code threshold estimates depend on the inference pipeline, including decoder and estimator choices. We compare decoders within a single LiDMaS+ workflow under Pauli-reference and digitized hybrid continuous-variable/discrete sweeps. In the Pauli-reference mode, the matching-style backend outperforms Union-Find and yields crossing median $p_c=0.0531$ (bootstrap interval $[0.0415,0.0572]$) and collapse fit $p_c=0.052$ ($\nu=1.35$). For the hybrid mode, a dense transition-window sweep at $d=3,5,7$ uses $\sigma\in[0.30,0.50]$ with step $0.01$ and $3000$ trials per point. After the initial exact-zero plateau is excluded from crossing localization, the matching-style backend gives interior crossing estimates $\sigma_c=0.4707$ for $(d=3,5)$ and $\sigma_c=0.3275$ for $(d=5,7)$; the latter lies in a low-LER region and remains estimator-sensitive. A targeted $d=9$ extension shows larger Union-Find LER at moderate-to-high $\sigma$ and matching-fallback rates up to $0.747$ at $\sigma=0.50$. In a $d=5$ neural-guidance sensitivity sweep, full learned reweighting reduces the sampled mean LER from $0.1773$ to $0.1663$ over $\sigma\in[0.35,0.55]$. These results show that estimator resolution and backend fallback diagnostics are part of an auditable decoder comparison.

quant-ph

DifGa: Differentiable Error Mitigation for Multi-Mode Gaussian and Non-Gaussian Noise in Quantum Photonic Circuits

We introduce DifGa, a fully differentiable error-mitigation framework for continuous-variable (CV) quantum photonic circuits operating under Gaussian loss and weak non-Gaussian noise. The approach is demonstrated using analytic simulations with the default.gaussian backend of PennyLane, where quantum states are represented by first and second moments and optimized end-to-end via automatic differentiation. Gaussian loss is modeled as a beam splitter interaction with an environmental vacuum mode of transmissivity $\eta \in [0.3,0.95]$, while non-Gaussian phase noise is incorporated through a differentiable Monte-Carlo mixture of random phase rotations with jitter amplitudes $\delta \in [0,0.7]$. The core architecture employs a multi-mode Gaussian circuit consisting of a signal, ancilla, and environment mode. Input states are prepared using squeezing and displacement operations with parameters $(r_s,\varphi_s,\alpha)=(0.60,0.30,0.80)$ and $(r_a,\varphi_a)=(0.40,0.10)$, followed by an entangling beam splitter with angles $(\theta,\phi)=(0.70,0.20)$. Error mitigation is achieved by appending a six-parameter trainable Gaussian recovery layer comprising local phase rotations and displacements, optimized by minimizing a quadratic loss on the signal-mode quadratures $\langle \hat{x}_0\rangle$ and $\langle \hat{p}_0\rangle$ using gradient descent with fixed learning rate $0.06$ and identical initialization across experiments. Under pure Gaussian loss, the optimized recovery suppresses reconstruction error to near machine precision ($<10^{-30}$) for moderate loss ($\eta \ge 0.5$). When non-Gaussian phase noise is present, noise-aware training using Monte Carlo averaging yields robust generalization, reducing error by more than an order of magnitude compared to Gaussian-trained recovery at large phase jitter. Runtime benchmarks confirm linear scaling with the number of Monte Carlo samples.

quant-ph

Opportunities and Challenges for Data Quality in the Era of Quantum Computing

In an era where data underpins decision-making across science, politics, and economics, ensuring high data quality is of paramount importance. Conventional computing algorithms for enhancing data quality, including anomaly detection, demand substantial computational resources, lengthy processing times, and extensive training datasets. This work aims to explore the potential advantages of quantum computing for enhancing data quality, with a particular focus on detection. We begin by examining quantum techniques that could replace key subroutines in conventional anomaly detection frameworks to mitigate their computational intensity. We then provide practical demonstrations of quantum-based anomaly detection methods, highlighting their capabilities. We present a technical implementation for detecting volatility regime changes in stock market data using quantum reservoir computing, which is a special type of quantum machine learning model. The experimental results indicate that quantum-based embeddings are a competitive alternative to classical ones in this particular example. Finally, we identify unresolved challenges and limitations in applying quantum computing to data quality tasks. Our findings open up new avenues for innovative research and commercial applications that aim to advance data quality through quantum technologies.

quant-ph

OptiMA: A Transaction-Based Framework with Throughput Optimization for Very Complex Multi-Agent Systems

In recent years, the research of multi-agent systems has taken a direction to explore larger and more complex models to fulfill sophisticated tasks. We point out two possible pitfalls that might be caused by increasing complexity; susceptibilities to faults, and performance bottlenecks. To prevent the former threat, we propose a transaction-based framework to design very complex multi-agent systems (VCMAS). To address the second threat, we offer to integrate transaction scheduling into the proposed framework. We implemented both of these ideas to develop the OptiMA framework and show that it is able to facilitate the execution of VCMAS with more than a hundred agents. We also demonstrate the effect of transaction scheduling on such a system by showing improvements up to more than 16\%. Furthermore, we also performed a theoretical analysis on the transaction scheduling problem and provided practical tools that can be used for future research on it.

cs.MA

QCardEst/QCardCorr: Quantum Cardinality Estimation and Correction

Cardinality estimation is an important part of query optimization in DBMS. We develop a Quantum Cardinality Estimation (QCardEst) approach using Quantum Machine Learning with a Hybrid Quantum-Classical Network. We define a compact encoding for turning SQL queries into a quantum state, which requires only qubits equal to the number of tables in the query. This allows the processing of a complete query with a single variational quantum circuit (VQC) on current hardware. In addition, we compare multiple classical post-processing layers to turn the probability vector output of VQC into a cardinality value. We introduce Quantum Cardinality Correction QCardCorr, which improves classical cardinality estimators by multiplying the output with a factor generated by a VQC to improve the cardinality estimation. With QCardCorr, we have an improvement over the standard PostgreSQL optimizer of 6.37 times for JOB-light and 8.66 times for STATS. For JOB-light we even outperform MSCN by a factor of 3.47.

quant-ph

Automated Archival Descriptions with Federated Intelligence of LLMs

Enforcing archival standards requires specialized expertise, and manually creating metadata descriptions for archival materials is a tedious and error-prone task. This work aims at exploring the potential of agentic AI and large language models (LLMs) in addressing the challenges of implementing a standardized archival description process. To this end, we introduce an agentic AI-driven system for automated generation of high-quality metadata descriptions of archival materials. We develop a federated optimization approach that unites the intelligence of multiple LLMs to construct optimal archival metadata. We also suggest methods to overcome the challenges associated with using LLMs for consistent metadata generation. To evaluate the feasibility and effectiveness of our techniques, we conducted extensive experiments using a real-world dataset of archival materials, which covers a variety of document types and formats. The evaluation results demonstrate the feasibility of our techniques and highlight the superior performance of the federated optimization approach compared to single-model solutions in metadata quality and reliability.

cs.AI

QCE'24 Tutorial: Quantum Annealing -- Emerging Exploration for Database Optimization

Quantum annealing is a meta-heuristic approach tailored to solve combinatorial optimization problems with quantum annealers. In this tutorial, we provide a fundamental and comprehensive introduction to quantum annealing and modern data management systems and show quantum annealing's potential benefits and applications in the realm of database optimization. We demonstrate how to apply quantum annealing for selected database optimization problems, which are critical challenges in many data management platforms. The demonstrations include solving join order optimization problems in relational databases, optimizing sophisticated transaction scheduling, and allocating virtual machines within cloud-based architectures with respect to sustainability metrics. On the one hand, the demonstrations show how to apply quantum annealing on key problems of database management systems (join order selection, transaction scheduling), and on the other hand, they show how quantum annealing can be integrated as a part of larger and dynamic optimization pipelines (virtual machine allocation). The goal of our tutorial is to provide a centralized and condensed source regarding theories and applications of quantum annealing technology for database researchers, practitioners, and everyone who wants to understand how to potentially optimize data management with quantum computing in practice. Besides, we identify the advantages, limitations, and potentials of quantum computing for future database and data management research.

quant-ph

Variables are a Curse in Software Vulnerability Prediction

Deep learning-based approaches for software vulnerability prediction currently mainly rely on the original text of software code as the feature of nodes in the graph of code and thus could learn a representation that is only specific to the code text, rather than the representation that depicts the 'intrinsic' functionality of a program hidden in the text representation. One curse that causes this problem is an infinite number of possibilities to name a variable. In order to lift the curse, in this work we introduce a new type of edge called name dependence, a type of abstract syntax graph based on the name dependence, and an efficient node representation method named 3-property encoding scheme. These techniques will allow us to remove the concrete variable names from code, and facilitate deep learning models to learn the functionality of software hidden in diverse code expressions. The experimental results show that the deep learning models built on these techniques outperform the ones based on existing approaches not only in the prediction of vulnerabilities but also in the memory need. The factor of memory usage reductions of our techniques can be up to the order of 30,000 in comparison to existing approaches.

cs.SE

Research Trends for the Interplay between Large Language Models and Knowledge Graphs

This survey investigates the synergistic relationship between Large Language Models (LLMs) and Knowledge Graphs (KGs), which is crucial for advancing AI's capabilities in understanding, reasoning, and language processing. It aims to address gaps in current research by exploring areas such as KG Question Answering, ontology generation, KG validation, and the enhancement of KG accuracy and consistency through LLMs. The paper further examines the roles of LLMs in generating descriptive texts and natural language queries for KGs. Through a structured analysis that includes categorizing LLM-KG interactions, examining methodologies, and investigating collaborative uses and potential biases, this study seeks to provide new insights into the combined potential of LLMs and KGs. It highlights the importance of their interaction for improving AI applications and outlines future research directions.

cs.AI

Are Large Language Models the New Interface for Data Pipelines?

A Language Model is a term that encompasses various types of models designed to understand and generate human communication. Large Language Models (LLMs) have gained significant attention due to their ability to process text with human-like fluency and coherence, making them valuable for a wide range of data-related tasks fashioned as pipelines. The capabilities of LLMs in natural language understanding and generation, combined with their scalability, versatility, and state-of-the-art performance, enable innovative applications across various AI-related fields, including eXplainable Artificial Intelligence (XAI), Automated Machine Learning (AutoML), and Knowledge Graphs (KG). Furthermore, we believe these models can extract valuable insights and make data-driven decisions at scale, a practice commonly referred to as Big Data Analytics (BDA). In this position paper, we provide some discussions in the direction of unlocking synergies among these technologies, which can lead to more powerful and intelligent AI solutions, driving improvements in data pipelines across a wide range of applications and domains integrating humans, computers, and knowledge.

cs.CL

Hype or Heuristic? Quantum Reinforcement Learning for Join Order Optimisation

Identifying optimal join orders (JOs) stands out as a key challenge in database research and engineering. Owing to the large search space, established classical methods rely on approximations and heuristics. Recent efforts have successfully explored reinforcement learning (RL) for JO. Likewise, quantum versions of RL have received considerable scientific attention. Yet, it is an open question if they can achieve sustainable, overall practical advantages with improved quantum processors. In this paper, we present a novel approach that uses quantum reinforcement learning (QRL) for JO based on a hybrid variational quantum ansatz. It is able to handle general bushy join trees instead of resorting to simpler left-deep variants as compared to approaches based on quantum(-inspired) optimisation, yet requires multiple orders of magnitudes fewer qubits, which is a scarce resource even for post-NISQ systems. Despite moderate circuit depth, the ansatz exceeds current NISQ capabilities, which requires an evaluation by numerical simulations. While QRL may not significantly outperform classical approaches in solving the JO problem with respect to result quality (albeit we see parity), we find a drastic reduction in required trainable parameters. This benefits practically relevant aspects ranging from shorter training times compared to classical RL, less involved classical optimisation passes, or better use of available training data, and fits data-stream and low-latency processing scenarios. Our comprehensive evaluation and careful discussion delivers a balanced perspective on possible practical quantum advantage, provides insights for future systemic approaches, and allows for quantitatively assessing trade-offs of quantum approaches for one of the most crucial problems of database management systems.

quant-ph

Exploring In-Context Learning Capabilities of Foundation Models for Generating Knowledge Graphs from Text

Knowledge graphs can represent information about the real-world using entities and their relations in a structured and semantically rich manner and they enable a variety of downstream applications such as question-answering, recommendation systems, semantic search, and advanced analytics. However, at the moment, building a knowledge graph involves a lot of manual effort and thus hinders their application in some situations and the automation of this process might benefit especially for small organizations. Automatically generating structured knowledge graphs from a large volume of natural language is still a challenging task and the research on sub-tasks such as named entity extraction, relation extraction, entity and relation linking, and knowledge graph construction aims to improve the state of the art of automatic construction and completion of knowledge graphs from text. The recent advancement of foundation models with billions of parameters trained in a self-supervised manner with large volumes of training data that can be adapted to a variety of downstream tasks has helped to demonstrate high performance on a large range of Natural Language Processing (NLP) tasks. In this context, one emerging paradigm is in-context learning where a language model is used as it is with a prompt that provides instructions and some examples to perform a task without changing the parameters of the model using traditional approaches such as fine-tuning. This way, no computing resources are needed for re-training/fine-tuning the models and the engineering effort is minimal. Thus, it would be beneficial to utilize such capabilities for generating knowledge graphs from text.

cs.CL