SearcharxivSearch

arXiv subjects

Jose Moreno

Publications and source records attributed to Jose Moreno.

4 recordsLinked to original sources

Guaranteeing Faithful Evidence Extraction in Speculative Retrieval-Augmented Generation

Large Language Models (LLMs) are increasingly used as interfaces for information retrieval, but they remain prone to hallucinations and faithfulness errors, in which the generated answers diverge from the retrieved evidence. While Retrieval-Augmented Generation (RAG) and recent hybrid or semi-extractive approaches mitigate this issue, they do not guarantee that quoted or extracted spans are verbatim from the retrieved context. This limitation can have severe consequences in safety-critical domains, where answers must exactly match certified documentation. We introduce Constrained Hybrid Decoding (CHyD), a novel faithfulness-first paradigm for speculative RAG. While traditional speculative decoding is optimized for inference speed, CHyD repurposes this architecture to ensure faithful verbatim evidence extraction when the extraction mode is correctly triggered. Our approach enforces hard decoding constraints that restrict generation to continuous spans present in the retrieved documents. This design provides a robust but straightforward guarantee: any explicitly quoted span in the output appears verbatim in the provided context. We evaluate our method across state-of-the-art LLMs on diverse abstractive, extractive, and semi-extractive QA benchmarks, including technical datasets motivated by aircraft maintenance. Results show that existing hybrid methods frequently hallucinate quoted spans, with exact extraction accuracy dropping below 40% in technical domains. In contrast, our approach achieves near-perfect extraction faithfulness regardless of the model used. Although enforcing hard constraints introduces a trade-off with fluency-oriented metrics, our method improves exact answer correctness and remains competitive overall, highlighting its suitability for safety-critical information retrieval applications.

cs.IR

Jointly Generating and Attributing Answers using Logits of Document-Identifier Tokens

Despite their impressive performances, Large Language Models (LLMs) remain prone to hallucination, which critically undermines their trustworthiness. While most of the previous work focused on tackling answer and attribution correctness, a recent line of work investigated faithfulness, with a focus on leveraging internal model signals to reflect a model's actual decision-making process while generating the answer. Nevertheless, these methods induce additional latency and have shown limitations in directly aligning token generation with attribution generation. In this paper, we introduce LoDIT, a method that jointly generates and faithfully attributes answers in RAG by leveraging specific token logits during generation. It consists of two steps: (1) marking the documents with specific token identifiers and then leveraging the logits of these tokens to estimate the contribution of each document to the answer during generation, and (2) aggregating these contributions into document attributions. Experiments on a trustworthiness-focused attributed text-generation benchmark, Trust-Align, show that LoDIT significantly outperforms state-of-the-art models on several metrics. Finally, an in-depth analysis of LoDIT shows both its efficiency in terms of latency and its robustness in different settings.

cs.CL

Complexity evaluation of network configurations and abstractions

Computer networks have been traditionally configured by humans using command-line interfaces. Some network abstractions have emerged in the last 10 years, but there is no easy way of comparing them to each other objectively. Therefore, there is no consensus in the industry of what direction modern network abstractions should take, and the adoption of these abstractions lags as a consequence. In this paper I propose a comparison framework using metrics derived from graph structures to evaluate the simplicity, efficiency, and effectiveness of different network abstraction models. The result of this comparison is that while some of the existing network abstractions are quite efficient to store network policy (such as the Kubernetes or the Cisco Application Centric Infrastructure models), others (notably public cloud) are still very infrastructure-centric and suffer from excessive complexity.

cs.NI

Research on pinches driven by SPPED 2 generator : hard X-ray and neutron emission in plasma focus configuration

SPEED2 is a generator based on Marx technology and was designed in the University of Dusseldorf. SPEED2 consists on 40 +/- Marx modules connected in parallel (4.1 mF equivalent Marx generator capacity, 300 kV, 4 MA in short circuit, 187 kJ, 400 ns rise time, dI/dt~1013 A/s). Currently the SPEED2 is operating at the Comision Chilena de Energia Nuclear, CCHEN, Chile, being the most powerful and energetic device for dense transient plasma in the Southern Hemisphere. Most of the previous works developed in SPEED2 at Dusseldorf were done in a plasma focus configuration for soft X-ray emission and the neutron emission from SPEED2 was not completely studied. The research program at CCHEN considers experiments in different pinch configurations (plasma focus, gas puffed plasma focus, gas embedded Z-pinch, wire arrays) at current of hundred of kiloamperes to mega-amperes, using the SPEED2 generator. The Chilean operation has begun implementing and developing diagnostics in a conventional plasma focus configuration operating in deuterium in order to characterize the neutron emission and the hard X-ray production. Silver activation counters, plastics CR39 and scintillator-photomultiplier detectors are used to characterize the neutron emission. Images of metallic plates with different thickness are obtained on commercial radiographic film, Agfa Curix ST-G2, in order to characterize an effective energy of the hard X-ray outside of the discharge .

physics.plasm-ph