SearcharxivSearch

arXiv subjects

Antonio Rago

Publications and source records attributed to Antonio Rago.

At least 19 recordsLinked to original sources

Contrastive Explanations in Quantitative Bipolar Argumentation Frameworks

Argumentation frameworks are useful tools for representing and reasoning with information in a variety of settings, e.g. in supplementing AI models as they perform classification tasks, with a notable benefit of providing additional explainability. In this paper, we introduce contrastive explanations for Quantitative Bipolar Argumentation Frameworks (QBAFs), one such formalism. Unlike most existing explanations for QBAFs, which explain the reasoning outcome of a single argument of interest (i.e. a topic argument), contrastive explanations explain the difference between two topic arguments. We introduce a general form of contrastive attribution functions (CAFs) and establish a set of general properties they should satisfy. We introduce CAFs based on removal, gradients and Shapley-values, and study their properties. Finally, to illustrate contrastive explanations, we demonstrate their usefulness in healthcare and bias identification settings.

cs.AI

A Theory of Post-hoc Debate Judgement

Debates have recently emerged as a useful methodology for agentic AI to improve performance as well as to aid explainability and user engagement. For example, LLM-empowered agents may debate internally (with themselves) and/or externally (with other agents). In many settings where debates are used, debates' outcomes and resulting outputs are determined post-hoc by external judges, often LLMs. In this paper we develop and test a novel theory of debate judgement applicable to all settings where agents engage in debates by providing pros and cons for their opinions therein. Specifically, we identify a number of formal properties that debate judgement may be required to satisfy in general, as concerns reproducibility, robustness, groundedness and explainability. Then, we explore their satisfaction formally and/or experimentally, for claim verification settings, for two specific alternative debate judgement methods: variants of the LLMs as a judge idea and formal semantics drawn from computational argumentation. We show that the two methods give similar accuracy performances but the former may lack formal guarantees that the latter brings. Overall, our study indicates argumentation semantics as an ideal candidate for principled judges in debate-driven AI.

cs.AI

Argumentation for Common Ground: Finding Zones of Possible Agreement between Individuals in Conflict

How can common ground between societies in conflict be identified when citizens' acceptability of peace agreements is shaped by contested narratives? Such acceptability is mediated not only by the clauses that agreements include or exclude, but crucially by citizens' subjective reasoning concerning agreements' clauses. In this paper, we leverage computational argumentation to introduce a novel approach to identifying mutually acceptable agreements among individuals in conflict, i.e. a Zone of Possible Agreement (ZOPA). First, we introduce a quantitative bipolar argumentation framework tailored to represent each side's reasoning about peace agreements. We then show how merging these frameworks can enable negotiators to identify peace agreements that are mutually acceptable. To evaluate our approach under conditions of real-world relevance, we focus on the Palestinian-Israeli conflict, where long-standing policy, practitioner and public interest underscores the demand for methods capable of analysing polarised public reasoning. We show how our framework identifies a ZOPA through theoretical analysis and preliminary experiments using survey data from both existing work and retrieved by a large language model. The results illustrate how argumentation can empower negotiators and conflict-resolution teams in mapping feasible ZOPAs grounded in citizens' reasoning.

cs.AI

Future Requirements of Lattice Field Theory Calculations on European High-Performance Computing Facilities

Lattice field theory provides a first-principles framework for studying properties of strongly interacting quantum field theories in elementary particle physics. Researchers in lattice field theory are also among the largest and most efficient users of high- performance computing resources in fundamental science. In this contribution, we outline the computational profile of lattice QCD, from gauge-field generation to large-scale measurements, and discuss the main hardware, software, and human resource requirements needed to sustain progress on current and future European HPC infrastructures.

hep-lat

Evaluating LLM-Driven Summarisation of Parliamentary Debates with Computational Argumentation

Understanding how policy is debated and justified in parliament is a fundamental aspect of the democratic process. However, the volume and complexity of such debates mean that outside audiences struggle to engage. Meanwhile, Large Language Models (LLMs) have been shown to enable automated summarisation at scale. While summaries of debates can make parliamentary procedures more accessible, evaluating whether these summaries faithfully communicate argumentative content remains challenging. Existing automated summarisation metrics have been shown to correlate poorly with human judgements of consistency (i.e., faithfulness or alignment between summary and source). In this work, we propose a formal framework for evaluating parliamentary debate summaries that grounds argument structures in the contested proposals up for debate. Our novel approach, driven by computational argumentation, focuses the evaluation on formal properties concerning the faithful preservation of the reasoning presented to justify or oppose policy outcomes. We demonstrate our methods using a case-study of debates from the European Parliament and associated LLM-driven summaries.

cs.CL

High-Performance Simulations of Higher Representations of Wilson Fermions

We present HiRep v2, an open-source software suite for high-performance lattice field theory simulations with dynamical Wilson fermions in higher representations of $SU(N_g)$ gauge groups. This new version fully supports GPU acceleration, optimizing both gauge configuration generation and measurements for NVIDIA and AMD GPUs. HiRep v2 integrates improved gauge and fermionic lattice actions, advanced inverters, and Monte Carlo algorithms, including (R)HMC with Hasenbusch acceleration. It exhibits excellent scalability across multiple GPUs and nodes with minimal efficiency loss, making it a robust tool for large-scale simulations in physics beyond the Standard Model.

hep-lat

From User Preferences to Base Score Extraction Functions in Gradual Argumentation (with Appendix)

Gradual argumentation is a field of symbolic AI which is attracting attention for its ability to support transparent and contestable AI systems. It is considered a useful tool in domains such as decision-making, recommendation, debate analysis, and others. The outcomes in such domains are usually dependent on the arguments' base scores, which must be selected carefully. Often, this selection process requires user expertise and may not always be straightforward. On the other hand, organising the arguments by preference could simplify the task. In this work, we introduce \emph{Base Score Extraction Functions}, which provide a mapping from users' preferences over arguments to base scores. These functions can be applied to the arguments of a \emph{Bipolar Argumentation Framework} (BAF), supplemented with preferences, to obtain a \emph{Quantitative Bipolar Argumentation Framework} (QBAF), allowing the use of well-established computational tools in gradual argumentation. We outline the desirable properties of base score extraction functions, discuss some design choices, and provide an algorithm for base score extraction. Our method incorporates an approximation of non-linearities in human preferences to allow for better approximation of the real ones. Finally, we evaluate our approach both theoretically and experimentally in a robotics setting, and offer recommendations for selecting appropriate gradual semantics in practice.

cs.AI

Towards an Argumentative Foundation for Evaluative AI

Evaluative AI (EAI) has been recently proposed as a way to support human decision-making, not by producing a single recommendation, but by presenting competing hypotheses together with evidence for and against each. In this position paper, we advocate (computational) argumentation as a particularly suitable paradigm to provide a formal, computable foundation for forms of EAI that are explainable and contestable, setting the ground for a long-term research agenda towards distributed and human-centred EAI systems.

cs.AI

Renormalized quark masses using gradient flow

We propose a new and simple method for determining the renormalized quark masses from lattice simulations. Renormalized quark masses are an important input to many phenomenological applications, including searching and modeling physics beyond the Standard Model. The non-perturbative renormalization is performed using gradient flow combined with the short-flow-time expansion that is improved by renormalization-group (RG) running to match to the $\overline{\text{MS}}$-scheme. Implementing the RG running perturbatively, we demonstrate this method works reliably at least up to the charm-quark mass and exhibits an easily-attainable ``windowing condition''. Using RBC/UKQCD's (2+1)-flavor Shamir domain-wall fermion ensembles with Iwasaki gauge action, we find $m_s^\overline{\text{MS}}(μ=2 \text{ GeV}) = 90(3)$ MeV and $m_c^\overline{\text{MS}}(μ=3 \text{ GeV}) = 972(16)$ MeV. These results predict the scale-independent ratio $m_c/m_s= 12.1(4)$. Generalization to other observables is possible, providing an efficient approach to determine non-perturbatively renormalized fermionic observables like form factors or bag parameters from lattice simulations.

hep-lat

Bag Parameters for Heavy Meson Lifetimes

We calculate the dimension-six $ΔQ=0$ four-quark matrix elements describing heavy-meson lifetime ratios using the gradient flow with its short flow-time expansion as a renormalization procedure. On six RBC/UKQCD 2+1-flavor domain-wall fermion ensembles, we determine flowed bag parameters for physical charm and strange quarks and match to the $\overline{\text{MS}}$ scheme with perturbative short flow-time expansion coefficients through next-to-next-to-leading order (NNLO). A multi-scale matching procedure using renormalization-group running improves the extrapolation to zero flow time. For the operators relevant to $τ(D_s)/τ(D^0)$ at the SU(3)$_{\rm F}$ symmetric point, we obtain $B_1^{\overline{\text{MS}}}(3\,{\rm GeV})=1.0524(97)$,$B_2^{\overline{\text{MS}}}(3\,{\rm GeV})=0.9621(70)$, $ε_1^{\overline{\text{MS}}}(3\,{\rm GeV})=-0.2275(76)$, and $ε_2^{\overline{\text{MS}}}(3\,{\rm GeV})=-0.0005(8)$ using a specific choice of evanescent operators. This is the first lattice-QCD determination of $ΔQ=0$ four-quark operators with a full error budget. It opens the path towards higher-precision predictions of heavy-meson lifetimes and similar quantities exhibiting operator mixing under renormalization.

hep-ph

Heavy-Meson Bag Parameters using Gradient Flow

We demonstrate the use of the gradient flow combined with the short flow-time expansion (GF+SFTX) as a renormalization procedure for four-quark operator matrix elements and associated bag parameters relevant to neutral heavy-meson mixing ($ΔQ=2$) and heavy-meson lifetimes ($ΔQ=0$). Using six RBC/UKQCD 2+1-flavor domain-wall fermion ensembles, we calculate for a charm-strange system with physical quark masses flowed bag parameters and match them to the $\overline{\text{MS}}$ scheme using perturbative SFTX coefficients up to next-to-next-to-leading order in QCD. We employ a multi-scale matching strategy and a renormalization-group improved flow-time evolution which allows for a reliable estimate of systematic uncertainties. For a fictitious neutral $D_s$ meson, we obtain the $ΔQ=2$ $\overline{\text{MS}}$ bag parameter ${\cal B}^{\overline{\text{MS}}}_1(3\,{\rm GeV})=0.7673(123)$, consistent with existing short-distance $D^0$ mixing determinations. For the $ΔQ=0$ lifetime-ratio operator basis, we find the $\overline{\text{MS}}$ results $B^{\overline{\text{MS}}}_1(3\,{\rm GeV})=1.0524(97)$, $B^{\overline{\text{MS}}}_2(3\,{\rm GeV})=0.9621(71)$, $ε^{\overline{\text{MS}}}_1(3\,{\rm GeV})=-0.2275(76)$, and $ε^{\overline{\text{MS}}}_2(3\,{\rm GeV})=-0.0005(8)$. We provide conversion formulae to re-express these results for an arbitrary choice of evanescent operators. These results demonstrate that GF+SFTX can deliver precise determinations of dimension-six four-quark operators and establish a framework for future lattice computations including more complex operator bases, where the challenge of power-divergent mixing is shifted to the continuum and handled in the SFTX.

hep-lat

Argumentative Human-AI Decision-Making: Toward AI Agents That Reason With Us, Not For Us

Computational argumentation offers formal frameworks for transparent, verifiable reasoning but has traditionally been limited by its reliance on domain-specific information and extensive feature engineering. In contrast, LLMs excel at processing unstructured text, yet their opaque nature makes their reasoning difficult to evaluate and trust. We argue that the convergence of these fields will lay the foundation for a new paradigm: Argumentative Human-AI Decision-Making. We analyze how the synergy of argumentation framework mining, argumentation framework synthesis, and argumentative reasoning enables agents that do not just justify decisions, but engage in dialectical processes where decisions are contestable and revisable -- reasoning with humans rather than for them. This convergence of computational argumentation and LLMs is essential for human-aware, trustworthy AI in high-stakes domains.

cs.AI

Synthesising Counterfactual Explanations via Label-Conditional Gaussian Mixture Variational Autoencoders

Counterfactual explanations (CEs) provide recourse recommendations for individuals affected by algorithmic decisions. A key challenge is generating CEs that are robust against various perturbation types (e.g. input and model perturbations) while simultaneously satisfying other desirable properties. These include plausibility, ensuring CEs reside on the data manifold, and diversity, providing multiple distinct recourse options for single inputs. Existing methods, however, mostly struggle to address these multifaceted requirements in a unified, model-agnostic manner. We address these limitations by proposing a novel generative framework. First, we introduce the Label-conditional Gaussian Mixture Variational Autoencoder (L-GMVAE), a model trained to learn a structured latent space where each class label is represented by a set of Gaussian components with diverse, prototypical centroids. Building on this, we present LAPACE (LAtent PAth Counterfactual Explanations), a model-agnostic algorithm that synthesises entire paths of CE points by interpolating from inputs' latent representations to those learned latent centroids. This approach inherently ensures robustness to input changes, as all paths for a given target class converge to the same fixed centroids. Furthermore, the generated paths provide a spectrum of recourse options, allowing users to navigate the trade-off between proximity and plausibility while also encouraging robustness against model changes. In addition, user-specified actionability constraints can also be easily incorporated via lightweight gradient optimisation through the L-GMVAE's decoder. Comprehensive experiments show that LAPACE is computationally efficient and achieves competitive performance across eight quantitative metrics.

cs.LG

Retrieval- and Argumentation-Enhanced Multi-Agent LLMs for Judgmental Forecasting (Extended Version with Supplementary Material)

Judgmental forecasting is the task of making predictions about future events based on human judgment. This task can be seen as a form of claim verification, where the claim corresponds to a future event and the task is to assess the plausibility of that event. In this paper, we propose a novel multi-agent framework for claim verification, whereby different agents may disagree on claim veracity and bring specific evidence for and against the claims, represented as quantitative bipolar argumentation frameworks (QBAFs). We then instantiate the framework for supporting claim verification, with a variety of agents realised with Large Language Models (LLMs): (1) ArgLLM agents, an existing approach for claim verification that generates and evaluates QBAFs; (2) RbAM agents, whereby LLM-empowered Relation-based Argument Mining (RbAM) from external sources is used to generate QBAFs; (3) RAG-ArgLLM agents, extending ArgLLM agents with a form of Retrieval-Augmented Generation (RAG) of arguments from external sources. Finally, we conduct experiments with two standard judgmental forecasting datasets, with instances of our framework with two or three agents, empowered by six different base LLMs. We observe that combining evidence from agents can improve forecasting accuracy, especially in the case of three agents, while providing an explainable combination of evidence for claim verification.

cs.AI

Meronymic Ontology Extraction via Large Language Models

Ontologies have become essential in today's digital age as a way of organising the vast amount of readily available unstructured text. In providing formal structure to this information, ontologies have immense value and application across various domains, e.g., e-commerce, where countless product listings necessitate proper product organisation. However, the manual construction of these ontologies is a time-consuming, expensive and laborious process. In this paper, we harness the recent advancements in large language models (LLMs) to develop a fully-automated method of extracting product ontologies, in the form of meronymies, from raw review texts. We demonstrate that the ontologies produced by our method surpass an existing, BERT-based baseline when evaluating using an LLM-as-a-judge. Our investigation provides the groundwork for LLMs to be used more generally in (product or otherwise) ontology extraction.

cs.CL

Representation Consistency for Accurate and Coherent LLM Answer Aggregation

Test-time scaling improves large language models' (LLMs) performance by allocating more compute budget during inference. To achieve this, existing methods often require intricate modifications to prompting and sampling strategies. In this work, we introduce representation consistency (RC), a test-time scaling method for aggregating answers drawn from multiple candidate responses of an LLM regardless of how they were generated, including variations in prompt phrasing and sampling strategy. RC enhances answer aggregation by not only considering the number of occurrences of each answer in the candidate response set, but also the consistency of the model's internal activations while generating the set of responses leading to each answer. These activations can be either dense (raw model activations) or sparse (encoded via pretrained sparse autoencoders). Our rationale is that if the model's representations of multiple responses converging on the same answer are highly variable, this answer is more likely to be the result of incoherent reasoning and should be down-weighted during aggregation. Importantly, our method only uses cached activations and lightweight similarity computations and requires no additional model queries. Through experiments with four open-source LLMs and four reasoning datasets, we validate the effectiveness of RC for improving task performance during inference, with consistent accuracy improvements (up to 4%) over strong test-time scaling baselines. We also show that consistency in the sparse activation signals aligns well with the common notion of coherent reasoning.

cs.CL

Moments of parton distributions functions of the pion from lattice QCD using gradient flow

We present a nonperturbative determination of the pion valence parton distribution function (PDF) moment ratios $\left\langle x^{n-1} \right\rangle / \left\langle x \right\rangle$ up to $n=6$, using the gradient flow in lattice QCD. As a testing ground, we employ SU($3$) isosymmetric gauge configurations generated by the OpenLat initiative with a pseudoscalar mass of $m_π\simeq 411~\text{MeV}$. Our analysis uses four lattice spacings and a nonperturbatively improved action, enabling full control over the continuum extrapolation, and the limit of vanishing flow time, $t\to0$. The flowed ratios exhibit O($a^2$) scaling across the ensembles, and the continuum-extrapolated results, matched to the $\overline {\text{MS}}$ scheme at $μ= 2$ GeV using next-to-next-to-leading order matching coefficients, show only mild residual flow-time dependence. The resulting ratios, computed with a relatively small number of configurations, are consistent with phenomenological expectations for the pion's valence distribution, with statistical uncertainties that are competitive with modern global fits. These findings demonstrate that the gradient flow provides an efficient and systematically improvable method to access partonic quantities from first principles. Future extensions of this work will target lighter pion masses toward the physical point, and applications to nucleon structure such as the proton PDFs and the gluon and sea-quark distributions.

hep-lat

Evaluating Uncertainty Quantification Methods in Argumentative Large Language Models

Research in uncertainty quantification (UQ) for large language models (LLMs) is increasingly important towards guaranteeing the reliability of this groundbreaking technology. We explore the integration of LLM UQ methods in argumentative LLMs (ArgLLMs), an explainable LLM framework for decision-making based on computational argumentation in which UQ plays a critical role. We conduct experiments to evaluate ArgLLMs' performance on claim verification tasks when using different LLM UQ methods, inherently performing an assessment of the UQ methods' effectiveness. Moreover, the experimental procedure itself is a novel way of evaluating the effectiveness of UQ methods, especially when intricate and potentially contentious statements are present. Our results demonstrate that, despite its simplicity, direct prompting is an effective UQ strategy in ArgLLMs, outperforming considerably more complex approaches.

cs.CL