SearcharxivSearch

arXiv subjects

Christian Buck

Publications and source records attributed to Christian Buck.

18 recordsLinked to original sources

AI-Assisted Scientific Assessment: A Case Study on Climate Change

The emerging paradigm of AI co-scientists focuses on tasks characterized by repeatable verification, where agents explore search spaces in 'guess and check' loops. This paradigm does not extend to problems where repeated evaluation is impossible and ground truth is established by the consensus synthesis of theory and existing evidence. We evaluate a Gemini-based AI environment designed to support collaborative scientific assessment, integrated into a standard scientific workflow. In collaboration with a diverse group of 13 scientists working in the field of climate science, we tested the system on a complex topic: the stability of the Atlantic Meridional Overturning Circulation (AMOC). Our results show that AI can accelerate the scientific workflow. The group produced a comprehensive synthesis of 79 papers through 104 revision cycles in just over 46 person-hours. AI contribution was significant: most AI-generated content was retained in the report. AI also helped maintain logical consistency and presentation quality. However, expert additions were crucial to ensure its acceptability: less than half of the report was produced by AI. Furthermore, substantial oversight was required to expand and elevate the content to rigorous scientific standards.

cs.CL

CLINB: A Climate Intelligence Benchmark for Foundational Models

Evaluating how Large Language Models (LLMs) handle complex, specialized knowledge remains a critical challenge. We address this through the lens of climate change by introducing CLINB, a benchmark that assesses models on open-ended, grounded, multimodal question answering tasks with clear requirements for knowledge quality and evidential support. CLINB relies on a dataset of real users' questions and evaluation rubrics curated by leading climate scientists. We implement and validate a model-based evaluation process and evaluate several frontier models. Our findings reveal a critical dichotomy. Frontier models demonstrate remarkable knowledge synthesis capabilities, often exhibiting PhD-level understanding and presentation quality. They outperform "hybrid" answers curated by domain experts assisted by weaker models. However, this performance is countered by failures in grounding. The quality of evidence varies, with substantial hallucination rates for references and images. We argue that bridging this gap between knowledge synthesis and verifiable attribution is essential for the deployment of AI in scientific workflows and that reliable, interpretable benchmarks like CLINB are needed to progress towards building trustworthy AI systems.

cs.AI

Assessing Large Language Models on Climate Information

As Large Language Models (LLMs) rise in popularity, it is necessary to assess their capability in critically relevant domains. We present a comprehensive evaluation framework, grounded in science communication research, to assess LLM responses to questions about climate change. Our framework emphasizes both presentational and epistemological adequacy, offering a fine-grained analysis of LLM generations spanning 8 dimensions and 30 issues. Our evaluation task is a real-world example of a growing number of challenging problems where AI can complement and lift human performance. We introduce a novel protocol for scalable oversight that relies on AI Assistance and raters with relevant education. We evaluate several recent LLMs on a set of diverse climate questions. Our results point to a significant gap between surface and epistemological qualities of LLMs in the realm of climate communication.

cs.CL

Decoding a Neural Retriever's Latent Space for Query Suggestion

Neural retrieval models have superseded classic bag-of-words methods such as BM25 as the retrieval framework of choice. However, neural systems lack the interpretability of bag-of-words models; it is not trivial to connect a query change to a change in the latent space that ultimately determines the retrieval results. To shed light on this embedding space, we learn a "query decoder" that, given a latent representation of a neural search engine, generates the corresponding query. We show that it is possible to decode a meaningful query from its latent representation and, when moving in the right direction in latent space, to decode a query that retrieves the relevant paragraph. In particular, the query decoder can be useful to understand "what should have been asked" to retrieve a particular paragraph from the collection. We employ the query decoder to generate a large synthetic dataset of query reformulations for MSMarco, leading to improved retrieval performance. On this data, we train a pseudo-relevance feedback (PRF) T5 model for the application of query suggestion that outperforms both query reformulation and PRF information retrieval baselines.

cs.CL

Zero-Shot Retrieval with Search Agents and Hybrid Environments

Learning to search is the task of building artificial agents that learn to autonomously use a search box to find information. So far, it has been shown that current language models can learn symbolic query reformulation policies, in combination with traditional term-based retrieval, but fall short of outperforming neural retrievers. We extend the previous learning to search setup to a hybrid environment, which accepts discrete query refinement operations, after a first-pass retrieval step via a dual encoder. Experiments on the BEIR task show that search agents, trained via behavioral cloning, outperform the underlying search system based on a combined dual encoder retriever and cross encoder reranker. Furthermore, we find that simple heuristic Hybrid Retrieval Environments (HRE) can improve baseline performance by several nDCG points. The search agent based on HRE (HARE) matches state-of-the-art performance, balanced in both zero-shot and in-domain evaluations, via interpretable actions, and at twice the speed.

cs.CL

Tomayto, Tomahto. Beyond Token-level Answer Equivalence for Question Answering Evaluation

The predictions of question answering (QA)systems are typically evaluated against manually annotated finite sets of one or more answers. This leads to a coverage limitation that results in underestimating the true performance of systems, and is typically addressed by extending over exact match (EM) with pre-defined rules or with the token-level F1 measure. In this paper, we present the first systematic conceptual and data-driven analysis to examine the shortcomings of token-level equivalence measures. To this end, we define the asymmetric notion of answer equivalence (AE), accepting answers that are equivalent to or improve over the reference, and publish over 23k human judgments for candidates produced by multiple QA systems on SQuAD. Through a careful analysis of this data, we reveal and quantify several concrete limitations of the F1 measure, such as a false impression of graduality, or missing dependence on the question. Since collecting AE annotations for each evaluated model is expensive, we learn a BERT matching (BEM) measure to approximate this task. Being a simpler task than QA, we find BEM to provide significantly better AE approximations than F1, and to more accurately reflect the performance of systems. Finally, we demonstrate the practical utility of AE and BEM on the concrete application of minimal accurate prediction sets, reducing the number of required answers by up to x2.6.

cs.CL

Boosting Search Engines with Interactive Agents

This paper presents first successful steps in designing search agents that learn meta-strategies for iterative query refinement in information-seeking tasks. Our approach uses machine reading to guide the selection of refinement terms from aggregated search results. Agents are then empowered with simple but effective search operators to exert fine-grained and transparent control over queries and search results. We develop a novel way of generating synthetic search sessions, which leverages the power of transformer-based language models through (self-)supervised learning. We also present a reinforcement learning agent with dynamically constrained actions that learns interactive search strategies from scratch. Our search agents obtain retrieval and answer quality performance comparable to recent neural methods, using only a traditional term-based BM25 ranking function and interpretable discrete reranking and filtering actions.

cs.CL

Meta Answering for Machine Reading

We investigate a framework for machine reading, inspired by real world information-seeking problems, where a meta question answering system interacts with a black box environment. The environment encapsulates a competitive machine reader based on BERT, providing candidate answers to questions, and possibly some context. To validate the realism of our formulation, we ask humans to play the role of a meta-answerer. With just a small snippet of text around an answer, humans can outperform the machine reader, improving recall. Similarly, a simple machine meta-answerer outperforms the environment, improving both precision and recall on the Natural Questions dataset. The system relies on joint training of answer scoring and the selection of conditioning information.

cs.CL

Novel Opaque Scintillator for Neutrino Detection

There is rising interest in organic scintillators with low scattering length for future neutrino detectors. Therefore, a new scintillator system was developed based on admixtures of paraffin wax in linear alkyl benzene. The transparency and viscosity of this gel-like material can be tuned by temperature adjustment. Whereas it is a colorless transparent liquid at temperatures around 40C it has a milky wax structure below 20C. The production and properties of such a scintillator as well as its advantages compared to transparent liquids are described.

physics.ins-det

Status of Light Sterile Neutrino Searches

A number of anomalous results in short-baseline oscillation may hint at the existence of one or more light sterile neutrino states in the eV mass range and have triggered a wave of new experimental efforts to search for a definite signature of oscillations between active and sterile neutrino states. The present paper aims to provide a comprehensive review on the status of light sterile neutrino searches in mid-2019: we discuss not only the basic experimental approaches and sensitivities of reactor, source, atmospheric, and accelerator neutrino oscillation experiments but also the complementary bounds arising from direct neutrino mass experiments and cosmological observations. Moreover, we review current results from global oscillation analyses that include the constraints set by running reactor and atmospheric neutrino experiments. They permit to set tighter bounds on the active-sterile oscillation parameters but as yet are not able to provide a definite conclusion on the existence of eV-scale sterile neutrinos.

hep-ex

Production and Properties of the Liquid Scintillators used in the Stereo Reactor Neutrino Experiment

The electron antineutrino spectrum in the Stereo reactor experiment (ILL Grenoble) is measured via the inverse beta decay signals in an organic liquid scintillator. The six target cells of the Stereo detector are filled with about 1800 litres of Gd-loaded liquid scintillator optimised for the requirements of the experiment. These target cells are surrounded by similar cells containing liquid scintillator without the Gd-loading. The development and characteristics of these scintillators are reported. In particular, the transparency, light production and pulse shape discrimination capabilities of the organic liquids are discussed.

physics.ins-det

Zero-Shot Dual Machine Translation

Neural Machine Translation (NMT) systems rely on large amounts of parallel data. This is a major challenge for low-resource languages. Building on recent work on unsupervised and semi-supervised methods, we present an approach that combines zero-shot and dual learning. The latter relies on reinforcement learning, to exploit the duality of the machine translation task, and requires only monolingual data for the target language pair. Experiments show that a zero-shot dual system, trained on English-French and English-Spanish, outperforms by large margins a standard NMT system in zero-shot translation performance on Spanish-French (both directions). The zero-shot dual method approaches the performance, within 2.2 BLEU points, of a comparable supervised setting. Our method can obtain improvements also on the setting where a small amount of parallel data for the zero-shot language pair is available. Adding Russian, to extend our experiments to jointly modeling 6 zero-shot translation directions, all directions improve between 4 and 15 BLEU points, again, reaching performance near that of the supervised setting.

cs.CL

Analyzing Language Learned by an Active Question Answering Agent

We analyze the language learned by an agent trained with reinforcement learning as a component of the ActiveQA system [Buck et al., 2017]. In ActiveQA, question answering is framed as a reinforcement learning task in which an agent sits between the user and a black box question-answering system. The agent learns to reformulate the user's questions to elicit the optimal answers. It probes the system with many versions of a question that are generated via a sequence-to-sequence question reformulation model, then aggregates the returned evidence to find the best answer. This process is an instance of \emph{machine-machine} communication. The question reformulation model must adapt its language to increase the quality of the answers returned, matching the language of the question answering system. We find that the agent does not learn transformations that align with semantic intuitions but discovers through learning classical information retrieval techniques such as tf-idf re-weighting and stemming.

cs.CL

Ask the Right Questions: Active Question Reformulation with Reinforcement Learning

We frame Question Answering (QA) as a Reinforcement Learning task, an approach that we call Active Question Answering. We propose an agent that sits between the user and a black box QA system and learns to reformulate questions to elicit the best possible answers. The agent probes the system with, potentially many, natural language reformulations of an initial question and aggregates the returned evidence to yield the best answer. The reformulation system is trained end-to-end to maximize answer quality using policy gradient. We evaluate on SearchQA, a dataset of complex questions extracted from Jeopardy!. The agent outperforms a state-of-the-art base model, playing the role of the environment, and other benchmarks. We also analyze the language that the agent has learned while interacting with the question answering system. We find that successful question reformulations look quite different from natural language paraphrases. The agent is able to discover non-trivial reformulation strategies that resemble classic information retrieval techniques such as term re-weighting (tf-idf) and stemming.

cs.CL

Sterile Neutrinos: Reactor Experiments

Nuclear reactors are strong, pure and well localized sources of electron antineutrinos with energies in the few MeV range. Therefore they provide a suitable environment to study neutrino properties, in particular neutrino oscillation parameters. Recent predictions of the expected antineutrino flux at nuclear reactors are about 6% higher than the average rate measured in different experiments. This discrepancy, known as the reactor antineutrino anomaly, is significant at the 2.5σ level. Several new experiments are searching for the origin of this observed neutrino deficit. One hypothesis to be tested is an oscillation to another neutrino state. In a three flavor model reactor neutrinos do not oscillate at baselines below 100 m. Hence, if such an oscillation is observed, it would imply the existence of at least one light sterile neutrino state not participating in weak interactions. Such a discovery would open the gate for new physics beyond the Standard Model.

hep-ex

Investigating the Spectral Anomaly with Different Reactor Antineutrino Experiments

The spectral shape of reactor antineutrinos measured in recent experiments shows anomalies in comparison to neutrino reference spectra. New precision measurements of the reactor neutrino spectra as well as more complete input in nuclear data bases are needed to resolve the observed discrepancies between models and experimental results. This article proposes the combination of experiments at reactors which are highly enriched in ${}^{235}$U with commercial reactors with typically lower enrichment to gain new insights into the origin of the anomalous neutrino spectrum. The presented method clarifies, if the spectral anomaly is either solely or not at all related to the predicted ${}^{235}$U spectrum. Considering the current improvements of the energy scale uncertainty of present-day experiments, a significance of three sigma and above can be reached. As an example, we discuss the option of a direct comparison of the measured shape in the currently running Double Chooz near detector and the upcoming Stereo experiment. A quantitative feasibility study emphasizes that a precise understanding of the energy scale systematics is a crucial prerequisite in recent and next generation experiments investigating the spectral anomaly.

hep-ex

Metal-loaded organic scintillators for neutrino physics

Organic liquid scintillators are used in many neutrino physics experiments of the past and present. In particular for low energy neutrinos when realtime and energy information are required, liquid scintillators have several advantages compared to other technologies. In many cases the organic liquid needs to be loaded with metal to enhance the neutrino signal over background events. Several metal loaded scintillators of the past suffered from chemical and optical instabilities, limiting the performance of these neutrino detectors. Different ways of metal loading are described in the article with a focus on recent techniques providing metal loaded scintillators that can be used under stable conditions for many years even in ton scale experiments. Applications of metal loaded scintillators in neutrino experiments are reviewed and the performance as well as the prospects of different scintillator types are compared.

physics.ins-det

Prototype scintillator cell for an In-based solar neutrino detector

We describe the work carried out at MPIK to design, model, build and characterize a prototype cell filled with a novel indium-loaded scintillator of interest for real-time low energy solar neutrino spectroscopy. First, light propagation in optical modules was studied with experiments and Monte Carlo simulations. Subsequently a 5 cm x 5 cm x 100 cm prototype detector was set up and the optical performances of several samples were measured. We first tested a benchmark PXE-based scintillator, which performed an attenuation length of ~ 4.2 m and a photo-electron yield of ~ 730 pe/MeV. Then we measured three In-loaded samples. At an In-loading of 44 g/l, an energy resolution of ~ 11.6 % and a spatial resolution of ~ 7 cm were attained for 477 keV recoil electrons. The long-range attenuation length in the cell was ~1.3 m and the estimated photo-electron yield ~ 200 pe/MeV. Light attenuation and relative light output of all tested samples could be reproduced reasonably well by MC. All optical properties of this system have remained stable over a period of > 1 y.

physics.ins-det