arXiv · 2504.21605
RDF-Based Structured Quality Assessment Representation of Multilingual LLM Evaluations
Abstract
Large Language Models (LLMs) increasingly serve as knowledge interfaces, yet systematically assessing their reliability with conflicting information remains difficult. We propose an RDF-based framework to assess multilingual LLM quality, focusing on knowledge conflicts. Our approach captures model responses across four distinct context conditions (complete, incomplete, conflicting, and no-context information) in German and English. This structured representation enables the comprehensive analysis of knowledge leakage-where models favor training data over provided context-error detection, and multilingual consistency. We demonstrate the framework through a fire safety domain experiment, revealing critical patterns in context prioritization and language-specific performance, and demonstrating that our vocabulary was sufficient to express every assessment facet encountered in the 28-question study.
Explore related subjects
Keep this discovery
Jonas Gwozdz, Andreas Both. 2025-04-30. RDF-Based Structured Quality Assessment Representation of Multilingual LLM Evaluations. https://arxiv.org/abs/2504.21605
Cite the original work for its findings. Save a collection to share your selection of sources.