SearcharxivSearch

arXiv subjects

Georgi Georgiev

Publications and source records attributed to Georgi Georgiev.

At least 19 recordsLinked to original sources

Overview of FinMMEval 2026 Task 1: Multilingual Financial Multiple-Choice Question Answering

FinMMEval 2026 Task 1 evaluates multilingual financial multiple-choice question answering in English, Chinese, Arabic, and Hindi. The task tests whether systems can select the correct answer to finance questions involving domain terminology, numerical interpretation, and conceptual financial reasoning across languages and scripts. The final-test set contains 800 questions, with 200 questions per language; gold answers were withheld during submission, and each language was ranked independently by accuracy. The final leaderboards contain 13 English, 11 Chinese, 11 Arabic, and 10 Hindi ranked submissions. Top accuracies range from 92.0% in Hindi to 97.5% in English and Arabic, with the same leading teams appearing near the top across all four languages. The documented systems used retrieval augmentation, direct answer-option scoring, language-specific prompting, selective self-consistency, confidence checks, and LLM-based review stages.

cs.CL

Overview of FinMMEval 2026 Task 2: Multilingual Financial Short-Answer Question Answering

FinMMEval 2026 Task 2 evaluates short-answer financial question answering over multilingual evidence. Each final-test item pairs an English question with financial statements and news in English, Chinese, Japanese, Spanish, and Greek. Participating systems submit one concise answer per item in JSONL format. The final-test set contains 256 items, split evenly between easy and expert tiers; each tier contains four question templates instantiated over 32 company-report groups. Gold answers were withheld during submission, and systems were ranked by macro-averaged item-level ROUGE-1 F1 against organizer-held reference answers. The final leaderboard includes 12 ranked submissions. The strongest systems are closely clustered, with the top four separated by less than one percentage point in ROUGE-1 F1. The submitted system papers document retrieval-augmented generation, cross-lingual evidence handling, structured prompting, answer compression, and validation strategies.

cs.CL

Nuclear electromagnetic moments by spin-precession methods

Nuclear moment studies carried out with spin-precession methods at and after the turn of the millennium are critically assessed. A period of about 30 years is covered, during which much of} the focus of nuclear structure research shifted from high-spin physics to studies of neutron-rich exotic nuclei. The formalism for the extraction of nuclear moments is described. The $\beta$-nuclear magnetic resonance/nuclear quadrupole resonance ($\beta$-NMR/NQR), the time-dependent perturbed angular distribution (TDPAD), the transient field, the recoil-in-vacuum (RIV), and the tilted-foils methods for measurements of nuclear magnetic dipole and electric quadrupole moments are described in detail, as well as the requirements for their application in studies of exotic nuclei. The impact of nuclear-moment measurements on the understanding of key topics of nuclear structure research is discussed. {Key results on short-lived states, mainly from transient-field measurements, are reviewed. Included are comparisons with large-basis shell model calculations, discussions on the nature of weakly-collective nuclei, insights into emerging collectivity away from closed shells, and electromagnetic properties of odd-$A$ rotors.} In the field of high-spin physics, research related to high-spin yrast and $\mathrm{K}$ isomers, superdeformation, magnetic, anti-magnetic, and chiral rotation is covered. In neutron-rich exotic nuclei, studies related to the $\mathrm{N=20}$, $\mathrm{N=28}$ and $\mathrm{N=40}$ ``islands of inversion'', the structure of nuclei around $^{68-78}$Ni and $^{132}$Sn, and in the $A \sim 100$ mass region are discussed.

nucl-ex

Towards better nuclear charge radii

Nuclear charge radii constitute a physical observable of growing significance across multiple subdisciplines of physics and related fields. Their determination relies on a combination of complementary experimental techniques and advanced theoretical frameworks. Current recommended values are informed by the outcomes of several independent working groups, each employing distinct methodological approaches and evaluation strategies. The present effort is directed toward a more precise and reliable extraction of charge radii, as well as the development of a modern, transparent, and methodologically robust compilation of recommended values.

nucl-ex

FinReporting: An Agentic Workflow for Localized Reporting of Cross-Jurisdiction Financial Disclosures

Financial reporting systems increasingly leverage Large Language Models (LLMs) to extract and summarize corporate disclosures. However, most existing approaches assume a single-market setting and overlook structural differences across jurisdictions. Variations in accounting taxonomies, tagging infrastructures (e.g., XBRL vs.\ PDF), and aggregation conventions introduce substantial challenges for semantic alignment and reliable verification. Here, we aim to bridge this gap. We present FinReporting, an agentic workflow for localized cross-jurisdiction financial reporting. The system constructs a unified canonical ontology spanning the income statement, balance sheet, and cash flow statement, and decomposes reporting into auditable stages, including filing acquisition, extraction, canonical mapping, and anomaly logging. Rather than treating LLMs as free-form generators, FinReporting employs them as constrained verifiers operating under explicit decision rules with evidence grounding. Evaluated on annual filings from the USA, Japan, and China, FinReporting improves consistency and reliability under heterogeneous reporting regimes. We further release an interactive demo that enables cross-market inspection and supports structured export of localized financial statements. Our demo is available at url{https://huggingface.co/spaces/BoomQ/FinReporting-Demo. A video describing our system is available at https://www.youtube.com/watch?v=f65jdEL31Kk.

cs.CL

The CLEF-2026 FinMMEval Lab: Multilingual and Multimodal Evaluation of Financial AI Systems

We present the setup and the tasks of the FinMMEval Lab at CLEF 2026, which introduces the first multilingual and multimodal evaluation framework for financial Large Language Models (LLMs). While recent advances in financial natural language processing have enabled automated analysis of market reports, regulatory documents, and investor communications, existing benchmarks remain largely monolingual, text-only, and limited to narrow subtasks. FinMMEval 2026 addresses this gap by offering three interconnected tasks that span financial understanding, reasoning, and decision-making: Financial Exam Question Answering, Multilingual Financial Question Answering (PolyFiQA), and Financial Decision Making. Together, these tasks provide a comprehensive evaluation suite that measures models' ability to reason, generalize, and act across diverse languages and modalities. The lab aims to promote the development of robust, transparent, and globally inclusive financial AI systems, with datasets and evaluation resources publicly released to support reproducible research.

cs.CL

Determination of the fifth Busy Beaver value

The Busy Beaver value $S(n)$ is the maximum number of steps that an $n$-state 2-symbol Turing machine can perform from the all-zero tape before halting. $S$ was historically introduced by Tibor Rad\'o in 1962 as one of the simplest examples of an uncomputable function. We prove that $S(5) = 47,176,870$ using the Coq proof assistant. The proof enumerates $181,385,789$ Turing machines with 5 states and, for each machine, decides whether it halts or not. Our result marks the first determination of a new Busy Beaver value in over 40 years and the first Busy Beaver value ever to be formally verified, attesting to the effectiveness of massively collaborative online research (bbchallenge$.$org).

cs.LO

FinChain: A Symbolic Benchmark for Verifiable Chain-of-Thought Financial Reasoning

Multi-step symbolic reasoning is essential for robust financial analysis; yet, current benchmarks largely overlook this capability. Existing datasets such as FinQA and ConvFinQA emphasize final numerical answers while neglecting the intermediate reasoning steps required for transparency and verification. To address this gap, we introduce FinChain, the first benchmark specifically designed for verifiable Chain-of-Thought evaluation in finance. FinChain spans 58 topics across 12 financial domains, each represented by parameterized symbolic templates with executable Python code that enable fully machine-verifiable reasoning and scalable, contamination-free data generation. To assess reasoning capacity, we propose CHAINEVAL, a dynamic alignment measure that jointly evaluates both the final-answer correctness and the step-level reasoning consistency. Our evaluation of 26 leading LLMs reveals that even frontier LLMs exhibit clear limitations in symbolic financial reasoning, while domain-adapted and math-enhanced fine-tuned models can substantially narrow this gap. Overall, FinChain exposes persistent weaknesses in multi-step financial reasoning and provides a foundation for developing trustworthy, interpretable, and verifiable financial AI. This project is available at https://github.com/mbzuai-nlp/finchain.git.

cs.CL

Legendre Functions and the Non-Integrability of a Hamiltonian System

In this paper we are studying the meromorphic integrability of a two-dimensional Hamiltonian system with a homogeneous potential of degree 6. The approach used in this work is the theory of the Ziglin-Moralez-Ruiz-Ramis-Simo. Within the scope of this theory, the study of such systems is reduced to determining the differential Galois group of a linear differential equation, obtained as a projection onto the tangent bundle of the phase curve of its non-equilibrium solution - Variation Equations (VE). In the case of Hamiltonian systems with homogeneous potentials, the variation equations are hypergeometric. If a standard approach is used to study such a system, it is necessary to calculate a Darboux point, which is not always easy. In this paper we can skip this difficulty by reducing VE to a Legendre equation. We use the results for commutativity of the Galois group of the associated Legendre equation for study a Hamiltonian system with a homogeneous polynomial potential of degree 6. The approach is different and answers are sought as to what exactly is happening in the gray areas of the classical results. For the full study, the second variations are used and conditions for a non-zero logarithmic term in their solutions are found. This is exactly the case when VE is solvable, but the unit component of the Galois group is not commutative.

math.DS

OpenFactCheck: A Unified Framework for Factuality Evaluation of LLMs

The increased use of large language models (LLMs) across a variety of real-world applications calls for automatic tools to check the factual accuracy of their outputs, as LLMs often hallucinate. This is difficult as it requires assessing the factuality of free-form open-domain responses. While there has been a lot of research on this topic, different papers use different evaluation benchmarks and measures, which makes them hard to compare and hampers future progress. To mitigate these issues, we developed OpenFactCheck, a unified framework, with three modules: (i) RESPONSEEVAL, which allows users to easily customize an automatic fact-checking system and to assess the factuality of all claims in an input document using that system, (ii) LLMEVAL, which assesses the overall factuality of an LLM, and (iii) CHECKEREVAL, a module to evaluate automatic fact-checking systems. OpenFactCheck is open-sourced (https://github.com/mbzuai-nlp/openfactcheck) and publicly released as a Python library (https://pypi.org/project/openfactcheck/) and also as a web service (http://app.openfactcheck.com). A video describing the system is available at https://youtu.be/-i9VKL0HleI.

cs.CL

OpenFactCheck: Building, Benchmarking Customized Fact-Checking Systems and Evaluating the Factuality of Claims and LLMs

The increased use of large language models (LLMs) across a variety of real-world applications calls for mechanisms to verify the factual accuracy of their outputs. Difficulties lie in assessing the factuality of free-form responses in open domains. Also, different papers use disparate evaluation benchmarks and measurements, which renders them hard to compare and hampers future progress. To mitigate these issues, we propose OpenFactCheck, a unified framework for building customized automatic fact-checking systems, benchmarking their accuracy, evaluating factuality of LLMs, and verifying claims in a document. OpenFactCheck consists of three modules: (i) CUSTCHECKER allows users to easily customize an automatic fact-checker and verify the factual correctness of documents and claims, (ii) LLMEVAL, a unified evaluation framework assesses LLM's factuality ability from various perspectives fairly, and (iii) CHECKEREVAL is an extensible solution for gauging the reliability of automatic fact-checkers' verification results using human-annotated datasets. Data and code are publicly available at https://github.com/yuxiaw/openfactcheck.

cs.CL

Factuality of Large Language Models: A Survey

Large language models (LLMs), especially when instruction-tuned for chat, have become part of our daily lives, freeing people from the process of searching, extracting, and integrating information from multiple sources by offering a straightforward answer to a variety of questions in a single place. Unfortunately, in many cases, LLM responses are factually incorrect, which limits their applicability in real-world scenarios. As a result, research on evaluating and improving the factuality of LLMs has attracted a lot of attention recently. In this survey, we critically analyze existing work with the aim to identify the major challenges and their associated causes, pointing out to potential solutions for improving the factuality of LLMs, and analyzing the obstacles to automated factuality evaluation for open-ended text generation. We further offer an outlook on where future research should go.

cs.CL

A deconvolution based signal reconstruction capable of piled-up pulse separation

This study provides a computationally effective deconvolution algorithm capable to reconstruct piled-up events in scintillating detector systems with high count rate where fully digitized waveforms are available. A fixed-point iteration algorithm is suggested and used to find properties of the signals which are later used during the signal preprocessing stage. The impulse response function is successfully extracted even from heavily piled-up event waveforms using an iterative approach. A methodology for pulse time and amplitude reconstruction is based on a deconvolution algorithm, which is described in details and some results are presented. The presented algorithms are meant to be general and might be successfully applied to other fields with minor to no modifications.

physics.ins-det

Status and Prospects of PADME

The Positron Annihilation to Dark Matter Experiment (PADME) was designed and constructed to search for dark photons ($A'$) in the process $e^+e^-\rightarrowγA'$, using the positron beam at the Beam Test Facility (BTF) at the National Laboratories of Frascati (LNF). Since the observation of an anomalous spectra in internal pair creation decays of nuclei seen by the collaboration at the ATOMKI institute, the PADME detector has been modified and a new data-taking run has been undertaken to probe the existance of the so-called ``X17" particle

hep-ex

Dark sector studies with the PADME experiment

The Positron Annihilation to Dark Matter Experiment (PADME) uses the positron beam of the DA$Φ$NE Beam-Test Facility, at the Laboratori Nazionali di Frascati (LNF) to search for a Dark Photon $A'$. The search technique studies the missing mass spectrum of single-photon final states in $e^+e^-\rightarrow A'γ$ annihilation in a positron-on-thin-target experiment. This approach facilitates searches for new particles such as long lived Axion-Like-Particles, protophobic X bosons and Dark Higgs. This talk illustrated the scientific program of the experiment and its first physics results. In particular, the measurement of the cross-section of the SM process $e^+e^-\rightarrow γγ$ at $\sqrt{s}$=21 MeV was shown.

hep-ex

Non-Integrability of the Trapped Ionic System II

In this paper we explore the two dimensional system describing trapped ionic system in the quadrapole field with a superposition of rationally symmetric hexapole and octopole fields for meromorphic integrability. We use the Lyapunov's and Ziglin-Morales-Ramis classical methods for the proofs. The inaccuracies from the previous such paper have been removed.

math.DS

A Model of Random Multiple Access in Unlicensed Spectrum Systems

We consider a classical multiple access system with a single transmission channel, finite number of users (users), and randomized transmission protocol (ALOHA). We assume that every user sends messages to the base station with various intensities. Due to the overlapping of messages during their sending there are restrictions on the time between the messages and a mathematical model adequate to the physical one becomes quite complicated. In this work, we propose a simplified mathematical model that is easy to analyze and takes into account the properties of real systems.

cs.IT