SearcharxivSearch

arXiv subjects

Matteo Greco

Publications and source records attributed to Matteo Greco.

4 recordsLinked to original sources

MultiGhostBench: A Multilingual Benchmark for Long-Form LLM-Generated Text Attribution under Distribution Shifts

While existing work on LLM authorship attribution (AA) has made progress, available benchmarks remain limited, often focusing on English, controlled settings, or relatively outdated models, with the few multilingual studies considering only relatively short texts. We introduce MultiGhostBench, a multilingual benchmark comprising 928 books generated by five recent LLMs across six languages and three scripts, with an average length of approximately 59K words per book. The benchmark supports evaluation under domain, author, and language shifts. Evaluation of representative AA methods shows that no single method consistently performs best across settings, and performance generally degrades under distribution shifts. Transformer-based detectors can retain generator-related information across languages, although transfer effectiveness varies by language pair, whereas statistical and fingerprint-based detectors are more language-dependent. We envision MultiGhostBench as a valuable resource for the development and evaluation of robust AA methods. The dataset and code can be found at https://github.com/GrecoMT/MultiGhostBench.

cs.CL

Searches for strong production of supersymmetric particles with the ATLAS detector

Supersymmetry (SUSY) provides elegant solutions to several open questions in the Standard Model, and searches for SUSY particles are an important component of the LHC physics program. Naturalness arguments favour supersymmetric partners of the gluons and third-generation quarks with masses light enough to be produced at the LHC. With increasing mass bounds on more classical Minimal Supersymmetric Standard Model (MSSM) scenarios other variations of supersymmetry, including non-minimal particle content, become increasingly interesting. This proceeding will present the latest results of searches conducted by the ATLAS experiment at LHC at center of mass energies of $\sqrt{s}=13$ and 13.6 TeV which target gluino and squark production, including stop, in a variety of decay modes.

hep-ex

False perspectives on human language: why statistics needs linguistics

A sharp tension exists about the nature of human language between two opposite parties: those who believe that statistical surface distributions, in particular using measures like surprisal, provide a better understanding of language processing, vs. those who believe that discrete hierarchical structures implementing linguistic information such as syntactic ones are a better tool. In this paper, we show that this dichotomy is a false one. Relying on the fact that statistical measures can be defined on the basis of either structural or non-structural models, we provide empirical evidence that only models of surprisal that reflect syntactic structure are able to account for language regularities.

cs.CL

Particle identification with the cluster counting technique for the IDEA drift chamber

IDEA (Innovative Detector for an Electron-positron Accelerator) is a general-purpose detector concept, designed to study electron-positron collisions in a wide energy range from a very large circular leptonic collider. Its drift chamber is designed to provide an efficient tracking, a high precision momentum measurement and an excellent particle identification by exploiting the application of the cluster counting technique. To investigate the potential of the cluster counting techniques on physics events, a simulation of the ionization clusters generation is needed, therefore we developed an algorithm which can use the energy deposit information provided by Geant4 toolkit to reproduce, in a fast and convenient way, the clusters number distribution and the cluster size distribution. The results obtained confirm that the cluster counting technique allows to reach a resolution 2 times better than the traditional dE/dx method. A beam test has been performed during November 2021 at CERN on the H8 to validate the simulations results, to define the limiting effects for a fully efficient cluster counting and to count the number of electron clusters released by an ionizing track at a fixed $\beta\gamma$ as a function of the track angle. The simulation and the beam test results will be described briefly in this issue.

hep-ex