Multi-Legal-Bench: When the Answer Is in the Input. Label Leakage in Legal Benchmarks Built from Court Registries
Court registries publish millions of decisions with structured metadata, which makes them an attractive source of labelled legal benchmarks: the court type, the form of the decision and its subject area come for free. We show that this convenience has a cost that registry-derived benchmarks rarely measure. We build Multi-Legal-Bench, which evaluates identical tasks on native court decisions from five national registries (France, the Netherlands, Poland, the Czech Republic and Lithuania), extending the Ukrainian UA-Legal-Bench, and audit what its cells actually measure. A keyword scan that uses no model reaches 96% on Dutch judgment-form classification against a 50% majority baseline, 84% on French court-type classification against 33%, and 66% on Polish judgment-form classification, where all nine models add at most seven points over it. Replacing the label names in the input by a mask brings the scan to the majority floor in all six affected cells, and eight of nine models lose accuracy significantly in most of them (39 of 54 model-cell pairs), by up to 48 points (Dutch judgment form: 100% to 52%); only Claude Sonnet 5 is essentially unaffected. Paired McNemar tests with Holm correction find a leader that beats every other model in only two of twelve cells, each a different model; but once the label names are masked, the same model leads both, and the spread between models widens in every masked cell. Part of the apparent parity between models is produced by the leakage itself. The audit also exposed a scoring defect in our own earlier release (answers written with diacritics were scored as wrong, understating Czech judgment-form accuracy by up to 90 points), which we correct and document. We recommend that every benchmark built from court registries report a no-model baseline and a masked control per cell. All data, prompts, predictions and scoring code are released.