arXiv · 2601.16946
Strategies for Span Labeling with Large Language Models
Abstract
Large language models (LLMs) are increasingly used for text analysis tasks, such as named entity recognition or error detection. Unlike encoder-based models, however, generative architectures lack an explicit mechanism to refer to specific parts of their input. This leads to a variety of ad-hoc prompting strategies for span labeling, often with inconsistent results. In this paper, we categorize these strategies into three families: tagging the input text, indexing numerical positions of spans, and matching span content. To address the limitations of content matching, we introduce LogitMatch, a new constrained decoding method that forces the model's output to align with valid input spans. We evaluate all methods across four diverse tasks. We find that while tagging remains a robust baseline, LogitMatch improves upon competitive matching-based methods by eliminating span matching issues and outperforms other strategies in some setups.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Danil Semin, Ondřej Dušek, Zdeněk Kasner. 2026-01-23. Strategies for Span Labeling with Large Language Models. https://arxiv.org/abs/2601.16946
Cite the original work for its findings. Save a collection to share your selection of sources.