arXiv · 2503.05298
Coreference as an indicator of context scope in multimodal narrative
Abstract
We demonstrate that large multimodal language models differ substantially from humans in the distribution of coreferential expressions in a visual storytelling task. We introduce a number of metrics to quantify the characteristics of coreferential patterns in both human- and machine-written texts. Humans distribute coreferential expressions in a way that maintains consistency across texts and images, interleaving references to different entities in a highly varied way. Machines are less able to track mixed references, despite achieving perceived improvements in generation quality. Materials, metrics, and code for our study are available at https://github.com/GU-CLASP/coreference-context-scope.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Nikolai Ilinykh, Shalom Lappin, Asad Sayeed, Sharid Loáiciga. 2025-03-07. Coreference as an indicator of context scope in multimodal narrative. https://arxiv.org/abs/2503.05298
Cite the original work for its findings. Save a collection to share your selection of sources.