TY - RPRT TI - Benchmarking Multi-Image Understanding in Vision and Language Models: Perception, Knowledge, Reasoning, and Multi-Hop Reasoning AU - Bingchen Zhao AU - Yongshuo Zong AU - Letian Zhang AU - Timothy Hospedales PY - 2024 UR - https://arxiv.org/abs/2406.12742 ID - 2406.12742 ER -