TY - RPRT TI - Benchmarking and Mechanistic Analysis of Vision-Language Models for Cross-Depiction Assembly Instruction Alignment AU - Zhuchenyang Liu AU - Yao Zhang AU - Yu Xiao PY - 2026 UR - https://arxiv.org/abs/2604.00913 ID - 2604.00913 ER -