arXiv · 2510.07665
Automatic Text Box Placement for Supporting Typographic Design
Abstract
In layout design for advertisements and web pages, balancing visual appeal and communication efficiency is crucial. This study examines automated text box placement in incomplete layouts, comparing a standard Transformer-based method, a small Vision and Language Model (Phi3.5-vision), a large pretrained VLM (Gemini), and an extended Transformer that processes multiple images. Evaluations on the Crello dataset show the standard Transformer-based models generally outperform VLM-based approaches, particularly when incorporating richer appearance information. However, all methods face challenges with very small text or densely populated layouts. These findings highlight the benefits of task-specific architectures and suggest avenues for further improvement in automated layout design.
Explore related subjects
Keep this discovery
Jun Muraoka, Daichi Haraguchi, Naoto Inoue, Wataru Shimoda, Kota Yamaguchi, Seiichi Uchida. 2025-10-09. Automatic Text Box Placement for Supporting Typographic Design. https://arxiv.org/abs/2510.07665
Cite the original work for its findings. Save a collection to share your selection of sources.