TY - RPRT TI - Aligning where to see and what to tell: image caption with region-based attention and scene factorization AU - Junqi Jin AU - Kun Fu AU - Runpeng Cui AU - Fei Sha AU - Changshui Zhang PY - 2015 UR - https://arxiv.org/abs/1506.06272 ID - 1506.06272 ER -