arXiv · 2212.14322
BagFormer: Better Cross-Modal Retrieval via bag-wise interaction
Abstract
In the field of cross-modal retrieval, single encoder models tend to perform better than dual encoder models, but they suffer from high latency and low throughput. In this paper, we present a dual encoder model called BagFormer that utilizes a cross modal interaction mechanism to improve recall performance without sacrificing latency and throughput. BagFormer achieves this through the use of bag-wise interactions, which allow for the transformation of text to a more appropriate granularity and the incorporation of entity knowledge into the model. Our experiments demonstrate that BagFormer is able to achieve results comparable to state-of-the-art single encoder models in cross-modal retrieval tasks, while also offering efficient training and inference with 20.72 times lower latency and 25.74 times higher throughput.
Explore related subjects
Keep this discovery
Haowen Hou, Xiaopeng Yan, Yigeng Zhang, Fengzong Lian, Zhanhui Kang. 2022-12-29. BagFormer: Better Cross-Modal Retrieval via bag-wise interaction. https://arxiv.org/abs/2212.14322
Cite the original work for its findings. Save a collection to share your selection of sources.