arXiv · 2104.14721
End-to-End Attention-based Image Captioning
Abstract
In this paper, we address the problem of image captioning specifically for molecular translation where the result would be a predicted chemical notation in InChI format for a given molecular structure. Current approaches mainly follow rule-based or CNN+RNN based methodology. However, they seem to underperform on noisy images and images with small number of distinguishable features. To overcome this, we propose an end-to-end transformer model. When compared to attention-based techniques, our proposed model outperforms on molecular datasets.
Explore related subjects
Keep this discovery
Carola Sundaramoorthy, Lin Ziwen Kelvin, Mahak Sarin, Shubham Gupta. 2021-04-30. End-to-End Attention-based Image Captioning. https://arxiv.org/abs/2104.14721
Cite the original work for its findings. Save a collection to share your selection of sources.