arXiv · 2402.05106
Image captioning for Brazilian Portuguese using GRIT model
Abstract
This work presents the early development of a model of image captioning for the Brazilian Portuguese language. We used the GRIT (Grid - and Region-based Image captioning Transformer) model to accomplish this work. GRIT is a Transformer-only neural architecture that effectively utilizes two visual features to generate better captions. The GRIT method emerged as a proposal to be a more efficient way to generate image captioning. In this work, we adapt the GRIT model to be trained in a Brazilian Portuguese dataset to have an image captioning method for the Brazilian Portuguese Language.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Rafael Silva de Alencar, William Alberto Cruz Castañeda, Marcellus Amadeus. 2024-02-07. Image captioning for Brazilian Portuguese using GRIT model. https://arxiv.org/abs/2402.05106
Cite the original work for its findings. Save a collection to share your selection of sources.