Bengali Visual Genome is a multimodal dataset consisting of text and images suitable for English-to-Bengali multimodal machine translation task and multimodal research. We have selected short English segments (captions) from Visual Genome along with associated images and manually translated them to Bengali by Bengali native speaker, taking the associated images into account.

The training set contains 29K segments. Further 1K and 1.6K segments are provided in a development and test sets, respectively, which follow the same (random) sampling from the original Hindi Visual Genome. Additionally, a challenge test set of 1400 segments was prepared for the WAT2022 multi-modal task. This challenge test set was created by searching for (particularly) ambiguous English words based on the embedding similarity and manually selecting those where the image helps to resolve the ambiguity. The surrounding words in the sentence however also often include sufficient cues to identify the correct meaning of the ambiguous word.


Dataset Sample

Sample items from the randomly selected segments (train/dtest/etest)

Image Image_ID X Y Width Height English Text Bengali Text
2323457 2323457 20 150 325 121 Many giraffes at a zoo একটি চিড়িয়াখানায় অনেক জিরাফ
2335684 2335684 61 191 437 182 Fruit stand outside market বাজারের বাইরে ফলের স্ট্যান্ড

Sample item from the challenge test set (chtest)

2372733 2372733 26 107 152 218 The tennis court is made up of sand and dirt টেনিস কোর্ট বালু এবং ময়লা দিয়ে গঠিত

Sample ambiguous word: court (the key different meanings are court of justice vs. tennis court). As the image example illustrates, the word alone is ambiguous but already the English caption may contain sufficient information to disambiguate the word; the image is thus not always necessary for correct translation even in our challenge test set.

Download Link


Paper and References

Please refer to the below papers:


Bengali Visual Genome: A Multimodal Dataset for Machine Translation and Image Captioning


[Reference Papers]

Multimodal Neural Machine Translation System for English to Bengali

How to cite

If you use this corpus, please cite the following paper:

How to cite

  title={Bengali Visual Genome: A Multimodal Dataset for Machine Translation and Image Captioning},
  author={Sen, Arghyadeep and Parida, Shantipriya and Kotwal, Ketan and Panda, Subhadarshi and Bojar, Ond{\v{r}}ej and Dash, Satya Ranjan},
  booktitle={Intelligent Data Engineering and Analytics},