Publiora

Menghubungkan ke Publiora...

Publiora

Transformer+transformer architecture for image captioning in Indonesian language

Wijaya, Bryan ChristoferSugiarto, Hendrik Santoso
IAES International Journal of Artificial Intelligence (IJ-AI) (Sinta 1)Vol. 0 No. 01 Juni 2025
DOI10.11591/ijai.v14.i3.pp2338-2346

Abstrak

Image captioning in Indonesian language poses a significant challenge due to the complex interplay between visual and linguistic comprehension, as well as the scarcity of publicly available datasets. Despite considerable advancements in this field, research specifically targeting the Indonesian language remains scarce. In this paper, we propose a novel image captioning model employing a transformer-based architecture for both the encoder and decoder components. Our model is trained and evaluated on the pre-translated Flickr30k dataset in the Indonesian language. We conduct a comparative analysis of various transformertransformer configurations and convolutional neural network (CNN)-recurrent neural network (RNN) architectures. Our findings highlight the superior performance of a vision transformer (ViT) as the visual encoder, combined with IndoBERT as the textual decoder. This architecture achieved a BLEU-4 score of 0.223 and a ROUGE-L score of 0.472.

Kata Kunci

Deep Learning

Cari jurnal yang tepat untuk naskah Anda

MatchMind AI mencocokkan abstrak naskah Anda dengan ribuan jurnal terakreditasi dan menampilkan rekomendasi terbaik beserta alasannya.

Coba MatchMind

Lihat profil lengkap jurnal ini

Waktu review, biaya APC, statistik sitasi, indeksasi Scopus, dan banyak lagi.

Buka IAES International Journal of Artificial Intelligence (IJ-AI)

Artikel ini juga tersedia di situs resmi jurnal.

Transformer+transformer architecture for image captioning in Indonesian language | IAES International Journal of Artificial Intelligence (IJ-AI) | Publiora