Publiora

Menghubungkan ke Publiora...

Publiora

Information Extraction from Makassar Culinary Images Using Vision Transformers and Cahya GPT-2 (Visual Question Answering Case Study

Tirta Chiantalia Sharief*Universitas Atma Jaya MakassarHazrianiUniversitas Atma Jaya MakassarSyamsulUniversitas Atma Jaya MakassarAnasUniversitas Atma Jaya MakassarYuyunUniversitas Atma Jaya Makassar
Indonesian Journal of Data and Science (Sinta 3)Vol. 6 No. 3 (2025)31 Desember 2025hal. 433-444
DOI10.56705/ijodas.v6i3.357

Abstrak

This study examines the development of a Visual Question Answering (VQA) system to extract information from images of Makassar culinary specialties by combining the Vision Transformer (ViT) and Cahya_GPT-2 models. The main objective is to integrate visual and natural language understanding so that computers can recognize visual objects (food images) and generate relevant text descriptions. The research method uses an experimental approach with a fine-tuning process of the pre-trained ViT model as a visual encoder and Cahya_GPT-2 as a text decoder. The dataset used includes images of Makassar culinary specialties such as Coto, Konro, Pisang Epe, Barongko, and Jalangkote with question and answer (QnA) annotations. Evaluation is carried out using the ROUGE metric to assess the semantic match between the model's answers and the actual answers. The results show that the developed multimodal model is able to accurately understand the image context with an average ROUGE-L score of 0.63, indicating a good level of closeness between the model's answers and the annotations. In conclusion, the combination of ViT and Cahya_GPT-2 can be an effective approach for natural language-based visual information extraction systems, especially in the Indonesian local culinary domain

Cari jurnal yang tepat untuk naskah Anda

MatchMind AI mencocokkan abstrak naskah Anda dengan ribuan jurnal terakreditasi dan menampilkan rekomendasi terbaik beserta alasannya.

Coba MatchMind

Lihat profil lengkap jurnal ini

Waktu review, biaya APC, statistik sitasi, indeksasi Scopus, dan banyak lagi.

Buka Indonesian Journal of Data and Science

Artikel ini juga tersedia di situs resmi jurnal.

Information Extraction from Makassar Culinary Images Using Vision Transformers and Cahya GPT-2 (Visual Question Answering Case Study | Indonesian Journal of Data and Science | Publiora