Publiora

Menghubungkan ke Publiora...

Publiora

Video captioning in Vietnamese using deep learning

Phuc, Dang ThiTrieu, Tran QuangTinh, Nguyen VanHieu, Dau Sy
International Journal of Electrical and Computer Engineering (IJECE) (Sinta 1)Vol. 0 No. 01 Juni 2022
DOI10.11591/ijece.v12i3.pp3092-3103

Abstrak

With the development of today's society, demand for applications using digital cameras jumps over year by year. However, analyzing large amounts of video data causes one of the most challenging issues. In addition to storing the data captured by the camera, intelligent systems are required to quickly analyze the data to correct important situations. In this paper, we use deep learning techniques to build automatic models that describe movements on video. To solve the problem, we use three deep learning models: sequence-to-sequence model based on recurrent neural network, sequence-to-sequence model with attention and transformer model. We evaluate the effectiveness of the approaches based on the results of three models. To train these models, we use microsoft research video description corpus (MSVD) dataset including 1970 videos and 85,550 captions translated into Vietnamese. In order to ensure the description of the content in Vietnamese, we also combine it with the natural language processing (NLP) model for Vietnamese.

Kata Kunci

attentionnatural language processingsequence-to-sequence modeltransformervideo caption

Cari jurnal yang tepat untuk naskah Anda

MatchMind AI mencocokkan abstrak naskah Anda dengan ribuan jurnal terakreditasi dan menampilkan rekomendasi terbaik beserta alasannya.

Coba MatchMind

Lihat profil lengkap jurnal ini

Waktu review, biaya APC, statistik sitasi, indeksasi Scopus, dan banyak lagi.

Buka International Journal of Electrical and Computer Engineering (IJECE)

Artikel ini juga tersedia di situs resmi jurnal.

Video captioning in Vietnamese using deep learning | International Journal of Electrical and Computer Engineering (IJECE) | Publiora