BanSpEmo: a Bangla audio dataset for speech emotion recognition and its baseline evaluation
Abstrak
Speech interfaces provide a natural and comfortable way for humans to communicate with machines. Recognizing emotions from acoustic signals is essential in audio and speech processing. Detection of emotion in speech is critical to the next generation of human-computer interaction (HCI) fields. However, a lack of large-scale datasets has hampered the progress of relevant research. In this study, we prepare BANSpEmo, a demanding Bangla speech emotion dataset consisting of 792 audio recordings totaling more than 1 hour and 23 minutes. The recordings feature 22 native speakers and each speaker uttered two sets of sentences representing six emotions: disgust, happiness, anger, sadness, surprise, and fear. The dataset consists of 12 Bangla sentences, each expressed in these six emotions. Furthermore, a series of investigations are carried out to assess the baseline performance of the support vector machine (SVM), logistic regression (LR), and multinomial Naive Bayes models on the BANSpEmo dataset presented in this study. The studies found that SVM performed best on this dataset, with an accuracy of 87.18%.
Kata Kunci
Cari jurnal yang tepat untuk naskah Anda
MatchMind AI mencocokkan abstrak naskah Anda dengan ribuan jurnal terakreditasi dan menampilkan rekomendasi terbaik beserta alasannya.
Coba MatchMindLihat profil lengkap jurnal ini
Waktu review, biaya APC, statistik sitasi, indeksasi Scopus, dan banyak lagi.
Buka Indonesian Journal of Electrical Engineering and Computer ScienceArtikel ini juga tersedia di situs resmi jurnal.
