Publiora

Menghubungkan ke Publiora...

Publiora

TunDC: a public benchmark dataset for sentiment analysis and language modeling in the Tunisian dialect

Khalil Boulahia, AhmedMars, Mourad
IAES International Journal of Artificial Intelligence (IJ-AI) (Sinta 1)Vol. 0 No. 01 April 2026
DOI10.11591/ijai.v15.i2.pp1891-1908

Abstrak

The development of natural language processing (NLP) applications has increasingly focused on dialectal variations of languages. The Tunisian dialect (TD), a widely spoken variant of Arabic, poses unique linguistic challenges due to its lack of standardized writing conventions and influences from multiple languages, including French, Italian, Turkish, and Berber. In this work, we introduce TunDC, a dataset of 20,044 labeled comments designed to advance NLP research on the TD. The dataset covers diverse linguistic forms (Arabic, Latin, and mixed scripts), and each comment was manually annotated for positive or negative sentiment by native speakers, achieving high inter-annotator agreement. To evaluate its effectiveness, we fine-tuned various models on TunDC. The bert-base-arabic-TunDC-mixed model achieved an accuracy of 0.84 and a macro-averaged F1-score of 0.83, demonstrating strong generalization across sentiment categories and writing systems. A stratified data-splitting strategy considering both sentiment and script type further improved accuracy by approximately 8% compared to standard splits. As a publicly available resource, TunDC contributes to the computational linguistics community, fostering advancements in language modeling and applications tailored to the TD.

Kata Kunci

Arabic datasetArtificial intelligenceFine-tuningLarge language modelLow-resource languageSentiment analysisTunisian dialect

Cari jurnal yang tepat untuk naskah Anda

MatchMind AI mencocokkan abstrak naskah Anda dengan ribuan jurnal terakreditasi dan menampilkan rekomendasi terbaik beserta alasannya.

Coba MatchMind

Lihat profil lengkap jurnal ini

Waktu review, biaya APC, statistik sitasi, indeksasi Scopus, dan banyak lagi.

Buka IAES International Journal of Artificial Intelligence (IJ-AI)

Artikel ini juga tersedia di situs resmi jurnal.

TunDC: a public benchmark dataset for sentiment analysis and language modeling in the Tunisian dialect | IAES International Journal of Artificial Intelligence (IJ-AI) | Publiora