Publiora

Menghubungkan ke Publiora...

Publiora

The effect of training set size in authorship attribution: application on short arabic texts

Al-Sarem, MohammedEmara, Abdel-Hamid
International Journal of Electrical and Computer Engineering (IJECE) (Sinta 1)Vol. 0 No. 01 Februari 2019
DOI10.11591/ijece.v9i1.pp652-659

Abstrak

Authorship attribution (AA) is a subfield of linguistics analysis, aiming to identify the original author among a set of candidate authors. Several research papers were published and several methods and models were developed for many languages. However, the number of related works for Arabic is limited. Moreover, investigating the impact of short words length and training set size is not well addressed. To the best of our knowledge, no published works or researches, in this direction or even in other languages, are available. Therefore, we propose to investigate this effect, taking into account different stylomatric combination. The Mahalanobis distance (MD), Linear Regression (LR), and Multilayer Perceptron (MP) are selected as AA classifiers. During the experiment, the training dataset size is increased and the accuracy of the classifiers is recorded. The results are quite interesting and show different classifiers behaviours. Combining word-based stylomatric features with n-grams provides the best accuracy reached in average 93%.

Kata Kunci

Computer and Informatics

Cari jurnal yang tepat untuk naskah Anda

MatchMind AI mencocokkan abstrak naskah Anda dengan ribuan jurnal terakreditasi dan menampilkan rekomendasi terbaik beserta alasannya.

Coba MatchMind

Lihat profil lengkap jurnal ini

Waktu review, biaya APC, statistik sitasi, indeksasi Scopus, dan banyak lagi.

Buka International Journal of Electrical and Computer Engineering (IJECE)

Artikel ini juga tersedia di situs resmi jurnal.

The effect of training set size in authorship attribution: application on short arabic texts | International Journal of Electrical and Computer Engineering (IJECE) | Publiora