Perhitungan Kemiripan Term Co-occurence Berdasarkan Cluster Dokumen Untuk Pengembangan Thesaurus Bahasa Arab
Abstrak
Generating automatically thesaurus is by calculating the similarity value term. To get the value of the similarity can be carried out with the co-occurence approach is to see the frequency of occurrence along these terms. The frequency of how much these terms occurence on documents corpus. Each of the documents contained in the corpus have content or topics vary. So the terms that are in the document a specific topic will have a different context with the terms of the document with other topics. Therefore, this paper proposes a new method of measurement term similarity with co-occurence based on cluster of documents on the generate of Arabic thesaurus. The documents will be in the corpus clustering to group by the proximity of the content of the document. To get the term similarity value calculation clusterweight by leveraging the value of inverse class frequency of each term to an existing cluster. Thesaurus is formed by looking at the value of the calculation result of the similarity term. Thesaurus formed by the proposed method succeeded in improving inter-term relevance is evidenced by the experimental results have a precision value of 63,3%, amounting to 78,6% recall and F-measure by 50%.
Cari jurnal yang tepat untuk naskah Anda
MatchMind AI mencocokkan abstrak naskah Anda dengan ribuan jurnal terakreditasi dan menampilkan rekomendasi terbaik beserta alasannya.
Coba MatchMindLihat profil lengkap jurnal ini
Waktu review, biaya APC, statistik sitasi, indeksasi Scopus, dan banyak lagi.
Buka JURNAL INFOTELArtikel ini juga tersedia di situs resmi jurnal.
