Ekstraksi informasi leksikal bahasa Kerinci dari kamus dwibahasa untuk klasifikasi Part-of-Speech otomatis
Abstrak
This study aims to develop initial linguistic resources for the Kerinci language by extracting lexical information from an Indonesian–Kerinci bilingual dictionary. The dictionary data, originally in digital document format (.docx), was converted into a structured dataset suitable for computational processing. Part-of-Speech (POS) classification was performed automatically using a rule-based approach, referring to the Indonesian equivalents that had been POS-tagged using reference lexicons. The results show that verbs and adjectives dominate the dictionary entries, differing from the common trend in bilingual dictionaries which are usually dominated by nouns. This approach proves effective in low-resource conditions and serves as a preliminary step toward the development of Kerinci NLP. The resulting POS-classified dataset has great potential for training POS taggers, supporting linguistic documentation, and contributing to language preservation through technology
Kata Kunci
Cari jurnal yang tepat untuk naskah Anda
MatchMind AI mencocokkan abstrak naskah Anda dengan ribuan jurnal terakreditasi dan menampilkan rekomendasi terbaik beserta alasannya.
Coba MatchMindLihat profil lengkap jurnal ini
Waktu review, biaya APC, statistik sitasi, indeksasi Scopus, dan banyak lagi.
Buka AITIArtikel ini juga tersedia di situs resmi jurnal.
