Publiora

Menghubungkan ke Publiora...

Publiora

Developing an effective focused crawler to retrieve data of Indian-origin scientists and utilizing text classification for comparative analysis

Gautam, ShivaniBhatia, RajeshJain, Shaily
International Journal of Electrical and Computer Engineering (IJECE) (Sinta 1)Vol. 0 No. 01 Oktober 2024
DOI10.11591/ijece.v14i5.pp5468-5480

Abstrak

This article presents the implementation of focused web crawling to retrieve data about scientists of Indian ancestry who are working in foreign nations. This study demonstrates the effectiveness of web scraping in obtaining large amounts of data from publicly available online pages. The objective is to construct a collection of data pertaining to Indian scientists who are now employed in national laboratories overseas. Collecting a vast quantity of data on the aforementioned Indian scientists through manual search is a pointless task. Therefore, this study proposes a detailed plan for a focused web crawler that can gather similar data. Subsequently, we present a comprehensive assessment of numerous classification models on this newly created dataset. Our assessments indicate that the random forest model surpasses the other supervised models. The empirical findings on large datasets demonstrated that the combination of random forest with synthetic minority oversampling technique (SMOTE) and k-fold cross-validation methods yielded better performance compared to K-nearest neighbors (KNN), support vector machine (SVM), and logistic regression (LR) for Indian origin scientists. Conversely, SMOTE with an 80-20 random split demonstrated superior performance on smaller datasets. Overall, the random forest classifier demonstrated the most favorable outcomes, attaining a micro-average area under curve (AUC) of 90%. The outcomes of our study provide a solid foundation for further investigation into classification of text of Indian origin scientists.

Kata Kunci

Computer ScienceComparative analysisFocused web crawlerIndian scientists database natural language processing text classificationInformation retrievalWeb scraping

Cari jurnal yang tepat untuk naskah Anda

MatchMind AI mencocokkan abstrak naskah Anda dengan ribuan jurnal terakreditasi dan menampilkan rekomendasi terbaik beserta alasannya.

Coba MatchMind

Lihat profil lengkap jurnal ini

Waktu review, biaya APC, statistik sitasi, indeksasi Scopus, dan banyak lagi.

Buka International Journal of Electrical and Computer Engineering (IJECE)

Artikel ini juga tersedia di situs resmi jurnal.

Developing an effective focused crawler to retrieve data of Indian-origin scientists and utilizing text classification for comparative analysis | International Journal of Electrical and Computer Engineering (IJECE) | Publiora