EFEKTIVITAS MODEL PEMROSESAN BAHASA ALAMI UNTUK ANALISIS SENTIMEN TEKS BAHASA INDONESIA PADA DOMAIN LAYANAN PUBLIK: SYSTEMATIC LITERATURE REVIEW 2021–2026
Abstract
Public sentiment toward government digital services in Indonesia is increasingly expressed through application reviews and social media, yet the choice of natural language processing (NLP) model for Indonesian-language text remains inconsistent across studies. This study aims to identify the most effective NLP models for Indonesian sentiment analysis published in 2021–2026, with a particular focus on the public service domain. A systematic literature review was conducted following the PRISMA 2020 statement and the Kitchenham and Charters guidelines, covering Scopus, IEEE Xplore, ScienceDirect, SpringerLink, ACL Anthology, arXiv, SINTA, Garuda, and Google Scholar. Eighteen primary studies met the inclusion criteria and were synthesised descriptively and thematically along five research questions. The results show a shift from classical machine learning (Naive Bayes, SVM) in 2021–2023 toward Indonesian-specific transformer models (IndoBERT, IndoBERTweet) and hybrid architectures in 2024–2025. In within-study comparisons, IndoBERT and its variants generally outperformed classical and recurrent models, reaching accuracy or F1-scores above 0.90 on public service datasets such as Mobile JKN and Coretax. Large language models are emerging but have not consistently surpassed fine-tuned encoders on standard sentiment benchmarks. The main gaps are reliance on automatic lexicon- or rating-based labelling, accuracy reporting on imbalanced data without macro-F1, and the scarcity of aspect-based and shared public service benchmarks. IndoBERT-family models are recommended as the default baseline, together with manually annotated public service corpora and macro-F1 reporting.
References
Ayomi, J. M., Vitianingsih, A. V., Kristyawan, Y., Maukar, A. L., & Widiartin, T. (2025). Sentiment analysis of user reviews for the PLN Mobile application using Naïve Bayes and long short-term memory. Journal of Information Systems and Informatics, 7(4), 3849–3873. https://doi.org/10.63158/journalisi.v7i4.1342
Cahyawijaya, S., Lovenia, H., Koto, F., Putri, R. A., Dave, E., Lee, J., … Fung, P. (2024). Cendol: Open instruction-tuned generative large language models for Indonesian languages. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 14899–14914). https://doi.org/10.18653/v1/2024.acl-long.796
Dermawan, S., & Ayunda, A. T. (2025). Sentiment analysis of Coretax on social media X using Naive Bayes, SVM, and LSTM for service improvement. Journal of Applied Informatics and Computing, 9(6), 3177–3190. https://doi.org/10.30871/jaic.v9i6.11063
Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (pp. 4171–4186). https://doi.org/10.18653/v1/N19-1423
Ellyanti, L., Ruldeviyani, Y., Pradana, L. E., & Harjanto, A. (2023). Sentiment analysis of Twitter users to the PeduliLindungi using Naïve Bayes algorithm. Jurnal RESTI (Rekayasa Sistem dan Teknologi Informasi), 7(2), 414–421. https://jurnal.iaii.or.id/index.php/RESTI/article/view/4684
Fauziah, Y., Yuwono, B. D., & Aribowo, A. S. (2021). Lexicon based sentiment analysis in Indonesia languages: A systematic literature review. RSF Conference Series: Engineering and Technology, 1(1), 363–367. https://doi.org/10.31098/cset.v1i1.397
Imaduddin, H., A'la, F. Y., & Nugroho, Y. S. (2023). Sentiment analysis in Indonesian healthcare applications using IndoBERT approach. International Journal of Advanced Computer Science and Applications, 14(8). https://doi.org/10.14569/IJACSA.2023.0140813
Jazuli, A., Widowati, W., & Kusumaningrum, R. (2025). Optimizing aspect-based sentiment analysis using BERT for comprehensive analysis of Indonesian student feedback. Applied Sciences, 15(1), 172. https://doi.org/10.3390/app15010172
Kitchenham, B., Brereton, O. P., Budgen, D., Turner, M., Bailey, J., & Linkman, S. (2009). Systematic literature reviews in software engineering – A systematic literature review. Information and Software Technology, 51(1), 7–15. https://doi.org/10.1016/j.infsof.2008.09.009
Kitchenham, B., & Charters, S. (2007). Guidelines for performing systematic literature reviews in software engineering (Technical Report EBSE-2007-01). Keele: Keele University and Durham University.
Komarudin, A., & Hilda, A. M. (2024). Analisis sentimen ulasan aplikasi Identitas Kependudukan Digital pada Play Store menggunakan metode Naïve Bayes. Computer Science (CO-SCIENCE), 4(1), 28–36. https://doi.org/10.31294/coscience.v4i1.2955
Koto, F., Lau, J. H., & Baldwin, T. (2021). IndoBERTweet: A pretrained language model for Indonesian Twitter with effective domain-specific vocabulary initialization. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (pp. 10660–10668). https://doi.org/10.18653/v1/2021.emnlp-main.833
Koto, F., Rahimi, A., Lau, J. H., & Baldwin, T. (2020). IndoLEM and IndoBERT: A benchmark dataset and pre-trained language model for Indonesian NLP. In Proceedings of the 28th International Conference on Computational Linguistics (pp. 757–770). https://doi.org/10.18653/v1/2020.coling-main.66
Koto, F., & Rahmaningtyas, G. Y. (2017). InSet lexicon: Evaluation of a word list for Indonesian sentiment analysis in microblogs. In 2017 International Conference on Asian Language Processing (IALP) (pp. 391–394). https://doi.org/10.1109/IALP.2017.8300625
Leong, W. Q., Ngui, J. G., Susanto, Y., Rengarajan, H., Sarveswaran, K., & Tjhi, W. C. (2023). BHASA: A holistic Southeast Asian linguistic and cultural evaluation suite for large language models (arXiv:2309.06085). https://arxiv.org/abs/2309.06085
Mandhasiya, D. G., Murfi, H., Bustamam, A., & Anki, P. (2022). Evaluation of machine learning performance based on BERT data representation with LSTM model to conduct sentiment analysis in Indonesian for predicting voices of social media users in the 2024 Indonesia presidential election. In 2022 5th International Conference on Information and Communications Technology (ICOIACT) (pp. 441–446). https://doi.org/10.1109/ICOIACT55506.2022.9972206
Maulana, R., Voutama, A., & Ridwan, T. (2023). Analisis sentimen ulasan aplikasi MyPertamina pada Google Play Store menggunakan algoritma NBC. Jurnal Teknologi Terpadu, 9(1), 42–48. https://doi.org/10.54914/jtt.v9i1.609
Nasution, A. H., Onan, A., Murakami, Y., Monika, W., & Hanafiah, A. (2025). Benchmarking open-source large language models for sentiment and emotion classification in Indonesian tweets. IEEE Access, 13, 94009–94025. https://doi.org/10.1109/ACCESS.2025.3574629
Owen, L., Tripathi, V., Kumar, A., & Ahmed, B. (2024). Komodo: A linguistic expedition into Indonesia's regional languages (arXiv:2403.09362). https://arxiv.org/abs/2403.09362
Page, M. J., McKenzie, J. E., Bossuyt, P. M., Boutron, I., Hoffmann, T. C., Mulrow, C. D., … Moher, D. (2021). The PRISMA 2020 statement: An updated guideline for reporting systematic reviews. BMJ, 372, n71. https://doi.org/10.1136/bmj.n71
Rizkia, A. S., Wufron, W., & Roji, F. F. (2025). Analisis sentimen Coretax: Perbandingan pelabelan data manual, transformers-based, dan lexicon-based pada performa IndoBERT. MALCOM: Indonesian Journal of Machine Learning and Computer Science, 5(3), 1037–1048. https://doi.org/10.57152/malcom.v5i3.2151
Setiawan, B. (2024). A review of sentiment analysis applications in Indonesia between 2023–2024. JIEET (Journal of Information Engineering and Educational Technology), 8(2), 71–83. https://doi.org/10.26740/jieet.v8n2.p71-83
Tamami, G., Triyanto, W. A., & Muzid, S. (2025). Sentiment analysis Mobile JKN reviews using SMOTE based LSTM. IJCCS (Indonesian Journal of Computing and Cybernetics Systems), 19(1), 13–24. https://doi.org/10.22146/ijccs.101910
Tarwoto, Nugroho, R., Azka, N., & Graha, W. S. R. (2025). Analisis sentimen ulasan aplikasi Mobile JKN di Google PlayStore menggunakan IndoBERT. Jurnal JTIK (Jurnal Teknologi Informasi dan Komunikasi), 9(2), 495–505. https://doi.org/10.35870/jtik.v9i2.3340
Wankhade, M., Rao, A. C. S., & Kulkarni, C. (2022). A survey on sentiment analysis methods, applications, and challenges. Artificial Intelligence Review, 55(7), 5731–5780. https://doi.org/10.1007/s10462-022-10144-1
Wilie, B., Vincentio, K., Winata, G. I., Cahyawijaya, S., Li, X., Lim, Z. Y., … Purwarianti, A. (2020). IndoNLU: Benchmark and resources for evaluating Indonesian natural language understanding. In Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing (pp. 843–857). Association for Computational Linguistics.
Winata, G. I., Aji, A. F., Cahyawijaya, S., Mahendra, R., Koto, F., Romadhony, A., … Ruder, S. (2023). NusaX: Multilingual parallel sentiment dataset for 10 Indonesian local languages. In Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics (pp. 815–834). https://doi.org/10.18653/v1/2023.eacl-main.57
Wohlin, C. (2014). Guidelines for snowballing in systematic literature studies and a replication in software engineering. In Proceedings of the 18th International Conference on Evaluation and Assessment in Software Engineering (Article 38). https://doi.org/10.1145/2601248.2601268
Yulianti, E., & Nissa, N. K. (2024). ABSA of Indonesian customer reviews using IndoBERT: Single-sentence and sentence-pair classification approaches. Bulletin of Electrical Engineering and Informatics, 13(5), 3579–3589. https://doi.org/10.11591/eei.v13i5.8032
Yunanto, I., & Yulianto, S. (2022). Twitter sentiment analysis PeduliLindungi application using Naïve Bayes and support vector machine. Jurnal Teknik Informatika (JUTIF), 3(4), 807–814. https://doi.org/10.20884/1.jutif.2022.3.4.292











