Perbandingan Model Deep Learning dan Transformer untuk Analisis Sentimen Ulasan Aplikasi Tokopedia
Keywords:
Analisis sentimen, Tokopedia, Weak Label, IndoBERT, Deep Learning, TransformerAbstract
Rating bintang pada ulasan aplikasi tidak selalu merepresentasikan polaritas teks. Penelitian ini membandingkan tujuh model machine learning, deep learning, dan Transformer pada 10.071 ulasan Tokopedia berbahasa Indonesia dengan label teks berbasis InSet serta rating sebagai weak-label pembanding. InSet mencakup 9,447 ulasan; label akhir terdiri atas 5,342 negatif, 1,019 netral, dan 3,710 positif. Sebanyak 4,360 label InSet berbeda dari pemetaan rating. Pada lima seed, Linear SVM memperoleh mean macro-F1 tertinggi 0,711 ± 0,000, diikuti Logistic Regression 0,709 ± 0,000. Pada seed 42, bootstrap 95% CI model terbaik adalah [0,680, 0,742], sedangkan perbedaannya dengan runner-up tidak signifikan menurut McNemar-Holm (p=0,5666). Hasil menunjukkan bahwa kesimpulan benchmark sensitif terhadap skema label dan baseline linear tetap kompetitif.
References
J. Dąbrowski, E. Letier, A. Perini, and A. Susi, “Analysing app reviews for software engineering: a systematic literature review,” Empir. Softw. Eng., vol. 27, no. 2, 2022, doi: 10.1007/s10664-021-10065-7.
X. Wang, T. Zhang, Y. Tan, W. Shang, and Y. Li, “How to effectively mine app reviews concerning software ecosystem? A survey of review characteristics,” Journal of Systems and Software, vol. 213, p. 112040, 2024, doi: 10.1016/j.jss.2024.112040.
P. Rasappan, M. Premkumar, G. Sinha, and K. Chandrasekaran, “Transforming sentiment analysis for e-commerce product reviews: Hybrid deep learning model with an innovative term weighting and feature selection,” Inf. Process. Manag., vol. 61, no. 3, p. 103654, May 2024, doi: 10.1016/j.ipm.2024.103654.
P. Ataei, S. Regula, D. Staegemann, and S. Malgaonkar, “Filtering useful app reviews using Naive Bayes: Which Naive Bayes?,” AI, vol. 5, no. 4, pp. 2237–2259, 2024, doi: 10.3390/ai5040110.
A. Almansour, R. Alotaibi, and H. Alharbi, “Text-rating review discrepancy (TRRD): an integrative review and implications for research,” Future Business Journal, vol. 8, no. 1, 2022, doi: 10.1186/s43093-022-00114-y.
A. Iqbal, R. Amin, J. Iqbal, R. Alroobaea, A. Binmahfoudh, and M. Hussain, “Sentiment Analysis of Consumer Reviews Using Deep Learning,” Sustainability, vol. 14, no. 17, p. 10844, 2022, doi: 10.3390/su141710844.
A. Koufakou, “Deep learning for opinion mining and topic classification of course reviews,” Educ. Inf. Technol. (Dordr)., vol. 29, no. 3, pp. 2973–2997, 2024, doi: 10.1007/s10639-023-11736-2.
Y. Mao, Q. Liu, and Y. Zhang, “Sentiment analysis methods, applications, and challenges: A systematic literature review,” Journal of King Saud University - Computer and Information Sciences, vol. 36, no. 4, p. 102048, 2024, doi: 10.1016/j.jksuci.2024.102048.
J. R. Jim, M. A. R. Talukder, P. Malakar, M. M. Kabir, K. Nur, and M. Mridha, “Recent advancements and challenges of NLP-based sentiment analysis: A state-of-the-art review,” Natural Language Processing Journal, vol. 6, p. 100059, 2024, doi: 10.1016/j.nlp.2024.100059.
Y. C. Hua, P. Denny, J. Wicker, and K. Taskova, “A systematic review of aspect-based sentiment analysis: domains, methods, and trends,” Artif. Intell. Rev., vol. 57, no. 11, 2024, doi: 10.1007/s10462-024-10906-z.
G. I. Winata, A. F. Aji, S. Cahyawijaya, R. Mahendra, F. Koto, and A. Romadhony, “NusaX: Multilingual Parallel Sentiment Dataset for 10 Indonesian Local Languages,” in in Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, Andreas Vlachos and Isabelle Augenstein, Eds., Association for Computational Linguistics, May 2023, pp. 815–834. doi: 10.18653/v1/2023.eacl-main.57.
S. Cahyawijaya, H. Lovenia, A. F. Aji, G. Winata, B. Wilie, and F. Koto, “NusaCrowd: Open Source Initiative for Indonesian NLP Resources,” in in Findings of the Association for Computational Linguistics: ACL, Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki, Eds., Association for Computational Linguistics, Jul. 2023, pp. 13745–13818. doi: 10.18653/v1/2023.findings-acl.868.
B. Wilie et al., “IndoNLU: Benchmark and Resources for Evaluating Indonesian Natural Language Understanding,” in Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing, Stroudsburg, PA, USA: Association for Computational Linguistics, 2020, pp. 843–857. doi: 10.18653/v1/2020.aacl-main.85.
Fajri Koto, Afshin Rahimi, Jey Han Lau, and Timothy Baldwin, “IndoLEM and IndoBERT: A Benchmark Dataset and Pre-trained Language Model for Indonesian NLP,” in in Proceedings of the 28th International Conference on Computational Linguistics, Donia Scott, Nuria Bel, and Chengqing Zong, Eds., International Committee on Computational Linguistics, Dec. 2020, pp. 757–770. doi: 10.18653/v1/2020.coling-main.66.
Y. Kim, “Convolutional Neural Networks for Sentence Classification,” in Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), Alessandro Moschitti, Bo Pang, and Walter Daelemans, Eds., Stroudsburg, PA, USA: Association for Computational Linguistics, Oct. 2014, pp. 1746–1751. doi: 10.3115/v1/D14-1181.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” in Proceedings of the 2019 Conference of the North, Jill Burstein, Christy Doran, and Thamar Solorio, Eds., Stroudsburg, PA, USA: Association for Computational Linguistics, Jun. 2019, pp. 4171–4186. doi: 10.18653/v1/N19-1423.
A. Conneau et al., “Unsupervised Cross-lingual Representation Learning at Scale,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel Tetreault, Eds., Stroudsburg, PA, USA: Association for Computational Linguistics, Jul. 2020, pp. 8440–8451. doi: 10.18653/v1/2020.acl-main.747.
Y. Liu et al., “RoBERTa: A Robustly Optimized BERT Pretraining Approach,” Jul. 2019, [Online]. Available: http://arxiv.org/abs/1907.11692
I. H. Setiawan, M. Rahardi, A. Aminuddin, and F. F. Abdulloh, “Sentiment Analysis of Tokopedia Application Reviews on Google Play Store Using BERT,” in in, 2024, pp. 242–247. doi: 10.1109/ICITSI65188.2024.10929357.
M. Idris, A. Rifai, and K. D. Tania, “Sentiment Analysis of Tokopedia App Reviews using Machine Learning and Word Embeddings,” sinkron, vol. 9, no. 1, pp. 210–219, 2025, doi: 10.33395/sinkron.v9i1.14278.
A. R. Aisy and G. Karyono, “Analisis Sentimen Terhadap Ulasan Google Play Store Aplikasi Lazada, Shopee, dan Tokopedia Menggunakan Algoritma IndoBERT,” Building of Informatics, vol. 7, no. 3, pp. 2127–2135, 2025, doi: 10.47065/bits.v7i3.8745.
R. Tajuddin, Y. Irawan, and R. R. Setiawan, “Classification of Sentiment Tokopedia and Shopee App Reviews on Google Playstore Using Naive Bayes,” bit-Tech, vol. 8, no. 2, pp. 2348–2357, 2025, doi: 10.32877/bt.v8i2.3246.
R. Wijianto, D. Pratmanto, A. Widayanto, and Ubaidilah, “Komparasi K-Nearest Neighbors (KNN) dan Naive Bayes pada Klasifikasi Sentimen Ulasan Aplikasi Tokopedia di Google Play Store,” Informatics and Computer Engineering Journal, vol. 5, no. 2, pp. 75–80, 2025, doi: 10.31294/icej.v5i2.8939.
F. D. Saputra and F. Budiman, “Comparison of Random Forest and LSTM for Tokopedia Sentiment Analysis,” Journal of Applied Informatics and Computing, vol. 10, no. 1, pp. 630–639, 2026, doi: 10.30871/jaic.v10i1.12042.
N. Alowidi, R. Alnanih, and N. Alsaleh, “Hybrid Deep Learning Approach for Automating App Review Classification: Advancing Usability Metrics Classification with an Aspect-Based Sentiment Analysis Framework,” Computers, vol. 82, no. 1, pp. 949–976, 2024, doi: 10.32604/cmc.2024.059351.
H. Murfi, Syamsyuriani, T. Gowandi, G. Ardaneswari, and S. Nurrohmah, “BERT-based combination of convolutional and recurrent neural network for indonesian sentiment analysis,” Appl. Soft Comput., vol. 151, p. 111112, 2023, doi: 10.1016/j.asoc.2023.111112.
Dhendra and V. Gayuh Utomo, “Benchmarking IndoBERT and Transformer Models for Sentiment Classification on Indonesian E-Government Service Reviews,” Jurnal Transformatika, vol. 23, no. 1, pp. 86–95, Jul. 2025, doi: 10.26623/transformatika.v23i1.12095.
C. Shaw, P. LaCasse, and L. Champagne, “Exploring emotion classification of indonesian tweets using large scale transfer learning via IndoBERT,” Soc. Netw. Anal. Min., vol. 15, no. 1, 2025, doi: 10.1007/s13278-025-01439-6.
Y. Wang, J. Wang, H. Zhang, X. Ming, and Q. Wang, “Better together: Automated app review analysis with deep multi-task learning,” Inf. Softw. Technol., vol. 177, p. 107597, 2024, doi: 10.1016/j.infsof.2024.107597.
N. Cassee, A. Agaronian, E. Constantinou, N. Novielli, and A. Serebrenik, “Transformers and meta-tokenization in sentiment analysis for software engineering,” Empir. Softw. Eng., vol. 29, no. 4, 2024, doi: 10.1007/s10664-024-10468-2.
F. Koto and G. Y. Rahmaningtyas, “Inset lexicon: Evaluation of a word list for Indonesian sentiment analysis in microblogs,” in 2017 International Conference on Asian Language Processing (IALP), Singapore: IEEE, Dec. 2017, pp. 391–394. doi: 10.1109/IALP.2017.8300625.
M. Sokolova and G. Lapalme, “A systematic analysis of performance measures for classification tasks,” Inf. Process. Manag., vol. 45, no. 4, pp. 427–437, Jul. 2009, doi: 10.1016/j.ipm.2009.03.002.
T. G. Dietterich, “Approximate Statistical Tests for Comparing Supervised Classification Learning Algorithms,” Neural Comput., vol. 10, no. 7, pp. 1895–1923, Oct. 1998, doi: 10.1162/089976698300017197.
Y. Bestgen, “Please, Don’t Forget the Difference and the Confidence Interval when Seeking for the State-of-the-Art Status,” in in Proceedings of the Thirteenth Language Resources and Evaluation Conference, European Language Resources Association, Jun. 2022, pp. 5956–5962. [Online]. Available: https://aclanthology.org/2022.lrec-1.640/
C. Guo, G. Pleiss, Y. Sun, and K. Q. Weinberger, “On Calibration of Modern Neural Networks,” in in Proceedings of the 34th International Conference on Machine Learning, 2017, pp. 1321–1330. [Online]. Available: https://proceedings.mlr.press/v70/guo17a.html
T. Pires, E. Schlinger, and D. Garrette, “How Multilingual is Multilingual BERT?,” in in Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Anna Korhonen, David Traum, and Lluís Màrquez, Eds., Association for Computational Linguistics, Jul. 2019, pp. 4996–5001. doi: 10.18653/v1/P19-1493.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Restu Jati Saputro, Ratri Kurniasari, Hana Nurdina

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.


