Development of Automated Essay Scoring Using Retrieval Augmented Generation in SAGE
DOI:
https://doi.org/10.31004/riggs.v5i2.8979Keywords:
Automated Scoring, Learning Management System, Retrieval Augmented Generation, Gemini API, Assessment AccuracyAbstract
Academic assessment through essay questions is a fundamental component of the educational ecosystem designed to measure students’ cognitive depth and critical reasoning. However, manual essay grading presents significant pedagogical and administrative challenges including high susceptibility to subjective bias and an overwhelming workload for educators. To address these critical issues this research develops the Smart Automated Grading Engine or SAGE, an advanced Learning Management System engineered to automate qualitative assessments. SAGE integrates large language models via the Gemini API with a robust Retrieval Augmented Generation architecture. By strictly grounding the artificial intelligence evaluation process in teacher curated reference documents and specific grading rubrics the system effectively neutralizes the risk of information hallucination. The system was empirically validated at SMAN 1 Blahbatuh involving 180 authentic essay responses from 36 eleventh grade students. The automated assessments were statistically compared against the manual evaluations of three expert history teachers. Comprehensive technical evaluations utilizing Black Box and White Box testing confirmed the platform absolute functional stability and architectural security. Crucially the accuracy testing demonstrated exceptional pedagogical reliability where the SAGE platform achieved a Quadratic Weighted Kappa coefficient of 0.9133 categorizing its performance as having almost perfect agreement. Furthermore, the system exhibited a remarkable precision rate of 94.44 percent within a stringent 10 point score tolerance. Ultimately the integration of this technology proves to be an effective objective and efficient solution capable of replicating human evaluation sharpness while significantly alleviating educator burnout.
Downloads
References
Y. Gao et al., “Retrieval-Augmented Generation for Large Language Models: A Survey,” Dec. 2023, [Online]. Available: http://arxiv.org/abs/2312.10997
M. Faseeh et al., “Hybrid Approach to Automated Essay Scoring: Integrating Deep Learning Embeddings with Handcrafted Linguistic Features for Improved Accuracy,” Mathematics, vol. 12, no. 21, Nov. 2024, doi: 10.3390/math12213416.
R. Bambang, E. Saputro, and R. Hikmawan, “A Rubric-Integrated Assessment System Using a Large Language Model for Automated Essay Evaluation in Secondary Vocational Schools,” Journal of Educational Sciences, vol. 10, no. 5, pp. 11–23, 2026, doi: 10.31258/jes.10.5.p.11-23.
I. M. Y. Wirawan, I. M. Yudana, and I. N. Natajaya, “EVALUASI PELAKSANAAN LEARNING MANAGEMENT SYSTEM (LMS) DI SEKOLAH PENGGERAK SMPK 1 HARAPAN DENPASAR,” JURNAL ADMINISTRASI PENDIDIKAN INDONESIA, vol. 13, no. 1, pp. 44–54, May 2022, doi: 10.23887/jurnal_ap.v13i1.957.
P. R. Darmayasa, I. B. N. Pascima, and K. Agustini, “Pengembangan Sistem Penilaian Keaktifan Belajar Otomatis Berbasis Learning Management System (LMS) Dengan Retrieval-Augmented Generation (RAG),” Kumpulan Artikel Mahasiswa Pendidikan Teknik Informatika (KARMAPATI), vol. 14, no. 3, 2025, doi: https://doi.org/10.23887/karmapati.v14i3.102876.
A. A. G. P. Prameswara, I. N. I. Wiradika, L. P. E. Damayanti, and I. P. G. P. Pastika, “Development of a Large Language Model-Based Diagnostic Assessment System for Detecting Students’ Initial Abilities,” Measurement in Educational Research, vol. X, pp. 1–11, 2022, doi: 10.33292/meter.v1i1.xxx.
Z. Li, Z. Wang, W. Wang, K. Hung, H. Xie, and F. L. Wang, “Retrieval-augmented generation for educational application: A systematic survey,” Computers and Education: Artificial Intelligence, vol. 8, p. 100417, Jun. 2025, doi: 10.1016/j.caeai.2025.100417.
Z. Jiang et al., “Active Retrieval Augmented Generation,” pp. 7969–7992, Oct. 2023, [Online]. Available: http://arxiv.org/abs/2305.06983
Sugiyono, Metode Penelitian Kuantitatif, Kualitatif, dan R&D. Bandung: Alfabeta, 2018.
S. Y. Luis, D. G. Reina, and S. T. Marín, “Towards a Retrieval-Augmented Generation Framework for Originality Evaluation in Projects-Based Learning Classrooms,” Educ. Sci. (Basel)., vol. 15, no. 6, p. 706, Jun. 2025, doi: 10.3390/educsci15060706.
S. Pambudi, Herman Dwi Surjono, Totok Sukardiyono, and Akhsin Nurlayli, “A Moodle-Based Digital Learning Approach to Enhance AI Literacy Competence in Non-STEM Programs,” Jurnal Nasional Pendidikan Teknik Informatika (JANAPATI), vol. 15, no. 1, pp. 183–194, Mar. 2026, doi: 10.23887/janapati.v15i1.105958.
I. K. Resika Arthana, N. Gunantara, M. Sudarma, and M. Sukarsa, “A Systematic Literature Review of Retrieval-Augmented Generation Implementation for Enhancing Large Language Models in Education,” Jurnal Nasional Pendidikan Teknik Informatika (JANAPATI), vol. 15, no. 1, pp. 91–109, Mar. 2026, doi: 10.23887/janapati.v15i1.112281.
R. V. Barenji, N. Salimi, and S. Khoshgoftar, “An LLM -Powered Assessment Retrieval-Augmented Generation (RAG) For Higher Education,” Jan. 2026, [Online]. Available: http://arxiv.org/abs/2601.06141
N. K. Nopiani, N. Sugihartini, I. Bagus, and N. Pascima, “PENGEMBANGAN KONTEN INTERAKTIF BERBASIS MODEL DISCOVERY LEARNING PADA MATA PELAJARAN ILMU PENYAKIT DAN PENUNJANG DIAGNOSTIK KELAS XI DI SMK NEGERI 4 NEGARA,” Kumpulan Artikel Mahasiswa Pendidikan Teknik Informatika (KARMAPATI), vol. 11, no. 1, 2022, doi: 10.23887/karmapati.v11i1.39827.
C. Pramartha, I. Koten, I. G. N. A. C. Putra, I. W. Supriana, and I. W. Arka, “Pengembangan Sistem Dokumentasi Melalui Pendekatan Ontologi untuk Praktek Budaya Bali,” Jurnal Nasional Pendidikan Teknik Informatika (JANAPATI), vol. 11, no. 3, pp. 259–268, Dec. 2022, doi: 10.23887/janapati.v11i3.53939.
E. Prasetio, L. B. H. Handoko, and K. Hastuti, “Improving Retrieval-Augmented Generation Performance Using the MAF-RAG Architecture, EVR–VOR Vector Retrieval, and Multi-Agent Fallback Reasoning,” Journal of Applied Informatics and Computing, vol. 10, no. 1, pp. 212–223, Feb. 2026, doi: 10.30871/jaic.v10i1.11738.
E. Kasneci et al., “ChatGPT for good? On opportunities and challenges of large language models for education,” Learn. Individ. Differ., vol. 103, p. 102274, Apr. 2023, doi: 10.1016/j.lindif.2023.102274.
J. R. Landis and G. G. Koch, “The Measurement of Observer Agreement for Categorical Data,” Biometrics, vol. 33, no. 1, p. 159, Mar. 1977, doi: 10.2307/2529310.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 I Gusti Nyoman Sapta Wiguna, Ida Bagus Nyoman Pascima, Luh Putu Eka Damayanthi

This work is licensed under a Creative Commons Attribution 4.0 International License.


















