MÔ HÌNH NGÔN NGỮ LỚN TRONG ĐÁNH GIÁ BÀI LUẬN TỰ ĐỘNG: TỔNG QUAN LÝ THUYẾT VÀ ĐỀ XUẤT KHUNG ỨNG DỤNG CHO GIÁO DỤC ĐẠI HỌC VIỆT NAM
DOI:
https://doi.org/10.59266/houjs.2026.1435Keywords:
chấm luận tự động, đánh giá theo rubric, giáo dục đại học, mô hình ngôn ngữ lớn, tiếng ViệtAbstract
Bài báo tổng quan đánh giá bài luận tự động (Automated Essay Scoring – AES) bằng mô hình ngôn ngữ lớn (LLMs) và đề xuất khung ứng dụng cho giáo dục đại học Việt Nam. Tổng quan tường thuật có cấu trúc phân tích 19 nguồn cốt lõi, trong đó có 6 công trình năm 2025–2026. Bằng chứng trực tiếp cho thấy mức đồng thuận LLM–người chấm thay đổi theo dữ liệu, mô hình, rubric và prompt; độ tin cậy không đồng nghĩa với độ giá trị. Khung LLM–Rubric–LMS–Human-in-the-Loop được đề xuất như một quy trình kỹ thuật–sư phạm chưa kiểm chứng trên bài luận tiếng Việt: LLM đề xuất điểm và phản hồi, giảng viên quyết định điểm cuối. Khung kèm rubric, prompt minh họa, cơ chế giảm thiểu rủi ro và kế hoạch kiểm định.
References
Bender, E. M., Gebru, T., McMillan-Major, A., & Shmitchell, S. (2021). On the dangers of stochastic parrots: Can language models be too big? Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, 610-623. https://doi.org/10.1145/3442188.3445922
Chu, S. Y., Kim, J. W., Wong, B., & Yi, M. Y.(2025). Rationale behind essay scores: Enhancing S-LLM’s multi-trait essay scoring with rationale generated by LLMs. Findings of the Association for Computational Linguistics: NAACL 2025, 5811-5829. https://doi.org/10.18653/v1/2025.findings-naacl.322
Hashemi, H., Eisner, J., Rosset, C., Van Durme, B., & Kedzie, C. (2024). LLM-RUBRIC: A multidimensional, calibrated approach to automated evaluation of natural language texts. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics, 13806-13834.
Huang, Y., & Wilson, J. (2025). Evaluating LLM-based automated essay scoring: Accuracy, fairness, and validity. Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress, 71-83. https://aclanthology.org/2025.aimecon-wip.9/
Kasneci, E., Sessler, K., Küchemann, S., Bannert, M., Dementieva, D., Fischer, F., Gasser, U., Groh, G., Günnemann, S., Hüllermeier, E., Krusche, S., Kutyniok, G., Michaeli, T., Nerdel, C., Pfeffer, J., Poquet, O., Sailer, M., Schmidt, A., Seidel, T., … Kasneci, G. (2023). ChatGPT for good? On opportunities and challenges of large language models for education. Learning and Individual Differences, 103, 102274. https://doi.org/10.1016/j.lindif.2023.102274
Kim, S., Shin, J., Cho, Y., Jang, J., Longpre, S., Lee, H., Yun, S., Shin, S., Kim, S., Thorne, J., & Seo, M. (2024). PROMETHEUS: Inducing fine-grained evaluation capability in language models. International Conference on Learning Representations (ICLR).
Liu, Y., Iter, D., Xu, Y., Wang, S., Xu, R., & Zhu, C. (2023). G-EVAL: NLG evaluation using GPT-4 with better human alignment. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP).
Mansour, W. A., Albatarni, S., Eltanbouly, S., & Elsayed, T. (2024). Can large language models automatically score the proficiency of written essays? Proceedings of the 2024
Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC- COLING 2024), 2777-2786.
Mizumoto, A., & Eguchi, M. (2023). Exploring the potential of using an AI language model for automated essay scoring. Research Methods in Applied Linguistics, 2(2), 100050. https://doi.org/10.1016/j.rmal.2023.100050
Mughal, N., Imran, A. S., Daudpota, S. M., Kastrati, Z., & Noor, W. (2026). Exploring potential of large language models for automated essay scoring in education. Discover Artificial Intelligence, 6, Article 166. https://doi.org/10.1007/s44163-026-01002-y
Nguyen, T.-H., Le, A.-C., & Nguyen, V.-C. (2024). ViLLM-Eval: A comprehensive evaluation suite for Vietnamese large language models [Preprint]. arXiv.
Pack, A., Barrett, A., & Escalante, J. (2024). Large language models and automated essay scoring of English language learner writing: Insights into validity and reliability. Computers and Education: Artificial Intelligence, 6, 100234. https://doi.org/10.1016/j.caeai.2024.100234
Shibata, T., & Miyamura, Y. (2025). LCES: Zero-shot automated essay scoring via pairwise comparisons using large language models. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 29988-30001. https://doi.org/10.18653/v1/2025.emnlp-main.1523
Su, J., Yan, Y., Fu, F., Han, Z., Ye, J., Liu, X., Huo, J., Zhou, H., & Hu, X. (2025). EssayJudge: A multi-granular benchmark for assessing automated essay scoring capabilities of multimodal large language models. Findings of the Association for Computational Linguistics: ACL 2025, 6363-6389. https://doi.org/10.18653/v1/2025.findings-acl.329
Tlili, A., Shehata, B., Adarkwah, M. A., Bozkurt, A., Hickey, D. T., Huang, R., & Agyemang, B. (2023). What if the devil is my guardian angel: ChatGPT as a case study of using chatbots in education. Smart Learning Environments, 10, Article 15. https://doi.org/10.1186/s40561-023-00237-x
Tran, M.-N., Nguyen, P.-V., Nguyen, L., & Dinh, D. (2024). ViGLUE: A Vietnamese general language understanding benchmark and analysis of Vietnamese language models. Findings of the Association for Computational Linguistics: NAACL 2024, 4174-4189. UNESCO. (2023). Guidance for generative AI in education and research. https://doi.org/10.54675/EWZM9535
Yoshida, L. (2024). The impact of example selection in few-shot prompting on automated essay scoring using GPT models. In Artificial Intelligence in Education (AIED 2024) (Lecture Notes in Computer Science, Vol. 14830, pp. 61-73). Springer. https://doi.org/10.1007/978-3-031-64315-6_5
Yoshida, L. (2025). Do we need a detailed rubric for automated essay scoring using large language models? In A. I. Cristea, E. Walker, Y. Lu, O. C. Santos, & S. Isotani (Eds.), Artificial intelligence in education (Lecture Notes in Computer Science, Vol. 15882, pp. 60-67). Springer. https://doi.org/10.1007/978-3-031-98465-5_8