PHÁT HIỆN VĂN BẢN TẠO SINH BỞI TRÍ TUỆ NHÂN TẠO TRONG ĐỒ ÁN TIẾNG VIỆT: THỰC NGHIỆM TRÊN PHOBERT VÀ XLM-ROBERTA VỚI CHIẾN LƯỢC PHÂN ĐOẠN TỪ
DOI:
https://doi.org/10.59266/houjs.2026.1423Keywords:
học đối lập đa tầng, ,ô hình ngôn ngữ lớn, PhoBERT, phân đoạn từ, văn bản tạo sinhAbstract
Các mô hình ngôn ngữ lớn (Large Language Models - LLMs) như ChatGPT, Gemini, DeepSeek, Claude AI đã thay đổi cách thức học tập và nghiên cứu của sinh viên. Việc lạm dụng văn bản tạo sinh thiếu kiểm chứng vào trong đồ án, luận văn, các nơi mà tính học thuật cần được đảm bảo và minh bạch rõ ràng đã và đang xuất hiện ngày một nhiều. Trong nghiên cứu này, chúng tôi thực nghiệm khung nghiên cứu, kế thừa từ khung công nghệ tiên tiến dùng để phân tích và phân loại văn bản (Fine-Grained AI-Generated Text Detection - FAID), một hệ thống phát hiện văn bản tạo sinh Tiếng Việt trong miền thông tin đồ án và khóa luận tốt nghiệp. Bài báo thực nghiệm với mô hình PhoBERT, thay thế cho mô hình đa ngôn ngữ XLM-RoBERTa. Bên cạnh đó, chúng tôi đề xuất 14 quy tắc học thuật thường xuất hiện trong các đồ án, luận văn để tiền xử lý văn bản Tiếng Việt với mục tiêu cải thiện hiệu quả của PhoBERT. Chúng tôi trích rút từ bộ dữ liệu của FAID để thu được 46,790 mẫu Tiếng Việt làm dữ liệu cho thực nghiệm. Kết quả cho thấy PhoBERT không vượt XLM‑RoBERTa về Accuracy/Macro F1; cấu hình tốt nhất trên hai chỉ số này vẫn là XLM‑RoBERTa không phân đoạn từ (Experiment1-EXP‑1: 95,76%/95,99%). Đóng góp của 14 quy tắc thể hiện rõ nhất ở việc hạ tỉ lệ cáo buộc sai (False Accusation Rate-FAR) xuống mức thấp nhất (8,73%) trong khi giữ Accuracy/Macro F1 tương đương, phù hợp với bối cảnh ưu tiên giảm nguy cơ buộc tội oan.
References
Abassy, M., Elozeiri, K., Aziz, A., Ta, M. N., Tomar, R. V., Adhikari, B., Ahmed, S.E. D., Wang, Y., Mohammed Afzal, O., Xie, Z., Mansurov, J., Artemova, E., Mikhailov, V., Xing, R., Geng, J., Iqbal, H., Mujahid, Z. M., Mahmoud, T., Tsvigun, A., … Nakov, P. (2024). LLM-DetectAIve: A Tool for Fine- Grained Machine-Generated Text Detection. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, 336-343. https://doi.org/10.18653/v1/2024.emnlp-demo.35
Bao, G., Rong, L., Zhao, Y., Zhou, Q., & Zhang, Y. (2025). Decoupling Content and Expression: Two-Dimensional Detection of AI-Generated Text (arXiv:2503.00258). arXiv. https://doi.org/10.48550/arXiv.2503.00258
Chen, Y., Kang, H., Zhai, V., Li, L., Singh, R., & Raj, B. (2023). Token Prediction as Implicit Classification to Identify LLM-Generated Text. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 13112-13120. https://doi.org/10.18653/v1/2023.emnlp-main.810
Conneau, A., Khandelwal, K., Goyal, N., Chaudhary, V., Wenzek, G., Guzmán, F., Grave, E., Ott, M., Zettlemoyer, L., & Stoyanov, V. (2020). Unsupervised Cross-lingual Representation Learning at Scale. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 8440-8451. https://doi.org/10.18653/v1/2020.acl-main.747
DeepSeek-AI, Guo, D., Yang, D., Zhang, H.,Song, J., Wang, P., Zhu, Q., Xu, R., Zhang, R., Ma, S., Bi, X., Zhang, X., Yu, X., Wu, Y., Wu, Z. F., Gou, Z., Shao, Z., Li, Z., Gao, Z., … Zhang, Z. (2025).
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. Nature, 645(8081), 633-638. https://doi.org/10.1038/s41586-025-09422-z
Grattafiori, A., Dubey, A., Jauhri, A., Pandey, A., Kadian,A.,Al-Dahle,A., Letman,A., Mathur, A., Schelten, A., Vaughan, A., Yang, A., Fan, A., Goyal, A., Hartshorn, A., Yang, A., Mitra, A., Sravankumar, A., Korenev, A., Hinsvark, A., … Ma, Z. (2024). The Llama 3 Herd of Models (arXiv:2407.21783). arXiv. https://doi.org/10.48550/arXiv.2407.21783
Guo, X., Zhang, S., He, Y., Zhang, T., Feng, W., Huang, H., & Ma, C. (2024). DeTeCtive: Detecting AI-generated Text via Multi- Level Contrastive Learning. Advances in Neural Information Processing Systems 37, 88320-88347. https://doi.org/10.52202/079017-2802
Hans, A., Schwarzschild, A., Cherepanova, V., Kazemi, H., Saha, A., Goldblum, M., Geiping, J., & Goldstein, T. (2024). Spotting LLMs With Binoculars: Zero- Shot Detection of Machine-Generated Text (arXiv:2401.12070). arXiv. https://doi.org/10.48550/arXiv.2401.12070
Nguyen, D. Q., & Nguyen, A. T. (2020). PhoBERT: Pre-trained language models for Vietnamese (Version 3). arXiv. https://doi.org/10.48550/ARXIV.2003.00744
OpenAI, Hurst, A., Lerer, A., Goucher, A. P., Perelman, A., Ramesh, A., Clark, A., Ostrow, A. J., Welihinda, A., Hayes, A., Radford, A., Mądry, A., Baker- Whitcomb, A., Beutel, A., Borzunov, A., Carney, A., Chow, A., Kirillov, A., Nichol, A., … Malkov, Y. (2024). GPT- 4o System Card (arXiv:2410.21276). arXiv. https://doi.org/10.48550/arXiv.2410.21276
Soni, R., Misra, R., & Mukopadhyay, S. (2025). DetectLLM: A Multimodal Fusion Approach for Detecting LLM-Generated Text. 2025 5th International Conference on AI- ML-Systems (AIMLSystems), 50-57. https://doi.org/10.1109/AIMLSystems67835.2025.11387030
Su, J., Zhuo, T., Wang, D., & Nakov, P. (2023). DetectLLM: Leveraging Log Rank Information for Zero-Shot Detection of Machine-Generated Text. Findings of the Association for Computational Linguistics: EMNLP 2023, 12395-12412. https://doi.org/10.18653/v1/2023.findings-emnlp.827
Ta, M. N., Van, D. C., Hoang, D.-A., Le-Anh, M., Nguyen, T., Nguyen, M. A. T., Wang, Y., Nakov, P., & Dinh, S. (2026). FAID: Fine-Grained AI-Generated Text Detection Using Multi-Task Auxiliary and Multi-Level Contrastive Learning (arXiv:2505.14271). arXiv. https://doi.org/10.48550/arXiv.2505.14271
Team, G., Anil, R., Borgeaud, S., Alayrac, J.-B., Yu, J., Soricut, R., Schalkwyk, J., Dai, A. M., Hauth, A., Millican, K., Silver, D., Johnson, M., Antonoglou, I., Schrittwieser, J., Glaese, A., Chen, J., Pitler, E., Lillicrap, T., Lazaridou, A., … Vinyals, O. (2025). Gemini: A Family of Highly Capable Multimodal Models (arXiv:2312.11805). arXiv. https://doi.org/10.48550/arXiv.2312.11805
Tran, Q.-D., Nguyen, V.-Q., Pham, Q.-H., Nguyen, K. B. T., & Do, T.-H. (2024). Vietnamese AI Generated Text Detection (arXiv:2405.03206). arXiv. https://doi.org/10.48550/arXiv.2405.03206
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017). Attention is All you Need. Advances in Neural Information Processing Systems 30 (NeurIPS 2017), 6000-6010. https://doi.org/10.48550/arXiv.1706.03762
Wang, P., Li, L., Ren, K., Jiang, B., Zhang, D., & Qiu, X.(2023). SeqXGPT: Sentence-Level AI-Generated Text Detection (Version 2). arXiv. https://doi.org/10.48550/ARXIV.2310.08903
Zhang, Q., Gao, C., Chen, D., Huang, Y., Huang, Y., Sun, Z., Zhang, S., Li, W., Fu, Z., Wan, Y., & Sun, L. (2024). LLM-as-a-Coauthor: Can Mixed Human-Written and Machine- Generated Text Be Detected? Findings of the Association for Computational Linguistics: NAACL 2024, 409-436. https://doi.org/10.18653/v1/2024.findings-naacl.29