PHÁT HIỆN VĂN BẢN TẠO SINH BỞI TRÍ TUỆ NHÂN TẠO TRONG ĐỒ ÁN TIẾNG VIỆT: THỰC NGHIỆM TRÊN PHOBERT VÀ XLM-ROBERTA VỚI CHIẾN LƯỢC PHÂN ĐOẠN TỪ

Authors

  • Nguyễn Thành Huy
  • Trần Duy Hùng

DOI:

https://doi.org/10.59266/houjs.2026.1423

Keywords:

học đối lập đa tầng, ,ô hình ngôn ngữ lớn, PhoBERT, phân đoạn từ, văn bản tạo sinh

Abstract

Các mô hình ngôn ngữ lớn (Large Language Models - LLMs) như ChatGPT, Gemini, DeepSeek, Claude AI đã thay đổi cách thức học tập và nghiên cứu của sinh viên. Việc lạm dụng văn bản tạo sinh thiếu kiểm chứng vào trong đồ án, luận văn, các nơi mà tính học thuật cần được đảm bảo và minh bạch rõ ràng đã và đang xuất hiện ngày một nhiều. Trong nghiên cứu này, chúng tôi thực nghiệm khung nghiên cứu, kế thừa từ khung công nghệ tiên tiến dùng để phân tích và phân loại văn bản (Fine-Grained AI-Generated Text Detection - FAID), một hệ thống phát hiện văn bản tạo sinh Tiếng Việt trong miền thông tin đồ án và khóa luận tốt nghiệp. Bài báo thực nghiệm với mô hình PhoBERT, thay thế cho mô hình đa ngôn ngữ XLM-RoBERTa. Bên cạnh đó, chúng tôi đề xuất 14 quy tắc học thuật thường xuất hiện trong các đồ án, luận văn để tiền xử lý văn bản Tiếng Việt với mục tiêu cải thiện hiệu quả của PhoBERT. Chúng tôi trích rút từ bộ dữ liệu của FAID để thu được 46,790 mẫu Tiếng Việt làm dữ liệu cho thực nghiệm. Kết quả cho thấy PhoBERT không vượt XLM‑RoBERTa về Accuracy/Macro F1; cấu hình tốt nhất trên hai chỉ số này vẫn là XLM‑RoBERTa không phân đoạn từ (Experiment1-EXP‑1: 95,76%/95,99%). Đóng góp của 14 quy tắc thể hiện rõ nhất ở việc hạ tỉ lệ cáo buộc sai (False Accusation Rate-FAR) xuống mức thấp nhất (8,73%) trong khi giữ Accuracy/Macro F1 tương đương, phù hợp với bối cảnh ưu tiên giảm nguy cơ buộc tội oan.

References

Abassy, M., Elozeiri, K., Aziz, A., Ta, M. N., Tomar, R. V., Adhikari, B., Ahmed, S.E. D., Wang, Y., Mohammed Afzal, O., Xie, Z., Mansurov, J., Artemova, E., Mikhailov, V., Xing, R., Geng, J., Iqbal, H., Mujahid, Z. M., Mahmoud, T., Tsvigun, A., … Nakov, P. (2024). LLM-DetectAIve: A Tool for Fine- Grained Machine-Generated Text Detection. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, 336-343. https://doi.org/10.18653/v1/2024.emnlp-demo.35

Bao, G., Rong, L., Zhao, Y., Zhou, Q., & Zhang, Y. (2025). Decoupling Content and Expression: Two-Dimensional Detection of AI-Generated Text (arXiv:2503.00258). arXiv. https://doi.org/10.48550/arXiv.2503.00258

Chen, Y., Kang, H., Zhai, V., Li, L., Singh, R., & Raj, B. (2023). Token Prediction as Implicit Classification to Identify LLM-Generated Text. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 13112-13120. https://doi.org/10.18653/v1/2023.emnlp-main.810

Conneau, A., Khandelwal, K., Goyal, N., Chaudhary, V., Wenzek, G., Guzmán, F., Grave, E., Ott, M., Zettlemoyer, L., & Stoyanov, V. (2020). Unsupervised Cross-lingual Representation Learning at Scale. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 8440-8451. https://doi.org/10.18653/v1/2020.acl-main.747

DeepSeek-AI, Guo, D., Yang, D., Zhang, H.,Song, J., Wang, P., Zhu, Q., Xu, R., Zhang, R., Ma, S., Bi, X., Zhang, X., Yu, X., Wu, Y., Wu, Z. F., Gou, Z., Shao, Z., Li, Z., Gao, Z., … Zhang, Z. (2025).

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. Nature, 645(8081), 633-638. https://doi.org/10.1038/s41586-025-09422-z

Grattafiori, A., Dubey, A., Jauhri, A., Pandey, A., Kadian,A.,Al-Dahle,A., Letman,A., Mathur, A., Schelten, A., Vaughan, A., Yang, A., Fan, A., Goyal, A., Hartshorn, A., Yang, A., Mitra, A., Sravankumar, A., Korenev, A., Hinsvark, A., … Ma, Z. (2024). The Llama 3 Herd of Models (arXiv:2407.21783). arXiv. https://doi.org/10.48550/arXiv.2407.21783

Guo, X., Zhang, S., He, Y., Zhang, T., Feng, W., Huang, H., & Ma, C. (2024). DeTeCtive: Detecting AI-generated Text via Multi- Level Contrastive Learning. Advances in Neural Information Processing Systems 37, 88320-88347. https://doi.org/10.52202/079017-2802

Hans, A., Schwarzschild, A., Cherepanova, V., Kazemi, H., Saha, A., Goldblum, M., Geiping, J., & Goldstein, T. (2024). Spotting LLMs With Binoculars: Zero- Shot Detection of Machine-Generated Text (arXiv:2401.12070). arXiv. https://doi.org/10.48550/arXiv.2401.12070

Nguyen, D. Q., & Nguyen, A. T. (2020). PhoBERT: Pre-trained language models for Vietnamese (Version 3). arXiv. https://doi.org/10.48550/ARXIV.2003.00744

OpenAI, Hurst, A., Lerer, A., Goucher, A. P., Perelman, A., Ramesh, A., Clark, A., Ostrow, A. J., Welihinda, A., Hayes, A., Radford, A., Mądry, A., Baker- Whitcomb, A., Beutel, A., Borzunov, A., Carney, A., Chow, A., Kirillov, A., Nichol, A., … Malkov, Y. (2024). GPT- 4o System Card (arXiv:2410.21276). arXiv. https://doi.org/10.48550/arXiv.2410.21276

Soni, R., Misra, R., & Mukopadhyay, S. (2025). DetectLLM: A Multimodal Fusion Approach for Detecting LLM-Generated Text. 2025 5th International Conference on AI- ML-Systems (AIMLSystems), 50-57. https://doi.org/10.1109/AIMLSystems67835.2025.11387030

Su, J., Zhuo, T., Wang, D., & Nakov, P. (2023). DetectLLM: Leveraging Log Rank Information for Zero-Shot Detection of Machine-Generated Text. Findings of the Association for Computational Linguistics: EMNLP 2023, 12395-12412. https://doi.org/10.18653/v1/2023.findings-emnlp.827

Ta, M. N., Van, D. C., Hoang, D.-A., Le-Anh, M., Nguyen, T., Nguyen, M. A. T., Wang, Y., Nakov, P., & Dinh, S. (2026). FAID: Fine-Grained AI-Generated Text Detection Using Multi-Task Auxiliary and Multi-Level Contrastive Learning (arXiv:2505.14271). arXiv. https://doi.org/10.48550/arXiv.2505.14271

Team, G., Anil, R., Borgeaud, S., Alayrac, J.-B., Yu, J., Soricut, R., Schalkwyk, J., Dai, A. M., Hauth, A., Millican, K., Silver, D., Johnson, M., Antonoglou, I., Schrittwieser, J., Glaese, A., Chen, J., Pitler, E., Lillicrap, T., Lazaridou, A., … Vinyals, O. (2025). Gemini: A Family of Highly Capable Multimodal Models (arXiv:2312.11805). arXiv. https://doi.org/10.48550/arXiv.2312.11805

Tran, Q.-D., Nguyen, V.-Q., Pham, Q.-H., Nguyen, K. B. T., & Do, T.-H. (2024). Vietnamese AI Generated Text Detection (arXiv:2405.03206). arXiv. https://doi.org/10.48550/arXiv.2405.03206

Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017). Attention is All you Need. Advances in Neural Information Processing Systems 30 (NeurIPS 2017), 6000-6010. https://doi.org/10.48550/arXiv.1706.03762

Wang, P., Li, L., Ren, K., Jiang, B., Zhang, D., & Qiu, X.(2023). SeqXGPT: Sentence-Level AI-Generated Text Detection (Version 2). arXiv. https://doi.org/10.48550/ARXIV.2310.08903

Zhang, Q., Gao, C., Chen, D., Huang, Y., Huang, Y., Sun, Z., Zhang, S., Li, W., Fu, Z., Wan, Y., & Sun, L. (2024). LLM-as-a-Coauthor: Can Mixed Human-Written and Machine- Generated Text Be Detected? Findings of the Association for Computational Linguistics: NAACL 2024, 409-436. https://doi.org/10.18653/v1/2024.findings-naacl.29

Loading...