TỰ ĐỘNG HÓA XỬ LÝ HỒ SƠ VÀ TÀI LIỆU TRONG QUẢN LÝ ĐÀO TẠO DỰA TRÊN NHẬN DẠNG KÝ TỰ QUANG HỌC, NHẬN DẠNG CHỮ VIẾT TAY VÀ MÔ HÌNH NGÔN NGỮ LỚN
DOI:
https://doi.org/10.59266/houjs.2026.1360Từ khóa:
mô hình ngôn ngữ lớn (LLM), nhận dạng ký tự quang học (OCR), nhận dạng chữ viết tay (HTR), quản lý đào tạo, xử lý hồ sơ, chuyển đổi sốTóm tắt
Nghiên cứu này đề xuất một quy trình tự động hóa xử lý hồ sơ, tài liệu nhằm đáp ứng yêu cầu chuyển đổi số trong quản trị đại học, đặc biệt ở tầng nghiệp vụ. Trên cơ sở tổng quan các nghiên cứu về nhận dạng ký tự quang học (OCR), nhận dạng chữ viết tay (HTR), Document AI và xử lý tài liệu có cấu trúc, nghiên cứu phân tích đặc thù nghiệp vụ nhập liệu, lưu trữ và đối chiếu dữ liệu đối với văn bằng, chứng chỉ, bảng điểm và danh sách thi trong môi trường giáo dục đại học. Từ đó, nhóm tác giả xây dựng một mô hình xử lý theo kiến trúc mô-đun, tích hợp nhận dạng văn bản, mô hình ngôn ngữ lớn (LLM), cơ chế hậu xử lý và khâu đối chiếu có giám sát của con người nhằm bảo đảm khả năng vận hành theo quy trình khép kín từ tiếp nhận tài liệu đến chuẩn hóa dữ liệu đầu ra. Kết quả thử nghiệm trên dữ liệu nghiệp vụ thực tế cho thấy mô hình có tiềm năng giảm đáng kể khối lượng nhập liệu thủ công, đồng thời nâng cao tính nhất quán, khả năng truy xuất và hiệu quả xử lý hồ sơ. Đóng góp chính của nghiên cứu không nằm ở việc đề xuất thuật toán nhận dạng mới, mà ở việc thiết kế và thử nghiệm một qui trình xử lý (pipeline) tích hợp đầu cuối (end-to-end) phù hợp với điều kiện quản lý đào tạo tại cơ sở giáo dục đại học Việt Nam.
Tài liệu tham khảo
Barchard, K. A., và Pace, L. A. (2011). Preventing human error: The impact of data entry methods on data accuracy and statistical results. Computers in Human Behavior, 27(5), 1834-1839. https://doi.org/10.1016/j.chb.2011.04.004
Burrows, S., Gurevych, I., và Stein, B. (2015). The Eras and Trends of Automatic Short Answer Grading. International Journal of Artificial Intelligence in Education, 25(1), 60-117. https://doi.org/10.1007/s40593-014-0026-8
Clausner, C., Pletschacher, S., và Antonacopoulos, A. (2011). Aletheia - An advanced document layout and text ground-truthing system for production environments. In 2011 International Conference on Document Analysis and Recognition (ICDAR) (pp. 48- 52). IEEE. https://www.primaresearch.org/www/assets/papers/ICDAR2011Clausner_Aletheia.pdf
Gemini Team Google. (2023). Gemini: Afamily of highly capable multimodal models (arXiv:2312.11805). arXiv. https://doi.org/10.48550/arXiv.2312.11805
Gemini Team Google. (2024). Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. arXiv:2403.05530. https://doi.org/10.48550/arXiv.2403.05530
Graves, A., Fernández, S., Gomez, F., và Schmidhuber, J. (2006). Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks. In Proceedings of the 23rd International Conference on Machine Learning (pp. 369-376). https://doi.org/10.1145/1143844.1143891
Huang, Y., Lv, T., Cui, L., Lu, Y., và Wei, F. (2022). LayoutLMv3: Pre-training for document AI with unified text and image masking. In Proceedings of the 30th ACM International Conference on Multimedia (MM ‘22) (pp. 4083-4091). ACM. https://doi.org/10.1145/3503161.3548112
Kim, G., Hong, T., Yim, M., Nam, J., Park, J., Yim, J., Hwang, W., Yun, S., Han, D., và Park, S. (2022). OCR-Free document understanding transformer. In S. Avidan, G. Brostow, M. Cissé, G. M. Farinella, và T. Hassner (Eds.), Computer Vision - ECCV 2022 (Lecture Notes in Computer Science, Vol. 13688, pp. 498-517). Springer. https://doi.org/10.1007/978-3-031-19815-1_29
Li, M., Lv, T., Chen, J., Cui, L., Lu, Y., Florencio, D., Zhang, C., Li, Z., và Wei, F. (2023). TrOCR: Transformer-based optical character recognition with pre- trained models. Proceedings of the AAAI Conference on Artificial Intelligence, 37(11), 13094-13102. https://doi.org/10.1609/aaai.v37i11.26538
Retsinas, G., Sfikas, G., Gatos, B., và Nikou, C. (2022). Best practices for a handwritten text recognition system. In Document Analysis Systems (DAS 2022) (Lecture Notes in Computer Science, Vol. 13237, pp. 247-259). https://doi.org/10.1007/978-3-031-06555-217
Shi, B., Bai, X., và Yao, C. (2017). An end- to-end trainable neural network for image-based sequence recognition and its application to scene text recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence, 39(11), 2298-2304. https://doi.org/10.1109/TPAMI.2016.2646371
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., và Polosukhin, I. (2017). Attention is all you need. In Advances in Neural Information Processing Systems (Vol. 30). https://proceedings.neurips.cc/paper_files/paper/2017/file/3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf