Intelligent Project Reimbursement Material Organization Mini-Program based on OCR Technology
Keywords:
OCR Technology, Intelligent Mini-Program, Project Reimbursement, Edge AI, AI in IndustryAbstract
The increasing volume and complexity of project reimbursement materials in universities have highlighted the limitations of traditional manual processing methods, which are time-consuming, error-prone, and inefficient. This paper presents an intelligent mini-program based on optical character recognition (OCR) technology designed to automate the organization and data extraction of reimbursement documents. The system integrates a WeChat mini-program front-end with a Flask back-end server and a MySQL database, employing a multi-strategy OCR scheduling mechanism that dynamically selects the most appropriate OCR engine (e.g., Tesseract for printed text, commercial engines for handwritten or low-quality images) based on document characteristics. Extracted data are cleaned and validated using regular expressions, and a finite state machine manages real-time processing status tracking. The system also generates standardized Excel reports compatible with university financial department requirements. Experimental results demonstrate that the proposed system achieves 97.3% character recognition accuracy and 94.7% field extraction accuracy, with an average response time of 3.2 seconds per document, outperforming Tesseract OCR and Abbyy FlexiCapture baselines. The mini-program significantly improves efficiency and accuracy in reimbursement material processing, reduces manual labor, and promotes digital transformation in university financial management. This research contributes to the application of AI-driven document intelligence in industry, providing a scalable and adaptable solution for automated financial workflows.
References
Mori, S., Suen, C. Y., & Yamamoto, K. (1992). Historical review of OCR research and development. Proceedings of the IEEE, 80(7), 1029-1058.
Smith, R. (2007). An overview of the Tesseract OCR engine. In Ninth International Conference on Document Analysis and Recognition (ICDAR 2007) (Vol. 2, pp. 629-633). IEEE.
Nguyen, T. T. H., Jatowt, A., Coustaty, M., & Doucet, A. (2022). Survey of post-OCR processing approaches. ACM Computing Surveys, 54(6), 1-37.
LeCun, Y., Bengio, Y., & Hinton, G. (2015). Deep learning. Nature, 521(7553), 436-444.
Graves, A., & Schmidhuber, J. (2009). Offline handwriting recognition with multidimensional recurrent neural networks. In Advances in Neural Information Processing Systems (Vol. 21, pp. 545-552). Curran Associates.
Li, M., & Wu, Y. (2020). Research on invoice recognition system based on OCR and deep learning. Journal of Physics: Conference Series, 1631(1), 012034.
Kumar, M., & Jindal, M. K. (2021). A systematic review on OCR techniques for handwritten documents. Archives of Computational Methods in Engineering, 28(5), 3511-3533.
Jaderberg, M., Simonyan, K., Vedaldi, A., & Zisserman, A. (2016). Reading text in the wild with convolutional neural networks. International Journal of Computer Vision, 116(1), 1-20.
Chandio, A. A., & Asikuzzaman, M. (2022). Automated invoice data extraction using OCR and NLP. In 2022 International Conference on Digital Image Processing (pp. 45-52). IEEE.
Zhang, Y., & Wang, L. (2019). Design and implementation of financial reimbursement system based on OCR. In 2019 IEEE 4th Advanced Information Technology (pp. 112-117). IEEE.
Jain, A. K., & Yu, D. (2019). Text recognition using deep learning. In Handbook of Pattern Recognition and Computer Vision (pp. 223-245). World Scientific.
Shi, B., Bai, X., & Yao, C. (2017). An end-to-end trainable neural network for image-based sequence recognition and its application to scene text recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence, 39(11), 2298-2304.
Liu, X., & Chen, Z. (2021). A lightweight OCR system for mobile devices. In 2021 IEEE International Conference on Mobile Computing (pp. 78-84). IEEE.
Deng, L., & Yu, D. (2014). Deep learning: methods and applications. Foundations and Trends in Signal Processing, 7(3-4), 197-387.
Sermanet, P., & LeCun, Y. (2011). Traffic sign recognition with multi-scale convolutional networks. In The 2011 International Joint Conference on Neural Networks (pp. 2809-2813). IEEE.
