Context-Aware Document Intelligence for Automated Compliance Evidence Extraction from Heterogeneous Financial Documents
Main Article Content
Abstract
Regulatory compliance activities across the financial sector produce a vast and growing amount of unstructured evidentiary material, such as loan agreements, non-disclosure agreements, annual reports, disclosure forms, scanned correspondence, and more — all of which auditors and regulators must review to assure compliance with disclosure, anti-money-laundering and data-protection requirements. Text-only Natural Language Processing pipelines are unable to handle this type of content because the facts relevant to compliance are often embedded in the format of a page, such as signature blocks, tabular disclosures and heading to clauses. This paper summarizes the results of two and a half years of work, including twenty-five papers from the field of transformer language modeling, the field of layout-aware document representation learning, the field of financial natural language processing and the field of applied regulatory-technology systems, all of which were published between 2015 and 2022. It explores how the performance of text-only encoders jumped from an F1 score of around sixty percent on benchmarks for form understanding to the layout-aware multimodal architectures that exceeded the ninety percent mark on the same benchmark, and reviews applied systems that lowered the manual anti-money-laundering review workload by more than fifty percent while maintaining detection rates above eighty percent. The review also outlines five benchmark datasets that test extraction tasks that are related to compliance and suggests a generalized six-stage extraction pipeline extracted from the reviewed architectures. The results show that layout-conditioned pretraining yields the highest improvements, not just a parameter scale, and that the human verification is a constant characteristic of all the deployed systems studied.
Article Details

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
References
[1] I. Anagnostopoulos, “Fintech and regtech: Impact on regulators and banks,” J. Econ. Bus., vol. 100, pp. 7–25, 2018, doi: 10.1016/j.jeconbus.2018.07.003.
[2] G. Gasparri, “Risks and opportunities of RegTech and SupTech developments,” Front. Artif. Intell., vol. 2, Art. no. 14, 2019, doi: 10.3389/frai.2019.00014.
[3] J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in Proc. 2019 Conf. North American Chapter Assoc. Comput. Linguistics: Human Language Technologies, vol. 1, 2019, pp. 4171–4186, doi: 10.18653/v1/N19-1423.
[4] Y. Xu, M. Li, L. Cui, S. Huang, F. Wei, and M. Zhou, “LayoutLM: Pre-training of text and layout for document image understanding,” in Proc. 26th ACM SIGKDD Int. Conf. Knowledge Discovery & Data Mining, 2020, pp. 1192–1200, doi: 10.1145/3394486.3403172.
[5] M. E. Peters et al., “Deep contextualized word representations,” in Proc. 2018 Conf. North American Chapter Assoc. Comput. Linguistics: Human Language Technologies, vol. 1, 2018, pp. 2227–2237, doi: 10.18653/v1/N18-1202.
[6] Y. Liu et al., “RoBERTa: A robustly optimized BERT pretraining approach,” arXiv:1907.11692, 2019, doi: 10.48550/arXiv.1907.11692.
[7] P. Rajpurkar, J. Zhang, K. Lopyrev, and P. Liang, “SQuAD: 100,000+ questions for machine comprehension of text,” in Proc. 2016 Conf. Empirical Methods in Natural Language Processing, 2016, pp. 2383–2392, doi: 10.18653/v1/D16-1264.
[8] A. Vaswani et al., “Attention is all you need,” arXiv:1706.03762, 2017, doi: 10.48550/arXiv.1706.03762.
[9] G. Jaume, H. K. Ekenel, and J.-P. Thiran, “FUNSD: A dataset for form understanding in noisy scanned documents,” in 2019 Int. Conf. Document Analysis and Recognition Workshops (ICDARW), vol. 2, 2019, pp. 1–6, doi: 10.1109/ICDARW.2019.10029.
[10] Y. Xu et al., “LayoutLMv2: Multi-modal pre-training for visually-rich document understanding,” in Proc. 59th Annu. Meeting Assoc. Comput. Linguistics and 11th Int. Joint Conf. Natural Language Processing, vol. 1, 2021, pp. 2579–2591, doi: 10.18653/v1/2021.acl-long.201.
[11] Y. Huang, T. Lv, L. Cui, Y. Lu, and F. Wei, “LayoutLMv3: Pre-training for document AI with unified text and image masking,” in Proc. 30th ACM Int. Conf. Multimedia, 2022, pp. 4083–4091, doi: 10.1145/3503161.3548112.
[12] G. Lample, M. Ballesteros, S. Subramanian, K. Kawakami, and C. Dyer, “Neural architectures for named entity recognition,” in Proc. 2016 Conf. North American Chapter Assoc. Comput. Linguistics: Human Language Technologies, 2016, pp. 260–270, doi: 10.18653/v1/N16-1030.
[13] J. Li, A. Sun, J. Han, and C. Li, “A survey on deep learning for named entity recognition,” IEEE Trans. Knowl. Data Eng., vol. 34, no. 1, pp. 50–70, 2022, doi: 10.1109/TKDE.2020.2981314.
[14] S. Francis, J. Van Landeghem, and M.-F. Moens, “Transfer learning for named entity recognition in financial and biomedical documents,” Information, vol. 10, no. 8, Art. no. 248, 2019, doi: 10.3390/info10080248.
[15] D. Araci, “FinBERT: Financial sentiment analysis with pre-trained language models,” arXiv:1908.10063, 2019, doi: 10.48550/arXiv.1908.10063.
[16] A. H. Huang, H. Wang, and Y. Yang, “FinBERT: A large language model for extracting information from financial text,” Contemp. Account. Res., vol. 40, no. 2, pp. 806–841, 2022, doi: 10.1111/1911-3846.12832.
[17] T. Stanisławek et al., “Kleister: Key information extraction datasets involving long documents with complex layouts,” in Document Analysis and Recognition – ICDAR 2021, 2021, pp. 564–579, doi: 10.1007/978-3-030-86549-8_36.
[18] D. Hendrycks, C. Burns, A. Chen, and S. Ball, “CUAD: An expert-annotated NLP dataset for legal contract review,” arXiv:2103.06268, 2021, doi: 10.48550/arXiv.2103.06268.
[19] M. Mathew, D. Karatzas, and C. V. Jawahar, “DocVQA: A dataset for VQA on document images,” in Proc. IEEE/CVF Winter Conf. Applications of Computer Vision (WACV), 2021, pp. 2200–2209, doi: 10.1109/WACV48630.2021.00225.
[20] A. W. Harley, A. Ufkes, and K. G. Derpanis, “Evaluation of deep convolutional nets for document image classification and retrieval,” in Proc. 2015 13th Int. Conf. Document Analysis and Recognition (ICDAR), 2015, pp. 991–995, doi: 10.1109/ICDAR.2015.7333910.
[21] P. Craja, A. Kim, and S. Lessmann, “Deep learning for detecting financial statement fraud,” Decis. Support Syst., vol. 139, Art. no. 113421, 2020, doi: 10.1016/j.dss.2020.113421.
[22] M. Jullum, A. Løland, R. B. Huseby, G. Ånonsen, and J. Lorentzen, “Detecting money laundering transactions with machine learning,” J. Money Laund. Control, vol. 23, no. 1, pp. 173–186, 2020, doi: 10.1108/JMLC-07-2019-0055.
[23] S. Liu, B. Zhao, R. Guo, G. Meng, F. Zhang, and M. Zhang, “Have you been properly notified? Automatic compliance analysis of privacy policy text with GDPR Article 13,” in Proc. Web Conf. 2021, 2021, pp. 2154–2164, doi: 10.1145/3442381.3450022.
[24] P. Riba, A. Dutta, L. Goldmann, A. Fornés, O. Ramos, and J. Lladós, “Table detection in invoice documents by graph neural networks,” in 2019 Int. Conf. Document Analysis and Recognition (ICDAR), 2019, pp. 122–127, doi: 10.1109/ICDAR.2019.00028.
[25] W. Yu, N. Lu, X. Qi, P. Gong, and R. Xiao, “PICK: Processing key information extraction from documents using improved graph learning-convolutional networks,” in 2020 25th Int. Conf. Pattern Recognition (ICPR), 2021, pp. 4363–4370, doi: 10.1109/ICPR48806.2021.9412927.