Unlocking Financial Statement Fraud Detection: Tracking Disclosure Changes via Representation Learning
Yue Yu, Zhen Wu, Yanni Han, Zhuoqun Li, Wenqi Wei
Abstract
The rapid dissemination of information through digital platforms has revolutionized the way we access and consume data, creating conditions that may lead to an increase in financial statement fraud, which jeopardizes the efficient functioning of capital markets. This paper propose a sophisticated representation learning method to detect financial statement fraud by tracking detailed changes in a firm’s Management Discussion and Analysis (MD&A) documents over time. Unlike traditional word frequency approaches, we start by aligning paragraphs between consecutive disclosures based on their representation-level similarities. Given the paragraph-embedding similarity, we categorize paragraphs into three types: added, deleted and matched. Next, we construct multivariate change trajectory representations based on fraud-relevant word categories, such as sentiment and uncertainties. Finally, we develop a fraud detection model using these word-level change trajectory representations. Extensive experiments on 24 years of financial report data, from 1995 to 2019, show that our representation learning approach significantly improves financial statement fraud detection performance across 7 different backbone machine learning models, consistently outperforming traditional word frequency-based approaches. Our method marks a new paradigm in feature engineering for financial statement fraud detection.
BibTeX
@inproceedings{icassp2025_unlockingfinanci,
title = {Unlocking Financial Statement Fraud Detection: Tracking Disclosure Changes via Representation Learning},
author = {Yue Yu and Zhen Wu and Yanni Han and Zhuoqun Li and Wenqi Wei},
booktitle = {ICASSP 2025},
year = {2025}
}