Source-linked AI summary
A Review of Financial Accounting Fraud Detection based on Data Mining Techniques
Anuj Sharma, Prabin Kumar Panigrahi
TL;DR
Financial accounting fraud detection remains difficult because internal auditing may miss fraud and financial data are large and complex. This paper reviews data-mining applications, proposes a detection framework, and finds classification techniques most extensively applied.
Problem
Financial accounting fraud detection remains difficult because internal auditing may fail to identify fraud and financial data are large and complex.
Method
The paper systematically reviews data-mining research, classifies applications into six classes, and proposes an expanded framework for fraud detection.
Results
Logistic models, neural networks, Bayesian belief networks, and decision trees have been applied most extensively, with reported accuracies reaching 95.1% for logistic regression.
Takeaways & Limitations
The proposed framework may help certified public accountants select suitable data and data-mining technologies for detecting fraud and provide a foundation for future research.
Takeaways & Limitations
The review searched only some online databases for articles published between 1992 and 2011 and considered only financial-accounting fraud articles.
Abstract
from arXiv · showhide
With an upsurge in financial accounting fraud in the current economic scenario experienced, financial accounting fraud detection (FAFD) has become an emerging topic of great importance for academic, research and industries. The failure of internal auditing system of the organization in identifying the accounting frauds has lead to use of specialized procedures to detect financial accounting fraud, collective known as forensic accounting. Data mining techniques are providing great aid in financial accounting fraud detection, since dealing with the large data volumes and complexities of financial data are big challenges for forensic accounting. This paper presents a comprehensive review of the literature on the application of data mining techniques for the detection of financial accounting fraud and proposes a framework for data mining techniques based accounting fraud detection. The systematic and comprehensive literature review of the data mining techniques applicable to financial accounting fraud detection may provide a foundation to future research in this field. The findings of this review show that data mining techniques like logistic models, neural networks, Bayesian belief network, and decision trees have been applied most extensively to provide primary solutions to the problems inherent in the detection and classification of fraudulent data.
1. INTRODUCTION
Financial accounting fraud detection has gained importance as traditional auditing struggles with complex, deceptive, and infrequent fraud. This paper reviews data-mining applications and proposes a framework to support fraud-detection practice and future research.
- Motivation: Financial accounting fraud is an increasingly serious problem, while traditional internal auditing procedures may be insufficient for effective detection.Auditors may lack fraud knowledge and experience, while insiders may intentionally deceive auditors.
- Motivation: Forensic accounting uses accounting, auditing, and investigative skills to detect fraud that internal auditing fails to identify.Its adoption follows failures of organizational internal-auditing systems.
- Data mining: Data mining can help auditors analyze large databases, identify statistically reliable patterns, and support fraud-detection decisions.It may help reconcile the effectiveness and efficiency of fraud detection.
- Contribution: The paper systematically reviews data-mining techniques for financial accounting fraud detection and proposes a framework to help certified public accountants select suitable data and technologies.The review is intended to provide a foundation for future research.
- Organization: The paper organizes prior work around data-mining classifications, applications, literature distributions, a detailed framework, and future research directions.The planned sections cover classification, applications, research distribution, framework details, and future directions.
2. CLASSIFICATION OF DATA MINING TECHNIQUES FOR FRAUD DETECTION
The framework organizes data-mining applications into six classes and reviews the principal algorithms used for financial accounting fraud detection. Logistic models, neural networks, Bayesian belief networks, and decision trees receive the most extensive attention among classification techniques.
- Framework structure: The framework comprises classification, clustering, prediction, outlier detection, regression, and visualization, supported by algorithmic approaches.These classes organize how data-mining methods extract relevant relationships for fraud detection.
- Classification: Classification predicts predefined categorical labels from training data to distinguish objects across classes.Common techniques include neural networks, Naïve Bayes, decision trees, and support vector machines.
- Prediction: Neural networks and logistic models are the most commonly used techniques for prediction, which estimates continuous-valued ordered future attributes.Prediction differs from classification because its target attribute is continuous-valued rather than categorical.
- Classification techniques: Logistic models, neural networks, Bayesian belief networks, and decision trees are the four classification techniques discussed most extensively.The review identifies these approaches as the principal techniques in the literature.
- Regression models: Logistic regression reached up to 95.1% detecting accuracy in reported accounting-fraud research.Regression-based models are mostly used in financial accounting fraud detection, with logistic regression being common.
- Neural networks: A neural-network model using financial ratios and qualitative variables was more effective than linear and quadratic discriminant analysis and logistic regression.The comparison evaluated alternative statistical methods against the neural-network fraud-detection model.
- Hybrid methods: Fuzzy neural networks reportedly outperformed traditional statistical models and neural-network models in prior studies.The literature also reports generalized adaptive neural-network architectures and adaptive logic networks for fraud detection.
- Bayesian belief networks: 90.3% of the validation sample was correctly classified by a Bayesian belief network, which outperformed neural network and decision-tree methods.The Bayesian belief network represents conditional independencies among random variables using a directed acyclic graph.
3. THE RESEARCH UNDER PROPOSED CLASSIFICATION FRAMEWORK
The reviewed literature is organized chronologically across major data-mining technique families for financial accounting fraud detection. The tables cover neural networks, regression models, fuzzy logic, expert systems, and genetic algorithms.
- Literature distribution: Tables 1 through 4 distribute the reviewed research papers according to publication year.The tables present the literature in chronological order.
- Neural networks: Table 1 covers neural-network research for financial accounting fraud detection.It presents the reviewed work in order of publication.
- Regression models: Table 2 covers regression-model research for financial accounting fraud detection.The table is part of the chronological literature review.
- Other techniques: Table 3 presents research on fuzzy logic, while Table 4 covers expert systems and genetic algorithms.Together, the tables summarize additional technique families used in financial accounting fraud detection.
4. DATA MINING BASED FRAMEWORK FOR FRAUD DETECTION
The paper proposes an expanded generic data-mining framework tailored to financial accounting fraud detection. It follows the general data-mining information flow while incorporating fraud-detection-specific characteristics.
- Framework process: The framework proceeds through feature selection, representation, data collection and management, pre-processing, data mining, post-processing, and performance evaluation.This sequence describes the information flow used for financial accounting fraud-detection techniques.
- Fraud-detection adaptation: The proposed framework expands the generic data-mining process to address specific characteristics of financial accounting fraud detection.The framework is presented as an adaptation of established data-mining information flow.
RESEARCH
The review examines data-mining applications for financial accounting fraud detection, classifying prior work across several technique families and proposing an expanded framework. It identifies limited use of outlier detection and visualization, data-access constraints, cost sensitivity, and bounded review coverage as important research considerations.
- Literature review: The review covers statistical tests, regression analysis, neural networks, decision trees, and Bayesian networks for financial accounting fraud detection.Regression analysis is described as widely used because of its explanatory ability, with Logit, Step-wise Logistic, UTADIS, and EGB2 among the models reported.
- Research gaps: Outlier detection and visualization have seen limited use despite suitability for identifying rare fraudulent patterns and presenting data anomalies.The review characterizes outlier detection as complex and visualization as potentially useful for identifying and quantifying fraud schemes.
- Research gaps: Obtaining sufficient research data, especially fraudulent financial statements, remains a major obstacle to financial accounting fraud detection research.The paper connects this challenge to the limited number of relevant journal articles and calls for closer practitioner–researcher collaboration.
- Research gaps: Few studies explicitly include cost sensitivity, although false negative errors are usually more costly than false positive errors.The review recommends that future data-mining research account for differing misclassification costs.
- Limitations: The review searched selected online databases for articles published between 1992 and 2011 and considered only financial-accounting-fraud articles.The authors state that the study is not exhaustive and could be expanded to other fraud domains.