Next Article in Journal
Comparing Meta-Learners for Estimating Heterogeneous Treatment Effects and Conducting Sensitivity Analyses
Previous Article in Journal
Prediction of Compressive Strength in Fine-Grained Soils Stabilized with a Combination of Various Stabilization Agents and Nano-SiO2 Using Machine Learning Algorithms
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

An Interpretable Financial Statement Fraud Detection Framework Enhanced by Temporal–Spatial Patterns

School of Accounting, Chongqing University of Technology, Chongqing 400054, China
*
Author to whom correspondence should be addressed.
Math. Comput. Appl. 2025, 30(6), 138; https://doi.org/10.3390/mca30060138
Submission received: 27 October 2025 / Revised: 4 December 2025 / Accepted: 10 December 2025 / Published: 15 December 2025

Abstract

In recent years, financial statement fraud schemes have evolved to become markedly more sophisticated and concealed, thereby posing severe threats to both social stability and economic health. Traditional detection methods, which rely primarily on fragmented corporate data, exhibit significant limitations in capturing the dynamic evolution and spatial diffusion characteristics of fraudulent behaviors over time and space. To address this issue, in this study, we undertake a thorough analysis of the intrinsic nature of fraud risk from a sociotechnical systems perspective and construct a multi-level indicator system to comprehensively quantify risk elements. Furthermore, recognizing the dynamic evolution nature and propagating characteristics of fraud risk, we propose a novel financial statement fraud detection framework to capture behavior patterns in temporal and spatial dimensions. Experiments on A-share-listed companies of high-risk industries in China demonstrate that the proposed framework significantly outperforms other mainstream machine learning and deep learning techniques. In addition, we open the “black box” of the detection framework and empirically validate fraud risk patterns with respect to social–technical elements by leveraging explainable AI techniques. Practically, the proposed framework and interpretable analysis are capable of providing precise early warnings and supervision.

1. Introduction

In contemporary business practice, enterprise organizations usually function as a complex sociotechnical system (STS), where human actors and technological tools work in collaboration to achieve specific objectives [1,2,3]. In this system, rapid variations in the external environment have significantly intensified pressures, such as unrealistic business expectations and fierce market competition, which can push individuals to commit fraud [4,5,6]. In addition, advanced information technology is increasingly exploited to conduct fraud, posing a great challenge for internal control and governance structure in detection and prevention [7,8,9,10,11]. In this way, instances of fraud often exhibit the following characteristics: strong motivation, concealment, and intelligence.
There are various types of fraud, among which we focus on financial statement fraud in this study. Financial statement fraud refers to deliberately engaging in the misleading manipulation of financial statements, violating accounting standards and regulations to obtain improper economic benefits [12,13,14]. In 2021, Shanghai Electric was exposed to financial fraud through its “private network communication business”, involving 15 listed companies and exceeding CNY 90 billion [15]. Fraudsters conduct private network communication business by circulating it among third parties without any substantive content [16]. Using this insidious, sophisticated, and systematic fraud scheme, financial statement fraud can evade traditional manual audits successfully. As one of the most prevalent types of occupational fraud in the world, it has caused severe adverse effects on capital markets and the broader socioeconomic environment [17,18,19]. Hence, it is necessary to construct an intelligent model for financial statement fraud detection (FSFD), which would function as an early warning mechanism for proactive fraud prediction, enabling timely intervention to mitigate losses and significantly supporting an organization’s capacity for resilience [20,21].
Preliminarily, most detection approaches employ financial data and corresponding financial indices as core carriers for FSFD [22]. In addition, many studies have shifted to exploring non-financial data, such as financial report data, governance data, and internal control data, for better detection performance [23,24,25,26,27]. However, financial statement fraud is not an abrupt occurrence but rather the outcome of long-term interactions among various factors within the enterprise STS. These data—both traditional financial data and newly incorporated non-financial data—predominantly represent static features of a certain period [28], and they have difficulty in capturing dynamically changing fraud risks. For instance, insidious fraudulent activities, such as “private network communication business”, are often conducted through circulation among third parties without substantive content, rendering traditional analyses based on static data largely ineffective.
With the rapid development of information technology, detection models have evolved beyond traditional rule-based systems and machine learning and deep learning techniques [29,30]. Although existing models have demonstrated continuously improved performance by extracting increasingly complex features, especially in deep learning, they lack the capability to systematically represent the holistic picture of fraud risks driven by the interplay of four sociotechnical elements.
In this study, multiple inputs from a sociotechnical perspective are studied. According to Leavitt [2], four key elements—namely, task, technology, actor, and structure—need to be considered in an enterprise system. Each element can be decomposed into a set of indicators, and each indicator can be quantified and taken as an input. Based on this framework, we attempt to resolve complex and dynamic fraud mechanisms and, furthermore, warn of fraud risk using these quantifiable inputs.
We propose a method for financial statement fraud detection enhanced by temporal–spatial patterns (FSFD-ETSP). Since four sociotechnical elements interact and evolve over time and risk behaviors spread among peers and supply chains, we integrate long short-term memory (LSTM) networks and clustering approaches to comprehensively capture the nuanced fraudulent patterns in both temporal and spatial dimensions. Specifically, LSTM can be used to analyze temporal changes among various inputs, extract specific shapes and features that repeatedly occur in time series, and be applied to characterize temporal patterns. Clustering algorithms can classify companies into groups. In each group, companies are categorized according to their similarity in the spatial dimension. In our framework, we assemble the attention mechanism to combine these two types of patterns, and transform them for FSFD through a fully connected layer network. Compared with other commonly used machine learning and deep learning models, the proposed framework exhibits higher detection performance with respect to real-world data.
Another achievement of this study is the examination of potential fraud risk suggested by corporate data through model interpretation. We apply the SHAP (SHapley Additive exPlanation) method to identify the most influential indicators in the model’s predictions. The resulting patterns provide holistic explanations consistent with an STS perspective on fraud risk, thereby ensuring the transparency of the detection process [31,32].
To improve the anticipation capability of organizational resilience [33], we introduce STS theory into fraud detection. Based on this, we then propose a specific temporal–spatial multi-input framework with SHAP-based interpretation in practice. Our work moves beyond traditional methods that rely on experience-based data accumulation and increasingly complex feature extraction, enabling a profound understanding of sophisticated fraud schemes within an enterprise STS, the construction of a novel systematic indicator system, and the capture of nuanced temporal–spatial patterns for superior detection. The key contributions of this study are as follows:
(1)
Beyond the traditional method of accumulating data for detection, we construct a multi-input set by mapping four core elements (task, technology, actor, and structure) within an enterprise STS, ensuring comprehensive coverage of all key factors where fraud may originate.
(2)
Considering the dynamic evolution and spatial diffusion of fraud risk within an enterprise STS, we innovatively combine supervised LSTM networks and an unsupervised clustering approach to extract both temporal and spatial patterns with an attention mechanism. With spatiotemporal patterns, the proposed FSFD-ETSP significantly outperforms traditional machine learning and other deep learning methods.
(3)
We introduce explainable AI techniques to analyze the most influential indicators and the associations among them, which align with the fraud risk pattern described by the enterprise STS.
The primary objectives of this study are:
(1)
To construct a multi-level indicator system as input to comprehensively cover potential risk sources according to STS theory;
(2)
To develop a temporal–spatial enhanced framework for FSFD by integrating LSTM networks and a clustering approach to effectively capture the nuanced behavior patterns of fraud risk;
(3)
To leverage the explainable AI techniques to ensure the trustworthiness and transparency of the proposed framework.
The rest of this article is organized as follows. In Section 2, we review the related literature, and in Section 3, we provide an overview of the framework and implementation details. In Section 4, we outline the experimental design, and in Section 5, we report on the results and further research. In Section 6, we summarize the thesis and propose future research directions.

2. Related Research

2.1. Conceptual Literature

As a systematic perspective on organizations [1], an STS can be viewed as an integrated whole, wherein the social subsystem and technical subsystem are inextricably intertwined [34,35]. The system usually comprises four core elements [2]: actors, social subsystem elements, and all stakeholders related to organizations; technology, technical subsystem elements, tools, equipment, and technical platforms that assist organizations; structures, social subsystem elements, and institutional arrangements, such as power distribution and management; and tasks, technical subsystem elements, and the specific goals of organizations.
Modern enterprises, as typical STSs, can be regarded as comprising the dynamic coupling of the two subsystems [36]. In this framework, financial or personal stresses, such as financial leverage and liquidity [37], can compel individuals to engage in fraudulent behavior. Actors, such as managers, find opportunities within the structural and technological elements to commit fraud, e.g., by exploiting weak ethical leadership and poor monitoring [20]. As for peers, suppliers, customers, and other company-level actors, unethical actions with respect to structural and technological elements can propagate and cause negative multiplier effects [38]. Based on these phenomena, we can conclude that the likelihood of intentional fraudulent behavior is continuously changing; that is, financial fraud risk is dynamic. One key manifestation of this risk is financial statement fraud [19].
As for the aforementioned risk, two solutions—a financial safety performance framework and a risk management system—are proposed in the enterprise STS, targeting social and technical subsystems, respectively [39]. The former solution, functioning as “a controller and decider”, normally comprises a performance evaluation mechanism for monitoring financial security and optimizing operations based on feedback [40]. The latter solution, serving as a critical tool for identifying potential fraud, can either collect company data in a timely manner (early warning mechanism) [41,42,43] or employ various models for statistical prediction (risk modeling) [44,45,46]. These two solutions are mutually reinforcing, with early warnings providing risk data and alert information, and the safety performance framework carrying out further actions to preserve financial security. Additionally, embedding a financial safety performance framework and an early warning system into an enterprise STS can effectively contribute to organizational resilience by enhancing early detection and risk insight. Organizational resilience refers to the ability of an organization to anticipate, withstand, respond to, and adapt to unexpected crises [47,48]. With these two solutions, the whole system can form a synergistic defense mechanism encompassing prevention, in-process control, and post-event improvement [42]. It is important to note that organizational resilience is inferred from the improved early warning and risk insight provided by the detection framework, rather than being observed directly. Accordingly, we can outline the conceptual framework in Figure 1.

2.2. Financial Statement Detection Methods

FSFD is an important task for auditors and accountants. In the early days, FSFD mainly relied on the empirical judgment of auditors [49]. With the development of artificial intelligence (AI), modern FSFD methods are mainly based on machine learning [18,50,51,52,53,54,55,56,57] and deep learning technologies [58,59]. The efficacy of these techniques is dependent on their ability to carry out discriminative feature extraction. For example, traditional machine learning methods, such as logistic regression [60,61,62,63], decision tree [64,65,66], and support vector machine [17,54,67,68,69], are capable of capturing linear relations among companies [70,71]. Meanwhile, deep learning methods have the potential to learn nonlinear and non-stationary correlations, rendering them state-of-the-art detection methods [72,73]. Table 1 provides a detailed description of the major body of research on FSFD.
According to Table 1, many studies were conducted primarily based on financial data, for which their lagging nature constrains predictive performance. Moreover, many other data types, such as textual data [66,69], are employed for FSFD, as they can complementarily provide real-time soft information regarding corporate behavior, managerial intent, and social interactions. By complementarily reducing information asymmetry, they demonstrate excellent detection performance. Nevertheless, these feature extraction approaches only focus on fragmented data, lacking a unified theoretical framework for feature integration and for revealing fraud mechanisms. This encourages further exploration from an STS perspective.

2.3. Deep Learning Models

Deep learning, with its powerful feature extraction and pattern recognition capabilities, has achieved significant breakthroughs in natural language processing [26]. By building multi-layer neural networks, deep learning models automatically learn high-dimensional feature representations, greatly enhancing prediction accuracy. Notably, convolutional neural networks (CNNs), recurrent neural networks (RNNs), and their variants, such as long short-term memory (LSTM) and gated recurrent units (GRUs), effectively address the long-term dependency problem through recurrent connections and gate control mechanisms.
CNN: The CNN architecture comprises stacked convolutional layers, fully connected layers, and an output layer. The convolutional layers extract localized features, while pooling layers perform downsampling to reduce dimensions. Dropout regularization is employed during training to prevent overfitting. This hierarchical design facilitates the progressive learning of multi-scale spatial representations, which are optimized via a loss function.
RNN: Unlike traditional feedforward neural networks, RNNs establish circular connections between hidden layers to capture long-term dependencies and dynamically update hidden states. However, they are prone to gradient vanishing or explosion when processing long sequences, limiting their effectiveness in long-term dependency tasks.
LSTM: LSTM is a variant of RNN. It can automatically learn the feature representations of MTS and use embeddings in end-to-end models. To control the storage unit, the LSTM requires three different gates.
Firstly, a forget gate is required to determine which information in the cell state should be discarded or retained. It outputs a value between 0 and 1 through the sigmoid activation function, indicating the proportion of information retained from the previous cell state. With respect to the values, 1 means “fully retained,” while 0 means “completely discarded.” The formula is as follows:
f t   =   σ W f h t 1 , X t   +   b f  
where   σ is the sigmoid activation function; W f and b f are the weight and bias of the forgetting gate; h t 1 is the hidden state of the previous time step; X t is the input of the current time step.
Secondly, an input gate is needed to determine new information that should be stored or updated in the cell state. It consists of two components: a sigmoid layer responsible for determining the updated extent (2) and a candidate cell state derived from the tanh activation function (3):
j t   =   σ W j · h t 1 , X t   +   b j
c ^ t   = tanh W c · h t 1 , X t + b c
where W j ,     W c ,     b j , and b c are the weights and biases of the input gates.
Furthermore, an output gate is needed to determine what information from the cell state should be output. Specifically, it produces a value between 0 and 1 via the sigmoid activation function. This value is subsequently multiplied by the candidate activation (the tanh function) to compute the new hidden state:
o t   =   σ W o · h t 1 , X t   +   b o
h t = o t · tanh c t
where W o and b o are the weight and bias of the output gate.
In addition to the gating mechanism, LSTM also introduces a cell state for storing long-term information. It combines the forgotten information in the cell state with the newly added information to obtain the updated cell state c t . The formula is as follows:
C t   =   ft   ·   C t 1   +   jt   ·   C ^ t
GRU: The GRU is a simplified version of LSTM. It combines the forget and input gates into an update gate to control what information is kept or discarded. It also combines the cell state and hidden state, reducing the number of parameters to two. The reset gate determines the retention of prior hidden state information, while the update gate adjusts the integration of new inputs into the current state. The resulting model is simpler than the standard LSTM model, and is hence becoming increasingly popular.
Attention mechanism: In neural networks, attention refers to a mechanism that simulates human cognitive attention [74]. The inspiration for this concept arises from the limitations of human attention, which cannot effectively process all significant information [26]. Instead, it can selectively concentrate on the most critical elements to acquire insights. Therefore, the core idea of the attention mechanism is to highlight the most relevant parts of input data rather than treating all parts equally.
Previous research has confirmed attention mechanisms to be a crucial component of predictive models, especially in the financial field, with their introduction enhancing the interpretability and predictive accuracy of models.
In this study, an attention mechanism is used to dynamically weight and integrate features according to their predictive importance. This approach amplifies impactful features to enhance model performance.

3. An Interpretable Financial Statement Fraud Detection Framework Enhanced by Temporal–Spatial Patterns (FSFD-ETSP)

In this study, we propose an interpretable financial statement fraud detection framework enhanced by temporal–spatial patterns (FSFD-ETSP), as shown in Figure 2. According to an STS, the framework firstly integrates diverse data as model inputs: operational and financial data for the technology element; audit and governance data for the structural element; social responsibility reports for the actor element; and key performance metrics for the task element. Then, we extract the behavior patterns of the spatial and temporal dimensions among peer firms using a clustering algorithm and LSTM networks, respectively. An attention mechanism is also incorporated to improve feature fusion by adaptively setting weights. Finally, the extracted features are fed into a fully connected (dense) layer for FSFD. In addition, we apply the SHAP method to provide holistic feature explanations with respect to FSFD, thereby validating fraud risks using an STS lens and ensuring detection transparency.

3.1. Multiple Inputs

From the perspective of an STS, an enterprise is a complex system in which task, technology, structure, and actor elements interact dynamically. Financial statement fraud is not caused by a single factor but results from the imbalance and maladaptive interactions of these four elements. To systematically analyze such risks, we have constructed a multi-level indicator system that is explicitly mapped onto the STS, as outlined in Table 2. The quantification of the four core elements corresponds to the following specific indicators. Within this system, diverse company data are organized through an STS lens: financial and operational data correspond to the technology element; audit and corporate governance data reflect the structural element; social responsibility reports pertain to the actor element; key operational metrics align with the task element.
The task element encompasses the firm’s objectives and external pressures. Here, unrealistic performance targets (X13), often stemming from supply chain pressures (X11) or business predicaments (X12), can create a powerful motive for fraud.
The technology element comprises tools and outputs for information processing and task achievement. With the misuse of technical tools, such as applying aggressive accounting treatment methods or directly beautifying reports, financial statements (X21) with financial information and their authenticity—dependent on the quality of accounting information (X22)—can be manipulated.
The structural element refers to institutional arrangements and the control environment, which determine fraud opportunities. A robust internal control system (X33) is the primary defense, while the firm’s equity (X34) and shareholder structures (X35) constitute the governance foundation. External audits, reflected in the audit opinion (X31) and audit fees (X32), represent the external assessment and supervision of this internal structure.
The actor element consists of decision-makers and executors within the system, whose behaviors directly precipitate fraud. Litigation X41 is a direct manifestation of conflicts among actors. A social responsibility report X42 can, to some extent, reflect the management’s strategic vision and capability. Furthermore, geographical location X43 and registered capital X44 help to analyze the behavioral patterns of actors at the company level, considering the existence of peer effects.
As multiple inputs for FSFD, this STS-based multi-level indicator system goes beyond examining factors in isolation. It enables a systemic diagnosis of fraud risk by providing a critical linkage between task, technology, structure, and actor elements. By analyzing these interactions, regulators, auditors, and researchers can predict fraud risks in a more precise and timely manner, which is crucial for developing robust and preemptive anti-fraud measures.

3.2. Spatial Behavior Pattern Extraction

From an STS perspective, fraud risk is not a sudden, isolated event but a dynamically evolving process within an enterprise organization. It originates from intense task pressure and escalates as actors exploit systemic vulnerabilities in structures and technology. These interconnections dynamically change and potentially culminate in a major fraud incident.
Simultaneously, we should note that, in reality, peer firms exhibit structural correlations that create channels for risk propagation. When fraudulent behavior emerges in one firm, it can spread through these connected channels, generating negative social multiplier effects. In this study, we employ K-means clustering—an unsupervised learning method—to group firms together and define “peer firms” based on feature similarity. Concerning the dynamic nature of fraud risk, we profile each company with its STS-based characteristics spanning all time points instead of at a single time point. Through the K-means approach, each resulting cluster represents a distinct configuration of these spatial sociotechnical characteristics. There is no doubt that companies exhibiting high-risk fraud behavior patterns are typically grouped together. This also provides crucial insights for targeted monitoring and early warning systems focused on high-risk STS profiles.
As shown in Figure 3a, the primary input is a tensor X of three ranks, namely, time, feature, and firm. Any single sample slice on the firm dimension represents a company’s profile over time. As illustrated in Figure 3b, for company c, its corresponding time–feature matrix is as follows:
X i , t c = x i , 1 , x i , 2 , , x i , t ( i [ 1 , j ] , t [ 1 , z ] )
where x i , t is the ith feature of company c at time t, resulting from the STS-based multi-level indicator system outlined in Table 2. Given the high dimensions of the feature vector for each company, the principal component analysis (PCA) approach—comprising the one-dimension reduction technique—is applied to reduce computational complexity and mitigate the curse of dimensionality while preserving as much information as possible. According to the PCA principle, an orthogonal projection matrix P can be derived from the computed principal components. Subsequently, as shown in formula (8) and Figure 3c, the time–feature matrix X i , t c of company c can be projected into a lower-dimensional space using matrix P:
Y i , t c = P X i , t c
where i [ 1 , j ] , i [ 1 , j ] , and j < j . Based on the low-dimensional time–feature matrix, we can construct a feature vector for each company. According to Figure 3d and formula (9), the time–feature matrix Y i , t c of company c can be flatted into a unified 1D vector Y c . The stretched vector integrates the STS-based characteristics of all time points.
Y c = [ y 1 , 1 , y 1 , 2 , , y 1 , t , , y i , 1 , y i , 2 , , y i , t , , y j , 1 , y j , 2 , , y j , t ]
In clustering, feature similarity can be measured using the average Euclidean distance. For any two companies ( c 1   and   c 2 ) , the degree of similarity can be expressed as follows:
d c 1 , c 2 = 1 T t = 1 T Y t c 1 Y t c 2 2
where 2 represents the L2 norm, T represents the total number of time points, and Y t c denotes the feature vector of company c at time t. The summation calculates the average Euclidean distance across all time points from 1 to T. Thus, formula (10) can compute the average distance between two companies over time.
As Figure 3e shows, through the K-means approach, the M clusters and the tensor X are transformed into M sets of the cluster matrix:
X = { V 1 , , V m }   ( m [ 1 , M ] )
where V m = { Y c } represents the matrix of cluster m, and Y c denotes company c in this cluster. This clustering approach is robust because its feature similarity measure assesses the holistic alignment of two firms’ STS elements throughout the observed period. It can analyze whether companies operating under similar task pressures—utilizing comparable technology—are governed by analogous structures and whether they are managed by actors with similar attributes over all time points rather than at any single time point. By analyzing patterns over all periods, companies sharing persistent and systemic spatial behavioral patterns are more likely to be grouped together. Finally, the matrices { V 1 , , V m } can be assembled into a single matrix G as the output, which is then used as the input for the next step (shown in Figure 3f).

3.3. Temporal Behavior Pattern Extraction

From an STS perspective, financial statement fraud is a dynamic process marked by the gradual accumulation of risk. To capture this temporal progression, we analyze multi-year corporate panel data using the most popular LSTM networks. LSTM networks are particularly adept at identifying long-range dependencies in sequential data, enabling the detection of subtle, evolving risk patterns that precede fraudulent activities.
In time series data, temporal behavior patterns are commonly observed as recurring shapes and features. For example, financial series can reflect industry traits, corporate strategies, and decision-making lagging effects. Deviations from existing patterns may indicate strategic misalignments or potential data manipulation. As shown in Figure 4, LSTM networks extract these temporal features through their internal cell states and network weights, learning representations of normal versus anomalous behavioral sequences. The practical significance of LSTM networks lies in their ability to detect abnormal changes during the years preceding fraud occurrence based on STS-based characteristics.
In this step, LSTM networks receive the matrix G as the input for extracting specific temporal behavior patterns and generate a low-dimensional representation H = h 1 , t , h 2 , t , , h m , t as the output. Here, h m , t represents the hidden layer state of the mth cluster at the tth time step.

3.4. Feature Weight Modification

In this framework, we integrate the spatial group features from the K-means clustering step with the temporal dynamic features extracted via LSTM networks to carry out comprehensive FSFD. The spatial features represent each company based on its cluster assignment, while the temporal features capture the dynamic evolutions of STS elements over time. The combined approach enables dynamic risk assessments. For improved integration, an attention mechanism is employed to optimally weight spatial and temporal features.
This mechanism takes the low-dimensional representations of temporal–spatial patterns, H = h 1 , t , h 2 , t , , h m , t ,   as inputs. These fused features are dynamically weighted based on their predictive relevance in order to obtain a global representation, which can simplify the model and accelerate calculations. Four steps are included in this process:
Step 1: Compute Query, Key, and Value separately for the low-dimensional representations, h m , t , of the temporal–spatial patterns in each cluster:
Query : Q m = h m , t · W q
Key : K m = h m , t · W k
Value : V m = h m , t · W u
where W q ,   W k , and   W u are the weight matrices.
Step 2: Calculate the attention score:
Dot   product   attention : attention _ scores mj = Q m   ·   K j T
where Q m and K j   are the Query and Key vectors of the temporal–spatial representation in the m-th and j-th clusters, respectively.
Scaling   dot   product   attention :   attention _ scores mk = Q m · K j T n
Softmax   normalization : α mj = soft   max ( attention _ scores mj ) = exp ( attention _ scores mj ) k = 1 t ( attention _ scores mk )
Here, α mj is the attention weight between the m-th and j-th clusters. Softmax ensures that the sum of the attention weights for each row is 1.
Step 3: Use the weighted sum to obtain the context vector:
attention _ output m = α mj · V j
where V j is the value vector in the j-th cluster.
Step 4: Carry out the temporal–spatial characterization of clusters:
S = 1 m i = 1 m attention _ output m
Finally, the attention mechanisms generate a global representation S as the output, which is used as the input for the next step.

3.5. Financial Statement Fraud Detection

In this subsection, we employ a fully connected (dense) layer as the prediction model. It receives global representation S from the attention mechanism as an input and generates the predicted value y ^ :
y ^   =   model.predict ( S )   =   W p S   +   b p
where W p   and b p are the weight matrix and bias vector of the fully connected layer, respectively.
In the training process, the Adam optimizer is used to adjust the learning rate of each parameter. Specifically, it minimizes the loss function L through a backpropagation algorithm, adjusting the model parameters W p and b p such that the predicted value y ^ is as close as possible to the true label y. Considering efficient gradients and faster convergence, we employ binary cross-entropy in the loss function to train our model for FSFD. The loss function is defined as follows:
L BCE = 1 N c = 1 N y c · log y ^ c + 1 y c · log 1 y ^ c  
where N is the number of companies, y ^ c is the predicted value of company c, and y c is the true label of company c.

3.6. Detection Result Explanation

Inspired by the Shapley value from cooperative game theory, this study uses the SHAP method to quantify indicator contributions for FSFD. It can provide a globally interpretable framework that transcends conventional indicator ranking. Quantifying both the direction and magnitude of indicator influences, it supports a logical interpretation of fraud risks through an STS perspective.
Financial statement fraud risks stem from maladaptive interactions among task, technology, structure, and actor elements. Positive SHAP values, visualized as red in dependency plots, indicate an increased likelihood of fraud risk. They may reflect task pressures, actor traits, or structural weaknesses that may result in fraudulent behaviors. Conversely, negative SHAP values, shown in blue, signal protective factors that reduce fraud risk. The SHAP analysis not only enhances model transparency but also decodes the essence of fraud risk from an STS perspective, thereby offering a systemic mechanism for risk generation and mitigation.
The principle of the SHAP method is as follows:
f x = 0 f + i = 1 M i f , x
where 0 f is the expected value of the model’s output, and the sum of i f , x matches the output f x   of the original tree-based model. i f , x is the contribution value of indicator x i , and its formula is as follows:
i f , x = 1 M s N i u S i u S C n 1 s
where u S   is the value of the characteristic subset S as the input. This equation represents the weighted average of the difference in contributions between all indicator subsets, regardless of whether indicator x i is included.
For each observation, we obtain the interpretation set E = i f 1 , x , i f 2 , x , , i f T , x for each indicator x i .

4. Experimental Design

4.1. Sample Data

Over the past 12 years (2010–2022), most financial statement fraud cases occurred in the manufacturing; agriculture; forestry, animal husbandry and fisheries; and information transmission, software, and IT service industries [11]. Concurrently, due to the effects of economic downturns and structural adjustments, China’s capital market has entered a period with a high incidence of financial fraud. Concerning the extreme imbalance between fraud and non-fraud companies across all industries, we confine our experiments to the above high-risk industries to improve sample quality. Through experiments, we aim to validate the effectiveness of the proposed FSFD framework, and we attempt to open the FSFD “black box” to reveal the essence of fraud risk.

4.2. Data Collection

With a comprehensive sampling approach, experimental data were collected from company research series in the CSMAR database (https://data.csmar.com/, accessed on 27 October 2025), covering basic finance, operation, audit, governance, social responsibility, and corporate information. The obtained data directly correspond to indicators across these dimensions on a one-to-one basis. The constructed dataset consists of 24,982 firm-year observations. We also manually extracted the violation information of those observations. Observations associated with reported violations are assigned a positive label (1), resulting in 11,247 positive instances, while the remaining 13,734 compliant instances are labeled as negative (0). This constitutes a 45.1% positive and 54.9% negative class distribution.

4.3. Data Preprocessing

We processed the data according to the following steps to prepare them for model training and evaluation:
(1)
First, we filled missing spaces with the mean values of their corresponding attributes. This method was selected to preserve the sample size and the dataset’s overall distribution.
(2)
Second, as financial statement fraud is detected based on annual reports, we set the year as the minimum time granularity and standardized all temporal data to a uniform yearly level. This step ensures consistency with the annual disclosure cycle.
(3)
We then standardized all indicators to have a mean of 0 and a standard deviation of 1 using Z-score normalization. This process removes the influence of different measurement scales and distribution ranges, which helped our model to learn effectively by transforming all indicators to the same statistical scale while preserving their original distribution shapes and relative relationships.
(4)
Finally, we randomly divided the processed data into training and testing sets at an 8:2 ratio. The training set was used for model training, hyperparameter tuning, and validation, while the testing test set was reserved solely for final evaluation to report the model’s performance.

4.4. Model Training and Hyperparameter Tuning

In our experiment, all algorithms were implemented in pycharm community edition 2023 on an AMD Ryzen 5 3500U with Radeon Vega Mobile Gfx (2.10 GHz). The training process of the proposed framework was iterated over 1000 epochs using the Adam optimizer with a learning rate of 0.001, which ensured stable and efficient convergence. In FSFD, we use corporate data from year t-2 to predict fraud occurrence in year t, which provides a two-year early-warning window. The model’s architecture employs bidirectional LSTM layers with 128 units, a global average pooling 1D attention mechanism, and dense layers comprising 64 and 1 unit with sigmoid activation. Regularization includes a dropout rate of 0.3 and L2 regularization of 0.001. Additional parameters include a batch size of 32; PCA with 95% variance retention; K-means clustering with 2–8 automatic cluster determinations; balanced class weighting; early stopping with a patience of 5; learning rate reduction with a patience of 3; and data preprocessing using fixed 3-step sequences.

4.5. Model Validation and Evaluation

To ensure the generalization ability of the proposed framework and mitigate possible overfitting, we adopted K-fold cross-validation on the training set. The original training data after the initial 8:2 split were partitioned into K = 5 folds with a fixed random seed of 42 to ensure reproducibility. For each unique fold, we carried out the following:
(1)
Treated the selected fold as a validation fold;
(2)
Trained the model on the remaining 4-fold subset;
(3)
Evaluated the model on the selected fold.
The final performance metrics, reported as the mean across all 5 folds, provide a reliable estimate of the model’s efficacy.
As illustrated in Table 3, we use a confusion matrix to evaluate detection performance. The matrix calculates five metrics, namely, accuracy, precision, recall, F1 score, and AUC (area under the ROC curve), as shown in Table 4.
TP and TN represent the number of companies correctly classified as fraudulent and non-fraudulent companies, respectively. FP and FN represent the number of companies falsely classified as fraudulent and non-fraudulent, respectively.

5. Results and Further Research

The analytical techniques adopted in this study are clearly outlined as follows. For performance evaluation, comparative analysis of key metrics served as our primary analytical technique. This involved using a set of widely recognized statistical indicators, namely, accuracy, precision, recall, F1 score, and auc, to assess all models on a test set. The models were then directly compared based on the results of these metrics to objectively determine the optimal performer. Furthermore, to enhance the interpretability of our framework, we incorporated SHAP method as a quantitative analytical technique from the field of explainable AI. This technique provides a rigorous, quantitative measure of feature importance and contribution for each indicator, thereby allowing for analyzing the framework’s decision-making process and exploring intrinsic patterns of fraud risk.
We firstly validate the effectiveness of LSTM networks and the K-means clustering approach in our proposed FSFD-ETSP framework. Then, we verified the performance of FSFD-ETSP with other classical machine learning and deep learning techniques. Finally, we applied the SHAP method to provide a holistic interpretation of the framework and reveal fraud risk patterns consistent with the enterprise STS framework. It is important to note that SHAP provides an interpretable breakdown of model predictions rather than inferring causality on its own.

5.1. Validate the Effectiveness of Temporal Behavior Patterns

Within an enterprise STS, a firm’s sequential data can reflect its operational status, and from the status, we can extract strategic shifts and anomalies (e.g., unusual revenue cycles or irregular expense) as temporal patterns for FSFD. In the proposed framework, we apply LSTM for temporal pattern extraction, and compare it with three other commonly used methods, namely, CNN, RNN, and GRU. The detection results from the test set are shown in Table 5. To ensure robust performance assessments, all reported metrics are computed using the weighted average method. This approach ensures that the performance contribution of each class is proportional to its size, providing a more reliable and aggregate view of the model’s effectiveness.
According to Table 5, LSTM outperforms the other three deep learning models across all metrics. This is primarily because LSTM excels at capturing long-term dependencies in time series, has strong memory capabilities, and allows for more flexible model tuning. In contrast, RNN struggles with these dependencies and often experiences gradient vanishing or explosion. GRU simplifies the LSTM structure, but this weakens its ability to handle long sequences. CNN is good at extracting local features, but it is generally less effective in processing global contextual information compared to LSTM. In summary, the proposed LSTM model performs best in temporal behavior pattern extraction procedures.

5.2. Validate the Effectiveness of Spatial Behavior Patterns

Since fraudulent behaviors can spread in an enterprise STS, companies exhibiting similar behavioral patterns are usually considered as “peer firms”. These structural correlations could comprise key factors affecting FSFD performance, which are usually ignored by current detection models. In this section, we compare detection performance using either temporal behavior patterns (T) alone or temporal–spatial patterns together (T+S) as input features. We conduct the entire experiment using four types of temporal behavior pattern extraction models, namely, LSTM, CNN, GRU, and RNN. The detection performance on the test set can be observed in Table 6. Similarly to Section 5.2, all reported metrics are computed using the weighted average method.
Using temporal–spatial behavior patterns provides an overall improvement with respect to all models, with a maximum increase of 0.1877 in accuracy, 0.2033 in precision, 0.2168 in recall, 0.2045 in F1 score, and 0.153 with respect to AUC. Among these models, LSTM exhibited superior performance, demonstrating the effectiveness of the proposed framework (FSFD-ETSP).
Models relying solely on temporal patterns can detect isolated enterprise violations; however, they ignore collective fraud signals. Spatial behavior patterns can capture structural correlations and offer a crucial complementary view among companies. For example, geographically proximate companies generally face the same regulatory pressures and policy incentives, which can exacerbate coordinated fraud. Industry peers exhibit correlated risk signals when adopting similar aggressive accounting practices. Consequently, these spatial behavior patterns usually expose collective fraud signals or shared external risks. Models incorporating temporal–spatial patterns outperform those that rely solely on temporal patterns.

5.3. Validate the Effectiveness of FSFD-ETSP

To test whether our proposed framework can enhance FSFD, we compared FSFD-ETSP with eight benchmark methods. Table 7 summarizes the weighted average results of all performance metrics on the test set.
Overall, FSFD-ETSP outperformed all benchmark methods in terms of all performance metrics, with an accuracy of 98.04%, precision of 97.69%, recall of 98%, F1 score of 97.83%, and AUC of 99.77%. FSFD-ETSP is a favorable candidate for FSFD, as it is able to extract complex patterns in both temporal and spatial dimensions. In FSFD-ETSP, PCA is first used to reduce the high dimensionality of corporate data while preserving critical information. The clustering step accounts for structural similarities and differences among firms, identifying subgroups with homogeneous risk profiles. A BiLSTM network then models the temporal sequences within each cluster to capture evolving risk patterns. An attention mechanism is also incorporated to improve feature fusion by adaptively setting weights, ensuring the model focuses on the most critical time periods. This PCA + clustering + BiLSTM + attention + SHAP-based architecture is appropriate as it systematically addresses the high dimensionality, heterogeneity, and temporal complexity inherent in financial fraud data. Simpler models often fail to simultaneously capture these intertwined spatial and temporal risk patterns. Furthermore, the inclusion of SHAP provides essential model interpretability, transcending a “black-box” prediction to offer reasonable insights into risk factors.
XGBoost and random forest ensemble weak classifiers achieve good performance, with accuracies of 92.48% and 87.85%, respectively. CNN, RNN, and GRU are deep learning models that can extract nonlinear patterns among data; however, they are inferior to ensemble methods. LR, SVM, and DT are traditional machine learning methods that have limitations when applied in complicated conditions, exhibiting the poorest detection performance.
As for CNN, RNN, and GRU, their performance varies, as can be observed in the results shown in Table 5 and Table 7. In Table 5, they are substitutes for LSTM networks and embedded to extract temporal patterns. As part of FSFD-ETSP, their excellent performance demonstrates the effectiveness of the framework. In contrast, the other deep learning models are applied independently as end-to-end classifiers to extract nonlinear patterns and carry out FSFD. They were trained and evaluated on the complete and raw input dataset without the specialized feature engineering and fusion modules present in FSFD-ETSP. As shown in Table 7, their performance is inferior to ensemble methods. This also identifies the importance of spatial and temporal patterns in FSFD.

5.4. Interpret the Detection Process in FSFD-ETSP

In this study, for better comparability across features, the SHAP values were first standardized (z-score normalization). We then averaged the absolute values of the final Shapley values to generate a global interpretation of FSFD-ETSP. Summing SHAP values across all samples provides a global indicator importance ranking, showing which indicator exerts the most influence on the overall model. These values can be positive or negative, indicating whether an indicator increases or decreases the likelihood of fraud. Red represents a positive contribution, increasing fraud risk; blue represents a negative contribution, reducing the risk. In Figure 5, the x-axis denotes the standardized SHAP value, indicating the feature’s marginal contribution to the predicted fraud risk, and the y-axis represents the feature value.
From an STS perspective, financial statement fraud arises from the complex interplay of four core elements, as shown in Figure 5. The analysis of the top 10 most important indicators validates this system, as they can be distinctly categorized into these four elements.
Task-related pressures—which comprise operational dependencies, ambitious earning targets, and intense performance reporting demands—often encourage fraudulent activities. This is evidenced by indicators such as customer concentration, which can motivate fraudulent activities. Actors play a central role within this system, with their decisions and incentives driving spatial behavior patterns. This is strongly supported by the dominance of geographical indicators, such as office latitude and the longitude of the registered address. Structural elements determine opportunities for fraud by defining governance structure. Indicators such as the chairman’s shareholding ratio, the top 10 shareholders, the total number of shareholders, and the number of senior executives reveal that governance mechanisms and organizational scale can create critical vulnerabilities in oversight. Technology serves as a catalyst for manipulation. Accounting indicators, such as capital reserve and undistributed profits, represent the technical tools and flexibilities in financial reporting that can be adjusted in response to performance pressures.
In conclusion, these indicators confirm that financial statement fraud is a systemic risk. The interconnectedness of tasks, actors, structures, and technology means that a vulnerability in one element can amplify weaknesses and propagate risks throughout the entire sociotechnical system.

5.5. Analyze the Intrinsic Associations of Fraud Risks Through Indicator Impact Patterns

We further carry out a deeper analysis of the specific relationships between the top 10 most important indicators and financial statement fraud using SHAP dependence plots. For better comparability across features, the SHAP values were standardized (z-score normalization) before generating the dependence plots and analysis. In Figure 6, Figure 7, Figure 8, Figure 9, Figure 10, Figure 11, Figure 12, Figure 13 and Figure 14, the x-axis represents the feature value, and the y-axis denotes the standardized SHAP value, indicating the feature’s marginal contribution to the predicted fraud risk.
(1)
Analysis of actor-related indicators
According to Figure 6 and Figure 7, actors’ behavior is significantly influenced by their geographic environment, which shapes fraud incentives via two primary aspects.
Figure 6 reveals a sharp spike in fraud risk within the 31.2° N–31.6° N latitude band—a major industrial zone and the core of the Yangtze River Delta. Here, SHAP values consistently exceed 0.05, peaking at 0.058 near 31.6° N, which indicates a pronounced “peer effect” in this high-risk region. When local firms engage in aggressive accounting and are rewarded by the market, they create powerful imitation incentives for their geographic peers.
Figure 7 shows persistently high risk within the coastal longitude range of 121.2° E–121.8° E, where SHAP values consistently exceed 0.045 and peak at 0.052 near 121.5°E. Here, actors face conflicting international and domestic institutional pressures. This environment fosters a specialized skillset in regulatory arbitrage, where firms learn to exploit loopholes across different regimes. These practices diffuse through local networks, forming a regional “accounting sub-culture”.
In essence, geography is not merely a location; it is a social and institutional environment that directly shapes how actors perceive opportunities and rationalize the manipulation of technology.
(2)
Analysis of technology-related indicators
Based on Figure 8 and Figure 9, technology-related indicators demonstrate that internal and external pressures can shape financial reporting behavior.
Figure 8 identifies an inverted U-shaped effect driven by internal profit allocation. When undistributed profits reach 15–25% of net assets with SHAP values increasing from 0.02 to 0.042, managers face conflicting objectives between profit retention and reinvestment, increasing incentives for accounting manipulation. Once this ratio exceeds 40%, evidenced by a drop in SHAP value to approximately 0.018, ample profit reserves and heightened external scrutiny jointly curb such behavior.
Figure 9 reveals that unusual capital accumulation often reflects technical adjustments aimed at meeting capital market expectations. Under weak internal controls, such adjustments are more likely to evolve into fraudulent reporting.
(3)
Analysis of task-related indicators
According to Figure 10, the accelerating growth pattern in SHAP values demonstrates that supply chain dependency drives technical manipulation. A high reliance on key customers creates significant operational pressure. This pressure compels management to employ technical methods to enhance reported performance.
(4)
Analysis of structure-related indicators
According to Figure 11, Figure 12, Figure 13 and Figure 14, the structural element determines opportunities for fraud. An analysis of the key structural features reveals a consistent pattern, where governance is most effective when power is balanced, and it is weakest when it is overly dispersed or concentrated.
Power distribution, measured using equity concentration, is a primary example. Below 35%, the SHAP value jumps from 0.038 to 0.048, and dispersed ownership results in ineffective monitoring. Above 65%, marked by a rebound in SHAP value from 0.021 to 0.034, dominant shareholders can overpower internal controls and manipulate technology for personal gain. Effective checks and balances are observed primarily within the moderate shareholding range of 35–65%, where the associated fraud risk remains low and stable, as evidenced by consistently low SHAP values around 0.02.
Decision-making dynamics, reflected in the number of senior executives, exhibit a similar nonlinear effect. Having fewer than five executives results in excessive power concentration, as indicated by an increase in SHAP value of approximately 0.019. On the other hand, teams with more than twelve executives result in high coordination costs and diffused responsibility, corresponding to a SHAP increase of about 0.021. A team of five to eight executives best balance efficiency with diverse oversight, as evidenced by consistently low SHAP values around 0.014.
Collective supervision, indicated by the total number of shareholders, also affects governance. Too few shareholders weaken oversight, while organizations with too many shareholders encourage free-rider behavior. Both scenarios reduce the incentive for individual shareholders to monitor management effectively.
In summary, the structure of an organization sets the foundation of its integrity. Imbalances in ownership or decision-making power undermine supervision. This creates critical vulnerabilities in the system, allowing actors to exploit technology under task-related pressures.

6. Conclusions

In this study, we constructed an STS-based multi-level indicator system for FSFD, providing the theoretical foundation and analytical framework for subsequent work. The effectiveness of current approaches can be attributed to pattern learning. With the influx of more data and the adoption of advanced technologies, these approaches are becoming increasingly capable of learning complex patterns. However, the employed data are normally fragmented and cannot provide a complete description of financial fraud, resulting in limited detection performance. Beyond traditional research, the proposed indicator system posits that fraud risk arises from long-term interactions among four core STS elements: task, technology, actor, and structure. Unlike data-driven approaches, the indicator system can capture fraudulent behaviors by accounting for all possible risk factors within an enterprise STS.
Furthermore, we developed a novel fraud detection framework, FSFD-ETSP, and validated its effectiveness with various baseline models. Concerning the dynamic changing and spreading nature of fraud risks in temporal and spatial dimensions, we integrated LSTM networks and the K-means approach to capture their temporal and spatial behavior patterns. Experiments on data from A-share listed companies demonstrate that firms exhibit explicit strategic shifts in the temporal dimension and peer contagion across the spatial dimension. Notably, the proposed FSFD-ETSP was found to clearly outperform other machine learning and deep learning techniques. The superior performance of FSFD-ETSP verifies the importance of spatial–temporal behavior patterns. Therefore, modeling these patterns is key to achieving early warning and supporting organizational resilience.
To gain deeper insights into the framework’s detection logic and validate theoretical hypotheses, we utilized the explainable SHAP technique to dissect key factors for FSFD and empirically validated the STS-based understanding of fraud risks. We opened the “black box” of the framework by analyzing the marginal contributions of the top 10 most important indicators with explainable AI techniques, specifically SHAP analysis. Moreover, we empirically validated the interaction patterns of four STS elements, associating the suggested potential patterns of fraud risks. These SHAP results not only reveal which indicators are most critical in FSFD but also align perfectly with our initial STS-based theoretical proposition on fraud risks.
In addition, the SHAP analysis carried out on each element in the enterprise STS provides some interesting insights regarding fraud risk. Inspired by our observations, we provide the following recommendations: (1) Dynamic resource allocation should be implemented. Fraud risks exhibit significant geographic concentration; for example, firms located at 31.2° N to 31.6° N and 121.2° E–121.8° E are more likely to exhibit fraudulent behaviors. Regulatory resources should be strategically concentrated in high-risk regions, rather than being evenly distributed. (2) A balanced governance structure should be built. Excessive concentration/decentralization of power distributions in a corporate setting could create opportunities for unethical manipulation; for example, if the top 10 shareholder indicator exceeds 65%, a great number of shareholders dominate the company and may manipulate internal control mechanisms for personal gain. In contrast, if the indicator falls below 35%, dispersive power usually results in ineffective monitoring. It is thus necessary to maintain reasonable ranges with respect to ownership concentration and size. Taken together, these observations suggest that indicators related to actor and structure elements are particularly crucial in prediction, which is in accordance with our STS-based understanding of fraud risk.
We acknowledge several limitations of this study. First, our data were sourced exclusively from high-risk industries of A-share listed companies in China, which may limit the generalizability of our findings to other economic contexts. We plan to incorporate data from other industries and countries to test the robustness of the proposed framework. Second, while the FSFD-ETSP framework demonstrates high performance, we did not conduct systematic ablation studies to explore the contribution of each component (e.g., LSTM and K-means). Future research could explore simpler or alternative architectures to confirm the necessity of our design.
Moreover, as our method integrates LSTM networks and the clustering approach, the proposed FSFD-ETSP framework requires higher computation power and storage capacity, reducing efficiencies in real-time applications. With respect to this limitation, we plan to utilize incremental data rather than complete data as inputs for FSFD, and we will attempt to integrate other techniques, such as incremental learning and model distillation, into the framework such that the FSFD solution can be used in resource-constrained environments. Finally, as noted in our conclusion, this study offers valuable insights with respect to indicator importance in FSFD; however, identifying the complex, nonlinear causal relationships among spatial and temporal features remains a challenge. In the future, we will employ various causal discovery algorithms—such as Peter–Clark-based algorithms—to identify causal relationships among multiple factors.

Author Contributions

Conceptualization, H.X.; methodology, J.J.; software, J.J.; validation, J.J.; formal analysis, H.X.; investigation, Q.W.; resources, Q.W.; data curation, J.J.; writing—original draft preparation, J.J., H.X. and Q.W.; writing—review and editing, J.J., H.X. and Q.W.; visualization, J.J. and Q.W.; supervision, H.X. and Q.W.; project administration, H.X.; funding acquisition, H.X. All authors have read and agreed to the published version of the manuscript.

Funding

This research was funded by National Social Science Fund of China, grant number 23CGL074.

Data Availability Statement

The datasets, code, and experimental records generated and analyzed during this study to develop and validate the interpretable financial statement fraud detection framework are publicly available in a github repository. The complete resources can be accessed via: https://github.com/1Freya1/An-Interpretable-Financial-Statement-Fraud-Detection-Framework-Enhanced-by-Temporal-Spatial-Patterns (accessed on 27 October 2025).

Acknowledgments

The authors gratefully acknowledge the funding support from the National Social Science Fund youth project (23CGL074).

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Trist, E.L.; Bamforth, K.W. Some social and psychological consequences of the longwall method of coal-getting: An examination of the psychological situation and defences of a work group in relation to the social structure and technological content of the work system. Hum. Relat. 1951, 4, 3–38. [Google Scholar] [CrossRef] [Scilit]
  2. Leavitt, H.J. Applied Organizational Change in Industry: Structural, Technological and Humanistic Approaches. In Handbook of Organization; Rand McNally & Company: Chicago, IL, USA, 1965; pp. 1144–1170. [Google Scholar]
  3. Abbas, R.; Michael, K. Socio-Technical Theory: A Review. In TheoryHub Book; TheoryHub: Newcastle upon Tyne, UK, 2025. [Google Scholar]
  4. Sun, Y.; Sun, X.; Wu, W. Who detects corporate fraud under the thriving of the new media? Evidence from Chinese-listed firms. ACC Financ. 2021, 61, 1313–1343. [Google Scholar] [CrossRef] [Scilit]
  5. Younas, M.W.; Li, S.; Maqsood, U.S.; Zahid, R.A. Leveraging artificial intelligence to detect and prevent corporate fraud in Chinese enterprises. Innov. Eur. J. Soc. Sci. Res. 2025, 1–29. [Google Scholar] [CrossRef] [Scilit]
  6. Kon, Z.S.; Lim, Y.H.; Choong, Y.O.; Paloosamy, J.R.; Low, B.T. The influence of pressure on intention to commit fraud: The mediating role of rationalization and opportunities. Asian J. Bus. Ethics 2024, 13, 175–195. [Google Scholar] [CrossRef] [Scilit]
  7. Huang, S.; Ye, Q.; Xu, S.; Ye, F. Analysis of Financial Fraud in Chinese Listed Companies from 2010 to 2019. Account. Mon. 2020, 14, 153–160. [Google Scholar]
  8. Zhu, X.; Ao, X.; Qin, Z.; Chang, Y.; Liu, Y.; He, Q.; Li, J. Intelligent financial fraud detection practices in post-pandemic era. Innovation 2021, 2, 100176. [Google Scholar] [CrossRef] [Scilit]
  9. Wang, Y.; Yu, M.; Gao, S. Gender diversity and financial statement fraud. J. Account. Public Policy 2022, 41, 106903. [Google Scholar] [CrossRef] [Scilit]
  10. Sun, N.; Salama, A.; Hussainey, K.; Habbash, M. Corporate environmental disclosure, corporate governance and earnings management. Manag. Audit. J. 2010, 25, 679–700. [Google Scholar] [CrossRef] [Scilit]
  11. Wang, B.; Bian, J.; Xu, Y.; Liu, D. The Role of Internal Control on Financial Fraud: Evidence of Pre-Fraud and Post-Punishment from Chinese Firms. Emerg. Mark. Financ. Trade 2025, 61, 1446–1459. [Google Scholar] [CrossRef] [Scilit]
  12. AICPA. AU-C Section 240: Consideration of Fraud in a Financial Statement Audit; American Institute of Certified Public Accountants: Durham, NC, USA, 2019. [Google Scholar]
  13. IAASB. International Standard on Auditing (ISA) 240: The Auditor’s Responsibilities Relating to Fraud in an Audit of Financial Statements; International Auditing and Assurance Standards Board: New York, NY, USA, 2020. [Google Scholar]
  14. ACFE. Occupational Fraud 2024: A Report to the Nations; Association of Certified Fraud Examiners: Austin, TX, USA, 2024; p. 5. [Google Scholar]
  15. Finance, S. The 90 Billion Yuan “Private Network Communication” Fraud Case Has Been Subject to Criminal Accountability for the First Time. 2025. Available online: https://finance.sina.cn/2025-07-14/detail-inffmtum6474546.d.html (accessed on 27 October 2025).
  16. Ye, Q.; Huang, S. New Characteristics and New Responses to Financial Fraud—Based on the Sample of Private Network Communication Fraud. Finance Account. Mon. 2024, 45, 3–8. [Google Scholar]
  17. Cecchini, M.; Aytug, H.; Koehler, G.J.; Pathak, P. Detecting management fraud in public companies. Manag. Sci. 2010, 56, 1146–1160. [Google Scholar] [CrossRef] [Scilit]
  18. Bao, Y.; Ke, B.; Li, B.; Yu, Y.J.; Zhang, J. Detecting accounting fraud in publicly traded US firms using a machine learning approach. J. Account. Res. 2020, 58, 199–235. [Google Scholar] [CrossRef] [Scilit]
  19. Aftabi, S.Z.; Ahmadi, A.; Farzi, S. Fraud detection in financial statements using data mining and GAN models. Expert. Syst. Appl. 2023, 227, 120144. [Google Scholar] [CrossRef] [Scilit]
  20. Sanusi, A.; Sanusi, I.; Yinusa, S.; Abugh, I. Fraud risk management and organizational resilience: An empirical analysis controlling for organizational culture. World J. Adv. Res. Rev. 2025, 26, 2446–2458. [Google Scholar] [CrossRef] [Scilit]
  21. Richtnér, A.; Löfsten, H. Managing in turbulence: How the capacity for resilience influences creativity. RD Manag. 2014, 44, 137–151. [Google Scholar] [CrossRef] [Scilit]
  22. Penman, S.H. Financial reporting quality: Is fair value a plus or a minus? Account. Bus. Res. 2007, 37, 33–44. [Google Scholar] [CrossRef] [Scilit]
  23. Kim, Y.J.; Baik, B.; Cho, S. Detecting financial misstatements with fraud intention using multi-class cost-sensitive learning. Expert. Syst. Appl. 2016, 62, 32–43. [Google Scholar] [CrossRef] [Scilit]
  24. Brown, N.C.; Crowley, R.M.; Elliott, W.B. What are you saying? Using topic to detect financial misreporting. J. Account. Res. 2020, 58, 237–291. [Google Scholar] [CrossRef] [Scilit]
  25. Li, J.; Li, N.; Xia, T.; Guo, J. Textual analysis and detection of financial fraud: Evidence from Chinese manufacturing firms. Econ. Model. 2023, 126, 106428. [Google Scholar] [CrossRef] [Scilit]
  26. Huang, Y.; Wang, Z.; Jiang, C. Diagnosis with incomplete multi-view data: A variational deep financial distress prediction method. Technol. Forecast. Soc. Change 2024, 201, 123269. [Google Scholar] [CrossRef] [Scilit]
  27. Wu, C.; Jiang, C.; Wang, Z.; Ding, Y. Predicting financial distress using current reports: A novel deep learning method based on user-response-guided attention. Decis. Support Syst. 2024, 179, 114176. [Google Scholar] [CrossRef] [Scilit]
  28. Stilinki, D.; Potter, K. Enhancing Fraud Detection Accuracy and Adaptability Through Dynamic Feature Engineering in NoSQL Databases. Available online: https://easychair.org/publications/preprint/F7X18 (accessed on 27 October 2025).
  29. Hilal, W.; Gadsden, S.A.; Yawney, J. Financial fraud: A review of anomaly detection techniques and recent advances. Expert. Syst. Appl. 2022, 193, 116429. [Google Scholar] [CrossRef] [Scilit]
  30. Motie, S.; Raahemi, B. Financial fraud detection using graph neural networks: A systematic review. Expert. Syst. Appl. 2024, 240, 122156. [Google Scholar] [CrossRef] [Scilit]
  31. Lin, K.; Gao, Y. Model interpretability of financial fraud detection by group SHAP. Expert. Syst. Appl. 2022, 210, 118354. [Google Scholar] [CrossRef] [Scilit]
  32. Zhou, Y.; Li, H.; Xiao, Z.; Qiu, J. A user-centered explainable artificial intelligence approach for financial fraud detection. Financ. Res. Lett. 2023, 58, 104309. [Google Scholar] [CrossRef] [Scilit]
  33. Madni, A.M.; Jackson, S. Towards a Conceptual Framework for Resilience Engineering. IEEE Syst. J. 2009, 3, 181–191. [Google Scholar] [CrossRef] [Scilit]
  34. Malatji, M.; Von Solms, S.; Marnewick, A. Socio-technical systems cybersecurity framework. Inf. Comput. Secur. 2019, 27, 233–272. [Google Scholar] [CrossRef] [Scilit]
  35. van Bruxvoort, X.; van Keulen, M. Framework for assessing ethical aspects of algorithms and their encompassing socio-technical system. Appl. Sci. 2021, 11, 11187. [Google Scholar] [CrossRef] [Scilit]
  36. Geels, F.W. Technological transitions as evolutionary reconfiguration processes: A multi-level perspective and a case-study. Res. Policy 2002, 31, 1257–1274. [Google Scholar] [CrossRef] [Scilit]
  37. Rahman, M.J.; Jie, X. Fraud detection using fraud triangle theory: Evidence from China. J. Financ. Crime 2024, 31, 101–118. [Google Scholar] [CrossRef] [Scilit]
  38. Xia, H.; Ma, H.; Cheng, P. PE-EDD: An efficient peer-effect-based financial fraud detection approach in publicly traded China firms. CAAI Trans. Intell. Technol. 2022, 7, 469–480. [Google Scholar] [CrossRef] [Scilit]
  39. O’Neill, S.; Flanagan, J.; Clarke, K. Safewash! Risk attenuation and the (Mis)reporting of corporate safety performance to investors. Saf. Sci. 2016, 83, 114–130. [Google Scholar] [CrossRef] [Scilit]
  40. Kalina, I.; Khurdei, V.; Shevchuk, V.; Vlasiuk, T.; Leonidov, I. Introduction of a corporate security risk management system: The experience of Poland. J. Risk Financ. Manag. 2022, 15, 335. [Google Scholar] [CrossRef] [Scilit]
  41. Wei, R.; Wong, E.Y.C.; Sun, M.; Wang, Z. Multidimensional financial metrics for corporate financial risk assessment and early warning mechanisms. J. Organ. End User Comput. 2024, 36, 1–23. [Google Scholar] [CrossRef] [Scilit]
  42. Koyuncugil, A.S.; Ozgulbas, N. Financial early warning system model and data mining application for risk detection. Expert. Syst. Appl. 2012, 39, 6238–6253. [Google Scholar] [CrossRef] [Scilit]
  43. Machado, M.R.; Chen, D.T.; Osterrieder, J.R. An analytical approach to credit risk assessment using machine learning models. Decis. Anal. J. 2025, 16, 100605. [Google Scholar] [CrossRef] [Scilit]
  44. Mostofi, F.; Toğan, V.; Ayözen, Y.E.; Tokdemir, O.B. Construction Safety Risk Model with Construction Accident Network: A Graph Convolutional Network Approach. Sustainability 2022, 14, 15906. [Google Scholar] [CrossRef] [Scilit]
  45. Mohammadfam, I.; Zarei, E. Safety risk modeling and major accidents analysis of hydrogen and natural gas releases: A comprehensive risk analysis framework. Int. J. Hydrog. Energy 2015, 40, 13653–13663. [Google Scholar] [CrossRef] [Scilit]
  46. Choe, S.; Leite, F. Transforming inherent safety risk in the construction Industry: A safety risk generation and control model. Saf. Sci. 2020, 124, 104594. [Google Scholar] [CrossRef] [Scilit]
  47. Su, W.; Junge, S. Unlocking the recipe for organizational resilience: A review and future research directions. Eur. Manag. J. 2023, 41, 1086–1105. [Google Scholar] [CrossRef] [Scilit]
  48. Kantur, D.; İşeri-Say, A. Organizational resilience: A conceptual integrative framework. J. Manag. Organ. 2012, 18, 762–773. [Google Scholar] [CrossRef] [Scilit]
  49. Loebbecke, J.K.; Eining, M.M.; Willingham, J.J. Auditors’ Experience with Material Irregularities: Frequency, Nature, and Detectability. Audit. J. Pract. Theory 1989, 9, 1. [Google Scholar]
  50. Abbasi, A.; Albrecht, C.; Vance, A.; Hansen, J. Metafraud: A meta-learning framework for detecting financial fraud. MIS Q. 2012, 36, 1293–1327. [Google Scholar] [CrossRef] [Scilit]
  51. An, B.; Suh, Y. Identifying financial statement fraud with decision rules obtained from Modified Random Forest. Data Technol. Appl. 2020, 54, 235–255. [Google Scholar] [CrossRef] [Scilit]
  52. Bertomeu, J.; Cheynel, E.; Floyd, E.; Pan, W. Using machine learning to detect misstatements. Rev. Account. Stud. 2021, 26, 468–519. [Google Scholar] [CrossRef] [Scilit]
  53. Achakzai, M.A.K.; Juan, P. Using machine learning Meta-Classifiers to detect financial frauds. Financ. Res. Lett. 2022, 48, 102915. [Google Scholar] [CrossRef] [Scilit]
  54. Ravisankar, P.; Ravi, V.; Rao, G.R.; Bose, I. Detection of financial statement fraud and feature selection using data mining techniques. Decis. Support Syst. 2011, 50, 491–500. [Google Scholar] [CrossRef] [Scilit]
  55. Kirkos, E.; Spathis, C.; Manolopoulos, Y. Data Mining techniques for the detection of fraudulent financial statements. Expert. Syst. Appl. 2007, 32, 995–1003. [Google Scholar] [CrossRef] [Scilit]
  56. Yao, J.; Zhang, J.; Wang, L. A financial statement fraud detection model based on hybrid data mining methods. In Proceedings of the 2018 International Conference on Artificial Intelligence and Big Data (ICAIBD), Chengdu, China, 26–28 May 2018; pp. 57–61. [Google Scholar]
  57. Kotsiantis, S.; Koumanakos, E.; Tzelepis, D.; Tampakas, V. Forecasting fraudulent financial statements using data mining. Int. J. Comput. Intell. 2006, 3, 104–110. [Google Scholar]
  58. Jan, C.-L. Detection of financial statement fraud using deep learning for sustainable development of capital markets under information asymmetry. Sustainability 2021, 13, 9879. [Google Scholar] [CrossRef] [Scilit]
  59. Wang, Y. Research on the model of preventing corporate financial fraud under the combination of deep learning and SHAP. Int. J. Adv. Comput. Sci. Appl. 2023, 14, 784–792. [Google Scholar] [CrossRef] [Scilit]
  60. Persons, O.S. Using financial statement data to identify factors associated with fraudulent financial reporting. J. Appl. Bus. Res. 1995, 11, 38. [Google Scholar] [CrossRef] [Scilit]
  61. Beneish, M.D. The detection of earnings manipulation. Financ. Anal. J. 1999, 55, 24–36. [Google Scholar] [CrossRef] [Scilit]
  62. Dechow, P.M.; Ge, W.; Larson, C.R.; Sloan, R.G. Predicting material accounting misstatements. Contemp. Account. Res. 2011, 28, 17–82. [Google Scholar] [CrossRef] [Scilit]
  63. Li, J.; Guo, C.; Lv, S.; Xie, Q.; Zheng, X. Financial fraud detection for Chinese listed firms: Does managers’ abnormal tone matter? Emerg. Mark. Rev. 2024, 62, 101170. [Google Scholar] [CrossRef] [Scilit]
  64. West, J.; Bhattacharya, M. Intelligent financial fraud detection: A comprehensive review. Comput. Secur. 2016, 57, 47–66. [Google Scholar] [CrossRef] [Scilit]
  65. Lee, C.-W.; Fu, M.-W.; Wang, C.-C.; Azis, M.I. Evaluating machine learning algorithms for financial fraud detection: Insights from Indonesia. Mathematics 2025, 13, 600. [Google Scholar] [CrossRef] [Scilit]
  66. Hajek, P.; Henriques, R. Mining corporate annual reports for intelligent detection of financial statement fraud–A comparative study of machine learning methods. Knowl. Based Syst. 2017, 128, 139–152. [Google Scholar] [CrossRef] [Scilit]
  67. West, J.; Bhattacharya, M. Mining financial statement fraud: An analysis of some experimental issues. In Proceedings of the 2015 IEEE 10th Conference on Industrial Electronics and Applications (ICIEA), Auckland, New Zealand, 15–17 June 2015; pp. 461–466. [Google Scholar]
  68. Zhang, X. Financial Data Anomaly Recognition Model Based on Improved Support Vector Machine. In Intelligent Computing Technology and Automation; IOS Press: Amsterdam, The Netherlands, 2024; pp. 924–933. [Google Scholar]
  69. Fu, B.; Tong, Y.; Li, Y.; Tang, Z.; Shang, Z.; Li, A. Financial Fraud Anomaly Detection of Listed Companies Based on Probabilistic Perspective Machine Learning Models. Procedia Comput. Sci. 2025, 266, 979–986. [Google Scholar] [CrossRef] [Scilit]
  70. Bengio, Y.; Courville, A.; Vincent, P. Representation learning: A review and new perspectives. IEEE Trans. Pattern Anal. Mach. Intell. 2013, 35, 1798–1828. [Google Scholar] [CrossRef] [Scilit]
  71. Lim, B.; Zohren, S. Time-series forecasting with deep learning: A survey. Philos. Trans. R. Soc. A 2021, 379, 20200209. [Google Scholar] [CrossRef] [Scilit]
  72. Guo, H.; Zhang, D.; Liu, S.; Wang, L.; Ding, Y. Bitcoin price forecasting: A perspective of underlying blockchain transactions. Decis. Support Syst. 2021, 151, 113650. [Google Scholar] [CrossRef] [Scilit]
  73. Zhong, C.; Du, W.; Xu, W.; Huang, Q.; Zhao, Y.; Wang, M. LSTM-ReGAT: A network-centric approach for cryptocurrency price trend prediction. Decis. Support Syst. 2023, 169, 113955. [Google Scholar] [CrossRef] [Scilit]
  74. Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, Ł.; Polosukhin, I. Attention is all you need. In Proceedings of the 31st International Conference on Neural Information Processing Systems, Long Beach, CA, USA, 4–9 December 2017. [Google Scholar]
Figure 1. Conceptual framework.
Figure 1. Conceptual framework.
Mca 30 00138 g001
Figure 2. Model framework. (a) Structure of LSTM. (b) Attention layer. (c) Output layer.
Figure 2. Model framework. (a) Structure of LSTM. (b) Attention layer. (c) Output layer.
Mca 30 00138 g002
Figure 3. Spatial behavioral pattern extraction procedure. (a) The primary input X. (b) A time–feature matrix X i , t c of a single corporate sample slice c. (c) A lower-dimensional time–feature matrix Y i , t c of company c through the PCA procedure. (d) The unified one-dimension vector Y c obtained through flattening. (e) M sets of the cluster matrix { V 1 , , V m }   after the K-means approach. (f) The assembled matrix G with { V 1 , , V m } .
Figure 3. Spatial behavioral pattern extraction procedure. (a) The primary input X. (b) A time–feature matrix X i , t c of a single corporate sample slice c. (c) A lower-dimensional time–feature matrix Y i , t c of company c through the PCA procedure. (d) The unified one-dimension vector Y c obtained through flattening. (e) M sets of the cluster matrix { V 1 , , V m }   after the K-means approach. (f) The assembled matrix G with { V 1 , , V m } .
Mca 30 00138 g003
Figure 4. Structure of LSTM. The previous hidden state h t 1   and current input X t are concatenated. This combined vector is fed to the three gates and the tanh layer. The forget gate σ decides what to remove from the old cell state c t 1 . The input gate σ and the tanh layer work together to create and select new candidate information. These updates are combined to form the new cell state c t . Finally, the output gate σ filters c t (via tanh) to produce the new hidden state h t .
Figure 4. Structure of LSTM. The previous hidden state h t 1   and current input X t are concatenated. This combined vector is fed to the three gates and the tanh layer. The forget gate σ decides what to remove from the old cell state c t 1 . The input gate σ and the tanh layer work together to create and select new candidate information. These updates are combined to form the new cell state c t . Finally, the output gate σ filters c t (via tanh) to produce the new hidden state h t .
Mca 30 00138 g004
Figure 5. Top 10 most important indicators.
Figure 5. Top 10 most important indicators.
Mca 30 00138 g005
Figure 6. (a) SHAP dependence plot of the latitude of the office without labels; (b) SHAP dependence plot of the latitude of the office with labels.
Figure 6. (a) SHAP dependence plot of the latitude of the office without labels; (b) SHAP dependence plot of the latitude of the office with labels.
Mca 30 00138 g006
Figure 7. (a) SHAP dependence plot of the longitude of the registered address without labels; (b) SHAP dependence plot of the longitude of the registered address with labels.
Figure 7. (a) SHAP dependence plot of the longitude of the registered address without labels; (b) SHAP dependence plot of the longitude of the registered address with labels.
Mca 30 00138 g007
Figure 8. (a) SHAP dependence plot of undistributed profit without labels; (b) SHAP dependence plot of undistributed profit with labels.
Figure 8. (a) SHAP dependence plot of undistributed profit without labels; (b) SHAP dependence plot of undistributed profit with labels.
Mca 30 00138 g008
Figure 9. (a) SHAP dependence plot of capital reserve without labels; (b) SHAP dependence plot of capital reserve with labels.
Figure 9. (a) SHAP dependence plot of capital reserve without labels; (b) SHAP dependence plot of capital reserve with labels.
Mca 30 00138 g009
Figure 10. (a) SHAP dependence plot of customer concentration without labels; (b) SHAP dependence plot of customer concentration with labels.
Figure 10. (a) SHAP dependence plot of customer concentration without labels; (b) SHAP dependence plot of customer concentration with labels.
Mca 30 00138 g010
Figure 11. (a) SHAP dependence plot of the top 10 shareholders without labels; (b) SHAP dependence plot of the top 10 shareholders with labels.
Figure 11. (a) SHAP dependence plot of the top 10 shareholders without labels; (b) SHAP dependence plot of the top 10 shareholders with labels.
Mca 30 00138 g011
Figure 12. (a) SHAP dependence plot of the number of senior executives without labels; (b) SHAP dependence plot of the number of senior executives with labels.
Figure 12. (a) SHAP dependence plot of the number of senior executives without labels; (b) SHAP dependence plot of the number of senior executives with labels.
Mca 30 00138 g012
Figure 13. (a) SHAP dependence plot of the total number of shareholders without labels; (b) SHAP dependence plot of the total number of shareholders with labels.
Figure 13. (a) SHAP dependence plot of the total number of shareholders without labels; (b) SHAP dependence plot of the total number of shareholders with labels.
Mca 30 00138 g013
Figure 14. (a) SHAP dependence plot of the chairman’s shareholding ratio without labels; (b) SHAP dependence plot of the chairman’s shareholding ratio with labels.
Figure 14. (a) SHAP dependence plot of the chairman’s shareholding ratio without labels; (b) SHAP dependence plot of the chairman’s shareholding ratio with labels.
Mca 30 00138 g014
Table 1. Summarized classification for CFF detection methods.
Table 1. Summarized classification for CFF detection methods.
Authors (Year)Detection ModelData SourceData InformationModel InputResults
Kotsiantis et al. [57]Hybrid decision support system using a stacking-based ensemble classifier164 Greek manufacturing firms (2001–2002)41 fraud and 123 non-fraud casesAudited financial statementsAccuracy: 95.1%
Kirkos et al. [55]DT, NN, BBN76 Greek manufacturing firms38 fraud and 38 non-fraud casesFinancial ratios from balance sheet and income statementAccuracy: 90.3% (BBN), 80% (NN), 73.6% (DT)
Cecchini et al. [17]SVM-FK205 fraudulent and 6427 non-fraudulent company-years from SEC AAERs and Compustat.205 fraud and 6427 non-fraud cases23 financial indicatorsAUC: 87.8% (SVM-FK)
Recall: 80% (SVM-FK)
Ravisankar et al. [54]MLFF, SVM, GP, LR, PNN202 Chinese companies101 fraud and 101 non-fraud casesBalance sheet, income statement, cash flow statement, statement of retained earningsAccuracy: 95.64% (PNN), 92.68% (GP)
Hajek and Henriques [66]BBN, DTNB, SVM, NN, DT, LR622 U.S. SEC AAERs (2005–2015)311 fraud and 311 non-fraud casesFinancial statements, analysts’ forecasts, and managerial commentsAccuracy: 90.32 ± 1.90 (BBN), 89.50 ± 1.91 (DTNB)
TP rate: 85.19 ± 4.26 (BBN), 87.21 ± 3.24 (DTNB)
Yao et al. [56]RF, SVM, DT, ANN, LR240 Chinese companies (2007–2016)120 fraud and 120 non-fraud casesFinancial and non-financial
information from annual report
Accuracy: 71.67% (SVM), 70.83% (ANN)
Bao et al. [18]RUSBoost, LR, SVMAll publicly listed U.S. firms (1991–2008)1171 fraud and 206,026 non-fraud casesFinancial statementsAUC: 72.5% (RUSBoost), 69% (Logit), 62.6% (SVM)
An and Suh [51]MRF, ANN, SVM, LR, DTCombined dataset from KIND system (1996–2018) 1591 fraud and 31,628 non-fraud casesKIND system and KIS
value database
Accuracy: 78.06% (MRF), 78.51% (ANN)
Bertomeu et al. [52]GBRT, RUSBoost, LRAudit analytics data (2001–2014)428 fraud and 11,323 non-fraud casesAudit analytics non-reliance restatement database and financial statementsAUC: 76.1% (RUSBoost), 72.8% (GBRT)
Jan et al. [58]RNN, LSTM153 Taiwanese listed companies (2001–2019)51 fraud and 102 non-fraud casesFinancial statementsAccuracy: 94.88% (LSTM), 87.18% (RNN)
AUC: 95.86% (LSTM), 91.30% (RNN)
Achakzai et al. [53]Meta-Classifiers (Stacked & Voting) combining LR, RF, RUSBoost, DT, SVM, MLP32173 firm-year observations of Chinese listed firms (2007–2019)2847 fraud and 29,326 non-fraud casesFinancial statementsAUC: 73.8% (Meta-classifiers), 66.4% (LR), 69.5% (DT)
Wang [59]XGBoostChinese listed companies from CSMAR (2010–2021)457 fraud and 3146 non-fraud cases Multiple indicators across governance, supervision, and financial dimensions.Accuracy: 85.9% (XGBoost)
AUC: 89.6% (XGBoost)
Zhang [68]SVM, BP, PSO-SVMEnterprise financial data sourced from WIND China Financial Database60 fraud and 60 non-fraud cases20 financial indicators and 8 non-financial indicatorsAccuracy: 85.64% (PSO-SVM)
Li et al. [63]LR, DT, NB, AdaBoost, BP, ensemble algorithm (RF, LightGBM, AdaBoost, and XGBoost)3547 Chinese listed firms (2012–2021) from CSMAR and WIND financial database6077 fraud and 6077 non-fraud cases207 financial, non-financial and textual indicators from MD&AAccuracy: 70.01% (RF), 70.14% (AdaBoost)
AUC: 70.00% (RF), 70.14 (AdaBoost)
Lee et al. [65]LR, KNN, SVM, DT, RFFirms listed on the Indonesia Stock Exchange (IDX) (2015–2023)2373 firm-year observationsBalance sheets and income statementsPrecision: 84% (RF), 74% (LR)
Recall: 67% (RF), 66% (LR)
Fu et al. [69]Probability-based One-Class SVM, DT, ensemble models (LightGBM, AdaBoost, and XGBoost)A-share listed companies (2011–2021) from CSMAR/67 feature multimodal index system of financial, non-financial and textual indicatorsAccuracy: 95% (DT), 85% (One-Class SVM)
AUC: 95% (DT), 84% (One-Class SVM)
Notes: ‘Year’ is the publishing year; ‘LR’ is short for Logistic Regression, ‘DT’ is short for Decision Tree; ‘NB’ is short for Native Bayes; ‘NN’ is short for Neural Networks; and ‘SVM’ is short for Support Vector Machine; ‘GA’ is short for Genetic Algorithm; ‘NLP’ is short for Natural Language Processing; ‘RF’ is short for Random Forest; ‘BBN’ is short for Bayesian Belief Network; ‘MFNN’ is short for Multilayer Feedforward Neural Network; ‘PNN’ is short for Probabilistic Neural Network; ‘DTNB’ is short for Decision Table/Naïve Bayes; ‘RUSBoost’ is short for Random Under-Sampling Boosting; ‘MRF’ is short for Modified Random Forest; ‘GBRT’ is short for Gradient Boosted Regression Trees, ‘ANN’ is short for Recurrent Neural Network; ‘LSTM’ is short for Long Short-Term Memory Network; ‘DBN’ is short for deepbeliefnetwork; ‘BP’ is short for BackPropagation; ‘FK’ is short for Financial Kernel, ‘GBT’ is short for Gradient Boosted Trees. SEC refers to the US Securities and Exchange Commission. AAERs refers to the Accounting and Auditing Enforcement Releases. CSMAR refers to China Stock Market & Accounting Research Database.
Table 2. An STS-based multi-level indicator system.
Table 2. An STS-based multi-level indicator system.
STS ElementsSubdivision DimensionCorresponding Indicators
X1TaskX11Supply ChainTop 5 supplier purchase volume, ratio of procurement value of the largest supplier to total procurement value, supplier concentration, supply chain concentration
X12Business PredicamentTop 5 customer sales, ratio of sales to total sales of the largest customer, customer concentration, profit per capital
X13Ambitious earnings targetFinancial indicators
X2TechnologyX21Financial StatementsAll 122 accounting subject items in balance sheets, profit sheets and cash flow sheets
X22The quality of accounting informationFinancial distress Z-score
X3StructureX31Audit opinionThe auditor is from the top 4 or top 8
X32Audit feeTotal audit fees
X33Internal controlThe share capital structure changes, the chairman of the board of directors and the general manager concurrently serve as the situation, committees, the internal control evaluation and audit report disclosure, effectiveness and deficiencies of internal control
X34EquityEquity nature, separation rate of two powers, the consistency of the working place of independent directors and companies, board, supervisory board and shareholders’ meetings, nature of property rights, the proportion of ownership, the separation rate of the two rights, the shareholding ratio and nature of the controlling shareholder, equity checks and balances, one control multiple situations, network centrality of independent directors
X35ShareholderNumber of directors, supervisors and senior executives, independent directors, shares held by directors, supervisors, senior management, regulator and senior executives, the total annual salary of directors and supervisors, total remuneration of directors, supervisors, executives, female supervisors, female directors, independent female directors, female senior managers, unpaid supervisors, directors and supervisors of unpaid remuneration, the first salary of the supervisor, remuneration of the first director, supervisors, senior management, top 10 shareholders, employees, retired employees, shareholding ratio of general manager, institutional investors and other financial institutions, total management compensation, employee density, excess employee rate, overseas or financial background, average age, proportion of men in management
X4ActorX41Litigation and arbitrationLitigation and arbitration information
X42Social responsibility reportThird-party organization verification, refer to the GRI sustainability reporting guidelines, the protection of shareholders, creditors, employees, suppliers, customer and consumers’ rights and interests, environmental and sustainable development, public relations and social welfare undertakings, the construction of the social responsibility system and the measures for improvement, the content of safety production, the company’s deficiencies, announcement disclosure intention,
X43Geographical locationLongitude and latitude of office, Longitude and latitude of registration, Province, Region,
X44registered capitalRegistered capital
Table 3. The Confusion matrix.
Table 3. The Confusion matrix.
PositiveNegative
PositiveTPFN
NegativeFPTN
Table 4. Specific of each metric.
Table 4. Specific of each metric.
MetricFormulaCode
Accuracy Accuracy = TP + TN TP + FP + FN + TNTP + TN 24
Precision Precision = TP TP + FP 25
Recall Recall = TP TP + FN 26
F1-Score F1-Score = 2 × Precision × Recall Precision + Recall 27
AUC TPR = TP TP + FN
FPR = FP FP + TN
28
29
Table 5. The detection performance of temporal behavior pattern extraction model. All methods utilized the same identical strategies listed in Section 4.4.
Table 5. The detection performance of temporal behavior pattern extraction model. All methods utilized the same identical strategies listed in Section 4.4.
MethodAccuracyPrecisionRecallF1 ScoreAUC
LSTM0.98040.97670.98000.97830.9977
CNN0.97180.97200.97180.97180.9934
GRU0.9474 0.9165 0.9715 0.9432 0.9945
RNN0.93370.93630.93370.93390.9876
Table 6. The detection performance of models with the features T and T+S. All methods utilized the same identical strategies listed in Section 4.4. ROC demonstrates the rate of change in model performance.
Table 6. The detection performance of models with the features T and T+S. All methods utilized the same identical strategies listed in Section 4.4. ROC demonstrates the rate of change in model performance.
FSFD-ETSPFeaturesAccuracyPrecisionRecallF1 ScoreAUC
LSTMT0.81350.77340.82840.80000.8909
T+S0.98040.97670.98000.97830.9977
ROC0.16690.20330.15160.17830.1068
CNNT0.84210.82640.82010.82320.9188
T+S0.97180.97200.97180.97180.9934
ROC0.12970.14560.15170.14860.0746
GRUT0.75970.72350.75470.73870.8415
T+S0.94740.91650.97150.94320.9945
ROC0.18770.1930.21680.20450.153
RNNT0.77750.74610.76670.75620.8638
T+S0.93370.93630.93370.93390.9876
ROC0.15620.19020.1670.17770.1238
Table 7. The detection performance of various methods. All methods utilized the same identical strategies listed in Section 4.4.
Table 7. The detection performance of various methods. All methods utilized the same identical strategies listed in Section 4.4.
MethodAccuracyPrecisionRecallF1 ScoreAUC
FSFD-ETSP0.98040.97670.98000.97830.9977
CNN0.71240.66020.74440.69980.7889
RNN0.76450.71620.78980.75120.8504
GRU0.79230.76810.77160.76980.8759
LR0.62400.57440.62230.59740.6719
SVM0.72440.68310.71880.70050.8002
DT0.71300.67510.69370.68430.7112
RF0.87850.90490.81600.85810.9534
XGB0.92480.92980.90090.91510.9784
Note: The specific hyperparameter settings for benchmark methods. LR: max_iter = 1000, solver = ‘liblinear’; SVM: kernel = ‘rbf’, C = 1.0, probability = True; DT: criterion = ‘gini’, max_depth = None; RF: n_estimators = 300, max_depth = None; XGB: n_estimators = 300, max_depth = 6, learning_rate = 0.1; CNN: Conv1D layer (32 filters, kernel_size = 2); GlobalAveragePooling, Dense(16), Dropout(0.3); GRU: 64 units, Dropout(0.2), Recurrent Dropout(0.2), Dense(32); RNN: SimpleRNN (64 units), Dropout(0.2), Recurrent Dropout(0.2), Dense(32); LSTM: Bidirectional architecture [128, 64 units], attention mechanism.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Xia, H.; Jiang, J.; Wang, Q. An Interpretable Financial Statement Fraud Detection Framework Enhanced by Temporal–Spatial Patterns. Math. Comput. Appl. 2025, 30, 138. https://doi.org/10.3390/mca30060138

AMA Style

Xia H, Jiang J, Wang Q. An Interpretable Financial Statement Fraud Detection Framework Enhanced by Temporal–Spatial Patterns. Mathematical and Computational Applications. 2025; 30(6):138. https://doi.org/10.3390/mca30060138

Chicago/Turabian Style

Xia, Hui, Jinhong Jiang, and Qin Wang. 2025. "An Interpretable Financial Statement Fraud Detection Framework Enhanced by Temporal–Spatial Patterns" Mathematical and Computational Applications 30, no. 6: 138. https://doi.org/10.3390/mca30060138

APA Style

Xia, H., Jiang, J., & Wang, Q. (2025). An Interpretable Financial Statement Fraud Detection Framework Enhanced by Temporal–Spatial Patterns. Mathematical and Computational Applications, 30(6), 138. https://doi.org/10.3390/mca30060138

Article Metrics

Back to TopTop