Skip to Content
ComputersComputers
  • Article
  • Open Access

27 June 2026

22 Pages

NAPO-SCVD: Noise-Aware Preference Reinforcement Large Language Model for Smart Contract Vulnerability Detection

,
,
,
and
1
Department of Network Security and Protection, Shanxi Police College, Taiyuan 030400, China
2
School of Software, North University of China, Taiyuan 030051, China
*
Author to whom correspondence should be addressed.

Abstract

As the core automated execution components of blockchain technology, smart contracts enable programmatic control over digital assets; however, their immutable characteristics and inherent logical vulnerabilities give rise to substantial security risks. Although smart contract vulnerability detection methods based on large language models (LLMs) have exhibited certain potential in vulnerability detection and explanation, the coarse-grained modeling of traditional binary preference optimization paradigms hinders the model ability to learn the priority of domain-specific requirements, frequently leading to extreme optimization at the cost of detection accuracy. Furthermore, existing approaches fail to consider non-ideal factors in real-world application scenarios and overlook noise interference induced by missing prompts, which results in inadequate detection stability and reliability, making them challenging to adapt to complex practical scenarios. To address these critical issues, this study proposes a Noise-Aware Preference Reinforcement Large Language Model for Smart Contract Vulnerability Detection (NAPO-SCVD). This method adopts a four-stage framework consisting of data construction, continuous pre-training, supervised fine-tuning, and noise-aware preference optimization. Specifically, it enhances the model’s comprehension of contract syntax and semantics through domain-specific pre-training, improves its detection and explanation capabilities using high-quality datasets, constructs deliberately guided biased explanations to simulate noisy samples, refines preference gradients, and strengthens the model’s anti-interference ability. Consequently, this approach achieves high-precision and high-reliability smart contract vulnerability detection, along with fine-grained explanations.

1. Introduction

Supported by its decentralized core architecture, blockchain technology has been successfully implemented and scaled in fields such as financial settlement, supply chain traceability, and government evidence preservation [1]. This technology establishes a secure and trustworthy distributed ledger system, which records transaction data through a consensus mechanism among distributed nodes. Combined with cryptographic algorithms, it ensures the integrity, traceability, and verifiability of all transactions at the technical foundation level [2]. Within the blockchain ecosystem, smart contracts act as the core automated execution components, which implement end-to-end programmatic control over cryptocurrencies and various digital assets based on predefined programmable logic. When the trigger conditions embedded in the contract are satisfied, the predefined instructions are automatically executed; both the execution results and the contract source code are permanently recorded in the blockchain ledger, without the need for any third-party intervention [3]. However, the immutable nature of smart contracts, coupled with inherent vulnerabilities in their logical architecture, gives rise to significant and difficult-to-mitigate security risks. In 2016, attackers exploited a recursive call vulnerability in the DAO smart contract to illegally transfer Ethereum tokens worth tens of millions of US dollars, leading to substantial asset losses and exerting a profound impact on the technical trust and development trajectory of the blockchain industry [4,5].
Currently, existing smart contract vulnerability detection methods can be categorized into three main types: traditional methods, deep learning-based methods, and large language model-based methods [6]. Traditional smart contract vulnerability detection approaches primarily include symbolic execution (e.g., Oyente, Mythril) and static analysis (e.g., Slither, SmartCheck). These methods rely heavily on predefined vulnerability patterns or rule sets, and their detection performance is largely constrained by the completeness of the predefined logic. In practical application scenarios characterized by increasingly complex contract logic and diversified business operations, these traditional methods are prone to problems such as false negatives and false positives, which lead to a significant degradation in overall detection performance [7].
In response to the inherent limitations of traditional methods, academic research has gradually shifted toward deep learning-based smart contract vulnerability detection approaches. Deep learning models have significantly improved the accuracy and cross-scenario generalization of smart contract vulnerability detection. Typical applications in this field include: constructing bidirectional long short-term memory (LSTM) networks integrated with attention mechanisms to detect specific types of vulnerabilities [8] and leveraging deep learning models to measure the semantic and structural similarity between the target contract and contracts with known vulnerabilities, thereby achieving transferable vulnerability identification [9]. Meanwhile, graph neural networks (GNNs), which can effectively capture the topological information of smart contracts, have emerged as a research hotspot in this field due to their strong alignment with the structured characteristics of contract code. By supplementing graph features with multimodal features such as expert knowledge and bytecode, the performance of smart contract vulnerability detection can be further enhanced [10,11]. However, these deep learning-based approaches impose stringent requirements on the scale, quality, and annotation completeness of the training datasets. Furthermore, the generalization ability of the models is constrained by the scope of scenarios covered in the training data, making it difficult for them to adapt to the vulnerability detection needs of smart contracts involving different data types and business models, which ultimately results in a lack of universality [12].
In recent years, large language models (LLMs) have emerged as a research hotspot in the field of smart contract vulnerability detection and explanation, owing to their powerful natural language understanding and logical reasoning capabilities. Chen et al. [3] and David et al. [13] systematically evaluated the detection capabilities of mainstream large language models using real-world smart contract vulnerability datasets, clarifying their technical potential and existing bottlenecks; Hu et al. [14] explored the logical reasoning capabilities of large language models, providing a theoretical foundation for identifying complex vulnerabilities; Wei et al. [15] proposed Fine-Tuned Language Model for Smart Contract Auditing (FTSmartAudit), which achieves competitive performance against mainstream existing methods in the field. The two-stage iAudit framework proposed by Ma et al. [16] enables both vulnerability detection and result interpretation simultaneously; however, this decoupled architecture carries the risk of inconsistencies between detection and interpretation results. Yu et al. [17] proposed a Reinforced Large Language Model for Smart Contract Vulnerability Detection (Smart-LLaMA-DPO), which employs a three-stage training process—comprising continuous pre-training, supervised fine-tuning, and direct preference optimization—to detect and localize vulnerabilities in smart contracts.A comprehensive comparison of mainstream smart contract vulnerability detection methods is illustrated in Table 1.
Table 1. Comparison of mainstream smart contract vulnerability detection methods.
Notably, the representative baseline Smart-LLaMA-DPO only adopts standard binary direct preference optimization (DPO) and ignores noise interference from incomplete prompts. Meanwhile, most existing robustness-enhanced LLM alignment methods are developed for general text tasks and fail to handle domain-specific noise in smart contract auditing. These are the critical limitations this study intends to address. The continuous pre-training and supervised fine-tuning adopted in this work are classic and widely used generic pipelines for LLM adaptation in code-related tasks, which inherit from prior research including Smart-LLaMA-DPO. Different from both binary preference learning and general robustness optimization solutions for large language models, our core innovations focus on breaking the constraints of binary preference modeling and enhancing model anti-noise ability tailored for practical noisy scenarios in smart contract auditing.
However, vulnerability detection methods for smart contracts based on large language models have the following limitations: (1) The traditional binary preference optimization paradigm (good/bad) suffers from coarse-grained preference modeling, failing to effectively capture the fine-grained, gradient-based preference structures required for expert-level vulnerability detection. This makes it difficult for the model to accurately learn the priority ranking of core domain requirements, preventing it from distinguishing the fundamental difference between ‘basic correctness is ensured but details are biased’ and ‘core detection logic is incorrect’. Consequently, the model falls into an optimization dilemma, where an excessive pursuit of interpretive completeness may sacrifice the accuracy of the core detection objective, making it difficult to consistently produce detection results that align with the task’s core requirements. (2) Existing methods do not fully account for non-ideal factors in real-world scenarios, particularly overlooking the issue of misleading explanations caused by partial prompt omissions. For example, a missing prompt for external call analysis will make the model fail to fully interpret reentrancy risks, and ignoring storage layout analysis will lead to unreliable delegatecall detection. In actual workflows, critical information may be missing from prompts fed into the model due to operational oversights or incomplete requirement specifications. Such incomplete prompts can easily drive the model to generate biased explanations that deviate from the actual vulnerability conditions. Current mainstream models lack the ability to assess the completeness of prompts; they cannot identify the validity of such biased explanations nor do they possess targeted correction mechanisms. Consequently, the models are highly susceptible to interference from noisy samples generated by missing prompts, resulting in a significant decline in detection stability and reliability, making them difficult to adapt to complex practical scenarios. Additionally, general LLMs tend to generate fluent but factually wrong explanations, which is extremely risky for asset-related smart contract auditing.
To address these challenges, we propose a method based on noise-aware preference reinforcement for large language models (NAPO-SCVD) for smart contract vulnerability detection. This method combines four stages: data construction, continuous pre-training, supervised fine-tuning, and noise-aware preference optimization. It utilizes large-scale smart contract code for domain-specific pre-training, enabling the large language model to better understand the syntax and semantics of smart contracts. Subsequently, a comprehensive smart contract vulnerability dataset containing detailed explanations and precise location information is constructed. This high-quality dataset is used for fine-tuning to enhance the large language model’s ability to detect vulnerabilities and provide explanations. Finally, a dataset comprising pairs of model outputs of varying quality is constructed. Each output pair in the dataset consists of three parts: a human expert preference section (high-quality explanations), a human expert non-preference section (low-quality explanations), and a human expert sub-preference section (deliberately biased explanations mimicking noise samples). By employing a designed loss function, the method encourages the large language model to increase the probability of preferred outputs, reduce the probability of non-preferred outputs, and resist interference from noise samples, thereby enhancing the model’s robustness in multimodal scenarios. The superiority of the proposed model is validated on a dataset covering four major vulnerability types: reentry, timestamp dependency, integer overflow/underflow, and delegation calls. The key innovations are as follows:
(1) We design a ternary preference dataset containing preferred, non-preferred and sub-preference samples. The sub-preference samples are manually crafted biased explanations to simulate missing-prompt noise in practical scenarios, breaking the limitation of traditional binary preference modeling. Combined with the newly designed NAPO loss function, the model can distinguish gradient differences among different-quality outputs and gain strong anti-noise capability.
(2) Different from existing DPO-based methods such as Smart-LLaMA-DPO, we propose a dynamic weighted fusion strategy for DPO loss and NAPO loss, which dynamically adjusts the optimization direction according to data quality, and further optimizes the model’s detection accuracy and explanation rationality on the basis of domain adaptation.
(3) We build a complete evaluation system including a prompt noise robustness test and dual-track explanation quality assessment. Extensive experiments on real-world smart contract datasets verify that NAPO-SCVD outperforms existing baselines in detection performance, explanation quality and noise resistance.

2. Problem Definition

NAPO-SCVD is an integrated framework that jointly performs four core tasks for smart contract security analysis: binary vulnerability detection, vulnerability type classification, vulnerability localization and explanation generation. This method assigns labels y ^ to each individual smart contract, where y ^ = 1 indicates that the smart contract has a vulnerability, and y ^ = 0 indicates that the smart contract is secure. It focuses on four types of vulnerabilities: reentrancy (RE) vulnerabilities, timestamp dependence (TD) vulnerabilities, integer overflow/underflow (IO), and delegatecall (DE) vulnerabilities [18]. Reentrancy (RE) vulnerabilities occur when a contract calls an external contract or transfers assets (Ether or tokens) before completing internal state changes, allowing an attacker to repeatedly call the vulnerable function and potentially withdraw funds multiple times. Timestamp dependence (TD) vulnerabilities occur when a smart contract relies on block timestamps for critical operations; if a miner manipulates the timestamp during random number generation or critical decision-making processes, it compromises contract integrity and causes financial losses. Integer overflow/underflow (IO) occurs when the result of an arithmetic operation exceeds the storage range of a variable. In an overflow scenario, the value wraps around the minimum value of the type; in an underflow scenario, the value wraps around the maximum value, leading to unexpected behaviors such as balance errors or runaway loops in the contract. Delegated Execution (DE) is a low-level function call mechanism that allows a contract to dynamically load code from other contracts. While this mechanism provides powerful upgradeability, it executes within the context of the called contract, potentially tampering with the calling contract’s storage [19].

3. Smart Contract Vulnerability Detection Using Noise-Aware Reinforcement Learning for Large Language Models

The proposed NAPO-SCVD model comprises four key stages: data construction, continuous pre-training, supervised fine-tuning, and noise-aware preference optimization. The specific model architecture is shown in Figure 1. First, a dataset is constructed. Building upon a pre-trained large language model, we integrate smart contract code with general-purpose data to perform continuous pre-training. Next, we utilize annotated vulnerability detection data for supervised fine-tuning. Finally, we perform noise-aware preference optimization by comparing the policy model with a reference model, resulting in a target model that balances robustness and accuracy.
Figure 1. Overall framework of the Noise-Aware Preference Reinforcement Large Language Model for Smart Contract Vulnerability Detection (NAPO-SCVD).

3.1. Data Construction

Continuous pre-training (CPT) Dataset: The continuous pre-training dataset is based on the dataset constructed by Yu et al. [17], following the process described below: Using Google BigQuery, Ethereum smart contract addresses containing at least one transaction were filtered out, and the corresponding contract source code was retrieved via the Etherscan API. To ensure data uniqueness, token-based similarity detection using the Jaccard index was employed. The smart contracts were decomposed into three components: core business logic, library code, and imported files. This approach effectively removed duplicate code while preserving unique implementation logic.
Supervised fine-tuning SFT data construction: The supervised fine-tuning data utilized the labeled dataset constructed by Yu et al. [17]. This dataset employed the open-source large language models Qwen2.5-72B-Instruct and Mistral-Large-Instruct-2407-123B as initial models, and designed dedicated prompt strategies for each vulnerability type (RE, TD, DE, IO) to generate explanations; subsequently, LLaMA 3.1-70B-Instruct was used as the evaluation model to assess the explanations generated by Qwen2.5 and Mistral based on three criteria: correctness (weight 0.6), thoroughness (weight 0.3), and clarity (weight 0.1), with each criterion scored on a 1–10 scale. Finally, the explanation with the highest weighted composite score (WCS = 0.6 × Correctness + 0.3 × thoroughness + 0.1 × Clarity) was selected for manual review. Clear annotation guidelines were established, and human reviewers manually annotated the vulnerabilities in the smart contracts along with their corresponding explanations in detail.
Noise-Aware Preference Optimization (NAPO) Data Construction: The dataset comprises a human expert preference section (high-quality explanations), a human expert non-preference section (low-quality explanations), and a human expert sub-preference section (deliberately biased explanations). The human expert preference section is derived from high-quality explanations in the SFT dataset; the human expert non-preference section consists of low-scoring outputs from LLaMA 3.1, where the reasoning or analytical logic is completely flawed. The human expert sub-preference section uses deliberately biased explanations; while maintaining basic correctness, it intentionally reduces analytical depth, simplifies logical derivations, ignores associated risks, and lowers the rigor and completeness of the explanations; RE identifies only external calls without analyzing execution order, reentry risks based on inheritance, or token transfer scenarios; TD only flags direct use of block.timestamp without examining the impact of state transitions or analyzing the randomness of timestamps in gaming/lottery applications; IO only flags simple overflow points without considering complex expressions, type conversions, or version-specific verification requirements; and DE only identifies basic delegatecall usage without analyzing the security of storage layouts or escalation mechanisms in the proxy pattern.
Annotation Protocol and Quality Control: A total of five professional annotators participated in data annotation. Three annotators have more than 3 years of practical experience in smart contract security auditing, and the other two are researchers specializing in smart contract vulnerability analysis. We used Cohen’s Kappa coefficient to measure inter-annotator agreement, and the average Kappa score reached 0.87, indicating excellent annotation consistency. All biased explanations for sub-preference samples were manually revised from high-quality explanations by experts, instead of being generated by LLMs or automatic tools. To guarantee that sub-preferred explanations remain basically correct with only partial biases, we implemented a two-round cross-review process: samples with wrong core judgments were eliminated. In addition, a double-blind design was adopted during annotation, and annotators were not informed of the model information for original explanation generation. To prevent data leakage, we split all data at the contract level with a ratio of 7:1:2 for training, validation and test sets. All samples from one original contract were assigned to a single dataset partition.

3.2. Continuous Pre-Training

During the continuous pre-training phase, the focus is on enhancing the large language model’s understanding of concepts specific to smart contract security. This process is guided by:
L C P T = − E x ~ D [ ∑ i = 1 n log P ( x i | x < i , c i ) ]
where D represents the smart contract continuous pre-training dataset, x i is a sequence of tokens derived from the contract, c i denotes the contextual information surrounding x i , and P ( x i | x < i , c i ) is the prediction probability of the large language model.
This formalized approach enables large language models to learn key elements: first, the syntax and semantics of smart contract understanding; second, the ability to prevent reentrancy attacks through a check-effect-interaction mechanism, combined with other security best practices, to help smart contract agents identify vulnerabilities; third, critical security functions (such as transfer, send, and call.value()) essential for fund transfers and inter-contract interactions, enabling smart contract agents to recognize their correct usage and potential risks; and fourth, the ability to perform contextual analysis to understand contract structures, function interactions, and state variables for vulnerability detection, maintaining general capabilities while avoiding catastrophic forgetting—the system integrates diverse data from mathematics, coding, and linguistics to enhance its generalization capabilities for novel contract patterns.

3.3. Supervised Fine-Tuning

This stage jointly optimizes two core tasks: vulnerability detection and explanation generation, based on the domain knowledge learned in continuous pre-training. The corresponding loss function is formulated as:
L S F T = 1 2 ( L d e t e c t + L e x p l a i n ) L d e t e c t = − ∑ ( x , y ) ∈ D v u l log P ( y | x ; θ ) L e x p l a i n = − ∑ ( x , e ) ∈ D e x p ∑ i = 1 | e | log P ( e i | x , e < i ; θ )
Here, x represents the input smart contract code, and y and e represent the target outputs for the detection task (vulnerability labels) and the generation task (explanations), respectively. The parameters θ include all trainable components of the large language model, such as the attention mechanism, feedforward network, and embedding layers. This balanced loss function design ensures that the large language model is effectively trained for both vulnerability detection and explanation generation tasks. The domain-specific knowledge acquired during the continuous pre-training phase serves as the critical foundation for this fine-tuning stage. The model is not only capable of understanding the complex context and nuances of smart contracts but also further optimizes its ability to generate explanations while enhancing its vulnerability detection capabilities.

3.4. Noise-Aware Preference Optimization

To address the issue of coarse-grained optimization in binary preference optimization, NAPO constructs a ternary gradient-based preference structure: the human expert preference component y w (high-quality explanations), the human expert non-preference component y l (low-quality explanations), and the human expert sub-preference component y w l (deliberately guided biased explanations—mimicking noise samples). The optimization objective is a weighted combination of the general preference loss (DPO) and the noise-aware preference optimization (NAPO) loss. The model can learn to identify biased explanations caused by suboptimal factors such as missing prompt words, thereby avoiding the model falling into extreme optimization traps. It accurately learns the core priorities of the task, enhances the model’s resistance to noise samples, and adapts to the complex demands of real-world application scenarios. The specific framework diagram is shown in Figure 2.
Figure 2. Noise-aware preference optimization framework.
Standard DPO offers a simplified method for learning human preferences without the need for explicit reward modeling or reinforcement learning. Under the KL-constrained optimization objective, the optimal policy π ∗ for the reward function r ∗ can be expressed as:
π ∗ ( y | x ) = 1 Z ( x ) π r e f ( y | x ) exp ( 1 β r ∗ ( x , y ) )
Specifically, this x represents a snippet of smart contract code, y represents a vulnerability analysis and explanation generated by a large language model, and π r e f represents a reference strategy, the final optimized model to be obtained π ∗ ( y | x ) , and a vulnerability analysis generated in accordance with expert preferences.
Rewrite the equation and express the reward function in terms of the optimal policy:
r ∗ ( x , y ) = β log π ∗ ( y | x ) π r e f ( y | x ) + β log Z ( x )
Reformulate the Bradley–Terry preference model in terms of preferences rather than rewards:
ψ ( x , y w , y l ) = β log π ( y w | x ) π r e f ( y w | x ) − β log π ( y l | x ) π r e f ( y l | x )
In this context, y 1 and y 2 indicate two different interpretations of the vulnerability.
DPO: By incorporating both the human expert’s preferred interpretations (high-quality explanations) and non-preferred interpretations (low-quality explanations) into the DPO loss function, this approach encourages the large language model to assign higher probabilities to the preferred interpretations and lower probabilities to the non-preferred ones, thereby effectively learning the expert’s preferences. By minimizing the loss, the policy model is directly optimized to align with the expert’s preferred interpretations of the smart contract:
L D P O ( π θ ; π r e f ) = − E ( x , y w , y l ) ~ D log σ ( β log π θ ( y w | x ) π r e f ( y w | x ) ) − β log π θ ( y l | x ) π r e f ( y l | x )
In particular, ( x , y w , y l ) are triples from dataset D, representing the input, the human expert’s preferred part (high-quality explanations), and the human expert’s non-preferred part (low-quality explanations), respectively; π θ is the policy model being optimized, initialized to π r e f , which learns to generate outputs that align with expert preferences through the training process.
NAPO: By incorporating both the human expert’s primary preferences (high-quality explanations) and secondary preferences (deliberately guided biased explanations—mimicking noisy samples) into the NAPO loss function, and combining the cross-entropy loss function from DPO with the mean absolute error (MAE) loss processed via the negative Box–Cox transformation, the method achieves rapid convergence of the cross-entropy loss function while preserving the noise robustness of the MAE. This enables large language models to strike a balance between learning high-quality responses and resisting noise interference, clearly defining the boundary between high-quality and acceptable results, and enhancing model robustness. The formula is as follows:
L N a P O ( x , y w , y l b ) = 1 q ( 1 − σ ( β log π θ ( y w | x ) π r e f ( y w | x ) − β log π θ ( y l b | x ) π r e f ( y l b | x ) ) q )
Here, the triplets ( x , y w , y l b ) in dataset D represent the input, the human expert preference component (high-quality explanations), and the human expert sub-preference component (deliberately biased explanations—simulated noise samples); π θ represents the policy model being optimized, initialized as π r e f ; and the dynamic coefficient q is adjusted based on the reward difference between high-quality interpretations and biased interpretations. When response discriminability is low (data is noisy), q increases, making the model rely more on the noise-resistant properties of the mean absolute error loss; when discriminability is high (data quality is good), q decreases, making the model rely more on the fast convergence of the cross-entropy loss function. The formula is as follows:
q = 1 − σ ( α ( β log π ( y w | x ) π r e f ( y w | x ) − β log π ( y l b | x ) π r e f ( y l b | x ) ) )
As shown in Equation (8), q is computed adaptively based on the model’s preference difference between high-quality and biased explanations. Its natural range is (0,1), and we clip it to [0.1,0.9] during training to ensure stability. Since q is not a fixed hyperparameter but an adaptive coefficient determined by the data distribution, the model maintains stable performance across the entire valid range without relying on manual tuning.
Final optimization objective: By dynamically weighting and fusing the general preference loss (DPO) and the noise-aware preference optimization (NAPO) loss, we use these weights to give high-quality explanations a greater weight in the loss function, while balancing the training directions of direct preference optimization and noise-aware preference optimization. The formula for the dynamic weights γ i is as follows:
γ i = ψ ( x , y w , y i ) ∑ i ψ ( x , y w , y i ) , i ∈ [ y l , y l b ]
The final loss function both aligns the model with general preferences and enhances the model’s robustness against noisy samples:
L = γ y l L D P O ( x , y w , y l ) + γ y l b L N a P O ( x , y w , y l b )

4. Experimental Results and Analysis

We evaluate the NAPO-SCVD method by examining the following research questions:
(1) RQ1: How does NAPO-SCVD perform in detecting four types of smart contract vulnerabilities compared to existing baseline methods?
(2) RQ2: How does each individual module affect the effectiveness of NAPO-SCVD?
(3) RQ3: How do the explanations generated by NAPO-SCVD perform in terms of accuracy, thoroughness, and clarity?
(4) RQ4: How does NAPO-SCVD perform in terms of the stability of explanation quality under the influence of missing prompts and noise interference?

4.1. Datasets

Continuous pre-training dataset: We use the dataset constructed by Yu et al. [17], which contains 286,397 instances from various domains, including general-purpose code, mathematics, and English and Chinese text.
Supervised fine-tuning dataset: We use the dataset constructed by Yu et al. [17], which includes real smart contracts from the SmartBugs dataset, multi-user vulnerabilities, Ethereum scanners, GitHub repositories, and blog posts; all data is from 2020 to 2024. The final dataset contains 3390 RE, 1167 TD, 1013 IO, and 698 delegatecall samples.
Noise-aware preference optimization Dataset: The NAPO dataset shares the same sources as the SFT dataset. It contains 270 RE, 227 TD, 260 IO, 265 DE instances, including high-quality explanations, low-quality explanations, and deliberately biased explanations.
Evaluation dataset: The dataset combines four vulnerability types (RE, TD, IO, DE). After processing, the dataset contains 3542 samples, with the distribution across vulnerability categories as follows: RE (116/470, 13.27%), TD (672/896, 25.30%), IO (354/1458, 41.16%), and DE (76/340, 9.60%). The numbers in parentheses represent (number of vulnerability samples/total number of samples, i.e., the proportion of samples of that type in the evaluation dataset). For the quality evaluation phase, considering the cost of manual review, we maintained the same ‘number of vulnerability samples/total number of samples’ ratio as in the optimized dataset and randomly sampled approximately 200 samples for each vulnerability type, resulting in a total of 1061 samples, distributed as follows: RE (58/235), TD (142/224), IO (59/243), and DE (38/170).

4.2. Experimental Setup

The comparative study includes baseline methods for smart contract vulnerability detection across four categories: rule-based methods, neural network-based methods, pre-trained model-based methods, and large language model-based methods. Rule-based methods include tools such as Mythril [20], Osiris [21], Oyente [22], Slither [23], Conkas [24], Smartian [24], sFuzz [25], Solhint [17], and Smartcheck [26]; neural network-based methods include GCN [27], TMP [28], AME [29], SMS [30], and DMT [30]; methods based on pre-trained models rely on pre-trained models such as CodeT5 [31] and fine-tuning techniques to identify vulnerabilities, including Peculiar [32] and PSCVFinder [33]; and methods based on large language models rely on large language models to identify vulnerabilities in smart contracts, including LLaMA 3.1-8B-Instruct [34], LLaMA 3.1-70B-Instruct [34], Qwen2.5-7B-Instruct [35], Qwen2.5-72B-Instruct [35], GPT-4o [36] (gpt-4o-2024-08-06), Claude-3.5-Sonnet [17] (claude-3-5-sonnet-20241022), GPTScan [17], GPTLens [17], and the fine-tuned methods FTSmartAudit [15], iAudit [16], Agent4Vul [37] and Smart-LLaMA-DPO [17]. All baseline methods were evaluated under consistent conditions on the same evaluation dataset.
Models were pre-trained, supervised fine-tuned, and noise-aware preference-optimized using LlamaFactory [38] and DeepSpeed [39], with cross-entropy loss and AdamW for parameter optimization. The backbone base model of NAPO-SCVD is LLaMA 3.1-70B-Instruct. All training and inference are performed on NVIDIA A100 80GB GPUs, with the maximum context length set to 1024 tokens. The total training time of the full pipeline is approximately 72 h. During pre-training, the batch size was set to 64, the number of training epochs to two, and the learning rate to 10−8. During supervised fine-tuning, the batch size was set to eight, the number of training epochs to three, and the learning rate to 10−5. For DPO training, the truncation length was set to 1024, the batch size to eight, the learning rate to 10−5, and the number of training epochs to 50. For fair evaluation, all baselines including closed-source models GPT-4o and Claude-3.5-Sonnet share unified inference settings: the temperature is fixed at 0.1 with greedy decoding. All experiments are repeated three times and average results are used. All models adopt the same context window during inference. Model performance was ultimately evaluated across two dimensions: vulnerability detection capability and explanation quality. For vulnerability detection, four standard metrics are used: precision measures the proportion of actual vulnerabilities among predicted positive results; recall indicates the proportion of detected vulnerabilities among all actual vulnerabilities; the F1 score provides a balanced metric through the harmonic mean of precision and recall; and accuracy reflects the overall correctness of the predicted results. Vulnerability explanation quality is assessed using three key metrics: correctness, thoroughness, and clarity. All open-source baseline models including Smart-LLaMA-DPO are fully reproduced by us following the official configurations, datasets and training pipelines released in their original paper; our NAPO-SCVD is improved based on the reproduced Smart-LLaMA-DPO framework with LLaMA 3.1-70B-Instruct as the shared backbone model. All experimental hyperparameters and settings for our model and comparison models follow the configurations reported in https://doi.org/10.5281/zenodo.15200616.
To further verify the reliability of performance improvements, we conduct paired t-tests and calculate 95% confidence intervals based on the results of three independent repeated experiments. The statistical results show that for all four vulnerability types, the performance of NAPO-SCVD is significantly better than Smart-LLaMA-DPO with p < 0.05. All values in Table 2 are reported as mean (standard deviation) calculated from three independent repeated experiments.
Table 2. Performance comparison of multiple metrics for smart contract vulnerability detection (reentrancy and timestamp dependency vulnerabilities).

4.3. Experimental Results

(1) RQ1: How does NAPO-SCVD perform in detecting four types of smart contract vulnerabilities compared to existing baseline methods? To answer this question, we compared the performance of the proposed method against baseline methods across four different vulnerability types. The results are shown in Table 2 and Table 3.
Table 3. Performance comparison of multiple metrics for smart contract vulnerability detection (overflow/underflow and delegatecall vulnerabilities).
NAPO-SCVD demonstrates superior performance across all vulnerability detection categories. We further provide confusion matrices for each vulnerability type to intuitively reveal the model’s performance on precision and recall, as well as its misclassification characteristics. By establishing a three-tier preference structure—“high-quality (completely correct, comprehensive, and clear) → suboptimal (basically correct but lacking in detail) → low-quality (incorrect or severely deficient)”—the model is able to accurately learn the core priorities of the task, prioritizing the correct identification of vulnerability types and locations, followed by the thoroughness and clarity of vulnerability explanations.
The confusion matrices shown in Figure 3 illustrate that NAPO-SCVD maintains high recall rates from 92.04% (RE) to 100.00% (DE) across four vulnerability categories. Timestamp dependency (TD) achieves the best overall performance with low false negatives (FN = 7) and false positives (FP = 21). For delegatecall (DE), the model obtains 100.00% recall and 75.00% precision: it detects all real vulnerabilities with zero false negatives (FN = 0) but generates 25 false positives (FP = 25). Overflow/underflow (IO) has 62 false positives, leaving room for improvement in identifying malicious integer anomalies.
Figure 3. Confusion matrices of NAPO-SCVD on four vulnerability types.
This precision–recall trade-off is an inherent characteristic of vulnerability detection tasks, and it is especially prominent in smart contract security auditing. In blockchain scenarios, once vulnerabilities are missed and exploited by attackers, it will lead to irreversible digital asset losses, user property damage and even systemic security risks for the entire project. Therefore, pursuing a high recall rate to ensure all potential vulnerabilities are fully captured is the core principle of practical auditing work.
Nevertheless, we recognize that excessive false positives cannot be overlooked, as they introduce tangible practical costs in real audit pipelines. Massive false alarms extend manual review time, weaken auditors’ trust in the detection tool, and easily trigger alert fatigue among security practitioners. This limitation is particularly evident for delegatecall detection, where recall reaches 100% while precision only hits 75%, producing a large volume of redundant candidate samples for manual screening. Although a certain number of false positive samples will increase the manual screening burden for auditors, such overhead is completely acceptable in actual deployment. Professional auditors can quickly distinguish and eliminate false alarms through targeted code review and vulnerability verification, which will not affect the overall auditing efficiency and practical application value of our model.
(2) RQ2: How does each individual module influence the effectiveness of NAPO-SCVD?
To address this question, we conducted ablation experiments comparing the proposed method’s performance when core modules were removed, designing multiple sets of experiments as follows: Using the complete NAPO-SCVD model—which integrates continuous pre-training (CPT), direct preference optimization (DPO), and noise-aware preference optimization (NAPO)—as the baseline model base, we constructed ablation variants by removing core modules individually or in combination: w/o NAPO (removing only the NAPO module while retaining CPT and DPO), w/o DPO + NAPO (removing the DPO and NAPO modules while retaining only CPT), w/o CPT (removing only the CPT module while retaining DPO and NAPO), and w/o CPT + DPO + NAPO (removing CPT, DPO, and NAPO). By comparing the performance differences in vulnerability detection metrics (accuracy, F1 score) between various ablation variants and the baseline model, we clarify the key contributions of the NAPO and CPT modules to model performance. The results are shown in Figure 4.
Figure 4. To compare the proposed method through ablation experiments on core modules, multiple sets of ablation experiments were designed: (a) ablation experiment accuracy; (b) ablation experiment F1 Score.
a. Continuous pre-training (CPT) is a core prerequisite for ensuring the model’s baseline detection performance. Compared to the baseline model, the detection accuracy and F1 score of the w/o CPT version (with the CPT module removed) both show a significant decline. This indicates that domain-specific pre-training using large-scale smart contract data effectively enhances the model’s understanding of Solidity syntax, contract semantics, and vulnerability features, thereby equipping it with key domain knowledge for vulnerability detection.
b. Noise-aware preference optimization (NAPO) is a key enhancement module for improving model performance. The performance of w/o NAPO (without the NAPO module) is lower than that of the baseline model, proving that incorporating both human expert preferences (high-quality explanations) and human expert sub-preferences (deliberately guided biased explanations—mimicking noise samples) into the NAPO loss function, through fine-grained preference modeling and noise-robust design, effectively aligns with expert-level detection standards, precisely capturing the core priorities of the task, thereby further improving the accuracy and stability of the model’s detection.
c. The combined optimization of DPO and NAPO offers complementary value: The results without DPO + NAPO demonstrate that the preference optimization mechanism (DPO + NAPO) can effectively translate human experts’ detection preferences and judgment logic into explicit optimization objectives that the model can learn. By establishing a foundational preference alignment framework via DPO and leveraging NAPO to achieve fine-grained optimization and noise resistance, it addresses the shortcomings of relying solely on domain-specific pre-training to precisely match actual detection requirements.
d. The synergy among multiple modules significantly enhances the overall performance of the vulnerability detection model: The model w/o CPT + DPO + NAPO (with all core modules removed) performs the worst, while the baseline model integrating CPT, DPO, and NAPO performs the best. The CPT module provides the ability to adapt domain-specific concepts for smart contract vulnerability detection; the DPO module establishes preference alignment between the foundation and human-interpretable preferences; and the NAPO module enhances noise robustness and preference granularity. The synergy among these three modules—combining domain knowledge, foundational alignment, and robust optimization—collectively underpins the outstanding detection performance of NAPO-SCVD.
(3) RQ3: How do the explanations generated by NAPO-SCVD perform in terms of accuracy, thoroughness, and clarity?
To systematically evaluate the quality of explanations for smart contract vulnerability detection, a dual-track verification and evaluation framework was developed. This framework comprises two modules: automated evaluation using large language models (LLMs) and evaluation by human experts. By combining subjective and objective assessments, the framework aims to overcome the limitations of relying on a single evaluation method. Both types of evaluation quantify explanation quality based on three core dimensions: correctness, thoroughness, and clarity. The correctness dimension focuses on whether the explanation accurately matches the vulnerability’s essential characteristics, technical principles, and trigger conditions, with no factual errors; the thoroughness dimension examines whether the explanation comprehensively covers the vulnerability’s root cause analysis, potential risks, scope of impact, and remediation strategies, with no critical information missing; and the clarity dimension measures whether the explanation has a clear logical structure, uses concise and accurate language, and is easily understandable by non-technical personnel.
A 1–4 point scale was adopted (1 indicates strongly disagree, 2 indicates mostly disagree, 3 indicates somewhat agree, and 4 indicates strongly agree), and the combined proportion of 3-point (somewhat agree) and 4-point (strongly agree) ratings was defined as the overall positive evaluation rate to measure the effectiveness of the model-generated explanations.
To clarify the scoring rule: in practical smart contract security auditing, explanations rated 3 (moderately acceptable) and 4 (excellent) can both meet the basic work requirements, so we merged these two scores to calculate the positive evaluation rate for an intuitive reflection of the overall qualified ratio.
For the LLM-based automatic evaluation, the evaluator LLM is fully blinded to the identity of the model that generated each explanation sample. All test samples are randomly shuffled before scoring to eliminate potential sequence bias.
Regarding the human evaluation protocol: a total of five professional experts participated in this evaluation, including three senior practitioners with more than three years of practical experience in smart contract security auditing and two researchers specializing in vulnerability analysis. Every sample is independently scored by all five evaluators without sample segmentation and assignment. When inconsistent scoring results appeared among evaluators, all experts held group discussions to re-assess relevant samples and reach a unified conclusion. Furthermore, we calculated the Cohen’s Kappa coefficient to evaluate inter-rater reliability, and the average Kappa value reached 0.87, which proves excellent consistency among all evaluators.
The experiment selected LLaMA 3.1-8B, FTSmartAudit, iAudit, and Smart-LLaMA-DPO as baseline models to conduct a multidimensional performance comparison with NAPO-SCVD. The results including positive evaluation rates, average scores and standard deviations are shown in Table 4. The experimental results indicate that NAPO-SCVD performs well in both automated large language model evaluation and human expert evaluation scenarios: in LLM evaluation, the model achieved overall positive evaluation rates of 88.88%, 90.76%, and 98.11% for correctness, thoroughness, and clarity, respectively, surpassing the baseline model’s rates of 86.62%, 90.10%, and 97.46%; in human evaluations, it also maintained a leading position, with positive evaluation rates of 83.51%, 86.15%, and 95.66% for correctness, thoroughness, and clarity, respectively, outperforming the baseline model’s rates of 81.15%, 83.88%, and 94.6%.
Table 4. Performance of each model in terms of accuracy, thoroughness, and clarity scores.
This performance improvement can be directly attributed to the model’s introduction of a ternary gradient preference structure that incorporates noise-aware preference optimization: by simultaneously modeling the human expert’s preferred (high-quality explanations), non-preferred (low-quality explanations), and secondary-preferred (medium-quality explanations) components, this structure constructs a refined quality gradient space. It effectively enhances the model’s ability to learn the gradient of explanation quality, generating more accurate, comprehensive, and easily understandable vulnerability explanations, thereby providing more reliable technical support for smart contract security audits.
(4) RQ4: How does NAPO-SCVD perform in terms of the stability of explanation quality under noise interference caused by missing prompts?
To systematically explore how prompt-induced noise affects the quality and stability of vulnerability explanations generated by NAPO-SCVD, we adopt the standard auditing prompts for smart contract vulnerability detection as the baseline reference. Instead of generating noisy inputs in a random manner, we design a hierarchical prompt degradation scheme to simulate real-world incomplete prompts encountered in practical auditing workflows. During the construction of noisy samples, we selectively remove domain-specific keywords closely related to vulnerability identification, such as external call, state update operation, .call() function, block.timestamp field, delegatecall invocation, and integer overflow and underflow conditions. Combined with the inherent identification characteristics of four mainstream smart contract vulnerabilities, namely reentrancy (RE), timestamp dependency (TD), integer overflow/underflow (IO), and delegatecall (DE), we establish a noise-free control group and three gradient-based prompt missing scenarios with increasing noise intensity. In total, we develop 18 distinct prompt variants covering all three noise levels. This multi-gradient experimental setup enables a comprehensive and quantitative assessment of the model’s robustness against different degrees of prompt corruption.
We categorize the degraded prompts into three hierarchical noise levels with well-defined rules: In the mild noise setting, all core functional descriptions, key operational logic and vulnerability-specific keywords are fully preserved. We only remove redundant auxiliary descriptive text, while retaining the complete specifications for explanation generation, so as to simulate slightly simplified prompts in daily auditing. In the moderate noise setting, the general scope of target vulnerabilities remains unambiguous. We eliminate part of the core identification keywords for each vulnerability category and retain only one representative feature per type, while keeping the basic requirements for vulnerability localization and analytical explanation. In the severe noise setting, all core vulnerability features and professional technical terms are completely removed, the definition scope of vulnerability types becomes ambiguous, and the explanation requirements are simplified to the most concise expressions, which corresponds to the extreme cases of severely incomplete auditing prompts in practice.
In this evaluation, we select two representative methods, Smart-LLaMA-DPO and iAudit, as the baseline control models. To achieve a comprehensive and objective assessment of explanation quality, we employ a dual evaluation framework integrating large language model automatic scoring and professional human expert review. Three core evaluation dimensions are adopted: correctness, thoroughness and clarity of generated explanations (Figure 5). Meanwhile, we introduce a quality stability coefficient to quantitatively characterize the model’s robustness under varying noise intensities. The detailed design of all types of prompts is illustrated in Figure 6, and Figure 7 presents the intuitive comparison of explanation outputs from different models using a typical reentrancy vulnerability case.
Figure 5. Experiments on the stability of explanation quality for models in noisy environments employing a dual-track validation paradigm. Explanation quality is quantified across three dimensions—accuracy, thoroughness, and clarity—and the model’s robustness under varying noise levels is characterized using a quality stability coefficient. (a) Automatic scoring results of Smart-LLaMA-DPO; (b) Automatic scoring results of NAPO-SCVD; (c) Automatic scoring results of iAudit; (d) Manual expert evaluation results of NAPO-SCVD; (e) Manual expert evaluation results of Smart-LLaMA-DPO; (f) Manual expert evaluation results of iAudit.
Figure 6. The detailed design of standard and degraded prompts at different noise levels.
Figure 7. The output results of the running instances in Smart-LLaMA-DPO (LLaMA) and NAPO-SCVD.
The experimental results reveal a clear performance gap between NAPO-SCVD and the two control models under noisy input conditions. When using standard noise-free prompts, all approaches deliver satisfactory explanation results. However, as the intensity of prompt noise gradually rises, the performance disparity becomes increasingly prominent. NAPO-SCVD maintains high-quality, logically rigorous and reliable analytical outputs across mild, moderate and severe noise levels. In sharp contrast, the two baseline models show obvious performance degradation: they produce incomplete analytical content under mild noise, generate biased and misleading conclusions under moderate noise, and even output totally irrelevant descriptions that deviate from vulnerability analysis when facing severe prompt corruption.
We define the quality stability coefficient as the ratio of the average positive evaluation rate calculated across all three noise levels to the positive evaluation rate obtained under the noise-free condition. A coefficient value closer to 1 indicates stronger anti-interference capability and better robustness. According to the automatic evaluation results from large language models, the stability coefficients of NAPO-SCVD for correctness, thoroughness and clarity reach 0.92, 0.89 and 0.96 respectively. The corresponding metrics from human expert evaluation are 0.91, 0.88 and 0.93. All coefficient values of our model are consistently higher than those of the compared baselines.
Such prominent robustness advantage can be attributed to the core design of our proposed noise-aware preference optimization (NAPO) module. The constructed ternary gradient preference structure empowers the model to effectively mitigate the negative impacts caused by missing core features in prompts. Even if the input prompts only provide partial guidance for vulnerability analysis, our model can still prioritize the correctness and clarity of explanations, and suppress the bias introduced by incomplete information.
On the contrary, the two control models inherently lack such anti-noise designs. Smart-LLaMA-DPO merely adopts a binary preference paradigm, which is unable to distinguish between sub-preference biased samples and completely incorrect outputs. The iAudit framework does not incorporate dedicated preference alignment and noise resistance mechanisms. Consequently, both methods suffer from dramatic deterioration in explanation quality when confronted with prompt noise.

5. Conclusions

To address the issues of coarse-grained modeling in binary preference optimization and insufficient noise resistance due to missing prompts in traditional large language models for smart contract vulnerability detection, we propose the NAPO-SCVD model, which employs a four-stage core process comprising data construction, continuous pre-training, supervised fine-tuning, and noise-aware preference optimization. The data construction phase integrates large-scale contract code with real vulnerability samples to build a ternary preference dataset comprising preferred, non-preferred, and sub-preferred (noise-mimicking) samples; the continuous pre-training phase focuses on domain-specific language modeling to enhance the model’s understanding of contract syntax, semantics, and vulnerability features; supervised fine-tuning solidifies the model’s foundational performance by balancing a dual-task loss function that integrates vulnerability detection and explanation generation; and in the core noise-aware preference optimization phase, the model innovatively combines the general preference loss of DPO with the noise-aware loss of NAPO, using dynamic coefficient and weight adjustments to enable the model to accurately learn task priorities and effectively filter out noise interference. Experimental results demonstrate that NAPO-SCVD achieves better performance than baseline models in terms of detection accuracy and the stability of explanation quality across four core vulnerability categories. Its multi-stage collaborative architecture and ternary preference modeling provide a more reliable technical solution for smart contract security auditing. Future work may further expand vulnerability coverage and explore integration pathways with dynamic code analysis technologies.

Author Contributions

D.X. contributed to methodology, formal analysis, investigation, data curation, visualization, writing—original draft, and writing—review and editing. W.S. contributed to conceptualization. B.Z. contributed to supervision, project administration, resources, and writing—review and editing. Y.L. contributed to validation. All authors have read and agreed to the published version of the manuscript. R.G. contributed to formal analysis.

Funding

This research was funded by the National Natural Science Foundation of China under the project “Research on Radio Frequency Fingerprint Identification of Low-Power Wide-Area Network Based on Multi-Feature Fusion and Convolutional Neural Networks” (Grant No. 62306207) and the Scientific and Technological Innovation Programs of Higher Education Institutions in Shanxi: Research on Smart Contract Vulnerability Detection Method Based on Multimodal Decoupled Feature Enhancement (Grant No. 2025L148).

Data Availability Statement

Data are available from the corresponding author upon reasonable request.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Zhuang, Y.; Liu, Z.; Qian, P.; Liu, Q.; Wang, X.; He, Q. Smart contract vulnerability detection using graph neural networks. In Proceedings of the 29th International Joint Conference on Artificial Intelligence, Yokohama, Japan, 7–15 January 2021; pp. 3283–3290. [Google Scholar]
  2. Jiang, B.; Liu, Y.; Chan, W.K. Contractfuzzer: Fuzzing smart contracts for vulnerability detection. In Proceedings of the 33rd ACM/IEEE International Conference on Automated Software Engineering, Montpellier, France, 3–7 September 2018; Association for Computing Machinery: New York, NY, USA, 2018; pp. 259–269. [Google Scholar]
  3. Chen, C.; Su, J.; Chen, J.; Wang, Y.; Bi, T.; Yu, J.; Lin, X.; Chen, T.; Zheng, Z. When ChatGPT meets smart contract vulnerability detection. ACM Trans. Softw. Eng. Methodol. 2025, 34, 1–30. [Google Scholar] [CrossRef] [Scilit]
  4. He, D.; Wu, R.; Li, X.; Chan, S.; Guizani, M. Detection of vulnerabilities of blockchain smart contracts. IEEE Internet Things J. 2023, 10, 12178–12185. [Google Scholar] [CrossRef] [Scilit]
  5. Chen, Y.; Sun, Z.; Gong, Z.; Hao, D. Improving smart contract security with contrastive learning. In Proceedings of the 46th IEEE/ACM International Conference on Software Engineering, Lisbon, Portugal, 14–20 April 2024; IEEE: New York, NY, USA, 2024; pp. 1–11. [Google Scholar]
  6. Gandhi, S.T. AI-driven deep learning approach for smart contract vulnerability detection. Int. J. Adv. Res. Comput. Sci. Technol. 2025, 8, 11540–11547. [Google Scholar]
  7. Li, J.; Lu, G.; Gao, Y.; Gao, F. A smart contract vulnerability detection method based on multimodal feature fusion and deep learning. Mathematics 2023, 11, 4823. [Google Scholar] [CrossRef] [Scilit]
  8. Luo, F.; Luo, R.; Chen, T.; Qiao, A.; He, Z.; Song, S.; Jiang, Y.; Li, S. Scvhunter: Smart contract vulnerability detection based on heterogeneous graph attention network. In Proceedings of the 46th IEEE/ACM International Conference on Software Engineering, Lisbon, Portugal, 14–20 April 2024; Association for Computing Machinery: New York, NY, USA, 2024; pp. 1–13. [Google Scholar]
  9. Hwang, S.-J.; Choi, S.-H.; Shin, J.; Choi, Y.-H. CodeNet: Code-targeted convolutional neural network for smart contract vulnerability detection. IEEE Access 2022, 10, 32595–32607. [Google Scholar] [CrossRef] [Scilit]
  10. Kasula, V.K.; Yadulla, A.R.; Yenugula, M.; Konda, B.; Alshboul, A. Enhancing vulnerability detection using transformer-based embeddings and graph neural networks. In Proceedings of the 2024 34th International Conference on Computer Theory and Applications (ICCTA), Alexandria, Egypt, 14–16 December 2024; IEEE: New York, NY, USA, 2024; pp. 177–182. [Google Scholar]
  11. Chen, D.; Feng, L.; Fan, Y.; Shang, S.; Wei, Z. Smart contract vulnerability detection based on semantic graph and residual graph convolutional networks with edge attention. J. Syst. Softw. 2023, 202, 111705. [Google Scholar] [CrossRef] [Scilit]
  12. Wang, T.; Zhao, X.; Zhang, J.T. TMF-Net: Multimodal smart contract vulnerability detection based on multiscale transformer fusion. Inf. Fusion 2025, 122, 103189. [Google Scholar] [CrossRef] [Scilit]
  13. Chaliasos, S.; Charalambous, M.A.; Zhou, L.; Galanopoulou, R.; Gervais, A.; Mitropoulos, D.; Livshits, B. Smart contract and defi security tools: Do they meet the needs of practitioners? In Proceedings of the 46th IEEE/ACM International Conference on Software Engineering, Lisbon, Portugal, 14–20 April 2024; Association for Computing Machinery: New York, NY, USA, 2024; pp. 1–13. [Google Scholar]
  14. Hu, S.; Huang, T.; İlHan, F.; Tekin, S.F.; Liu, L. Large language model-powered smart contract vulnerability detection: New perspectives. In Proceedings of the 2023 5th IEEE International Conference on Trust, Privacy and Security in Intelligent Systems and Applications (TPS-ISA), Atlanta, GA, USA, 1–4 November 2023; IEEE: New York, NY, USA, 2023; pp. 297–306. [Google Scholar]
  15. Wei, Z.; Sun, J.; Zhang, Z.; Zhang, X.; Li, M. Leveraging fine-tuned language models for efficient and accurate smart contract auditing. arXiv 2024, arXiv:2410.13918. [Google Scholar]
  16. Ma, W.; Wu, D.; Sun, Y.; Wang, T.; Liu, S.; Zhang, J.; Xue, Y.; Liu, Y. Combining fine-tuning and LLM-based agents for intuitive smart contract auditing with justifications. arXiv 2024, arXiv:2403.16073. [Google Scholar]
  17. Yu, L.; Huang, Z.; Yuan, H.; Cheng, S.; Yang, L.; Zhang, F.; Shen, C.; Ma, J.; Zhang, J.; Lu, J.; et al. Smart-LLaMA-DPO: Reinforced large language model for explainable smart contract vulnerability detection. Proc. ACM Softw. Eng. 2025, 2, 182–205. [Google Scholar] [CrossRef] [Scilit]
  18. Sun, X.; Tu, L.; Zhang, J.; Cai, J.; Li, B.; Wang, Y. ASSBert: Active and semi-supervised BERT for smart contract vulnerability detection. J. Inf. Secur. Appl. 2023, 73, 103423. [Google Scholar] [CrossRef] [Scilit]
  19. Ashizawa, N.; Yanai, N.; Cruz, J.P.; Okamura, S. Eth2vec: Learning contract-wide code representations for vulnerability detection on Ethereum smart contracts. In Proceedings of the 3rd ACM International Symposium on Blockchain and Secure Critical Infrastructure, Virtual, 7 June 2021; Association for Computing Machinery: New York, NY, USA, 2021; pp. 47–59. [Google Scholar]
  20. Sharma, N.; Sharma, S. A Survey of Mythril, a Smart Contract Security Analysis Tool for EVM Bytecode. Indian J. Nat. Sci. 2022, 13, 51003–51010. [Google Scholar]
  21. Torres, C.F.; Schütte, J.; State, R. Osiris: Hunting for integer bugs in Ethereum smart contracts. In Proceedings of the 34th Annual Computer Security Applications Conference, San Juan, PR, USA, 3–7 December 2018; Association for Computing Machinery: New York, NY, USA, 2018; pp. 664–676. [Google Scholar]
  22. Luu, L.; Chu, D.-H.; Olickel, H.; Saxena, P.; Hobor, A. Making smart contracts smarter. In Proceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security, Hofburg Palace, VIE, Austria, 24–28 October 2016; Association for Computing Machinery: New York, NY, USA, 2016; pp. 254–269. [Google Scholar]
  23. Feist, J.; Grieco, G.; Groce, A. Slither: A static analysis framework for smart contracts. In Proceedings of the 2019 IEEE/ACM 2nd International Workshop on Emerging Trends in Software Engineering for Blockchain (WETSEB), Montreal, QC, Canada, 27 May 2019; IEEE: New York, NY, USA, 2019; pp. 8–15. [Google Scholar]
  24. Choi, J.; Kim, D.; Kim, S.; Grieco, G.; Groce, A.; Kil Cha, S. Smartian: Enhancing smart contract fuzzing with static and dynamic data-flow analyses. In Proceedings of the 2021 36th IEEE/ACM International Conference on Automated Software Engineering (ASE), Melbourne, VIC, Australia, 15–19 November 2021; IEEE: New York, NY, USA, 2021; pp. 227–239. [Google Scholar]
  25. Nguyen, T.D.; Pham, L.H.; Sun, J.; Lin, Y.; Minh, Q.T. sFuzz: An efficient adaptive fuzzer for Solidity smart contracts. In Proceedings of the ACM/IEEE 42nd International Conference on Software Engineering, Seoul, Republic of Korea, 27 June–19 July 2020; Association for Computing Machinery: New York, NY, USA, 2020; pp. 778–788. [Google Scholar]
  26. Tikhomirov, S.; Voskresenskaya, E.; Ivanitskiy, I.; Takhaviev, R.; Marchenko, E.; Alexandrov, Y. SmartCheck: Static analysis of Ethereum smart contracts. In Proceedings of the 1st International Workshop on Emerging Trends in Software Engineering for Blockchain, Gothenburg, Sweden, 27 May 2018; Association for Computing Machinery: New York, NY, USA, 2018; pp. 9–16. [Google Scholar]
  27. Kipf, T.N.; Welling, M. Semi-supervised classification with graph convolutional networks. arXiv 2016, arXiv:1609.02907. [Google Scholar]
  28. Zenggang, X.; Qiangqiang, L.; Gang, Z.; Xuemin, Z.; Hao, C.; Yuan, L.; Jing, L. A multimodal-based approach for smart contract vulnerability detection. J. Signal Process. Syst. 2026, 98, 17. [Google Scholar] [CrossRef] [Scilit]
  29. Liu, Z.; Qian, P.; Wang, X.; Zhu, L.; He, Q.; Ji, S. Smart contract vulnerability detection: From pure neural network to interpretable graph feature and expert pattern fusion. arXiv 2021, arXiv:2106.09282. [Google Scholar]
  30. Qian, P.; Liu, Z.; Yin, Y.; He, Q. Cross-modality mutual learning for enhancing smart contract vulnerability detection on bytecode. In Proceedings of the ACM Web Conference, Austin, TX, USA, 30 April–4 May 2023; Association for Computing Machinery: New York, NY, USA, 2023; pp. 2220–2229. [Google Scholar]
  31. Wang, Y.; Wang, W.; Joty, S.; Hoi, S.C. CodeT5: Identifier-aware unified pre-trained encoder-decoder models for code understanding and generation. arXiv 2021, arXiv:2109.00859. [Google Scholar]
  32. Wu, H.; Zhang, Z.; Wang, S.; Lei, Y.; Lin, B.; Qin, Y.; Zhang, H.; Mao, X. Peculiar: Smart contract vulnerability detection based on crucial data flow graph and pre-training techniques. In Proceedings of the 2021 IEEE 32nd International Symposium on Software Reliability Engineering (ISSRE), Wuhan, China, 25–28 October 2021; IEEE: New York, NY, USA, 2021; pp. 378–389. [Google Scholar]
  33. Yu, L.; Lu, J.; Liu, X.; Yang, L.; Zhang, F.; Ma, J. PSCVFinder: A prompt-tuning based framework for smart contract vulnerability detection. In Proceedings of the 2023 IEEE 34th International Symposium on Software Reliability Engineering (ISSRE), Florence, Italy, 9–12 October 2023; IEEE: New York, NY, USA, 2023; pp. 556–567. [Google Scholar]
  34. Dubey, A.; Jauhri, A.; Pandey, A.; Kadian, A.; Al-Dahle, A.; Letman, A.; Mathur, A.; Schelten, A.; Vaughan, A.; Yang, A.; et al. The Llama 3 herd of models. arXiv 2024, arXiv:2407.21783. [Google Scholar]
  35. Qwen Team. Qwen2 technical report. arXiv 2024, arXiv:2407.10671. [Google Scholar]
  36. Hirano, Y.; Hanaoka, S.; Nakao, T.; Miki, S.; Kikuchi, T.; Nakamura, Y.; Nomura, Y.; Yoshikawa, T.; Abe, O. GPT-4 Turbo with Vision fails to outperform text-only GPT. Jpn. J. Radiol. 2024, 42, 918–926. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  37. Jie, W.; Qiu, W.; Yang, H.; Guo, M.; Huang, X.; Lei, T.; Zhang, Q.; Zheng, H.; Zheng, Z. Agent4Vul: Multimodal LLM agents for smart contract vulnerability detection. Sci. China Inf. Sci. 2025, 68, 160101. [Google Scholar] [CrossRef] [Scilit]
  38. Zheng, Y.; Zhang, R.; Zhang, J.; YeYanhan, Y.; Luo, Z. LLaMAFactory: Unified efficient fine-tuning of 100+ language models. arXiv 2024, arXiv:2403.13372. [Google Scholar]
  39. Rasley, J.; Rajbhandari, S.; Ruwase, O.; He, Y. DeepSpeed: System optimizations enable training deep learning models with over 100 billion parameters. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, Virtual, 6–10 July 2020; Association for Computing Machinery: New York, NY, USA, 2020; pp. 3505–3506. [Google Scholar]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Article Metrics

Citations

Article Access Statistics

Multiple requests from the same IP address are counted as one view.