1. Introduction
Steelmaking is a dominant steel production method that utilizes hot metal and scrap steel as primary raw materials, characterized by high efficiency and large-scale production. The primary material inputs include hot metal, scrap steel, slag-forming agents (such as lime and magnesia balls), and oxygen. High-speed oxygen is blown into the converter through an oxygen lance, where high-concentration elements like carbon and silicon in the hot metal are oxidized, and harmful elements such as phosphorus and sulfur are removed, thereby producing molten steel with the desired temperature and composition, as well as generating slag and flue gas [
1]. Steel scrap is regarded as one of the key raw materials in steelmaking, and the utilization of scrap steel is of great significance for lowering carbon emissions. For a single type of steel scrap, its yield is defined as the proportion of steel converted into molten steel to its total weight, measuring the efficiency of scrap conversion into molten steel during the smelting process. The scrap steel yield rate is an important indicator of scrap resource utilization efficiency in the steelmaking process [
2,
3,
4,
5]. The average scrap steel yield rate for different types of scrap is a comprehensive metric, calculated as the weighted average of the individual yields, reflecting the overall scrap utilization efficiency. The identification of scrap steel yield rate plays a significant role in the recycling efficiency of scrap resources and the green and low-carbon transformation of the steel industry. With the advancement of China’s strategic research into the high-quality recycling of scrap resources under the “double carbon” background, improving scrap steel yield rates is of great significance for promoting the green and low-carbon development of the steel industry [
6].
However, there is limited research specifically focused on the identification of scrap steel yield rate. The yield rate of a single scrap type is defined as the mass ratio of molten steel converted from this scrap to its total feeding mass, and the average scrap yield of one whole heat is the weighted average value of all single scrap yields. Predicting scrap steel yield rate in steelmaking is challenging due to the wide variety of scrap types and numerous process factors. Firstly, the influencing factors are complex and diverse, including scrap charging patterns, hot metal temperature, endpoint temperature, oxygen flow rate, and oxygen blowing volume, as well as the contents of elements like carbon, sulfur, manganese, and phosphorus in hot metal. Identifying scrap steel yield rates through mechanism analysis is difficult due to the high dimensionality and non-linearity of these factors. Secondly, due to the continuity and high efficiency of the steelmaking process, it is impractical for online measurement for each heat. The complex furnace working conditions cannot be simulated in a laboratory, where it is difficult to fully replicate the high temperatures, high pressures, complex chemical reactions, and large-scale material flow in actual furnace working conditions. Additionally, the iron content in scrap cannot be equated to the true yield. The iron content only reflects the proportion of iron in the scrap raw materials, while yield involves the proportion of steel actually converted into steel products during the complex process. Scrap steel yield rates are mainly affected by iron oxidation, scrap melting losses, and splashing losses. Therefore, in the actual smelting process, the scrap steel yield rate cannot be equal to the iron content in the scrap steel.
Existing identification models for scrap steel yield rates are primarily categorized into mechanism-driven and data-driven approaches. Mechanism-driven methods primarily predict the scrap steel yield rate during steelmaking through static and dynamic mechanism models formed over a long period in the metallurgical field, and they possess a certain degree of interpretability. Existing studies mostly focus on mechanism analysis of the scrap melting process in converters. Wang [
7] numerically calculated the melting behavior of scrap steel in a full-scale converter under different blowing conditions using a three-dimensional coupling model for turbulent multiphase flow, melting, and solute transport. Singha [
8] developed a thermodynamic calculation model to describe the steelmaking process involving scrap and lime dissolution. Deng et al. [
9] established a mathematical model to study the melting mechanism of scrap in a dephosphorization converter, considering the coupling effects of the carbon content in scrap steel, the temperature of hot metal, and the temperature of preheating scrap on the scrap melting rate.
In terms of scrap steel yield rate identification methods, mechanism-driven research has mainly focused on laboratory measurement methods. The scrap steel yield rate is measured using an intermediate frequency furnace proposed by Ding Jian [
10]. However, this method relies on preprocessed samples with a certain composition. However, the composition of scrap steel is diverse in the actual steel-making process. Moreover, only the weight ratio of molten steel to raw scrap is calculated without considering the loss by smelting, such as chemical oxidation loss, smoke emission, and slag entrainment. The comparative analysis method combining heat balance and material balance was proposed by He Ruifei [
11]. But the establishment of balance equations relies on ideal metallurgical assumptions, making it difficult to quantify the impact of dynamic factors such as uneven molten pool temperature and fluctuating oxygen supply under industrial operating conditions; some key parameters in the equations need to be estimated through experience, and the accumulation of errors will significantly reduce the accuracy of yield rate identification. The adaptive algorithm for scrap steel yield rate (ASY-DCO) proposed by Gan Junjie [
12] realizes adaptive adjustment of the yield rate through dynamic cycle optimization based on the yield rate measured in the laboratory. But the initial reference value of the algorithm is derived from laboratory measurement results, and the deviation between laboratory data and industrial reality will be directly transmitted to the optimization process, leading to an inaccurate basis for adaptive adjustment, and the adjustment strategy only focuses on the cyclic correction of yield rate values without adapting to the core influencing factors of scrap steel yield rate, making it difficult to truly reflect the actual utilization efficiency of scrap steel in industrial scenarios.
Mechanism calculation methods utilize long-established mechanisms and expert knowledge in the metallurgical field and have excellent interpretability. However, since mechanism formulas contain a large number of unmeasurable key parameters, the identification of unmeasurable parameters is a key challenge affecting the accuracy of mechanism methods.
Data-driven scrap steel yield rate identification methods, centered on machine learning algorithms, have attracted certain attention in recent years. Zhang Chaojie [
13] proposed a method for predicting the deviation coefficient of scrap steel yield rate in steelmaking based on the Random Forest (RF) algorithm and the multi-layer neural network (MLNN), realizing the identification of scrap steel yield rate by leveraging the fitting and learning capabilities of the algorithms. Erik Sandberg et al. [
14] evaluated scrap steel properties using data from the Electric Arc Furnace (EAF) steelmaking process and constructed identification models for molten steel chemical composition, power consumption, and scrap steel yield rate using Partial Least Squares (PLS). However, these data-driven methods face significant limitations. Neural networks, often regarded as “black boxes”, lack interpretability regarding the relationship between input factors and identifications, limiting their acceptance in industrial operations. The generalization ability of some algorithms relies on a large amount of high-quality data, resulting in the inability to reproduce the theoretical hit rate of the model in actual on-site applications, thereby leading to low hit rates and poor scalability.
While existing studies have made some progress in scrap steel yield rate identification, critical limitations persist. The scrap steel yield rate in steelmaking is inherently affected by multi-dimensional and non-linear factors. The above methods fail to integrate metallurgical mechanisms with neural networks, resulting in models lacking interpretability and the capability for online identification.
Considering the severe heterogeneity of scrap feeding patterns across different heats, a single global model cannot fit diversified furnace conditions accurately. A knowledge graph can quantitatively describe the feeding similarity between heats through topologically weighted edges, providing an unsupervised clustering carrier for subsequent batch community division. Therefore, this paper constructs a scrap charging knowledge graph and adopts the LPA algorithm to divide heats with analogous feeding features into independent sub-communities for separate modeling. Thus, there is a critical need for the development of an efficient and online scrap steel yield rate identification method, and this method must be highly adaptable to the complex and changing furnace working conditions. In this paper, a mechanism and data joint-driven model for scrap steel yield rate identification is designed, and identification precision is further enhanced by community division of the knowledge graph. In
Section 2, a knowledge graph for scrap charging patterns and the label propagation algorithm (LPA) is used to identify communities for simulating the heats with similar material input characteristics. The mechanism formula for iron mass balance is incorporated into a multi-layer neural network as the loss function so that the difference between input and output iron mass can be minimized. SHAP is employed to evaluate and quantify the influence of various factors on scrap steel yield rates. In
Section 3, the experimental design and results are presented. Finally,
Section 4 is for the presentation of the experimental conclusions and the future outlook of this study.
4. Summary and Prospect
This paper developed a method for identifying scrap steel yield rates that integrates a mechanism and data joint-driven approach. By constructing the scrap charging pattern knowledge graph, using the LPA algorithm to find community, and establishing a physics-informed neural network model based on the iron balance mechanism, the identification of the hot metal yield rate and the scrap steel yield rate is achieved. The SHAP analysis method is utilized in this paper to enhance the interpretability of the model from the perspective of feature contribution. The proposed mechanism and data joint-driven identification model for scrap steel yield rates based on knowledge graph community detection achieves the highest identification precision, with an RMSE of 3.20 tons and an MAE of 2.62 tons. Compared with the baseline mechanism and data joint-driven identification model for scrap steel yield rates (without community detection), the proposed method reduces the MAE by 26.6% (from 3.57 t to 2.62 t) and significantly improves the hit rate within ±5 tons by 13.97 percentage points (reaching 87.55%). These improvements validate the effectiveness of the community-based divide-and-conquer strategy in handling complex charging patterns. It has important application value for promoting the green and low-carbon transformation of the iron and steel industry. Although the physics-informed neural network can achieve a high accuracy in identifying scrap steel yield rates, there are still some limitations and challenges: This model is trained based on ten types of scrap steel involved in the historical data of a certain iron and steel plant. Therefore, when applied to identify the scrap steel yield rate of new or previously unencountered scrap steel, transfer learning or retraining of the model may be required. The model does not quantitatively decouple error transmission from different types of measuring instruments, and the prediction fluctuation caused by heterogeneous sensor precision is not independently modeled, which will be optimized in follow-up research.
In future research, we will further optimize the model structure to enhance identification accuracy and stability, so as to better serve the actual steel production. For heats with abnormal conditions (e.g., over-blowing or insufficient oxygen blowing), targeted incremental learning will be conducted in the future after more abnormal heats are collected, given the limited amount of data currently available for such abnormal conditions. At the same time, more production link data will be considered for incorporation into the model to achieve comprehensive optimization and intelligent management and control of the steel production process.