Next Article in Journal
Coupled Transformation Processes of Cr-Adsorbed Schwertmannite and Chromium Redistribution Controlled by Ca(II) Speciation
Previous Article in Journal
Study on Drill String Vibration Characteristics and Structural Optimization During Wellbore Quality Design for Shale Gas and Oil Wells
Previous Article in Special Issue
Machine Learning-Based Prediction of Optimum Design Parameters for Axially Symmetric Cylindrical Reinforced Concrete Walls
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

A Deep Learning Approach to Predicting the Durability of Limestone Aggregates Using the FT-Transformer

1
Geological Engineering Department, Engineering Faculty, Istanbul University-Cerrahpasa, Istanbul 34098, Türkiye
2
Computer Engineering Department, Engineering Faculty, Istanbul University-Cerrahpasa, Istanbul 34098, Türkiye
*
Author to whom correspondence should be addressed.
Processes 2026, 14(8), 1257; https://doi.org/10.3390/pr14081257
Submission received: 12 February 2026 / Revised: 29 March 2026 / Accepted: 7 April 2026 / Published: 15 April 2026
(This article belongs to the Special Issue Machine Learning Models for Sustainable Composite Materials)

Abstract

The durability of limestone aggregates is a critical factor affecting the long-term performance of concrete, particularly under aggressive environmental conditions. However, conventional durability tests such as the magnesium sulfate soundness test are time-consuming and labor-intensive. In this study, a transformer-based deep learning model, namely the FT-Transformer, was employed to predict the magnesium sulfate soundness loss of limestone aggregates using standard aggregate index properties. The dataset used in the modeling stage consisted of 108 limestone aggregate samples, each characterized by particle density, water absorption, Los Angeles fragmentation, and magnesium sulfate soundness loss. Although four standardized laboratory tests were conducted for each sample, yielding 432 individual test results in total, the prediction dataset comprised 108 complete observations. The predictive performance of the FT-Transformer was evaluated and compared with Linear Regression, Polynomial Regression, and Support Vector Regression models. Under the single-split evaluation, the FT-Transformer achieved a test R 2 value of 0.6473 and a test MSE value of 0.0212. In addition, a repeated random-split statistical analysis demonstrated that the FT-Transformer achieved better average predictive performance than Linear Regression across 50 repeated train, validation and test partitions. These findings indicate that transformer-based tabular learning can provide an effective and practically applicable framework for aggregate durability prediction and may support preliminary material assessment and engineering decision-making processes.

1. Introduction

Aggregates, composed of a blend of fine and coarse particles such as gravel, sand, and crushed stone, play a fundamental role in various construction applications. High-quality aggregates are expected to possess sufficient strength, favorable engineering characteristics, and resistance to environmental effects [1]. The quality of aggregates directly affects both the cost efficiency and mechanical performance of concrete, with durability being one of the most critical parameters [2].
The durability of aggregates is primarily governed by their resistance to freeze–thaw cycles and physical degradation [3]. Freeze–thaw resistance becomes especially critical in cold climates, where repeated temperature fluctuations significantly influence concrete durability and mechanical performance [4]. Aggregates lacking adequate freeze–thaw resistance tend to deteriorate over time, resulting in cracking and structural damage [5]. High durability against freeze–thaw effects is therefore essential to prolong the service life of infrastructure elements such as roads, bridges, and buildings exposed to severe environmental conditions [6]. In contrast, the use of low-quality aggregates accelerates deterioration and increases maintenance and repair costs, while durable aggregates enhance long-term cost efficiency. Consequently, freeze–thaw testing constitutes a fundamental procedure to verify that aggregates meet durability requirements, particularly in regions with significant temperature variations.
The freezing and thawing behavior of aggregates can be evaluated through durability tests that employ chemical solutions, typically saturated sodium or magnesium sulfate solutions, to accelerate the disintegration process [7]. Sulfate durability tests generate crystallization and/or hydration pressures within aggregate pores, which may cause substantial deterioration in their mechanical properties [8]. Aggregate resistance to disintegration is commonly assessed according to ASTM C88 [9] or EN 1367-2 [10] standards. These protocols evaluate the ability of aggregates to withstand five cycles for concrete aggregates and ten cycles for ballast aggregates. Although such durability tests are relatively straightforward, they remain time-consuming and costly, motivating the search for alternative predictive methods to evaluate aggregate durability.
Recent studies in rock and construction materials have shown that durability- and degradation-related behavior is controlled by multiple interacting physical parameters and often exhibits highly nonlinear characteristics. Accordingly, advanced data-driven approaches have become increasingly important for capturing complex engineering responses and improving predictive performance in materials-related applications [11,12]. In parallel with these developments, recent research has also continued to emphasize the importance of multi-factor material characterization and performance-oriented evaluation in geomaterials and construction materials.
Previous studies focusing specifically on predicting the durability of natural aggregates are relatively limited in the literature [6,13]. In contrast, numerous studies have investigated the application of artificial intelligence techniques in areas related to concrete durability and strength. These works primarily focus on predictive modeling, optimization of material properties, and structural performance evaluation. Several researchers have studied freeze–thaw and weathering-related durability of concrete [4,14,15], including service life prediction under repeated cyclic loading [16,17,18]. Machine learning approaches have also been extensively applied to recycled aggregate concrete, covering compressive strength prediction [19,20,21], sustainable mix design [22,23], and aggregate composition analysis [24,25]. Recently, machine learning methodologies have attracted considerable attention due to their strong predictive capabilities and ability to process large datasets efficiently, making them suitable for complex engineering prediction tasks.
Machine learning algorithms such as Support Vector Machines (SVMs), Artificial Neural Networks (ANNs), Decision Trees (DTs), Random Forests (RFs), and Gradient Boosting (GB) methods have been widely used in durability prediction studies. More recently, deep learning models have also been employed for structured engineering datasets because of their ability to learn nonlinear relationships among multiple variables. Among these approaches, transformer-based models have emerged as promising alternatives for tabular prediction problems. In particular, the FT-Transformer was developed as a transformer-based architecture specifically adapted for tabular data, where input features are represented as tokens and their interactions are modeled through self-attention mechanisms [26]. This design enables the model to capture complex inter-feature dependencies effectively and makes it especially suitable for engineering datasets characterized by coupled and nonlinear material behavior.
In the present study, the FT-Transformer architecture is employed to predict the magnesium sulfate soundness loss of limestone aggregates using Los Angeles fragmentation, water absorption, and particle density as input variables. These parameters are widely recognized indicators of aggregate durability and are commonly used in aggregate characterization [27,28]. Since magnesium sulfate soundness loss is closely related to the internal structure, porosity, and degradation resistance of aggregates, the use of a transformer-based architecture offers a flexible and powerful framework for modeling the complex relationships among these variables.
The principal novelty of this study lies in its focus on natural limestone aggregate samples, which play a crucial role in determining concrete quality, and in the application of a transformer-based deep learning model to predict aggregate durability from standard laboratory index properties. Laboratory durability tests often face challenges arising from experimental variability, long testing durations, and manual measurement errors. Therefore, this research proposes a predictive modeling framework as an efficient and innovative alternative to conventional testing procedures. The developed models aim to reduce potential experimental uncertainty while providing faster, more practical, and more reliable estimations of aggregate durability.

2. Materials and Methods

2.1. Materials

This study forms part of an ongoing research program on the durability of limestone aggregates, particularly in view of the limited number of studies that address the prediction of aggregate durability properties using artificial intelligence-based approaches. Limestone aggregates are widely used in concrete production due to their favorable adhesion properties with cementitious materials [29]. For this reason, limestone aggregates obtained from quarries in Istanbul and different regions of Turkey were selected as the material group of this study.
To improve the representativeness and traceability of the dataset, the sampling program was designed to include geographically distributed quarry sites from both Istanbul and other major limestone-producing regions of Turkey. In Istanbul, samples were collected from quarries located in Çatalca–Muratbey, Silivri–Danamandıra, Çatalca–İhsaniye, Çatalca–Aydınlar, Sultangazi–Cebeci, Sarıyer–Gümüşdere, Çekmeköy–Ömerli, Çekmeköy–Hüseyinli, Şile–Üvezli, Şile–Ahmetli, Şile–Yaylalı, and Şile–Tekeköy. Additional samples were obtained from quarries in Kırklareli–Kofcaz, Kırklareli–Vize, İzmit–Kirazpınar, İzmit–Hereke, Sakarya–Ferizli, Bursa–Orhangazi, Bursa–Nilüfer, and İzmir–Karaburun. Since some of these areas contain more than one active quarry, the sampling framework was established to capture the geographical diversity of limestone aggregate sources used in construction practice.
All sampled materials were derived from limestone quarries. From a mineralogical point of view, limestone is composed predominantly of calcium carbonate (CaCO3), mainly in the form of calcite, although minor variations in composition may occur depending on the geological characteristics of the source formation. In this respect, while all samples belong to the same principal rock group, differences in local geological setting may lead to variations in texture, grain structure, and minor mineral content, which may in turn influence aggregate performance. Therefore, the inclusion of samples from different quarry regions was intended to reflect the natural variability in limestone aggregates within a consistent lithological framework.
The sampling criteria were defined to ensure both material consistency and experimental reliability. Accordingly, only representative aggregate samples obtained from unweathered or slightly weathered limestone zones were included in the study, whereas heavily weathered, highly fractured, visibly altered, or contaminated materials were excluded. Field sampling was carried out at different times over a three-year period, and each sample was recorded according to its source location before laboratory testing in order to improve the traceability of the experimental workflow.
Within the scope of the experimental program, the aggregate samples were subjected to particle density, water absorption, Los Angeles fragmentation, and magnesium sulfate soundness tests. A total of 108 limestone aggregate samples were used in the study, and each sample was characterized by these four standardized laboratory tests. Therefore, although 432 individual test results were obtained in total, these measurements correspond to 108 complete sample records rather than 432 independent aggregate samples. Accordingly, the prediction dataset used in the modeling stage consisted of 108 observations. The sampling procedures, laboratory test methods, and calculation stages were carried out in accordance with the relevant European standards, including EN 1367-2 [10], EN 1097-6 [30], and EN 1097-2 [31]. Thus, methodological consistency and compliance with internationally recognized testing procedures were ensured. The results obtained from the experimental studies are presented in Table 1.

2.2. Methods

In this study, magnesium sulfate soundness loss was considered as the target variable because it is a widely accepted indicator of aggregate durability against weathering-related deterioration. The predictor variables used in the modeling stage were Los Angeles fragmentation, water absorption, and particle density. These parameters were selected because they are widely recognized indicators of aggregate durability and are commonly used in aggregate characterization [27,28]. The Los Angeles fragmentation test measures the resistance of aggregates to abrasion and mechanical degradation, which is closely related to their structural integrity. Water absorption is associated with the porosity of the aggregates; in general, higher water absorption indicates a more porous internal structure, which may reduce durability and increase susceptibility to weathering-related deterioration. Particle density reflects the compactness and internal structure of the aggregate material and is often associated with its mechanical performance. Since magnesium sulfate soundness loss is influenced by the internal structure, porosity, and degradation resistance of aggregates, these three variables were considered to provide physically meaningful and complementary information for the prediction models.
In addition, these parameters are among the standard aggregate properties commonly used in engineering applications and were available for all samples in the dataset through standardized laboratory testing procedures. Therefore, their use allowed the modeling process to be established on a consistent and complete experimental basis. Although other physical, chemical, or petrographic parameters may also affect the durability behavior of limestone aggregates, such variables were not included in the present study because they were not uniformly available for the entire sample set. The inclusion of incomplete or non-standardized variables could reduce dataset consistency and adversely affect the comparability and reliability of the prediction results. Accordingly, the feature selection strategy adopted in this study was based on the combined consideration of physical relevance, data completeness, and standard test availability.
The dataset consisting of 108 observations was randomly divided into training, validation, and test sets prior to model development. Specifically, 80% of the data were used for training, 10% for validation, and the remaining 10% for testing. The training set was used for model fitting, the validation set was used to monitor model performance during training and to evaluate convergence behavior, and the test set was reserved for final performance assessment on previously unseen data.
In addition to the single-split evaluation, a repeated random-split analysis was performed to assess the robustness of the comparative model performance. For this purpose, the dataset was randomly partitioned 50 times into training, validation, and test subsets using the same 80%/10%/10% configuration, and the FT-Transformer and Linear Regression models were evaluated on identical splits. The resulting test metrics were then compared by paired statistical analyses.

2.2.1. FT-Transformer Architecture

Among transformer-based deep learning models developed for structured datasets, the FT-Transformer has emerged as a prominent architecture specifically adapted for tabular data analysis [26]. Unlike the original Transformer model, which was initially proposed for sequence modeling tasks in natural language processing, the FT-Transformer was designed to process numerical and categorical tabular features in a unified framework (Figure 1). In this architecture, each input feature is represented as a token and then passed through stacked self-attention layers, enabling the model to learn feature-wise representations and capture complex dependencies among variables more effectively. The mathematical expressions given in Equations (1)–(8) represent the standard formulation of feature embedding, self-attention, multi-head attention, and encoder transformation adapted from the Transformer and FT-Transformer frameworks reported in the literature, and they are presented here to describe the theoretical basis of the implemented model.
Unlike conventional fully connected neural networks, the FT-Transformer first tokenizes numerical features and then processes them through Transformer encoder layers (Figure 2), enabling improved representation learning among features. This tokenization strategy allows each input variable to be treated as an individual feature token, while the self-attention mechanism captures inter-feature dependencies in a flexible manner. Such a structure is especially advantageous for tabular engineering datasets, where the target variable is often controlled by multiple interacting material parameters rather than by isolated effects.
In the present study, this property is particularly relevant because magnesium sulfate soundness loss is influenced by aggregate characteristics such as Los Angeles fragmentation, water absorption, and particle density, which may exhibit coupled and nonlinear relationships. By learning these interactions through attention, the FT-Transformer provides a more expressive framework than conventional feed-forward architectures for modeling durability-related aggregate behavior. In addition, previous studies have shown that FT-Transformer and related transformer-based tabular models can achieve strong predictive performance and competitive generalization ability on structured datasets [26].
The implemented FT-Transformer model was tested in this study using a dataset consisting of 108 limestone aggregate samples. Los Angeles fragmentation, water absorption, and particle density were used as input variables, whereas magnesium sulfate soundness loss was defined as the target output. The model was trained and evaluated on separate training, validation, and test sets, and its predictive performance was compared with benchmark models including Linear Regression, Polynomial Regression, and Support Vector Regression.
Given an input matrix X containing n samples and d features as defined in Equation (1):
X = [ x 1 , x 2 , , x d ] , x i R n
each feature x i is projected into a high-dimensional embedding space as follows:
z i = W i x i + b i , W i R d e × 1 , b i R d e
where d e is the embedding dimension, and z i R d e denotes the embedded representation of the i-th feature. The complete embedded sequence is then expressed as
Z = [ z 1 , z 2 , , z d ] R d × d e
This embedded representation is passed through a multi-head self-attention mechanism, where query (Q), key (K), and value (V) matrices are computed as
Q = Z W Q , K = Z W K , V = Z W V
where W Q , W K , and W V are learnable projection matrices. The self-attention operation is formulated as
Attention ( Q , K , V ) = softmax Q K T d k V
where d k is the dimensionality of the key vectors. This mechanism enables the model to assign context-dependent importance weights to each feature interaction. In multi-head attention, this operation is performed in parallel over multiple subspaces:
MultiHead ( Z ) = Concat ( h e a d 1 , h e a d 2 , , h e a d h ) W O
where each attention head is defined as
h e a d j = Attention ( Q j , K j , V j )
and h denotes the number of heads.
The output of the Transformer encoder is passed through residual connections, layer normalization, and feed-forward layers to obtain a final latent representation:
H = LayerNorm ( H + MultiHead ( H ) )
H = LayerNorm ( H + FFN ( H ) )
where FFN ( · ) denotes the position-wise feed-forward network. Finally, the aggregated representation is used to predict the magnesium sulfate soundness loss as the target output:
y ^ = f θ ( X )
where f θ denotes the FT-Transformer model parameterized by θ , and y ^ represents the predicted soundness loss value.

2.2.2. Training Procedure and Optimization

The FT-Transformer model was trained using the Adam optimizer with a learning rate of 10 3 . The mean squared error (MSE) loss function was used during training. The model was trained for 100 epochs, and the validation set was monitored throughout the training process to track convergence behavior and generalization performance. The batch size was set to 8.

2.2.3. Benchmark Models and Evaluation Metrics

For comparative purposes, Linear Regression, second- and third-degree Polynomial Regression, and Support Vector Regression (SVR) models were also implemented. Model performance was evaluated using the coefficient of determination ( R 2 ) and mean squared error (MSE). In addition to the single-split results, repeated random-split comparisons between the FT-Transformer and Linear Regression were statistically assessed by paired t-tests and Wilcoxon signed-rank tests based on repeated test metrics.

3. Results

The dataset employed in this study consists of 108 observations describing quantitative physical and mechanical properties of limestone aggregates. The magnesium sulfate soundness loss (%) was considered as the dependent variable, while three independent variables were used for prediction: Los Angeles fragmentation coefficient (%), water absorption (%), and particle density (Mg/m3). These parameters exhibit different numerical ranges; magnesium sulfate values vary between 0.8% and 12.1%, whereas particle density values lie within a relatively narrow range of 2.58–2.91 Mg/m3. The dataset contains no missing values, allowing all observations to be fully utilized during model training and evaluation.
The FT-Transformer model was trained for 100 epochs, and the corresponding training and validation loss curves are presented in Figure 3. A rapid decrease in loss is observed during the initial epochs, followed by a gradual stabilization, indicating that the model effectively learned the underlying relationships in the dataset. The close agreement between the training and validation loss trends suggests stable generalization performance throughout the optimization process. The absence of a pronounced divergence between these curves indicates that the model maintained a balanced learning behavior during training. Minor fluctuations observed in later epochs are expected in stochastic optimization and do not change the overall convergence pattern.
Figure 4 shows the relationship between the actual and predicted magnesium sulfate soundness loss values for the test set. Most data points are concentrated around the 45 reference line, indicating a high level of agreement between measured and predicted values. This pattern confirms that the FT-Transformer successfully captures the dominant trend of the target variable. Although a limited number of samples exhibit larger deviations, the overall distribution demonstrates that the model provides a reliable approximation of the experimental results across the test data.
A sample-wise comparison of actual and predicted values is presented in Figure 5. The predicted values generally follow the measured values closely, indicating that the model preserves the variation pattern of the test samples. At the same time, several local mismatches can be observed, where the model either underestimates or overestimates the target response. These deviations may be associated with the intrinsic heterogeneity of limestone aggregates and the fact that aggregate durability is influenced by multiple interacting physical characteristics.
Comparative analyses were conducted using Linear Regression, second- and third-degree Polynomial Regression, and Support Vector Regression (SVR). Table 2 summarizes training and test R 2 values along with mean squared error (MSE) metrics for each model. Under the adopted single-split setting, Linear Regression achieved a test R 2 score of 0.6340, slightly lower than that obtained by the FT-Transformer (0.6473). The FT-Transformer also demonstrated stronger training performance ( R 2 = 0.7735 ) than Linear Regression ( R 2 = 0.6252 ). The second-degree Polynomial Regression showed limited improvement in generalization, while the third-degree Polynomial Regression suffered from strong overfitting, resulting in a test R 2 value of 0.2382. The SVR model also showed lower generalization capability, with a test R 2 value of 0.5368.
To examine whether the observed model performance was dependent on a specific data partition, an additional repeated random-split analysis was conducted. In this procedure, the dataset was randomly divided 50 times into training, validation, and test subsets under the same configuration, and the FT-Transformer and Linear Regression models were evaluated on identical partitions. The results showed that the FT-Transformer outperformed Linear Regression on average across repeated splits. Specifically, the mean test R 2 values were 0.4308 ± 0.4314 for the FT-Transformer and 0.3406 ± 0.4686 for Linear Regression, whereas the mean test MSE values were 0.0262 ± 0.0116 and 0.0306 ± 0.0143, respectively. Paired statistical comparisons further confirmed that these differences were significant (paired t-test: p = 0.0103 for test R 2 , p = 0.0070 for test MSE; Wilcoxon signed-rank test: p = 0.0224 for test R 2 , p = 0.0191 for test MSE). These findings indicate that the better predictive performance of the FT-Transformer was not limited to a single train–test partition and was supported across repeated data splits.
The comparative performance obtained in this study is also consistent with the general trend reported in previous studies on durability-related prediction problems in construction materials. Earlier research has shown that machine learning approaches can successfully model complex relationships between material properties and durability indicators in concrete-related applications [4,14,15]. Similar findings have been reported for freeze–thaw resistance and life prediction of concrete [16,17,18], as well as for strength and mix design of recycled aggregate concrete [19,20,21]. Further studies on recycled aggregates have confirmed the applicability of data-driven methods for material characterization and performance evaluation tasks [22,23,24]. In particular, studies based on conventional machine learning algorithms have demonstrated that standard material parameters can provide useful predictive information for durability estimation, although model performance may vary depending on the complexity of the dataset and the nature of the selected variables.
Within this context, the results of the present study indicate that the FT-Transformer provides competitive and robust predictive performance for magnesium sulfate soundness loss prediction in natural limestone aggregates. This finding is in line with the broader literature showing that advanced learning-based models are effective in capturing nonlinear interactions among material properties. At the same time, the benchmark results obtained from Linear Regression, Polynomial Regression, and Support Vector Regression also confirm that the selected input variables contain meaningful predictive information, although their ability to represent complex relationships is more limited than that of the FT-Transformer. Therefore, the present results not only support previous findings regarding the usefulness of data-driven methods in durability prediction, but also extend the literature by demonstrating the applicability of a transformer-based tabular model to natural aggregate durability assessment.
Figure 6 compares the cumulative values obtained from the experimental data and the prediction models. In the lower range of the cumulative response, all models follow the actual trend relatively closely, indicating that the general behavior of samples with lower magnesium sulfate soundness loss can be represented with acceptable accuracy. In the intermediate range, the FT-Transformer and Linear Regression models remain closer to the actual cumulative curve, whereas the polynomial models begin to show larger deviations. In the upper range, where the cumulative response increases more sharply, the differences between the actual values and the model predictions become more pronounced. This region is particularly important because it corresponds to samples with relatively higher soundness loss values, which are more critical from a durability assessment perspective. Among the evaluated models, the FT-Transformer follows the actual cumulative trend more closely in this critical upper range, whereas the higher-degree polynomial models exhibit the largest departures. These observations indicate that the FT-Transformer provides a more robust representation of the overall response pattern, especially in the range associated with more severe durability loss.

4. Discussion

The results demonstrate that the FT-Transformer model effectively captures relationships among aggregate physical properties and provides reliable durability predictions. Although Linear Regression also yielded competitive performance, the additional repeated random-split analysis showed that the FT-Transformer achieved better average test performance under the adopted experimental setting. Therefore, the present findings support the view that the FT-Transformer offers not only competitive but also statistically supported predictive advantages for this structured durability dataset.
Some prediction deviations observed in scatter and sample-wise comparisons may stem from the intrinsic heterogeneity of aggregate physical characteristics. Variations in mineral composition, microstructure, pore structure, and local weathering conditions may introduce additional variability that is not fully described by the selected input parameters. Even when samples belong to the same principal rock group, local differences in internal structure and material integrity may lead to variations in durability behavior that are difficult to capture completely using a limited number of predictor variables. Consequently, incorporating additional descriptors, such as petrographic characteristics or pore structure information, may further enhance predictive accuracy.
Another issue that should be considered is the behavior of the models for extreme values. Samples with unusually high or low magnesium sulfate soundness loss may be more difficult to predict accurately because such observations are generally less represented in the dataset and may reflect more complex material responses. In tabular prediction problems, models typically perform better in regions where the data are more densely distributed, whereas predictive accuracy may decrease near the lower and upper boundaries of the response range. Therefore, part of the observed prediction deviation may be associated with the limited representation of extreme cases in the available dataset.
The size of the dataset is another important factor when interpreting the results. Although the dataset used in this study was sufficient to develop and compare the proposed models, a larger and more diverse dataset would likely improve the robustness and generalizability of the findings. In particular, increasing the number of samples from different quarry sources and including a broader range of durability responses may help the models learn rare patterns more effectively and reduce uncertainty in prediction.
The repeated random-split analysis also provides an important perspective on the comparative behavior of the models. Although the numerical difference observed under a single split was limited, the repeated evaluation showed that the FT-Transformer achieved better average test metrics than Linear Regression across multiple data partitions, and these differences were statistically significant. This finding suggests that the predictive advantage of the FT-Transformer should not be interpreted as a coincidence related to one particular split, but rather as a more robust pattern observed across repeated train–validation–test configurations.
The comparative evaluation also highlights the risk of model overfitting when higher-degree polynomial models are employed. While training performance increases with model complexity, generalization capability decreases, as observed in the third-degree Polynomial Regression results. In contrast, the FT-Transformer maintains a favorable balance between model complexity and predictive stability, allowing effective learning without substantial overfitting. Nevertheless, the possibility of overfitting should still be acknowledged, especially when deep learning models are applied to structured datasets of moderate size. In the present study, this risk was reduced by comparing the FT-Transformer with benchmark machine learning models, by evaluating model performance on separate validation and test data, and by conducting repeated random-split analyses across multiple data partitions. However, despite these precautions, the possibility of partial overfitting cannot be completely ruled out and should be considered in future studies involving larger datasets and external validation scenarios.
Furthermore, visual cumulative analyses confirm that the FT-Transformer better approximates the overall distribution of durability values compared with conventional methods. This property is particularly important in engineering applications where accurate representation of both low- and high-risk aggregate behavior is necessary.
Overall, the findings indicate that transformer-based deep learning approaches offer promising advantages for durability prediction problems involving structured tabular data. In particular, the repeated random-split statistical analysis showed that the FT-Transformer achieved better average predictive performance than Linear Regression, and that this advantage was statistically significant across multiple data partitions. At the same time, the results should still be interpreted together with the natural heterogeneity of the material, the limited representation of extreme values, the moderate dataset size, and the potential influence of overfitting. The proposed approach is expected to deliver even stronger performance when applied to larger datasets or when additional features representing aggregate mineralogy, pore structure, or environmental exposure conditions are incorporated.

5. Conclusions

In this study, the magnesium sulfate soundness loss of limestone aggregates was predicted using Los Angeles fragmentation, water absorption, and particle density as input variables through the FT-Transformer and benchmark regression models. The results showed that the FT-Transformer provided competitive predictions and achieved good agreement with the experimental data, demonstrating its capability to model the complex relationships among aggregate properties in structured tabular datasets.
In addition to the single-split evaluation, a repeated random-split statistical analysis was carried out in order to examine the robustness of the comparative model performance. The results showed that the FT-Transformer achieved better average test performance than Linear Regression across 50 repeated train–validation–test partitions, with higher mean test R 2 and lower mean test MSE values. Paired statistical tests further confirmed that these differences were significant. These findings indicate that the predictive advantage of the FT-Transformer was not limited to a single favorable data split.
From an engineering perspective, the proposed model has the potential to support rapid and practical durability assessment in situations where conventional magnesium sulfate soundness testing is time-consuming, labor-intensive, or costly. In this respect, the model may assist engineers and material practitioners in preliminary aggregate evaluation, material screening, and decision-making processes for concrete production and infrastructure applications.
Nevertheless, the results should be interpreted considering the natural heterogeneity of limestone aggregates and the limitations of the available dataset. Future studies should focus on expanding the sample set, including a wider range of aggregate sources and durability levels, and incorporating additional variables such as petrographic, chemical, and pore structure characteristics. Moreover, the proposed framework should be validated on different rock types and under broader environmental conditions to further assess its generalizability and practical applicability.
Overall, the present study demonstrates that the FT-Transformer is a promising data-driven approach for predicting the durability performance of limestone aggregates and that transformer-based tabular learning can provide statistically supported advantages for aggregate durability modeling under the adopted experimental setting.

Author Contributions

Conceptualization, M.Y. and A.T.; methodology, M.E.I.; software, M.E.I.; validation, M.Y., M.E.I. and A.T.; formal analysis, M.E.I.; investigation, M.Y.; resources, M.Y.; data curation, N.V.; writing—original draft preparation, M.E.I.; writing—review and editing, M.Y. and A.T.; visualization, M.E.I.; supervision, A.T.; project administration, M.Y. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Data Availability Statement

The original contributions presented in this study are included in the article. Further inquiries can be directed to the corresponding author.

Acknowledgments

The authors would like to thank Istanbul University-Cerrahpasa for providing laboratory infrastructure and technical support during the experimental studies. The authors reviewed and edited the generated content and take full responsibility for the final version of the manuscript.

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

The following abbreviations are used in this manuscript:
FT-TransformerFeature Tokenization Transformer
LALos Angeles fragmentation coefficient
SVRSupport Vector Regression
MSEMean Squared Error
R2Coefficient of Determination
MSAMulti-Head Self-Attention
FFNFeed-Forward Network

References

  1. Marinković, S.; Radonjanin, V.; Malešev, M.; Ignjatović, I. Comparative environmental assessment of natural and recycled aggregate concrete. Waste Manag. 2010, 30, 2255–2264. [Google Scholar] [CrossRef]
  2. Shaban, W.M.; Su, J.Y.H.; Mo, K.H.; Li, L.; Xie, J. Quality improvement techniques for recycled concrete aggregate: A review. J. Adv. Concr. Technol. 2019, 17, 151–167. [Google Scholar] [CrossRef]
  3. Williams, S.G.; Cunningham, J.B. Evaluation of Aggregate Durability Performance Test Procedures; Final Report TRC0905; Arkansas State Highway and Transportation Department: Little Rock, AR, USA, 2012.
  4. Huang, D.; Feng, Y.; Xia, Q.; Tian, J.; Li, X. Research on mechanical properties and durability of early frozen concrete: A review. Constr. Build. Mater. 2024, 425, 135988. [Google Scholar] [CrossRef]
  5. Wang, R.; Hu, Z.; Li, Y.; Wang, K.; Zhang, H. Review on the deterioration and approaches to enhance the durability of concrete in the freeze–thaw environment. Constr. Build. Mater. 2022, 321, 126371. [Google Scholar] [CrossRef]
  6. Kahraman, E.; Ozdemir, A.C. The prediction of durability to freeze–thaw of limestone aggregates using machine-learning techniques. Constr. Build. Mater. 2022, 324, 126678. [Google Scholar] [CrossRef]
  7. McNally, G.H. Soil and Rock Construction Materials; E & FN Spon: London, UK, 1998. [Google Scholar]
  8. Thomas, M.D.A.; Folliard, K.J. Concrete aggregates and the durability of concrete. In Durability of Concrete and Cement Composites; Page, C.L., Page, M.M., Eds.; Woodhead Publishing and CRC Press: Boca Raton, FL, USA, 2007; pp. 247–281. [Google Scholar]
  9. ASTM C88; Standard Test Method for Soundness of Aggregates by Use of Sodium Sulfate or Magnesium Sulfate. ASTM International: West Conshohocken, PA, USA, 2007.
  10. EN 1367-2; Tests for Thermal and Weathering Properties of Aggregates—Part 2: Magnesium Sulfate Test. European Committee for Standardization: Brussels, Belgium, 2011.
  11. Armaghani, D.J.; Asteris, P.G.; Cavaleri, L. A comparative study of various machine learning methods for rock failure prediction in presence of intact rock properties. Eng. Comput. 2023, 39, 2465–2485. [Google Scholar] [CrossRef]
  12. Asteris, P.G.; Skentou, A.D.; Bardhan, A.; Samui, P.; Pilakoutas, K. Predicting concrete compressive strength using hybrid ensembling of surrogate machine learning models. Cem. Concr. Res. 2022, 145, 106449. [Google Scholar] [CrossRef]
  13. Alavi Nezhad Khalil Abad, S.V.; Yilmaz, M.; Jahed Armaghani, D.; Tugrul, A. Prediction of the durability of limestone aggregates using computational techniques. Neural Comput. Appl. 2018, 29, 423–433. [Google Scholar] [CrossRef]
  14. Medina, C.; De Rojas, M.I.S.; Frías, M. Freeze–thaw durability of recycled concrete containing ceramic aggregate. J. Clean. Prod. 2013, 40, 151–160. [Google Scholar] [CrossRef]
  15. Yu, H.; Ma, H.; Yan, K. An equation for determining freeze–thaw fatigue damage in concrete and a model for predicting the service life. Constr. Build. Mater. 2017, 137, 104–116. [Google Scholar] [CrossRef]
  16. Gong, L.; Bu, Y.; Xu, T.; Zhao, X.; Yu, X.; Liang, Y. Research on freeze resistance and life prediction of desert sand–crushed stone fine aggregate concrete. Case Stud. Constr. Mater. 2024, 21, e03896. [Google Scholar] [CrossRef]
  17. Li, Y.; Liu, Y.; Guo, H.; Li, Y. Investigation of the freeze–thaw deterioration behavior of hydraulic concrete under various curing temperatures. J. Build. Eng. 2024, 95, 110247. [Google Scholar] [CrossRef]
  18. Luo, D.; Qiao, X.; Niu, D. A predictive model for freeze–thaw concrete durability index utilizing DeepLabv3+ with machine learning. Constr. Build. Mater. 2025, 459, 139788. [Google Scholar] [CrossRef]
  19. Mohamad Ali Ridho, B.K.A.; Ngamkhanong, C.; Wu, Y.; Kaewunruen, S. Recycled aggregates concrete compressive strength prediction using artificial neural networks. Infrastructures 2021, 6, 17. [Google Scholar] [CrossRef]
  20. Pan, X.; Xiao, Y.; Suhail, S.A.; Ahmad, W.; Murali, G.; Salmi, A.; Mohamed, A. Use of artificial intelligence methods for predicting the strength of recycled aggregate concrete. Materials 2022, 15, 4194. [Google Scholar] [CrossRef]
  21. Yuan, X.; Tian, Y.; Ahmad, W.; Ahmad, A.; Usanova, K.I.; Mohamed, A.M.; Khallaf, R. Machine learning prediction models to evaluate the strength of recycled aggregate concrete. Materials 2022, 15, 2823. [Google Scholar] [CrossRef]
  22. Mohammadi Golafshani, E.; Kim, T.; Behnood, A.; Ngo, T.; Kashani, A. Sustainable mix design of recycled aggregate concrete using artificial intelligence. J. Clean. Prod. 2024, 442, 140994. [Google Scholar] [CrossRef]
  23. Hoong, J.D.L.H.; Lux, J.; Mahieux, P.Y.; Turcry, P.; Ait-Mokhtar, A. Determination of recycled aggregate composition using deep learning-based image analysis. Autom. Constr. 2020, 116, 103204. [Google Scholar] [CrossRef]
  24. Owais, M.; Idriss, L.K. Modeling green recycled aggregate concrete using machine learning. Constr. Build. Mater. 2024, 440, 137393. [Google Scholar] [CrossRef]
  25. Ulucan, M.; Yıldırım, G.; Alatas, B.; Alyamaç, K.E. Comparing machine learning regression models for recycled aggregate concrete strength prediction. Fırat Üniversitesi Mühendislik Bilimleri Dergisi 2024, 36, 563–580. [Google Scholar] [CrossRef]
  26. Gorishniy, Y.; Rubachev, I.; Khrulkov, V.; Babenko, A. Revisiting deep learning models for tabular data. Adv. Neural Inf. Process. Syst. 2021, 34, 18932–18943. [Google Scholar]
  27. Köken, E. Assessment of Los Angeles abrasion value (LAAV) and magnesium sulphate soundness (Mwl) of rock aggregates using gene expression programming and artificial neural networks. Arch. Min. Sci. 2022, 67, 401–422. [Google Scholar] [CrossRef]
  28. Strzałkowski, P.; Köken, E.; Sousa, L. Guidelines for natural stone products in connection with European standards. Materials 2023, 16, 6885. [Google Scholar] [CrossRef]
  29. Carlos, A.; Masumi, I.; Hiroaki, M.; Maki, M.; Takahisa, O. The effects of limestone aggregate on concrete properties. Constr. Build. Mater. 2010, 24, 2363–2368. [Google Scholar] [CrossRef]
  30. EN 1097-6; Determination of Particle Density and Water Absorption of Aggregates. European Committee for Standardization: Brussels, Belgium, 2022.
  31. EN 1097-2; Determination of Resistance to Fragmentation. European Committee for Standardization: Brussels, Belgium, 2020.
Figure 1. FT-Transformer architecture used for aggregate durability prediction.
Figure 1. FT-Transformer architecture used for aggregate durability prediction.
Processes 14 01257 g001
Figure 2. Transformer architecture.
Figure 2. Transformer architecture.
Processes 14 01257 g002
Figure 3. Training and validation loss curves of the FT-Transformer model.
Figure 3. Training and validation loss curves of the FT-Transformer model.
Processes 14 01257 g003
Figure 4. Actual versus predicted values for the FT-Transformer model.
Figure 4. Actual versus predicted values for the FT-Transformer model.
Processes 14 01257 g004
Figure 5. Sample-wise comparison of actual and predicted values.
Figure 5. Sample-wise comparison of actual and predicted values.
Processes 14 01257 g005
Figure 6. Cumulative comparison of actual values and model predictions, with the low-, intermediate-, and high-response ranges indicated for clearer interpretation.
Figure 6. Cumulative comparison of actual values and model predictions, with the low-, intermediate-, and high-response ranges indicated for clearer interpretation.
Processes 14 01257 g006
Table 1. Statistical summary of aggregate test results.
Table 1. Statistical summary of aggregate test results.
Aggregate TestNMinMaxMeanStd. Dev.
Particle density (Mg/m3), EN 1097-6 [30]1082.582.912.700.04
Water absorption (%), EN 1097-6 [30]1080.061.400.460.35
Los Angeles coefficient (500 cycles) (%), EN 1097-2 [31]108103018.883.48
Aggregate durability, magnesium sulfate value (%), EN 1367-2 [10]1080.812.14.182.76
Table 2. Evaluation metrics for each model.
Table 2. Evaluation metrics for each model.
ModelTrain R 2 Train MSETest R 2 Test MSE
FT-Transformer0.77350.01340.64730.0212
Linear Regression0.62520.02210.63400.0215
Polynomial (2nd deg.)0.70760.01730.58040.0247
Polynomial (3rd deg.)0.75100.01470.23820.0448
SVR0.65510.02030.53680.0272
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Yılmaz, M.; Isenkul, M.E.; Vural, N.; Tugrul, A. A Deep Learning Approach to Predicting the Durability of Limestone Aggregates Using the FT-Transformer. Processes 2026, 14, 1257. https://doi.org/10.3390/pr14081257

AMA Style

Yılmaz M, Isenkul ME, Vural N, Tugrul A. A Deep Learning Approach to Predicting the Durability of Limestone Aggregates Using the FT-Transformer. Processes. 2026; 14(8):1257. https://doi.org/10.3390/pr14081257

Chicago/Turabian Style

Yılmaz, Murat, M. Erdem Isenkul, Nil Vural, and Atiye Tugrul. 2026. "A Deep Learning Approach to Predicting the Durability of Limestone Aggregates Using the FT-Transformer" Processes 14, no. 8: 1257. https://doi.org/10.3390/pr14081257

APA Style

Yılmaz, M., Isenkul, M. E., Vural, N., & Tugrul, A. (2026). A Deep Learning Approach to Predicting the Durability of Limestone Aggregates Using the FT-Transformer. Processes, 14(8), 1257. https://doi.org/10.3390/pr14081257

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop