Next Article in Journal
Coptis chinensis Extract-Loaded Mouthwash: Antimicrobial Efficacy, Biocompatibility, and Clinical Benefits for Periodontal Health
Previous Article in Journal
Software Fault Localization Approach with Coverage Matrix Optimization Boosted by LLM-Based Code Naturalness
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

A Data-Driven and Explainable AI Framework for Quantitative Analysis of Research Trends in Timber Seismic Engineering

1
Department of Building Materials and Components, Building Research Institute, Tachihara-1, Tsukuba 3050802, Japan
2
Department of Architectural Research, National Institute for Land and Infrastructure Management, Tachihara-1, Tsukuba 3050802, Japan
*
Author to whom correspondence should be addressed.
Appl. Sci. 2026, 16(9), 4418; https://doi.org/10.3390/app16094418
Submission received: 12 April 2026 / Revised: 29 April 2026 / Accepted: 29 April 2026 / Published: 30 April 2026
(This article belongs to the Section Civil Engineering)

Abstract

This study presents a data-driven and explainable artificial intelligence (XAI) framework for quantitatively analyzing research trends in the seismic performance of timber structures. Unlike conventional bibliometric approaches based on descriptive statistics, the framework integrates large-scale literature mining, natural language processing, topic modeling, network analysis, and SHAP-based machine learning to enable structural and temporal interpretation. A dataset of 248 journal articles from OpenAlex was processed through a unified pipeline, including domain-specific filtering, text preprocessing, and temporal balancing. Topic modeling identified eight research themes spanning traditional component-level mechanics and emerging areas such as cross-laminated timber (CLT), hybrid systems, and performance-based design. Network analysis revealed a highly interconnected structure centered on key concepts such as shear walls, connections, stiffness, and cyclic behavior. SHAP-based analysis further showed that research evolution follows a layered and cumulative pattern rather than simple topic replacement: classical themes remain foundational, while newer concepts such as CLT and structural capacity have become increasingly influential. The proposed framework provides a reproducible and scalable method for quantitatively mapping research structures and temporal dynamics in timber seismic engineering.

1. Introduction

Timber structures have been widely used in residential buildings, particularly in seismic regions such as Japan, due to their sustainability, lightweight characteristics, and construction efficiency [1]. However, ensuring the seismic safety of wooden houses remains a critical challenge, as their structural response is strongly influenced by the material properties, connection behavior, and structural configuration. Over the past decades, extensive research has been conducted to investigate the seismic performance of timber structures through experimental testing, numerical modeling, and analytical approaches.
In parallel with these developments, the broader field of structural engineering has increasingly emphasized sustainability and low-carbon material design. Recent studies on advanced construction materials, such as engineered cementitious composites incorporating industrial by-products (e.g., lithium slag), have demonstrated that it is possible to simultaneously enhance mechanical performance and reduce environmental impact [2]. These findings reflect a growing trend toward integrating structural performance with environmental sustainability. Although such studies are not directly focused on timber structures, they provide an important contextual perspective. In particular, the increasing attention to hybrid systems, composite materials, and performance-oriented design suggests that future timber seismic research may also evolve toward incorporating sustainability-driven material innovations alongside traditional structural considerations.
Early studies on timber seismic design emphasized the importance of performance-based methodologies. Filiatrault and Folz [3] proposed a displacement-based design framework for wood-frame buildings, highlighting the limitations of conventional force-based approaches.
Subsequent research expanded these concepts into more comprehensive performance-based design (PBD) frameworks for timber structures, incorporating limit states associated with damage, serviceability, and collapse prevention. These studies systematically investigated the relationships between structural components, deformation capacities, and seismic demand, leading to improved design methodologies for timber wall systems and connections [4]. In particular, advancements in nonlinear analysis and component-based modeling have enabled a more accurate representation of the seismic behavior of timber structures, bridging the gap between simplified design approaches and actual structural performance. More recently, a recent review [5] identified these developments and identified emerging challenges in timber seismic design, particularly in the context of engineered wood products and evolving design standards.
A significant portion of the literature has focused on the behavior of lateral-load-resisting components, including shear walls, panels, and connections. Previous studies [6,7,8] have shown that shear walls play a dominant role in controlling stiffness, strength, and energy dissipation capacity in timber buildings. In the context of cross-laminated timber (CLT), it has been demonstrated that the global seismic response is largely governed by connection behavior rather than panel deformation. In addition, system-level investigations, including multi-story building analyses, have contributed to a better understanding of overall seismic performance and structural interactions.
Recent research, such as X. Estrella et al.’s study [9], has also explored seismic performance factors and their implications for design codes. Their study, combining experimental data and nonlinear analysis, proposed refined response modification factors for timber systems, thereby improving the reliability of practical design approaches. Furthermore, increasing attention has been given to resilient design strategies, such as self-centering and energy-dissipating connections, which aim to reduce residual deformation and enhance post-earthquake functionality [8]. Retrofit techniques for existing timber structures have also been developed to improve seismic resistance, particularly in regions with aging building stock [10].
In parallel, research on traditional and heritage timber structures has highlighted additional complexities associated with material degradation, construction techniques, and preservation constraints [11,12]. Moreover, probabilistic approaches, including fragility analysis and damage prediction, have been increasingly applied to assess seismic risk in timber housing systems [13]. These studies demonstrate that timber seismic research encompasses a wide range of topics, from component-level behavior to system-level performance and risk assessment.
In recent years, the application of artificial intelligence and data-driven approaches in timber engineering has also attracted increasing attention. A comprehensive review by Namba [14] summarized the current state of artificial intelligence applications in wooden structures, highlighting their potential in structural assessment, damage detection, and performance prediction [15,16,17]. In particular, SHAP (Shapley Additive Explanations), which is in one explainable AI (XAI), provides a theoretically grounded framework for quantifying the contribution of individual variables to model outputs, enabling a more transparent and quantitative interpretation of complex data-driven analyses [18,19,20,21]. Although SHAP has been successfully applied in structural engineering to identify influential parameters and interpret predictive models, its application to the quantitative analysis of research trends remains largely unexplored. This trend indicates a gradual transition toward integrating traditional structural engineering methodologies with advanced computational techniques.
Despite the growing body of literature, existing studies remain fragmented across different research domains, making it difficult to obtain a comprehensive understanding of overall research trends and knowledge structures. Conventional literature reviews are often qualitative and limited in scope, which restricts their ability to systematically analyze large volumes of publications. This limitation is also reflected in the diversity of design frameworks and analytical approaches adopted in timber engineering, which are often governed by region-specific standards such as the American Wood Council (NDS) and Eurocode 5, as well as classical analytical models such as the European Yield Model proposed by Johansen [22,23,24]. Furthermore, due to the inherent material variability of wood, many studies rely on simplified or empirical formulations rather than fully quantitative frameworks, as discussed in foundational works by Bodig and Jayne [25].
To address these challenges, bibliometric analysis and text mining techniques have emerged as powerful tools for extracting knowledge from large collections of scientific publications [26,27,28]. These approaches enable the systematic and data-driven exploration of research trends, topic structures, and temporal evolution. However, their application to timber seismic research remains limited, and few studies have attempted to integrate advanced machine learning techniques with bibliometric analysis to interpret the evolution of the field.
Therefore, this study aims to provide a systematic and quantitative characterization of research trends in the seismic performance of timber structures by integrating large-scale literature mining, natural language processing, topic modeling, network analysis, and explainable artificial intelligence. In addition, SHAP [29] was introduced as a quantitative analysis method to evaluate the contribution of variables. While previous studies have applied bibliometric and science-mapping approaches to timber-related research, they have often focused on broader domains or relied on a limited set of analytical techniques. In contrast, the present study specifically targets timber seismic engineering and combines multiple data-driven methods within a unified and reproducible framework.
Understanding the evolution of research trends in timber seismic engineering requires a systematic and data-driven analytical framework capable of processing large-scale bibliographic data. In this study, a unified framework integrating literature mining, natural language processing, topic modeling, network analysis, and SHAP-based explainable artificial intelligence is developed. This framework enables the identification of dominant research themes, as well as the quantitative analysis of their interrelationships and temporal evolution. Furthermore, SHAP (Shapley Additive Explanations) is incorporated to quantify the contribution of individual keywords, providing an interpretable understanding of research trends. The main contributions of this study are summarized as follows: (1) A systematic and quantitative framework is developed to analyze large-scale literature on timber seismic research; (2) Major research themes and their interrelationships are quantitatively identified using topic modeling and network analysis; (3) The temporal evolution of the field is evaluated using data-driven metrics; (4) an XAI, “SHAP”, is applied to interpret the contribution of keywords to research trends.
Overall, this study provides a reproducible and data-driven framework for understanding the evolving research landscape of timber seismic performance, offering complementary insights to existing bibliometric and qualitative review studies.

2. Methodology

2.1. Overall Workflow

This study proposes a comprehensive data-driven framework to analyze research trends in timber seismic engineering by integrating large-scale literature mining, natural language processing (NLP), topic modeling, network analysis, and explainable machine learning techniques. The overall workflow of the proposed methodology is illustrated in Figure 1.
This study demonstrates that the evolution of timber seismic research is not characterized by simple topic replacement, but by a layered and cumulative expansion of research themes in which classical mechanical studies continue to underpin emerging system-level and resilience-oriented approaches.
First, bibliographic data were collected from the OpenAlex [30] database using multiple domain-specific search queries. Second, a rule-based filtering process was applied to ensure the relevance of publications to timber structures and seismic performance. Third, textual data were preprocessed and transformed into structured representations suitable for analysis. To mitigate temporal bias, the dataset was balanced by limiting the number of publications per year. Subsequently, topic modeling using Latent Dirichlet Allocation (LDA) was conducted to extract latent research themes. A keyword co-occurrence network was constructed to identify relationships among key concepts. Furthermore, a machine learning model based on XGBoost (Extreme Gradient Boosting), an ensemble learning method based on boosting that sequentially improves model performance [31], was developed to predict publication year from textual features. Finally, SHAP [29] was applied to interpret the contribution of individual keywords to temporal trends. In addition to the above framework, it is important to clarify the relationship between the proposed methodology and conventional bibliometric analysis tools. Widely used software such as VOSviewer [32] has been extensively applied in civil engineering research to visualize bibliometric networks, including co-authorship, co-citation, and keyword co-occurrence relationships. These tools are highly effective for exploring the structural organization of research fields and identifying major clusters through intuitive graphical representations.
However, such approaches primarily focus on descriptive visualization and do not inherently provide quantitative measures for interpreting the contribution of individual factors to research trends. In contrast, the methodology proposed in this study integrates topic modeling and explainable machine learning techniques, enabling not only the identification of latent topic structures, but also the quantitative interpretation of their temporal evolution. In particular, the application of SHAP to the XGBoost model allows for the explicit evaluation of the importance of individual keywords in shaping research trends. Therefore, while conventional bibliometric tools serve as effective exploratory instruments, the proposed framework offers complementary capabilities by providing a more analytical and interpretable understanding of the underlying mechanisms driving the evolution of timber seismic research. The detailed procedures and implementation of each analytical step are described in the following sections.

2.2. Data Collection from OpenAlex

A large-scale dataset of scientific publications was constructed using the OpenAlex API [30], which provides structured and openly accessible bibliographic data in Python [33]. To comprehensively cover the domain of timber seismic research, multiple search queries were designed, including terms related to timber structures, earthquake response, and cyclic loading behavior. To ensure computational reproducibility, the logical structure of the search queries was explicitly defined using Boolean operators. The core query was formulated as follows:
  • (“timber” OR “wood” OR “wooden structure” OR “cross-laminated timber” OR “CLT” OR “glulam”) AND (“seismic” OR “earthquake” OR “seismic response” OR “cyclic loading” OR “hysteresis”)
These queries were implemented through the OpenAlex API endpoint [30] with filter parameters applied to restrict publication year (1980–2026), document type (journal articles), and language (English). An example API request is given as:
Pagination was handled iteratively to retrieve the complete set of records, and all results were aggregated into a unified dataset. For each publication, metadata including title, abstract, keywords, authors, affiliations, publication year, and citation information were retrieved. The abstract text was reconstructed from the inverted index provided by OpenAlex. Duplicate records were removed based on DOI and title matching to ensure data consistency.
To ensure domain-specific relevance, a rule-based filtering approach was implemented. Publications were retained only if they contained at least one timber-related term (e.g., timber, wood, cross-laminated timber) and at least one seismic-related term (e.g., earthquake, cyclic loading, hysteresis) within the title or abstract. In addition, a set of optional structural engineering terms (e.g., shear wall, connection, drift, stiffness) was used for diagnostic purposes to evaluate structural relevance.
Furthermore, to improve the quality of the vocabulary used in subsequent text mining and topic modeling, additional domain-specific filtering was introduced. Publications containing noise-related terms associated with unrelated disciplines (e.g., acoustics, paleontology, forensic science, business, biomedical fields, geology, and geochemistry) were excluded. In parallel, weakly informative and overly generic academic terms (e.g., engineering, science, study, analysis, method, model, and data) were removed from the vocabulary during preprocessing to enhance the interpretability of keyword frequency and topic distributions. This combined filtering strategy ensured that the final dataset consisted exclusively of publications relevant to timber structural systems subjected to seismic actions while also improving the clarity and robustness of the extracted textual features.
To ensure full transparency and reproducibility of the data collection process, the number of records at each stage was explicitly tracked. The initial dataset was obtained from multiple search queries, and the retrieved records were aggregated into a unified dataset. Duplicate records were removed based on DOI and title matching. Subsequently, rule-based filtering was applied to exclude irrelevant publications. The number of records removed at each filtering stage, including noise exclusion and domain-specific filtering, was recorded. Finally, a temporal balancing procedure was applied to obtain the final analytical dataset. A detailed summary of the number of records at each stage is provided in Table 1. All query scripts and preprocessing procedures are openly accessible via the link provided in the appendix, ensuring full reproducibility.

2.3. Text Preprocessing

The textual data used in this study were constructed by concatenating the title, abstract, and keywords of each publication. The combined text was subjected to several preprocessing steps to standardize and clean the data.
First, all text was converted to lowercase. URLs and non-alphanumeric characters were removed to eliminate noise. The text was then tokenized using regular expression-based parsing. Stopwords were removed using a combination of standard English stopwords (from NLTK and Scikit-learn), domain-generic terms (e.g., study, method, analysis), and core query terms (e.g., timber, seismic) to avoid trivial dominance of frequently occurring but non-informative words. Only tokens with a minimum length of three characters were retained. Additionally, documents containing fewer than five valid tokens were excluded to ensure sufficient textual information for analysis.

2.4. Temporal Balancing

To address the imbalance in publication frequency over time, a temporal balancing strategy was applied after relevance screening, noise removal, and text preprocessing. Because the number of publications has increased markedly in recent years, direct analysis of the preprocessed but unbalanced dataset would place disproportionate emphasis on recent topics and vocabulary patterns. To mitigate recency bias, up to 40 publications per year were retained via random sampling to construct a temporally balanced dataset. A fixed random seed was used to reduce stochastic variability, and the sampling procedure is explicitly documented.
To evaluate the effect of this procedure, the yearly publication distributions before and after balancing were compared, as shown in Figure 2. The unbalanced dataset shows a sharp increase in recent years, whereas the balanced dataset provides a more even representation across years. Although this procedure reduced the absolute number of recent publications, subsequent robustness checks confirmed that the overall thematic structure and temporal evolution remained broadly consistent between the unbalanced and balanced datasets. This indicates that temporal balancing mitigates recency bias without fundamentally altering the underlying structure of the research landscape.

2.5. Document-Term Matrix Construction and Topic Modeling

A document-term matrix (DTM) was constructed from the preprocessed and temporally balanced dataset using the Bag-of-Words representation implemented with CountVectorizer. The maximum number of features was limited to 350 in order to control dimensionality. Terms appearing in fewer than three documents were excluded to remove rare noise terms, whereas terms appearing in more than 75% of documents were excluded to eliminate overly common terms. Both unigrams and bigrams were included so that meaningful phrases could be retained in addition to single-word expressions. The resulting DTM provided the structured numerical representation used for topic modeling, network analysis, and machine-learning-based temporal interpretation.
Latent Dirichlet Allocation (LDA) was employed to identify latent research topics within the preprocessed and temporally balanced dataset. Each document was represented as a probabilistic distribution over topics, and each topic was characterized by a distribution of keywords. The model was trained using the batch learning methods with a maximum of 30 iterations.
The number of topics (k) was determined based on both interpretability and quantitative evaluation. As shown in Figure 3, the topic coherence (UMass) reached its highest value at k = 4 and decreased monotonically as the number of topics increased, with the lowest coherence observed around k = 10. This trend indicates that increasing k leads to a gradual reduction in semantic coherence.
However, although smaller values of k yielded higher coherence scores, they resulted in overly coarse topic structures that merged distinct research themes into a limited number of broad categories. In contrast, larger values of k enable finer-grained thematic separation, which is essential for distinguishing important subdomains such as connection behavior, shear wall systems, and emerging engineered wood technologies (e.g., CLT and hybrid systems).
Considering this trade-off between coherence and interpretability, k = 8 was selected as a balanced and practically meaningful configuration. At this value, the model sacrifices a moderate degree of coherence but achieves a clearer and more informative decomposition of the research landscape.
To further assess robustness, the stability of the topic structure was evaluated using multiple random seeds. The consistency of dominant topic assignments was quantified using normalized mutual information (NMI). Although some overlap between closely related topics remains, the overall thematic structure was found to be stable across different model runs. These results support the reliability of the selected topic configuration for subsequent analysis and interpretation.

2.6. Co-Occurrence Network Analysis

A keyword co-occurrence network was constructed from the preprocessed and temporally balanced dataset to capture relationships among terms. In this network, nodes represent keywords, and edges represent co-occurrence within the same document. Edge weights correspond to the frequency of co-occurrence. To reduce weak or potentially noisy relationships, edges with weights below four were removed, and isolated nodes were excluded from the final network.
To examine the structural characteristics of the network, graph-level metrics such as density and average node degree were calculated. In addition, node-level measures, including degree and betweenness centrality, were computed to identify influential keywords within the research landscape. These analyses were used to clarify the conceptual core of the field and the connectivity among major research themes.

2.7. Temporal Trend Analysis

Temporal trends were analyzed by calculating the number of publications per year for both the unbalanced and the preprocessed and temporally balanced datasets. In addition, each document in the balanced dataset was assigned a dominant topic according to the LDA results, and topic frequencies were aggregated by publication year. This procedure enabled the visualization of shifts in thematic emphasis over time and supported the interpretation of long-term research evolution.
To complement these topic-based trends, period-wise word frequency comparisons were also performed by grouping publications into broad temporal intervals. This additional analysis provided a descriptive view of how salient terminology changed across historical stages of timber seismic research.

2.8. Machine Learning Model

To quantitatively analyze temporal patterns in the literature, a regression model was developed using XGBoost. The target variable was the publication year, which served as a proxy for temporal position within the evolution of the field, while the explanatory variables were derived from the document-term matrix.
Textual features were generated using CountVectorizer with the same settings described above (max_features = 350, min_df = 3, max_df = 0.75, ngram_range = (1,2)). These features were used as input to an XGBRegressor configured with n_estimators = 300, max_depth = 4, learning_rate = 0.05, subsample = 0.9, colsample_bytree = 0.9, and objective = “reg:squarederror”.
Before interpreting the model using SHAP, predictive performance was evaluated quantitatively. Repeated random holdout validation was conducted using multiple random seeds, and regression performance was assessed by the mean absolute error (MAE), root mean squared error (RMSE), and coefficient of determination (R2). In addition, robustness was examined using alternative split strategies, including k-fold cross-validation and a grouped split based on publication year. These evaluations were introduced to confirm that the model captured meaningful temporal patterns rather than artifacts of a single train–test split.
To interpret the trained model, SHAP (SHapley Additive exPlanations) was applied. SHAP values quantify the contribution of each textual feature to the predicted publication year, thereby enabling transparent interpretation of which terms are associated with earlier or later stages of development in timber seismic research. Furthermore, SHAP-based results obtained from the balanced dataset were compared with those from a weighted analysis on the unbalanced dataset to assess the robustness of keyword importance patterns.
In this study, the final analytical dataset refers to the preprocessed and temporally balanced dataset obtained after relevance screening, noise removal, vocabulary cleaning, and year-wise subsampling. To ensure transparency and reproducibility, detailed descriptions of the data collection, preprocessing procedures, and supplementary analyses are provided in the Supplementary Materials.

3. Results and Discussion

3.1. Predictive Performance and Temporal Generalization

Figure 4 presents the relationship between the true publication year and the predicted year obtained from the regression model under two validation strategies: (a) random holdout and (b) group-wise splitting by publication year. Overall, the predicted values exhibit a general alignment with the diagonal line, indicating that the model captures the temporal trend of publication years to a certain extent. However, several important patterns can be observed. First, a strong concentration of data points is visible in recent years (particularly after 2020), where the predictions tend to be more accurate and closely follow the diagonal. In contrast, for earlier years (approximately 2005–2015), the predictions show larger dispersion and a tendency toward systematic overestimation, with predicted years biased toward more recent values. This behavior likely reflects both the temporal imbalance of the dataset (i.e., a higher proportion of recent publications) and the evolution of terminology and research topics over time. Second, the difference between the two validation strategies is notable. In the random holdout setting (Figure 4a), training and test samples are randomly mixed across time, allowing the model to implicitly learn temporal patterns from future data. As a result, the predictions appear more tightly aligned with the diagonal, potentially leading to an optimistic estimation of model performance. In contrast, the group-wise split by year (Figure 4b) enforces a stricter and more realistic evaluation by preventing temporal leakage. Under this setting, the scatter becomes more pronounced, particularly for older publication years, highlighting the difficulty of temporal extrapolation. This indicates that the model’s predictive capability is sensitive to distribution shifts over time. These results suggest that while the model is capable of capturing general temporal trends in the literature, its performance is strongly influenced by dataset imbalance and temporal domain shift. Therefore, validation strategies that account for temporal structure are essential for obtaining a reliable assessment of predictive performance.
Figure 5 illustrates the distribution of residuals (true year − predicted year) under two validation strategies: (a) random holdout and (b) group-wise splitting by publication year. In both cases, the residuals are centered approximately around zero, indicating that the model does not exhibit a strong global bias in prediction. However, the shape and spread of the distributions reveal important differences in model behavior.
In the random holdout setting (Figure 5a), the residuals are relatively concentrated near zero, with a moderately symmetric distribution. This suggests that the model achieves stable performance when training and test samples are randomly mixed, benefiting from the implicit inclusion of temporal information across the dataset. Nevertheless, several extreme negative residuals can be observed, indicating occasional large overestimations of publication year.
In contrast, the group-wise split by year (Figure 5b) shows a noticeably wider distribution of residuals, with heavier tails on both sides. This indicates increased variability in prediction errors when temporal leakage is prevented. In particular, the presence of large negative residuals suggests that the model tends to overpredict publication years for older papers, while positive residuals indicate underprediction for some recent cases. Such patterns reflect the challenges of temporal generalization and the influence of shifting research trends over time. Overall, the comparison demonstrates that while the model maintains reasonable central tendency in both settings, its error distribution becomes significantly more dispersed under temporally consistent validation. This highlights the importance of evaluating models under realistic temporal constraints, as random holdout may underestimate the true uncertainty associated with predictions in evolving research domains.
Table 2 compares the prediction performance metrics (MAE, RMSE, and R2). The results indicate that the prediction errors (MAE and RMSE) increase under the group-wise splitting by publication year, suggesting a degradation in model performance when evaluated under a more realistic temporal setting. This implies that the random holdout strategy may lead to optimistic performance estimates due to potential information leakage or similarity between the training and test samples. Interestingly, the R2 value is higher in the year-based split despite larger errors. This can be attributed to increased variance in the target variable across different years, making the explained variance ratio appear larger. Nevertheless, the overall R2 values remain low, indicating limited explanatory power of the model. These findings highlight the importance of appropriate data splitting strategies for robust evaluation and suggest that temporal or group-wise splitting provides a more realistic assessment of model generalization. The discrepancy between low absolute errors and low R2 suggests that the model may be biased toward predicting the mean of the target variable, rather than learning meaningful patterns. This highlights the need for improved feature representation or model structure.

3.2. High-Frequency Keywords

The word-frequency analysis highlights the dominant vocabulary within the balanced corpus. As shown in Figure 6, the most frequent term is “structural”, followed by “design”, “frame”, “stiffness”, and “connections”, along with other high-ranking terms such as “energy”, “joints”, “CLT”, and “dissipation”. Several important observations can be drawn from this distribution. First, the prominence of terms such as “frame”, “stiffness”, “wall”, and “lateral” indicates that structural performance under lateral loading remains a central concern in timber seismic research. Rather than focusing solely on specific components, the vocabulary suggests a strong emphasis on system-level behavior, including global stiffness characteristics and load-resisting mechanisms. Second, the high frequency of “connections”, “connection”, “joints”, “loading”, “cyclic”, and “damage” highlights the continued importance of component-level mechanical behavior. This reflects the well-established understanding that the seismic performance of timber structures is governed by connection behavior, particularly in terms of energy dissipation, nonlinear response, and failure mechanisms. Third, the notable presence of “CLT” indicates the increasing importance of engineered wood products in recent research. Unlike traditional terms such as “wall” or “frame”, the prominence of “CLT” reflects a shift toward modern timber systems and industrialized construction methods. Furthermore, the appearance of terms such as “steel”, “composite”, and “masonry” suggests that the research scope extends beyond pure timber systems to include hybrid structures, multi-material systems, and comparative studies. Overall, the word-frequency distribution suggests that the field is structured around three major themes: (1) system-level structural behavior under lateral loading, (2) connection and joint mechanics governing seismic performance, and (3) the growing role of engineered timber and hybrid structural systems. These findings are consistent with the evolving research landscape, where traditional mechanics-based studies are increasingly integrated with modern materials and design approaches.

3.3. Keyword Co-Occurrence Network

Figure 7 presents a reduced co-occurrence network composed of the top 30 most frequent keywords. By limiting the network to the most representative terms, the visualization more effectively highlights the essential structural relationships and underlying patterns within the research field. This simplification enables a more precise understanding of the structural composition and evolving trends in timber seismic performance research.
The resulting network exhibits a highly dense and strongly interconnected topology, indicating that the field is not divided into isolated subdomains but is instead characterized by a high degree of conceptual integration. At the core of the network, keywords such as wall, walls, connections, joints, stiffness, frame, and lateral form a tightly coupled cluster. This central structure clearly reflects the fundamental role of lateral load-resisting systems in timber structures, particularly emphasizing the interaction between shear walls and connection behavior. These components directly govern key performance metrics, including stiffness, strength, and deformation capacity under seismic loading. A distinct yet strongly integrated substructure emerges around connection-related behavior, where terms such as joints, connections, steel, energy, dissipation, and cyclic are closely linked. This cluster highlights the critical importance of hysteretic behavior and energy dissipation mechanisms in seismic design. The presence of steel within this group further suggests an increasing research focus on hybrid connection systems, in which steel components are incorporated to enhance mechanical performance and reliability.
Material-related keywords, including CLT, composite, and materials, are also well embedded within the central network. This indicates that advances in engineered wood products and hybrid material systems are not treated as independent topics but are deeply integrated with structural performance considerations. In particular, the positioning of these terms suggests that material innovation is directly contributing to improvements in seismic resistance.
In addition, the network demonstrates a strong experimental orientation. Keywords such as test, tests, loading, and cyclic are widely distributed and highly connected, underscoring the central role of experimental validation in this field. This reflects the continued reliance on cyclic loading tests and shaking table experiments to verify analytical models and to capture complex nonlinear behaviors. In contrast, peripheral nodes such as masonry, geology, and construction exhibit relatively weaker connectivity. These terms likely represent interdisciplinary extensions or comparative studies, including research that contrasts timber systems with other structural materials or incorporates site-specific geotechnical conditions.
Overall, the absence of clearly separated clusters provides strong evidence that timber seismic research constitutes a highly cohesive and integrated domain. Structural behavior, connection mechanics, material innovation, and experimental validation are tightly interwoven, forming a unified research framework. This structural coherence distinguishes the field from more fragmented research areas and underscores the necessity of holistic, system-level approaches in advancing timber seismic design.

3.4. Topic Modeling Results

The LDA analysis extracted eight interpretable topics, each representing a distinct thematic direction within timber seismic research. Table 3 and Table 4 list the clustered topics and their associated keywords extracted from the LDA analysis.
While both datasets yielded broadly similar thematic structures, important differences can be observed in terms of topic clarity, balance, and interpretability. Across both datasets, several core research themes consistently emerged. These included (i) connection and joint mechanics, characterized by terms such as “connections”, “joints”, “dissipation”, and “cyclic”; (ii) structural system behavior under lateral loading, represented by “frame”, “stiffness”, “wall”, and “lateral”; and (iii) engineered wood systems, particularly cross-laminated timber (CLT), indicated by terms such as “clt”, “cross laminated”, and “panels”. These recurring themes confirm that the fundamental structure of timber seismic research is robust and largely independent of sampling strategy. However, the comparison between the unbalanced and balanced datasets revealed several notable differences.
In the unbalanced dataset, certain topics are strongly influenced by dominant and frequently occurring research areas. For example, topics related to finite element analysis (“finite element”, “nonlinear”, “dynamic”) and traditional shear wall systems appear prominently, reflecting the natural distribution of the literature where well-established research areas occupy a large proportion. In addition, some topics exhibit partial redundancy, such as multiple topics related to CLT and wall systems, suggesting that high-frequency domains may be overrepresented and further subdivided by the LDA model. In contrast, the balanced dataset produces more evenly distributed and semantically distinct topics. For instance, specific themes such as traditional joinery (“mortise”, “tenon”), energy dissipation mechanisms (“hysteretic”, “energy dissipation”), and hybrid or emerging materials (“bamboo”, “veneer”, “composite”) become more clearly identifiable. Furthermore, system-level topics such as modal behavior (“mode”, “modal”, “diaphragm”) and performance-based design concepts are more explicitly separated. This indicates that temporal balancing reduces the dominance of recent or highly published topics, allowing less frequent but still important research areas to be more clearly extracted. Another important observation is that the balanced dataset mitigates topic redundancy. While the unbalanced dataset tends to produce overlapping themes (e.g., multiple CLT-related or wall-related topics), the balanced dataset better differentiates between component-level behavior, system-level response, and material-specific studies. This suggests that balancing improves the interpretability and granularity of topic modeling results.
Nevertheless, some limitations remain. As observed in both datasets, low-information or domain-irrelevant tokens such as “in” or “table” still appear in certain topics, indicating that further refinement of domain-specific preprocessing is necessary. This is consistent with earlier observations that standard stop word removal alone is insufficient for specialized engineering corpora. In addition, minor overlaps between related topics persist, reflecting the inherent continuity of research themes and the limitations of unsupervised topic modeling.
Overall, the comparison demonstrates that while the core thematic structure of timber seismic research is stable, the balanced dataset provides a more reliable and interpretable representation of the research landscape. By reducing temporal and publication bias, it enables a clearer identification of both dominant and emerging topics, thereby offering a more robust basis for subsequent trend analysis.

3.5. Topic Evolution over Time

Figure 8 illustrates the temporal evolution of the identified topics based on the balanced dataset. The figure reveals a clear and non-uniform growth pattern in timber seismic research, with a particularly rapid increase in publications after approximately 2020.
First, the overall publication volume remained relatively low and stable before 2018, with only minor fluctuations across topics. This suggests that during this period, research activity was limited and concentrated in a small number of traditional areas. In contrast, a sharp increase in the number of publications was observed after 2020, indicating a significant expansion of the research field. Second, distinct differences in the growth patterns of individual topics could be identified. Topics related to structural systems and design frameworks (e.g., Topics 6 and 8) showed a substantial increase in recent years, suggesting a shift toward system-level and performance-oriented research. Similarly, topics associated with engineered wood products and hybrid systems (e.g., Topics 2 and 5) exhibited noticeable growth, reflecting the rising importance of modern materials such as CLT and composite structures.
In contrast, more traditional topics, such as connection mechanics and fundamental structural behavior, remained present throughout the entire period but did not exhibit the same level of rapid growth. This indicates that while these foundational areas continue to underpin the field, the relative focus of research has shifted toward more complex and integrated systems. Another notable observation was the increasing diversification of topics in recent years. After 2020, multiple topics simultaneously exhibited growth, resulting in a more balanced and distributed research landscape. This suggests that the field is no longer dominated by a small number of core themes but has evolved into a multi-dimensional research domain encompassing materials, structural systems, and performance-based design.
Overall, the temporal analysis indicates that timber seismic research has transitioned from a relatively narrow, component-focused discipline into a broader and more integrated field. The recent surge in publications and the diversification of topics reflect the combined influence of technological advancements, the increased adoption of engineered wood products, and growing interest in sustainable and resilient structural systems.

3.6. SHAP-Based Interpretation of Temporal Trends

The SHAP summary beeswarm plot shown in Figure 9 provides a quantitative and interpretable representation of the temporal dynamics of the research field. Unlike frequency-based analyses, this figure not only reveals which keywords are influential in predicting publication year, but also how their contributions vary depending on their relative importance within individual documents.
In this plot, each point corresponds to a single publication. The horizontal axis represents the SHAP value, which quantifies the contribution of a given keyword to the predicted publication year. Negative SHAP values indicate an association with earlier publications, whereas positive values indicate a contribution toward more recent years. The color of each point represents the feature value (i.e., the relative frequency of the keyword within the document), with blue indicating low frequency and red indicating high frequency. The vertical dispersion within each row reflects the beeswarm arrangement, enabling visualization of both the density and variability without overlap.
The most influential keywords identified in Figure 10 include dowel, presented, hysteretic, post, development, properties, carbon, composite material, loading, physics, nonlinear, shear wall, bracket, terms, curves, moment, mathematics, capacity, hysteresis, and composite. Notably, these terms were not necessarily the most frequent in the corpus, confirming that temporal significance is not determined by frequency alone but by the discriminative power of keywords across different periods.
A clear temporal pattern emerges from the distribution of SHAP values. Keywords such as dowel and presented exhibited predominantly negative SHAP values, particularly when their feature values were high. This indicates that documents in which these terms appear frequently are strongly associated with earlier publication years. Similarly, hysteretic and post also showed a tendency toward negative contributions, suggesting that earlier research placed greater emphasis on connection-level behavior, experimental reporting, and fundamental hysteretic response.
In contrast, terms such as development, carbon, and capacity displayed predominantly positive SHAP values, especially for high feature values. This indicates that these keywords are characteristic of more recent publications. In particular, carbon suggests a growing research focus on sustainability and carbon-related performance, while capacity and development reflect an increasing emphasis on structural performance, design advancement, and practical implementation. The keywords composite and composite material also showed a tendency toward positive SHAP values, indicating a recent expansion of interest in hybrid and engineered material systems.
Several keywords, including loading, nonlinear, shear wall, and moment, exhibited SHAP value distributions centered near zero or spanning both negative and positive regions. This suggests that these concepts have remained consistently relevant across different time periods. Rather than being tied to a specific phase, they represent fundamental concepts that persist throughout the evolution of the field, albeit with changing contexts and levels of emphasis.
Overall, the SHAP beeswarm plot demonstrates that the temporal evolution of timber seismic research is characterized by a layered and cumulative progression rather than a simple shift from old to new topics. Early research themes, particularly those related to connection mechanics and hysteretic behavior, continue to underpin the field, while newer emphases—such as structural capacity, composite systems, and sustainability considerations—have emerged and gained increasing importance. This result provides strong evidence that the development of the field is not discontinuous but integrative, with new research directions building upon and extending established knowledge frameworks.
Figure 10 presents representative SHAP dependency plots for selected keywords, illustrating how the contribution of each keyword to the predicted publication year varies as a function of its occurrence within individual documents. Unlike the summary beeswarm plot, these dependency plots enable a more detailed examination of feature-specific behavior and potential nonlinear relationships. The keyword dowel (Figure 10a) exhibited a clear negative association with publication year. High occurrences of dowel were consistently linked to strongly negative SHAP values, indicating that documents emphasizing dowel-type fasteners are more likely to correspond to earlier studies. As the frequency decreased, the SHAP values approached zero, suggesting a diminishing temporal influence. This pattern reflects the foundational role of dowel-type connection mechanics in earlier timber engineering research. The keyword hysteretic (Figure 10b) also showed a predominantly negative SHAP distribution, particularly at low to moderate frequencies. This indicates that studies focusing on hysteretic behavior are generally associated with earlier phases of research. However, compared to dowel, the distribution was more dispersed, and SHAP values gradually approached zero as the feature value increased. This suggests that while hysteretic analysis remains relevant, its role has become more integrated into broader analytical frameworks rather than serving as a primary distinguishing feature of recent studies. For the keyword post (Figure 10c), a similar negative trend was observed. Most data points were concentrated in the negative SHAP region, indicating that the presence of post is more strongly associated with earlier publications. The relatively wide spread of SHAP values, particularly at low frequencies, suggests variability in how this term contributes to temporal prediction. This may reflect the diverse contexts in which post is used, including both traditional post-and-beam systems and more general structural descriptions.
An important observation across all three keywords is that high feature values (i.e., frequent occurrence within a document) tend to correspond to more negative SHAP values. This consistent pattern reinforces the interpretation that these terms are characteristic of earlier research stages. At the same time, the gradual convergence of SHAP values toward zero at lower frequencies indicates that their discriminative power diminishes when they are not a dominant feature of the document.
Overall, the dependency plots confirm and refine the trends observed in the SHAP summary analysis. Early-stage timber seismic research is strongly characterized by keywords related to connection mechanics (dowel), fundamental cyclic behavior (hysteretic), and traditional structural systems (post). In contrast, the reduced influence of these terms in recent publications suggests a shift toward more system-level, material-oriented, and performance-based research themes. Importantly, this transition does not imply the disappearance of earlier topics, but rather their incorporation into a broader and more integrated research framework.

3.7. Integrated Interpretation

When the results of the frequency analysis, topic modeling, co-occurrence network analysis, and SHAP-based interpretation are considered collectively, a coherent and consistent picture of the evolution of timber seismic research emerges. The field can be interpreted as having progressed through three broad stages, characterized by a gradual shift in focus while maintaining strong continuity in fundamental concepts.
The first stage is characterized by relatively limited publication volume and a primary focus on fundamental mechanical behavior, including joints, nails, stiffness, slip, and lateral resistance in wall systems. During this period, research was largely centered on component-level mechanics, particularly the behavior of nailed light-frame shear walls and connections. These studies established the foundational understanding necessary for subsequent developments in seismic design. The second stage reflects a transition toward thematic diversification and increasing system-level consideration. Research expanded to include global structural behavior, dynamic testing, diaphragms, and multi-story response. This phase represents a critical integration of component-level mechanics with structural-scale seismic performance, supported by studies on the cyclic behavior of connections, moment-resisting timber frames, and shaking-table experiments. The third and current stage is marked by rapid growth and a pronounced shift toward advanced and integrated systems. Research increasingly emphasizes engineered wood products—particularly cross-laminated timber (CLT)—as well as hybrid structural systems, post-tensioned, and energy-dissipative design strategies, and performance- or resilience-based evaluation. These studies focus on system-level behavior, including global response, damage progression, and collapse performance, reflecting the maturation of the field. Importantly, this evolution does not represent a sequence of discrete paradigm shifts, but rather a layered and cumulative expansion of research themes. Fundamental topics such as connection mechanics, cyclic behavior, and shear resistance remain central, as confirmed by the dense and highly interconnected structure of the keyword co-occurrence network. At the same time, the SHAP analysis demonstrates that temporal evolution is driven by diagnostically significant terms—such as slip, damper, buckling, capacity, and application—rather than by frequency alone. These keywords correspond to critical mechanical and design concepts, reinforcing the validity of the data-driven interpretation.
Overall, the results indicate that contemporary timber seismic research has evolved from a component-focused discipline into a system-oriented and performance-driven field. The integration of material innovation, structural systems, and design methodologies reflects the increasing demand for resilient and sustainable timber structures. This progression highlights the necessity of holistic approaches that bridge traditional mechanics with emerging technologies and system-level performance considerations.

4. Conclusions

This study established a data-driven and explainable artificial intelligence (XAI) framework for analyzing research trends in the seismic performance of timber structures by integrating large-scale literature mining, natural language processing, topic modeling, network analysis, and SHAP-based machine learning.
Unlike conventional bibliometric approaches, the proposed framework enables not only the identification of dominant research themes, but also the quantitative and interpretable characterization of their temporal evolution. In particular, the integration of topic modeling with SHAP provides a novel analytical perspective, allowing the relative importance of keywords to be evaluated in terms of their contribution to temporal trends. This extends traditional frequency- and network-based analyses toward a more diagnostic, reproducible, and mechanistically interpretable framework.
From a practical standpoint, the proposed approach offers a scalable and generalizable tool for systematically mapping the evolution of complex engineering research domains. It enables the identification of emerging research directions, clarifies the relationships between component-level mechanics and system-level structural behavior, and supports evidence-based decision-making for future research planning. Moreover, the framework is transferable to other areas of structural and civil engineering where large and heterogeneous bodies of literature must be synthesized in a quantitative and transparent manner.
However, the results should be interpreted within the scope of the adopted methodology. The findings are derived from bibliometric and textual representations of the literature and therefore reflect patterns of research focus rather than direct measures of physical performance or experimental validation. As such, they provide insight into the evolution of scientific discourse rather than definitive conclusions regarding structural behavior.
Several limitations should be acknowledged. The dataset was limited to English-language journal articles, excluding conference proceedings that often play a significant role in engineering research. In addition, the analysis relied on predefined search queries and abstract-level text, which may introduce selection bias and restrict semantic depth. The results are also sensitive to preprocessing decisions, including stopword removal and domain-specific vocabulary filtering, as well as to modeling choices such as the number of topics in the LDA framework. Furthermore, the temporal balancing strategy, while effective in mitigating recency bias, may partially distort the natural growth pattern of the literature. Finally, the use of a single bibliographic database may limit the overall completeness of the dataset.
Future research should address these limitations by incorporating multiple data sources, including conference proceedings and full-text datasets, to enhance coverage and semantic richness. In addition, systematic robustness analyses—such as sensitivity tests for preprocessing strategies, feature selection, and topic modeling parameters—are necessary to further strengthen the reliability of the framework. The integration of more advanced language models and embedding-based representations also represents a promising direction for improving semantic interpretation and capturing deeper contextual relationships within the literature.

Supplementary Materials

The following supporting information can be downloaded at: https://www.mdpi.com/article/10.3390/app16094418/s1.

Author Contributions

T.N. and Y.S.: Methodology, investigation, writing—review and editing. All authors have read and agreed to the published version of the manuscript.

Funding

This research was funded by JSPS KAKENHI, Grant Number 25K23510.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

Some or all data, models, or code that support the findings of this study are available from the corresponding author upon reasonable request.

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

The following abbreviations are used in this manuscript:
AIArtificial Intelligence
XAIExplainable Artificial Intelligence
SHAPSHapley Additive exPlanations

References

  1. Win, L.S.Y.; Isoda, H.; Nakagawa, T.; Shinohara, M. Comparative analysis of CO2 emissions during construction of Japanese wooden buildings by construction method. In Summaries of Technical Papers of Annual Meeting, Architectural Institute of Japan (Kyushu); Architectural Institute of Japan: Tokyo, Japan, 2025; pp. 2343–2344. (In Japanese) [Google Scholar]
  2. Bai, M.; Song, L.; Xiao, Q.; Sun, J.; Liu, H. Development of low-carbon engineered cementitious composites incorporating lithium slag: Mechanical properties, microstructure evolution, and life cycle assessment. Constr. Build. Mater. 2026, 520, 145983. [Google Scholar] [CrossRef] [Scilit]
  3. Filiatrault, A.; Folz, B.T. Performance-based seismic design of wood frame buildings. J. Struct. Eng. 2002, 128, 39–47. [Google Scholar] [CrossRef] [Scilit]
  4. Seim, W.; Hummel, J.; Vogt, T. Earthquake design of timber structures: Remarks on force-based design procedures for different wall systems. Eng. Struct. 2014, 76, 124–137. [Google Scholar] [CrossRef] [Scilit]
  5. Stepinac, I.; Šušteršič, I.; Gavrić, D.; Rajčić, V. Seismic design of timber buildings: Highlighted challenges and future trends. Appl. Sci. 2020, 10, 1380. [Google Scholar] [CrossRef] [Scilit]
  6. Gavrić, I.; Ceccotti, A.; Fragiacomo, M. Cyclic behavior of cross-laminated timber wall systems. J. Struct. Eng. 2015, 141, 04015034. [Google Scholar] [CrossRef] [Scilit]
  7. Franco, L.; Fragiacomo, M.; Casagrande, R. Strategies for modeling of cross-laminated timber panels under cyclic loads. Eng. Struct. 2019, 198, 109476. [Google Scholar] [CrossRef] [Scilit]
  8. Zhang, X.; Isoda, H.; Sumida, K.; Araki, Y. Seismic performance of three-story cross-laminated timber structures. J. Struct. Eng. 2021, 147, 04020319. [Google Scholar] [CrossRef] [Scilit]
  9. Estrella, X.; Erochko, J.; Tesfamariam, S. Seismic performance factors for timber buildings with wood-frame shear walls. Eng. Struct. 2021, 248, 113185. [Google Scholar] [CrossRef] [Scilit]
  10. Hashemi, A.; Bagheri, H.; Yousef-Beik, M.M.; Mohammadi-Darani, F.; Zarnani, P.; Quenneville, B. Enhanced seismic performance of timber structures using resilient connections. J. Struct. Eng. 2020, 146, 04020180. [Google Scholar] [CrossRef] [Scilit]
  11. Hirakawa, T.; Uematsu, A.; Fukushima, Y.; Adachi, Y.; Kikuta, K. Seismic retrofit technique using plywood and common nails for timber frame structures. Buildings 2022, 12, 1029. [Google Scholar] [CrossRef] [Scilit]
  12. Tan, W.; Wang, J.; Liu, Z. Seismic performance of a single-story timber-framed masonry structure strengthened with fiber-reinforced cement mortar. Materials 2024, 17, 3644. [Google Scholar] [CrossRef] [Scilit]
  13. Shabani, A.; Alinejad, A.; Teymouri, M.; Costa, A.N.; Shabani, M.; Kioumarsi, M. Seismic vulnerability assessment and strengthening of heritage timber buildings: A review. Buildings 2021, 11, 661. [Google Scholar] [CrossRef] [Scilit]
  14. Namba, T. A review on artificial intelligence application for wooden structure. Artif. Intell. Data Sci. 2025, 6, 624–631. (In Japanese) [Google Scholar] [CrossRef]
  15. Mangalathu, S.; Hwang, S.-H.; Jeon, J.-S. Failure mode and effects analysis of RC members based on a machine-learning-based Shapley additive explanations (SHAP) approach. Eng. Struct. 2020, 219, 110927. [Google Scholar] [CrossRef] [Scilit]
  16. Malaga-Chuquitaype, E.J.C.; Chawgien, K. Interpretable machine learning models for the estimation of seismic drifts in CLT buildings. J. Build. Eng. 2023, 70, 106365. [Google Scholar] [CrossRef] [Scilit]
  17. Namba, T. Fundamental validation of an AI-based impact analysis framework for structural elements in wooden structures. Appl. Sci. 2026, 16, 915. [Google Scholar] [CrossRef] [Scilit]
  18. Oh, B.K.; Kim, J. Optimal architecture of a convolutional neural network to estimate structural responses for safety evaluation of the structures. Measurement 2021, 177, 109313. [Google Scholar] [CrossRef] [Scilit]
  19. Lee, S.C.; Park, S.K.; Lee, B.H. Development of the approximate analytical model for the stub-girder system using neural networks. Comput. Struct. 2001, 79, 1013–1025. [Google Scholar] [CrossRef] [Scilit]
  20. Kang, M.C.; Yoo, D.Y.; Gupta, R. Machine learning-based prediction for compressive and flexural strengths of steel fiber-reinforced concrete. Constr. Build. Mater. 2021, 266, 121117. [Google Scholar] [CrossRef] [Scilit]
  21. Namba, T. Machine learning surrogate for seismic response of a wooden house: A comparison of SHAP, Sobol, and Morris sensitivity analyses. Appl. Sci. 2026, 16, 3201. [Google Scholar] [CrossRef] [Scilit]
  22. Johansen, K.W. Yield theory of wood connections. Int. Assoc. Bridge Struct. Eng. Publ. 1949, 9, 249–262. [Google Scholar]
  23. American Wood Council. National Design Specification (NDS) for Wood Construction; American Wood Council: Leesburg, VA, USA, 2018. [Google Scholar]
  24. EN 1995-1-1; Eurocode 5: Design of Timber Structures—Part 1-1: General—Common Rules and Rules for Buildings. European Committee for Standardization (CEN): Brussels, Belgium, 2004.
  25. Bodig, J.; Jayne, B.A. Mechanics of Wood and Wood Composites; Van Nostrand Reinhold: New York, NY, USA, 1982. [Google Scholar]
  26. Xu, N.; Zhou, X.; Guo, C.; Xiao, B.; Wei, F.; Hu, Y. Text mining applications in the construction industry: Current status, research gaps, and prospects. Sustainability 2022, 14, 16846. [Google Scholar] [CrossRef] [Scilit]
  27. Jin, R.; Zou, P.X.W.; Piroozfar, P.; Wood, H.; Yang, Y.; Yan, L.; Han, Y. A science mapping approach-based review of construction safety research. Saf. Sci. 2019, 113, 285–297. [Google Scholar] [CrossRef] [Scilit]
  28. Wenzel, A.; Guindos, P.; Carpio, M. Using timber in mid-rise and tall buildings to construct our cities: A science mapping study. Sustainability 2025, 17, 1928. [Google Scholar] [CrossRef] [Scilit]
  29. Lundberg, S.M.; Lee, S.-I. A unified approach to interpreting model predictions. In Advances in Neural Information Processing Systems 30; Neural Information Processing Systems Foundation, Inc.: San Diego, CA, USA, 2017. [Google Scholar]
  30. Priem, J.; Piwowar, H.; Orr, R. OpenAlex: A Fully Open Index of Scholarly Works, Authors, Venues, Institutions, and Concepts. 2022. Available online: https://openalex.org (accessed on 10 April 2026).
  31. Chen, T.; Guestrin, C. XGBoost: A scalable tree boosting system. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, San Francisco, CA, USA, 13–17 August 2016; pp. 785–794. [Google Scholar]
  32. van Eck, N.J.; Waltman, L. Software survey: VOSviewer, a computer program for bibliometric mapping. Scientometrics 2010, 84, 523–538. [Google Scholar] [CrossRef] [Scilit]
  33. Pedregosa, F.; Varoquaux, G.; Gramfort, A.; Michel, V.; Thirion, B.; Grisel, O.; Blondel, M.; Prettenhofer, P.; Weiss, R.; Dubourg, V.; et al. Scikit-learn: Machine learning in Python. J. Mach. Learn. Res. 2011, 12, 2825–2830. [Google Scholar]
Figure 1. Flowchart of large-scale literature mining.
Figure 1. Flowchart of large-scale literature mining.
Applsci 16 04418 g001
Figure 2. Year distribution before and after balancing.
Figure 2. Year distribution before and after balancing.
Applsci 16 04418 g002
Figure 3. Topic coherence sensitivity.
Figure 3. Topic coherence sensitivity.
Applsci 16 04418 g003
Figure 4. Comparison between true and predicted publication years under two validation strategies: (a) random holdout and (b) group-wise splitting by publication year.
Figure 4. Comparison between true and predicted publication years under two validation strategies: (a) random holdout and (b) group-wise splitting by publication year.
Applsci 16 04418 g004
Figure 5. Residual distributions (true year − predicted year) under two validation strategies: (a) random holdout and (b) group-wise splitting by publication year.
Figure 5. Residual distributions (true year − predicted year) under two validation strategies: (a) random holdout and (b) group-wise splitting by publication year.
Applsci 16 04418 g005
Figure 6. Top 30 word frequencies (refined balanced set).
Figure 6. Top 30 word frequencies (refined balanced set).
Applsci 16 04418 g006
Figure 7. Top 30 keywords co-occurrence network.
Figure 7. Top 30 keywords co-occurrence network.
Applsci 16 04418 g007
Figure 8. Topic composition by year (balanced refined set).
Figure 8. Topic composition by year (balanced refined set).
Applsci 16 04418 g008
Figure 9. SHAP summary plot.
Figure 9. SHAP summary plot.
Applsci 16 04418 g009
Figure 10. SHAP dependency plots; (a) dowel; (b) hysteretic; (c) post.
Figure 10. SHAP dependency plots; (a) dowel; (b) hysteretic; (c) post.
Applsci 16 04418 g010
Table 1. Dataset screening and balancing summary.
Table 1. Dataset screening and balancing summary.
StageDescriptionNumber of Records
Initial retrievalRecords collected from all queries9779
After deduplicationDuplicate records removed5892
After relevance filteringTimber + seismic filter applied937
After noise removalNon-relevant domains removed418
After preprocessingShort/invalid documents removed418
Final balanced datasetMax 40 per year applied248
Table 2. Comparison of prediction performance metrics (MAE, RMSE, and R2).
Table 2. Comparison of prediction performance metrics (MAE, RMSE, and R2).
Random HoldoutGroup-Wise Splitting by Publication Year
MAE0.000530.00072
RMSE3.654.26
R20.00970.1002
Table 3. Topic and top words in the unbalanced dataset.
Table 3. Topic and top words in the unbalanced dataset.
TopicTop Words
1element, finite, finite element, stiffness, dynamic, nonlinear, parametric, elements, damper, displacement
2joints, joint, tenon, steel, mortise, dissipation, composite, loading, mortise tenon, energy
3clt, laminated, cross, cross laminated, panels, laminated clt, clt panels, connections, connection, existing
4masonry, damage, wall, roof, assessment, retrofit, walls, retrofitting, collapse, design
5stiffness, load, ductility, capacity, wall, composite, lateral, plane, in, walls
6design, construction, based, concrete, materials, systems, timber, energy, high, architectural
7connection, connections, mass, design, rocking, damage, post, wall, bamboo, modular
8frame, frames, buckling, steel, brace, networking, frame networking, mass, column, light
Table 4. Topic and top words in the balanced dataset.
Table 4. Topic and top words in the balanced dataset.
TopicTop Words
1masonry, damage, post, plane, walls, in, assessment, tensioned, collapse, table
2clt, cross, laminated, cross laminated, connection, energy, panels, design, construction, connections
3load, steel, composite, stiffness, bearing, design, materials, properties, element, material
4mass, design, buckling, mode, diaphragm, lateral, modal, frame, large, restrained
5laminated, bamboo, veneer, frame, high, lumber, column, design, rise, concrete
6dissipation, stiffness, connections, energy, steel, energy dissipation, hysteretic, cyclic, joint, moment
7tenon, mortise, mortise tenon, joints, friction, reinforced, damper, tenon joints, dampers, reinforcement
8frame, design, based, walls, systems, time, framework, light, shear, concrete
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Namba, T.; Sakai, Y. A Data-Driven and Explainable AI Framework for Quantitative Analysis of Research Trends in Timber Seismic Engineering. Appl. Sci. 2026, 16, 4418. https://doi.org/10.3390/app16094418

AMA Style

Namba T, Sakai Y. A Data-Driven and Explainable AI Framework for Quantitative Analysis of Research Trends in Timber Seismic Engineering. Applied Sciences. 2026; 16(9):4418. https://doi.org/10.3390/app16094418

Chicago/Turabian Style

Namba, Tokikatsu, and Yuta Sakai. 2026. "A Data-Driven and Explainable AI Framework for Quantitative Analysis of Research Trends in Timber Seismic Engineering" Applied Sciences 16, no. 9: 4418. https://doi.org/10.3390/app16094418

APA Style

Namba, T., & Sakai, Y. (2026). A Data-Driven and Explainable AI Framework for Quantitative Analysis of Research Trends in Timber Seismic Engineering. Applied Sciences, 16(9), 4418. https://doi.org/10.3390/app16094418

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop