Skip to Content
BuildingsBuildings
  • Article
  • Open Access

16 September 2026

Preliminary Application Research on Deep Learning Based on a Topic Focusing on the Identification of Modern Cultural and Educational Architectural Styles in Hunan, China

,
,
,
,
,
,
and
College of Architecture, Changsha University of Science & Technology, Changsha 410076, China
*
Author to whom correspondence should be addressed.

Abstract

Modern cultural and educational buildings are important material carriers of regional modernization, educational transformation, and Sino-Western cultural exchange. In Hunan, these buildings exhibit complex stylistic features shaped by the interaction of Western architectural languages, Chinese revival forms, modern decorative expressions, and early modern architectural tendencies. This study aims to explore a subject-focused deep learning framework for auxiliary-style recognition of modern educational and cultural buildings in Hunan and to investigate the feasibility of transferring stylistic representations learned from non-Hunan samples to the Hunan target domain under cross-regional conditions. A six-class style classification system with operational morphological criteria was established, and a geographically isolated cross-regional target-domain evaluation framework was adopted. Modern public, educational, and cultural building samples from outside Hunan Province were used for training and validation, while Hunan samples were retained as a fixed target-domain test set. Several representative backbone networks were used as references for model selection, and Swin Transformer was adopted as the unified backbone. The effects of L-Softmax, multi-scale feature fusion, global channel–spatial attention, PFP/PA, GBVS/AGL subject guidance, and Grad-CAM-based background constraints were compared under the unified cross-regional target-domain evaluation protocol. EMA was used only as an auxiliary training stabilization strategy. Model performance was evaluated using image-level accuracy, building-level accuracy based on multi-view majority voting, balanced accuracy, macro-F1, bootstrap confidence intervals, confusion matrices, and representative boundary cases. On the fixed Hunan target-domain test set containing 186 images from 37 building instances, the highest image-level accuracy reached 58.60%, while the highest building-level accuracy reached 64.86%. Considering the strong class imbalance and limited number of building instances in several categories, these results were further interpreted together with balanced metrics and bootstrap confidence intervals rather than accuracy alone. The A1–A8 ablation experiments indicated that the A7 variant incorporating GBVS/AGL achieved a relatively balanced performance between image-level and building-level evaluations under the current target-domain condition. The proposed framework provides an exploratory auxiliary approach for architectural heritage surveys, digital documentation, and style-boundary analysis, rather than a fully generalized architectural style recognition system, while offering methodological reference for future component-level quantitative research on Si-no-Western architectural integration.

1. Introduction

Modern educational and cultural buildings are important material witnesses to local modernization, educational transformation, and Sino-Western architectural and cultural exchange in China. The development of modern Chinese architecture was not a linear process in which one style simply replaced another; rather, it unfolded through the complex interaction of imported Western architectural forms, the continuation of indigenous architectural traditions, the emergence of modern professional architectural education, and the dissemination of modernist ideas [1,2]. Existing studies have argued that research on modern Chinese architectural history should move beyond the accumulation of individual cases toward more integrated analyses of architectural types, regional characteristics, and theoretical issues [3,4]. At the same time, the stylistic evolution of modern architecture should be understood in relation to architectural form, institutional background, and socio-cultural context [5]. Modern educational and cultural buildings therefore possess significance not only for educational and social history, but also for architectural typology, stylistic history, and heritage conservation [6,7].
Concepts such as “Yangshi”, “style”, and “architectural style” in modern architecture were not fixed visual labels but historically constructed concepts that emerged and were continually reinterpreted within modern architectural discourse [6]. Accordingly, modern educational and cultural buildings constitute important objects for research on architectural typology, stylistic history, and modern architectural heritage conservation [7,8]. Veranda-style architecture, as one of the important early imported forms in modern Chinese architecture, reflects the transmission, adaptation, and translation of Western architectural forms within the Chinese local environment [9]. Architectural debates during and around the period of the War of Resistance against Japanese Aggression further reshaped the relationships among national forms, modern structures, and architectural education [10,11]. The transformation of traditional ideas and the construction of theories of national architecture within modern architectural education also provide an important background for understanding traditional revival and the transition toward modern architecture in modern educational and cultural buildings [12].
Since the late Qing and early Republican periods, Hunan has developed a number of representative school buildings, missionary educational facilities, and public educational buildings. Compared with coastal treaty-port cities, the modernization process and educational reforms in Hunan followed a different historical rhythm, and both local social modernization and institutional transformation in education influenced the development of modern educational and cultural buildings [13,14]. Morphologically, these buildings often exhibit a complex stylistic spectrum. Their facades may incorporate Western architectural languages such as verandas, classical orders, pediments, arches, and eclectic composition, while also integrating palace-style roofs, traditional motifs, national-form revival, modern decorative expression, and transitional features associated with early modernism. Recent studies on the digital conservation of modern educational architecture have also demonstrated the practical need to identify, document, and preserve such buildings through digital methods [15]. However, because architectural styles are closely associated with regional historical contexts, construction traditions, and image acquisition environments, models trained on samples from other regions may encounter domain differences when applied to Hunan buildings. Therefore, evaluating cross-regional transferability rather than performance under randomly mixed datasets is essential for understanding the practical applicability of deep learning-based architectural style recognition.
Traditional architectural style assessment primarily relies on field surveys, historical archives, and expert judgment, all of which can provide detailed and in-depth architectural–historical interpretation. However, in large-scale heritage surveys, image-based documentation, cross-regional comparison, and rapid preliminary screening, conventional methods often face limitations in efficiency and reproducibility. Early computer-vision studies demonstrated the feasibility of automatic architectural style recognition and visual-pattern discovery from facade imagery [16,17]. These studies provided an important methodological basis for the subsequent application of deep learning to architectural image classification. It should be emphasized that this study does not seek to replace architectural–historical judgment with deep learning. Instead, deep learning is used as an auxiliary tool for preliminary style recognition, sample documentation, and the identification of boundary cases that warrant further expert review.
However, the automatic style recognition of modern educational and cultural buildings in Hunan is not merely a conventional image-classification problem. First, some style categories are sparsely represented in the Hunan target domain, particularly Chinese revival palace-style architecture, Chinese revival modern-style architecture, and the “New Architecture” of the modern period, namely early modern architecture. The limited availability of such building instances requires model evaluation to be conducted under small-sample and imbalanced-class conditions. Under these circumstances, overall accuracy alone may not sufficiently reflect model behavior because dominant categories can disproportionately influence performance. Therefore, balanced evaluation metrics and uncertainty estimation are required to provide a more reliable interpretation of model performance. Second, clear boundary ambiguity exists among several style categories. Confusion may arise between Western eclectic architecture and Chinese revival hybrid architecture, between Chinese revival palace-style architecture and Chinese revival hybrid architecture, and between Chinese revival modern-style architecture and early modern architecture because of the coexistence of Chinese and Western components, the simplification of decorative languages, or similarities in massing organization. Third, field photographs frequently contain non-building elements such as trees, sky, roads, pedestrians, signs, and surrounding buildings. Models may incorrectly use such background information as a basis for style judgment, thereby affecting the stability of cross-regional recognition. In addition, architectural style categories derived from historical interpretation need to be translated into reproducible morphological criteria for computational recognition. Therefore, this study establishes operational annotation rules based on roof morphology, facade composition, decorative systems, compositional organization, and the relationship between Chinese and Western architectural elements.
Against this background, this study conducts a subject-focused deep learning exploration for the style recognition of modern educational and cultural buildings in Hunan. It addresses three main research questions: (1) Can a model trained on samples from outside Hunan be transferred to the target domain of modern educational and cultural buildings in Hunan? (2) Under cross-regional target-domain conditions, can subject-focused mechanisms and background-suppression strategies improve the stability of building-level recognition? (3) Can model confusion patterns provide architectural evidence for understanding transitional boundaries within the stylistic spectrum of modern architecture? The purpose of this study is not to demonstrate that the model has already achieved mature, high-precision classification. Rather, it examines whether subject-focused and background-suppression mechanisms can improve the stability of building-level classification under conditions characterized by regional differences between the training and Hunan target domains, ambiguous stylistic boundaries, and complex image backgrounds.
In summary, the main contributions of this study are fourfold. First, a cross-regional target-domain evaluation framework is established for the style recognition of modern educational and cultural buildings. Rather than relying on randomly mixed datasets, the framework separates source-domain and Hunan target-domain samples at the building-instance level to examine model transferability under regional differences. Second, several representative backbone networks and candidate mechanisms, including PFP/PA, MSFF, GCSA, GBVS/AGL, Grad-CAM-based constraints, and L-Softmax, are comparatively evaluated to investigate subject-focused and background-suppression strategies under complex architectural-image conditions. These experiments are intended to examine the relative contribution of different mechanisms rather than to establish a universally optimal classification model. Third, model performance is evaluated at both the image and building levels. Multi-view majority voting is used to aggregate image-level predictions from the same building into a final building-level classification, thereby aligning the evaluation more closely with architectural practice, in which the style of a building instance rather than that of an individual photograph is ultimately judged. Fourth, confusion patterns involving Categories 02/04, 03/04, and 05/06 are interpreted in relation to architectural transition boundaries, linking computational recognition results with architectural–historical analysis.

2. Literature Review

Research on modern Chinese architecture provides the historical and conceptual basis for the six-class style system used in this study. Synthetic histories describe the successive interaction of imported forms, revivalist programs, and emerging modernism [1,2]. They also locate these changes within institutional reform, architectural education, and Sino-Western exchange [3]. Methodological and periodization studies emphasize that modern Chinese architecture should be understood through broader typological, regional, and historical frameworks rather than through isolated cases or chronological labels alone [4,5]. The terminology itself was historically translated and reconstructed in Chinese architectural discourse, so labels such as “style” are analytical categories rather than self-evident visual Categories [6]. Reviews of modern Chinese architectural history and studies of modern architectural heritage conservation further demonstrate the need to connect historical research with heritage recognition and conservation practice [7,8]. Within this framework, veranda-style architecture and Western eclectic architecture represent early processes of importation, climatic adaptation, and localization [9]. Architectural discourse and practice during the wartime period further reconfigured the relationships among national forms, modern construction, and public or institutional architecture [10,11]. Modern architectural education also promoted the creative transformation of traditional concepts and contributed to the construction of theories of national architecture [12]. Chinese revival modern-style architecture and early modern architecture then marked a partial shift toward simplified facades, functional planning, and modernist concepts [3,12]. These Categories therefore overlap in elements and chronology [1,3]. Given the historically constructed and evolving nature of architectural-style terminology [6], overlaps among these categories are treated in this study as historically interpretable boundary conditions rather than being assumed to represent annotation noise alone.
Hunan further complicates these boundaries because its modernization and educational reforms followed a distinct regional trajectory [13,14]. Consequently, modern educational and cultural buildings in Hunan often combine Western forms, local cultural expression, national-form strategies, and early modernist features rather than reproducing a single canonical type. Digital-conservation research on modern educational buildings shows why such sites require systematic image-based documentation as well as architectural–historical assessment [15]. This regional hybridity creates a demanding target domain for models trained on architectural images from other cities or building populations.
Automated architectural-style recognition has progressed from manually designed descriptors and visual-pattern discovery to convolutional neural network (CNN)-based representation learning. Early studies tested recognition across architectural categories and mined recurring visual patterns from facade imagery [16,17]. Deep CNNs were subsequently used to recognize residential architectural styles [18]. Related studies inferred architectural age and style at urban scale or determined a building’s style with convolutional networks [19,20]. Chinese studies also applied deep learning to building-style recognition and modern historical buildings [21,22]. Other work addressed traditional architectural types and regional dwellings [23,24]. Together, these studies establish technical feasibility but also show that performance is closely tied to category definition, image source, and geographic scope.
Research in cultural heritage broadens this evidence base beyond architectural style classification alone. Cross-disciplinary surveys map machine-learning applications across heritage tasks and evaluate their use in documenting and analyzing historical landmarks [25,26]. Transfer learning and data augmentation have been applied to cultural heritage building classification [27]. A separate study classified cultural heritage buildings in Athens using deep learning [28]. Street-view architectural-style dataset construction and region-specific decorative-pattern recognition further illustrate the importance of image sampling and geographic context in architectural-vision tasks [29,30]. Interpretable classification of traditional settlements further demonstrates the importance of dataset construction and model interpretation in architectural-style recognition [31]. This cross-journal evidence supports the use of deep learning as an auxiliary documentation and classification tool, but not as a substitute for architectural–historical judgment.
Architectural style is expressed at multiple spatial scales. Facade-element detection incorporating prior knowledge demonstrates that explicit architectural components can be directly incorporated into computational facade analysis [32], while hierarchical analysis of traditional architectural styles combines environmental, building-level, and landscape-level evidence [33]. Studies of vernacular and regional architecture similarly emphasize roofs, openings, materials, ornament, facade composition, and settlement context [34,35]. Intelligent facade-style recognition in real-world architectural images also demonstrates that architectural subjects are commonly embedded within complex visual contexts [36]. Quantitative facade assessment and vision-language modeling further extend architectural image analysis toward richer visual and semantic representations [37,38]. More generally, however, contextual cues may become shortcut features when they correlate with training labels without representing the underlying architectural style, a phenomenon widely discussed in the machine-learning literature on shortcut learning [39].
This concern aligns with the general machine-learning literature on shortcut learning and domain generalization. Deep models can exploit easy but non-causal cues that perform well within a source dataset and fail after a change in background, location, or acquisition conditions [39]. Domain-generalization research therefore focuses on performance under unseen distributions rather than random splits alone [40]. WILDS demonstrates the practical severity of distribution shift across real-world datasets [41], while comparative evaluation shows that carefully implemented empirical risk minimization remains a strong baseline against which domain-generalization methods should be tested [42]. For Hunan heritage buildings, these findings justify explicit cross-regional target-domain evaluation and separate analysis of minority Categories and hybrid styles. They also motivate building-instance-level aggregation, because multiple views of the same building should not be treated as independent evidence about different buildings.
Interpretability and subject-focused modeling provide complementary safeguards. Grad-CAM can localize image regions that contribute strongly to model predictions [43]. In this study, such heatmaps are therefore interpreted as indicators of model sensitivity rather than as direct evidence of architectural reasoning. Swin Transformer provides hierarchical multi-scale representations through shifted windows [44], and a recent historical-building study has combined it with global channel-spatial attention [45]. Related work has applied Swin Transformer to regional ethnic-building facades [46]. Transformer-based remote-sensing studies demonstrate the modeling of building extents and spatial context [47,48]. Taken together, these studies support testing transformer and attention mechanisms for heritage recognition, while requiring localization checks and expert interpretation to determine whether gains arise from the building subject rather than the background. Accordingly, the unresolved issue is not whether deep learning can classify architectural images, but whether it can provide reliable and interpretable assistance for fine-grained, regionally hybrid heritage categories under cross-regional shift. Architectural–historical studies provide the typological rationale and explain boundary cases, whereas computer-vision studies provide models for multiscale representation, subject localization, and interpretation. Few studies connect these traditions within a single evaluation framework that combines operational architectural-style criteria, geographically isolated target-domain testing, subject-focused analysis, image-level and building-level classification, and architectural interpretation of confusion patterns. This study addresses this gap while treating model outputs as decision support for, rather than replacements for, architectural–historical expertise.

3. Materials and Methods

3.1. Experimental Framework and Cross-Regional Target-Domain Validation

This study establishes a cross-regional target-domain evaluation framework for the style recognition of modern educational and cultural buildings in Hunan (Figure 1). Rather than maximizing classification performance through randomly mixed datasets, the framework is designed to examine whether stylistic representations learned from geographically separated source-domain samples can be transferred to the Hunan target domain.
Figure 1. Overall research framework, from geographically isolated source-target domain construction to subject-focused model comparison and architectural interpretation.
The workflow consists of seven interconnected stages. First, architectural images are collected and screened to construct a building-instance-level image database, ensuring that multiple views of the same building remain grouped during dataset organization. Second, a six-class style annotation system is established based on architectural–historical knowledge and operational morphological criteria. Third, source-domain and Hunan target-domain datasets are geographically separated at the building-instance level to avoid data leakage and to simulate cross-regional application conditions. Fourth, Swin Transformer is adopted as the unified backbone for constructing candidate subject-focused models. Fifth, A1–A8 ablation experiments are conducted under identical experimental settings to examine the contribution of different subject-focused, feature-enhancement, and background-suppression mechanisms. Sixth, model performance is evaluated at both image-level and building-level scales, where multi-view majority voting is used to obtain building-instance predictions. Seventh, confusion patterns, representative boundary cases, and attention visualization results are analyzed to investigate stylistic ambiguity and the relationship between computational recognition and architectural interpretation (Figure 1).
The central purpose of this framework is not to obtain high classification accuracy through random dataset splitting, but to examine whether stylistic knowledge learned from architectural samples outside Hunan can be transferred to the target domain of modern educational and cultural buildings in Hunan. The framework is therefore designed not as a conventional random-split classification benchmark, but as an evaluation system for cross-regional generalization, subject recognition, and architectural interpretation (Figure 1). In addition, considering the limited number of instances in several Hunan style categories, subsequent evaluations incorporate balanced metrics and uncertainty estimation to avoid relying solely on overall accuracy.
Because architectural style recognition is highly dependent on regional construction traditions, image acquisition environments, and stylistic distributions, the separation between non-Hunan source-domain samples and Hunan target-domain samples introduces potential domain differences. Therefore, this study further incorporates domain-shift quantification and in-distribution comparison in Section 4.2 to examine whether the observed performance reflects cross-regional generalization difficulty rather than random variation.

3.2. Dataset Construction

This study uses the building instance as the basic unit of data organization and adopts a cross-regional target-domain evaluation strategy based on explicit source-domain and target-domain separation. The source-domain dataset consists of modern public, educational, and cultural building samples from multiple regions outside Hunan Province and is used for model training and validation (Table 1). All Hunan samples are retained as a fixed external target-domain test set to examine whether architectural style representations learned from geographically distinct source-domain samples can be transferred to the regional context of Hunan (Figure 2). Images showing different views of the same building are always assigned to the same building instance and kept within the same data subset, thereby preventing similar images of the same building from appearing in both the training and testing stages and reducing the risk of data leakage.
Table 1. Dataset composition by image count and building-instance count under the source-domain and Hunan target-domain evaluation protocol.
Figure 2. Dataset composition statistics for the source-domain training/validation set and the fixed Hunan target-domain test set. Green circles indicate small-sample target-domain categories (03, 05, and 06), which require cautious interpretation in subsequent performance analysis.
The images were mainly collected from publicly available online sources, historical building image archives, public architectural heritage information, and field survey photographs of Hunan collected by the research team. Data screening was based on two basic criteria: the identity of the building instance had to be verifiable, and its stylistic characteristics had to be sufficiently interpretable. Priority was given to images that clearly showed the main facade, overall massing, or key stylistic components of the building. Images showing different views of the same building are always assigned to the same building instance and kept within the same data subset. This building-instance-level partition prevents multi-view images of the same building from appearing across source-domain training and target-domain evaluation sets, thereby reducing potential data leakage and overly optimistic performance estimation. Multiple photographs of the same building were assigned to a single building instance ID, regardless of differences in shooting angle, shooting time, or partial field of view. The collection and screening process followed consistent inclusion criteria across both source-domain and target-domain samples, including verifiable building identity, interpretable architectural characteristics, and sufficient visual information for stylistic assessment.
A unified preprocessing procedure was applied to all images. Each input image was resized to 224 × 224 pixels and normalized according to the input requirements of the pretrained visual models. During training, online data augmentation was applied, including random rotation, scaling, brightness perturbation, contrast perturbation, and random occlusion, to improve the model’s robustness to variations in shooting angle, lighting conditions, and partial occlusion. These augmentation operations were applied only during the training stage and did not alter the original evaluation conditions of the fixed Hunan target-domain test set.
The source domain training/validation set and the Hunan target domain test set together comprise 1436 images and 343 building instances, while the Hunan fixed-target domain test set includes 186 images and 37 building instances (Figure 2). The source-domain dataset contains 1436 images from 343 building instances, while the fixed Hunan target-domain test set contains 186 images from 37 building instances. A substantial class imbalance exists across the six categories in the fixed Hunan target-domain test set. In particular, Categories 03, 05, and 06 contain only 1, 1, and 2 building instances, respectively. Therefore, the recognition results of these categories cannot support stable class-level statistical inference and are interpreted primarily through representative cases rather than as evidence of generalized category performance. These three categories are therefore highlighted with green circles and are treated separately as small-sample target-domain categories in the subsequent results analysis (Figure 2). Their individual classification outcomes are not interpreted as evidence of stable class-level generalization.
Geographic isolation serves as the fundamental premise of the cross-regional target-domain evaluation protocol. This separation is designed to introduce realistic regional differences rather than simply divide data into training and testing subsets. It visually illustrates the separation between the source domain and target domain: blue-annotated areas represent the source domain training/validation datasets outside Hunan Province, while Hunan is highlighted in orange as a dedicated target domain testing region (Figure 3). Six representative source-domain samples from these blue-annotated areas demonstrate the range of building types and visual styles encountered during model training (Figure 3). The purpose of this design is not to perform random image splitting within the same geographical region, but to examine whether style features learned from non-Hunan samples can be transferred to modern educational and cultural buildings in Hunan, which are characterized by distinct regional historical backgrounds and architectural expressions.
Figure 3. Geographically isolated source–target domain design and representative source-domain samples. Blue-highlighted regions provide the non-Hunan source-domain training/validation data, whereas Hunan is retained as the fixed external target-domain test region for evaluating cross-regional generalization.
This cross-regional partitioning reduces the possibility of inflated performance caused by the same building appearing in both training and testing stages. It also creates a more challenging evaluation condition in which the model must address potential distribution differences arising from regional architectural traditions, image acquisition conditions, background environments, and stylistic variations. At the same time, it substantially increases the difficulty of the classification task. The model must not only distinguish among the six architectural style categories, but also address domain shifts caused by regional differences, variations in photographic conditions, differences in background environments, and stylistic continuity. Accordingly, this study places greater emphasis on transfer performance, building-level classification, and multi-view prediction consistency in the Hunan target domain, rather than simply pursuing high classification accuracy under random-split conditions (Figure 3).

3.3. Six-Class Style System and Operational Annotation Criteria

This study adopts a six-class-style classification system for modern educational and cultural buildings (Table 2). The system is not based on a mechanical chronological division. Rather, it is derived from discussions of modern Chinese architectural forms in the history of Chinese architecture and is further synthesized with research on architectural typology, stylistic evolution, and the transition toward modern architecture in the history of modern Chinese architecture [1,2]. Existing studies have indicated that the principal developmental trajectories and periodization of modern Chinese architecture should be understood through the combined consideration of architectural form, institutional background, and socio-cultural context [4]. Research on modern architectural history has also emphasized that individual buildings, regional differences, and typological characteristics should be examined within a broader historical framework [5]. Moreover, concepts such as “Yangshi”, “style”, and “architectural style” in modern architectural discourse were historically constructed rather than fixed visual labels [8]. Accordingly, the classification system used in this study emphasizes the integrated relationships among the sources of architectural form, compositional organization, roof and facade components, decorative language, and the degree of integration between Chinese and Western architectural elements (Table 2 and Figure 4). To improve reproducibility, each category was further defined through operational morphological criteria, including primary criteria, supporting criteria, and boundary criteria. The complete annotation protocol and category-level decision rules are provided in Supplementary Table S2.
Table 2. Six-class architectural style system and summarized operational annotation criteria.
Figure 4. Representative samples of the six architectural style categories in the fixed Hunan target-domain test set, illustrating category-specific characteristics, potential boundary similarities, and special cases requiring historical interpretation. Red notes indicate special interpretive remarks based on historical records and stylistic continuity.
Unlike purely chronological or building-type classifications commonly used in conventional architectural history studies, this study focuses on the visual stylistic spectrum formed by modern educational and cultural buildings under the influence of Sino-Western cultural exchange (Table 2). The categories are therefore not treated as rigid chronological labels, but as morphological and historical style groups defined by the dominant relationships among architectural components. Category 01, veranda-style architecture, and Category 02, Western eclectic architecture, mainly reflect the transmission, adaptation, and translation of Western architectural forms in China during an earlier stage. In particular, veranda-style architecture demonstrates the adaptation of imported architectural forms to local climatic and construction conditions in China [9] (Figure 4). Category 03, Chinese revival palace-style architecture, and Category 04, Chinese revival hybrid architecture, reflect the reorganization of traditional roofs, facade order, and national symbols in modern public, educational, and cultural buildings. Related architectural discussions also indicate that the relationships among national forms, modern structures, and public architectural expression were continuously reconstructed in modern China [10,11] (Figure 4). Category 05, Chinese revival modern-style architecture, and Category 06, the “New Architecture” of the modern period, namely early modern architecture, further reflect the gradual transition of architectural expression from symbolic national forms toward simplified facades, functional organization, and modernist concepts [12] (Figure 4). The detailed morphological distinctions between these categories are summarized in Supplementary Table S2. Unlike the training and validation samples drawn from multiple regions across China, the representative building instances in the fixed Hunan target-domain test set illustrate the specific manifestations and potential boundaries of the six-class style system within the Hunan context (Figure 4). Categories 01 and 02 respectively present representative examples of veranda-style architecture and Western eclectic architecture in Hunan, while Categories 03 to 06 further demonstrate the continuous transformation among traditional revival, hybrid expression, modern-style translation, and early modern architecture. By juxtaposing different buildings within the same category, or supplementary views of the same building, Figure 4 also illustrates that architectural style recognition cannot rely solely on a single image, but should instead consider multiple views of a building instance together with its overall morphological relationships (Figure 4).
It should be noted that category assignment in this study is based primarily on stylistic genealogy and morphological continuity rather than on construction date alone. A limited number of buildings completed after 1949 were included when historical evidence and architectural morphology indicated continuity with modern architectural traditions. The inclusion decisions considered design period, architectural language, formal composition, and the continuity of stylistic expression. Detailed information regarding post-1949 cases, historical evidence, and inclusion rationale is provided in Supplementary Table S3.
For example, some buildings completed after 1949 continue the formal logic of Chinese revival national-form expression or Sino-Western hybrid composition. These cases were therefore interpreted as examples of stylistic continuity rather than chronological exceptions. Such treatment avoids reducing architectural style recognition to a simple construction-date classification and better reflects the historically continuous transformation of modern Chinese architecture. Although the six style categories exhibit relatively clear morphological distinctions, unavoidable overlaps also exist within the stylistic spectrum. In particular, Category 02, Western eclectic architecture, and Category 04, Chinese revival hybrid architecture, as well as Category 05, Chinese revival modern-style architecture, and Category 06, the “New Architecture” of the modern period, namely early modern architecture, often cannot be distinguished on the basis of a single architectural component. Their classification requires a comprehensive assessment of roof form, facade proportions, decorative systems, compositional order, and architectural–historical context. A similar boundary also exists between Categories 03 and 04 within traditional revival architecture. Their distinction is more clearly reflected in the dominant role of palace-style roofs, the overall monumental composition, and the organizational relationship between traditional and modern components. Therefore, the subsequent discussion of the 02/04, 03/04, and 05/06 boundary pairs is grounded in the historical continuity and morphological transitionality inherent in this stylistic system. The operational distinctions used to separate these boundary categories are summarized in Supplementary Table S2.
Modern educational and cultural buildings in Hunan exhibit pronounced characteristics of regional translation. Some buildings simultaneously incorporate Western classical composition, traditional Chinese roofs or decorative symbols, modern decorative language, and early modernist expression [13,14]. Therefore, category assignment in this study does not rely on isolated architectural components, but instead integrates overall composition, relationships among architectural elements, and architectural–historical context. This characteristic also provides an important basis for the subsequent cross-regional target-domain testing and stylistic boundary analysis.

3.4. Model Framework and Candidate Modules

This study adopts Swin Transformer as the main backbone network (Figure 5). The selection of this backbone was not based solely on a preference for a particular model, but was informed by the feature requirements of architectural image recognition, methodological trends in existing architectural heritage recognition studies, and comparisons with several representative standard backbone networks. Images of modern educational and cultural buildings contain visual cues at multiple hierarchical levels, including overall facade proportions, roof forms, axial relationships, window rhythms, column systems, and decorative patterns. The model therefore needs to capture both local architectural components and overall compositional relationships. Compared with conventional CNNs, which primarily rely on local convolutional receptive fields, transformer-based models are better suited to modeling long-range spatial relationships and multi-scale visual features in architectural facades. Compared with the standard ViT architecture, Swin Transformer employs hierarchical window-based attention and shifted-window mechanisms, enabling it to preserve local feature modeling while enhancing information exchange across windows. This architectural design provides a potentially suitable basis for architectural image recognition tasks in which both local components and overall composition need to be considered. All baseline models were evaluated on the same six-class architectural style recognition task, and their image-level accuracy and building-level accuracy were reported on the fixed Hunan target-domain test set containing 186 images from 37 building instances (Figure 5). The pretraining status and final performance of the baseline models are reported in Section 4.1.
Figure 5. Subject-focused Swin Transformer framework and candidate modules evaluated in the study.
To provide a clear reference for backbone selection, several representative standard backbone networks were introduced as baselines, including EfficientNet-B0, DeiT-Base Distilled (DeiT), ViT-Base, and Swin Transformer-Tiny (Table 3). These models represent, respectively, a lightweight and efficient CNN, representative Vision Transformer architectures, and a hierarchical window-attention transformer. The purpose of this baseline setting was not to provide an exhaustive comparison of all CNN and transformer architectures, but to establish a reasonable reference for selecting Swin Transformer as the unified backbone and for subsequently evaluating subject-focused mechanisms built upon it (Table 3). Additional external baselines involving domain-transfer or foundation-model approaches are evaluated separately in Section 4.1 to further examine the relative difficulty of the cross-regional target-domain recognition task. These comparisons are not included in the A1–A8 ablation framework because the purpose of A1–A8 is to isolate the contribution of subject-focused modules under a unified backbone. It should be noted that ViT-Base was trained without pretrained weights because pretrained weights were unavailable under the corresponding experimental conditions. Its results are therefore treated only as an exploratory reference under a non-pretrained setting and are not used as strong evidence regarding the overall performance of the ViT architecture (Table 3).
Table 3. Preliminary backbone screening and rationale for selecting Swin Transformer under the fixed Hunan target-domain evaluation protocol.
The representative backbone results show that model performance was not fully consistent across image-level and building-level evaluation. Pure DeiT-Base Distilled achieved a higher image-level accuracy than Pure Swin Transformer-Tiny, while both models obtained the same building-level accuracy. This indicates that image-level classification performance and building-level classification based on multi-view majority voting are not necessarily synchronized. Considering the task characteristics of modern educational and cultural buildings, which require simultaneous interpretation of local architectural components and overall facade composition, together with the application of Swin Transformer in previous architectural heritage recognition studies, this study ultimately adopted Swin Transformer as the unified backbone network. On this basis, the subsequent experiments examined whether subject-focused and background-suppression mechanisms could improve building-level classification performance in the Hunan target domain (Figure 5).
Swin Transformer constructs multi-scale visual representations through a hierarchical window-based attention mechanism and enhances cross-window information exchange through shifted windows. This architecture is well suited to visual recognition tasks in which overall facade composition and local architectural components coexist. For images of modern educational and cultural buildings, it can attend to local elements such as roofs, eaves, windows, columns, and decorative patterns while retaining information about overall facade proportions, axial relationships, and massing organization. Compared with conventional convolutional networks, this structure is better suited to jointly modeling local architectural components and overall compositional relationships. This capability is particularly relevant for distinguishing closely related categories such as Western eclectic architecture, Chinese revival hybrid architecture, and early modern architecture (Figure 5).
On the basis of the backbone network, candidate modules were introduced to investigate subject-focused feature learning, boundary enhancement, and training stability. The GCSA module recalibrates features at both the channel and spatial levels, guiding the model toward semantically discriminative channels and spatial regions. In architectural images, this mechanism is designed to encourage stronger responses to the main facade, roof boundaries, and decorative components while reducing potential interference from background elements such as sky, trees, utility wires, and temporary attachments. Previous historical building recognition studies have also indicated that combining Swin Transformer with global channel–spatial attention can enhance the extraction of multi-level visual features from historical buildings (Figure 5).
L-Softmax was introduced to enlarge the angular margin between Categories, with the aim of creating clearer feature-space boundaries among adjacent or transitional style categories such as 02/04, 03/04, and 05/06. MSFF was used to fuse facade features at different scales, allowing the model to integrate information from roofs, windows, decorative elements, and overall massing. PFP/PA was introduced to further guide the model toward building-subject regions. GBVS/AGL combines saliency guidance and attention guidance to reduce interference from trees, sky, and surrounding buildings. Grad-CAM-based constraints use activation-map information as an auxiliary regularization signal to reduce excessive attention responses in background regions. These constraints are intended to encourage subject-focused feature learning, while the resulting attention patterns are interpreted only as supportive evidence of model behavior rather than direct proof of architectural reasoning (Figure 5).
EMA was used to maintain a moving average of model parameters and thereby smooth the training process, reducing fluctuations caused by mini-batch updates and potentially improving stability under cross-regional target-domain testing. Therefore, EMA-related observations are discussed only as training stability information and are not used to support conclusions regarding module effectiveness. However, EMA was not treated as an independent ablation variable in the A1–A8 experimental matrix. Its related results were used only as auxiliary observations of training stability and were not included as a basis for the main quantitative conclusions of this study (Figure 5).

3.5. Experimental Design

To investigate the contribution of different candidate mechanisms under a unified backbone network and controlled experimental conditions, eight ablation experiments (A1–A8) were designed (Table 4). The purpose of these experiments was not to identify a universally optimal model, but to analyze how different subject-focused, feature-enhancement, and boundary-optimization strategies influence cross-regional target-domain recognition. All experiments followed the same source-domain training and validation protocol and were finally evaluated on the same fixed Hunan target-domain test set, which contained 186 images from 37 building instances. The model first performed image-level classification for each building image. The image-level predictions from multiple views of the same building instance were then aggregated through multi-view majority voting to obtain the final building-level classification result. The detailed evaluation metrics are described in Section 3.6.
Table 4. Ablation experiment settings of A1–A8 under the fixed Hunan target-domain evaluation protocol.
The technical workflow of this study can be summarized as data preparation–feature extraction–candidate mechanism evaluation–training optimization–evaluation and analysis (Figure 5). Specifically, the data preparation stage included building-instance-level organization, image screening, resizing and normalization, and data augmentation. During feature extraction, Swin Transformer was used to obtain multi-scale facade representations. During module enhancement, candidate modules and constraints, including GCSA, MSFF, PFP/PA, GBVS/AGL, and Grad-CAM-based constraints, were comparatively evaluated. During training optimization, L-Softmax was introduced as a boundary-enhancement loss, while class-balancing strategies were used to reduce bias caused by uneven category distributions. During evaluation and analysis, image-level results, building-level results, and class-confusion patterns were jointly examined.
To ensure comparability among the A1–A8 ablation experiments, a unified training, validation, and evaluation protocol was adopted. The source-domain data were divided into training and validation sets at the building-instance level rather than at the individual-image level, with 20% of the building instances assigned to the validation set. All multi-view images belonging to the same building were always retained within the same data subset, thereby preventing similar images of the same building from appearing simultaneously in the training and validation sets and reducing the risk of information leakage. The same building-instance-level training/validation split was used for all A1–A8 experiments. Because the fixed Hunan target-domain set contains a limited number of building instances, final performance comparisons are supplemented with balanced metrics and bootstrap-based uncertainty estimation described in Section 3.6. Only the activation states of the corresponding candidate modules were changed, thereby controlling experimental variables beyond the module combinations as far as possible.
All input images were resized to 224 × 224 pixels, the batch size was set to 48, the maximum number of training epochs was set to 200, and the random seed was fixed at 42. The fixed random seed was used to ensure procedural consistency among A1–A8 experiments. However, uncertainty estimation of final target-domain performance is additionally considered through bootstrap analysis rather than relying solely on a single accuracy value. All models were optimized using the AdamW optimizer with the same differential learning-rate strategy. The learning rate was set to 1 × 10−5 for the pretrained Swin Transformer backbone, 2 × 10−5 for the classification head, and 4 × 10−5 for newly introduced modules. This configuration allowed the pretrained backbone to be fine-tuned with a relatively small learning rate while enabling the newly introduced classification layers and functional modules to adapt with comparatively larger learning rates. The Hunan samples were not involved in model training, validation, hyperparameter tuning, early stopping decisions, or best-model selection. Instead, they were retained throughout the entire experimental process as a fixed target-domain test set and were used only for the final evaluation of whether architectural style features learned from the non-Hunan source domain could be transferred to modern educational and cultural buildings in Hunan.
Class-balanced weighting was used to address the uneven distribution of samples across the six categories in the training set. Its basic principle is to assign greater loss weights to minority Categories, thereby reducing excessive model bias toward categories with larger sample sizes. This mechanism was introduced to mitigate category imbalance during source-domain training. However, it does not compensate for the limited number of independent building instances in the Hunan target domain. Therefore, minority-category results are interpreted cautiously and are further examined through balanced metrics, confusion patterns, and representative cases. However, class weighting cannot substitute for the inclusion of additional independent building instances. Therefore, high accuracy obtained from a limited number of minority-class samples is not interpreted in this study as evidence of stable class-level generalization. Instead, the results are evaluated jointly through building instances, confusion matrices, and interpretability analyses.
All experiments used the same learning-rate scheduling strategy, consisting of an initial linear warm-up followed by cosine annealing with restarts. Because different module combinations may exhibit different convergence rates, an early-stopping mechanism was adopted with a patience of 30 epochs. Training was terminated if validation performance did not improve for 30 consecutive epochs. During training, the training loss, validation loss, training accuracy, and validation accuracy were recorded for each epoch. The model parameters corresponding to the highest source-domain validation accuracy were selected as the best model, ensuring that Hunan target-domain information was not involved in model selection.
After best-model selection, the selected model was evaluated once under the unified protocol on the fixed Hunan target-domain test set. Image-level accuracy and building-level accuracy were reported separately. Building-level classification was obtained by aggregating the image-level predictions from multiple views of the same building through multi-view majority voting. By strictly separating model training and selection from target-domain testing, this study sought to minimize the evaluation bias caused by the use of Hunan test information during model optimization and to reduce potential evaluation bias and allow comparisons among A1–A8 to more directly reflect differences associated with the candidate module combinations.

3.6. Evaluation Metrics and Statistical Reliability Assessment

To avoid relying solely on overall accuracy, which may be affected by category imbalance and may obscure differences at the building-instance level, this study evaluates model performance using complementary metrics at both image and building levels. The evaluation framework includes image-level accuracy, building-level voting accuracy, balanced accuracy, macro-F1, bootstrap confidence intervals, confusion matrices, and boundary-pair statistics. Image-level evaluation measures the model’s classification ability for individual architectural photographs, while building-level voting integrates predictions from multiple images belonging to the same building instance, which is more consistent with the objective of architectural style attribution. Confusion matrices and boundary-pair statistics are further used to analyze misclassification patterns among stylistically adjacent categories, particularly the 02/04, 03/04, and 05/06 boundary pairs
Let N i m g denote the number of test images in the fixed Hunan test set. For the (i)-th image x i , the predicted class is denoted as y ˆ i , and the true class is denoted as ( y i ). Image-level accuracy is defined as per Equation (1):
A c c i m g = 1 N i m g i = 1 N i m g I y ^ i = y i ,
where N i m g is the total number of test images, and I is the indicator function. It equals 1 when the predicted class is the same as the true class and 0 otherwise. This metric reflects the overall classification ability of the model for individual architectural photographs.
However, overall accuracy may be influenced by uneven category distributions because categories with larger sample sizes contribute more strongly to the final value. Therefore, balanced accuracy is additionally calculated to evaluate whether the model maintains comparable recognition ability across different style categories. Balanced accuracy is defined as Equation (2):
B a l a n c e d A c c u r a c y = 1 K k = 1 K R e c a l l k
where K represents the number of style categories and R e c a l l k represents the recall value of category k . This metric provides a category-balanced evaluation by assigning equal importance to each style category.
In addition, macro-F1 is reported to further evaluate category-level performance. Unlike weighted F1-score, macro-F1 assigns equal importance to each category and therefore reduces the influence of dominant categories with larger sample numbers. This is defined as per Equation (3):
M a c r o F 1 = 1 K k = 1 K F 1 k
where F 1 is defined as per Equation (4):
F 1 k = 2 P k R k P k + R k
where P k and R k   denote the precision and recall of category, respectively. As one building usually contains multiple images from different viewpoints, shooting times, or local components, the prediction result of a single image cannot be regarded as fully equivalent to the final style attribution of the building instance. Therefore, this study further adopts a building-level majority voting method. Let B b denote the image set corresponding to the (b)-th building instance. The voted prediction of this building is defined as Equation (5):
Y ^ b = m o d e y ^ i x i B b
where x i denotes an image belonging to building instance B b , mode denotes the majority voting function, and ( y ^ b ) represents the final predicted class of the (b)-th building after voting. Therefore, Equation (2) describes the aggregation process of generating building-level classification results from image-level predictions. When two or more categories achieve the same maximum number of votes, this paper further calculates the average prediction probability for all viewing angles of the building across candidate categories and selects the category with the highest average prediction probability as the final building-level classification result.
Let N b l d denote the total number of building instances in the fixed Hunan test set, and let Y b denote the ground-truth class of the (b)-th building. Building-level accuracy is defined as per Equation (6):
A c c b l d = 1 N b l d b = 1 N b l d I Y ˆ b = Y b
Because the fixed Hunan target-domain test set contains only 37 building instances, a single accuracy value may not fully represent the uncertainty of model performance. Therefore, bootstrap resampling is adopted to estimate confidence intervals for the main evaluation metrics.
Specifically, building instances in the Hunan target-domain test set are repeatedly sampled with replacement, and the corresponding evaluation metrics are recalculated in each iteration. A total of 1000 bootstrap iterations are performed. The 95% confidence interval is obtained from the 2.5th and 97.5th percentiles of the bootstrap distribution:
C I 95 % = P 2.5 , P 97.5
where P 2.5 and P 97.5 represent the lower and upper percentile boundaries of the bootstrap distribution, respectively.
Building-level accuracy can reduce accidental misclassification caused by viewpoint variation, local occlusion, lighting changes, or missing stylistic components in a single image. It is therefore more consistent with the practical need of architectural research to identify the style of a building instance. This study reports both image-level accuracy and building-level accuracy. The difference between them is used as an auxiliary indicator of building-level voting stability, as shown in Equation (8):
G a i n b l d i m g = A c c b l d A c c i m g
When G a i n b l d i m g > 0 , it indicates that the building-level accuracy exceeds the image-level accuracy, suggesting improved instance-level classification performance after multi-view aggregation; when G a i n b l d i m g < 0 , it signifies that the building-level accuracy is lower than the image-level accuracy, indicating that the good performance in single-image classification has not translated into higher building-instance classification accuracy. As the image-level and building-level accuracies correspond to different statistical units, this difference serves only as a descriptive auxiliary indicator and should not be regarded as a direct measure of multi-view prediction consistency.
In addition to accuracy, this study uses confusion matrices to analyze the direction of misclassification among categories. Let C p q denote the number of images whose true class is p but whose predicted class is q. The confusion matrix can be represented as Equation (9):
C = C p q , p , q 1,2 , , K = 6
where (K = 6), corresponding to the six style categories of modern educational and cultural buildings used in this study. The confusion matrix is used not only to observe the overall misclassification structure, but also to identify boundary categories with architectural–historical significance.
For each key boundary pair p , q , this study further calculates its bidirectional confusion count, as shown in Equation (10):
B p / q = C p q + C q p
where B p / q denotes the total boundary confusion between categories (p) and (q). This study focuses on the 02/04, 03/04, and 05/06 boundary pairs because they respectively correspond to the transitional relationships between Western eclectic architecture and Chinese revival hybrid architecture, Chinese revival palace-style architecture and Chinese revival hybrid architecture, and Chinese revival modern-style architecture and early modern “New Architecture”, respectively.
It should be noted that although precision, recall, and F1-score are common classification evaluation metrics, the limited number of building instances in categories 03,05, and 06 within the Hunan test set means that the correct or incorrect classification of individual instances can significantly impact category-level metrics. Therefore, this study does not interpret individual category-level precision, recall, or F1-score values from extremely small categories as independent evidence of stable class-level generalization. Instead, macro-level metrics, bootstrap uncertainty estimation, image-level versus building-level comparisons, multi-view aggregation results, confusion matrices, and representative cases are jointly considered to avoid overinterpretation under small-sample conditions. This methodology helps avoid overinterpretation of unstable category-level metrics under small-sample conditions while better aligning with the emphasis on boundary samples and representative cases in architectural heritage style interpretation.
To improve the interpretability of Grad-CAM analysis and avoid relying solely on subjective visual inspection, this study further established an architectural diagnostic-region annotation procedure. The diagnostic regions were defined according to the morphological criteria summarized in Supplementary Table S2, including roof morphology, facade composition, opening rhythm, veranda and column systems, decorative components, and Sino-Western compositional relationships. Based on established architectural–historical studies of modern Chinese architecture, two researchers with architectural backgrounds independently annotated the diagnostic regions of representative cases. Any disagreements were discussed and resolved to obtain the final annotation results.
The manually annotated regions were subsequently compared with the predicted-class Grad-CAM attention maps. Let R denote the manually annotated architectural diagnostic region and A x , y denote the normalized Grad-CAM activation intensity at spatial location x , y . The ROI attention energy was calculated as the proportion of total Grad-CAM activation distributed within the expert-defined diagnostic region in Equation (11):
E R O I = x , y R A x , y x , y I A x , y
where I represents the entire image domain. A higher ROI attention energy indicates that a larger proportion of model attention is concentrated on architecturally meaningful regions.
To further evaluate spatial correspondence, the binary activation regions generated from the top 20% Grad-CAM responses were compared with manually annotated ROIs. Intersection-over-union (IoU) was calculated as per Equation (12):
I o U = R G R G
where G represents the high-activation region extracted from the Grad-CAM map.
In addition, activation coverage and enrichment ratio were calculated to examine whether diagnostic regions contained disproportionately higher activation compared with non-diagnostic areas in Equations (13) and (14):
C o v e r a g e = R G R
E R = E R O I R / I
These indicators were used not to prove that the model possesses architectural cognition, but to examine whether model attention patterns were spatially consistent with expert-defined stylistic evidence.

4. Results

4.1. Representative Backbone Baselines and Overall Target-Domain Classification Performance

Before conducting the A1–A8 ablation experiments on subject-focused modules, several representative standard backbone networks were first introduced as baseline references to examine their basic performance on the fixed Hunan target-domain test set (Table 5). The purpose of this baseline setting was not to provide an exhaustive comparison of all CNN and transformer architectures, but to establish an experimental reference for the subsequent selection of Swin Transformer as the unified backbone network. All baseline models addressed the same six-class style classification task for modern educational and cultural buildings and were evaluated under the same fixed Hunan target-domain testing protocol, using a test set containing 186 images from 37 building instances. In addition to image-level accuracy and building-level accuracy based on multi-view majority voting, balanced accuracy, macro-F1, and bootstrap confidence intervals were calculated for the selected main model to further evaluate performance stability under the imbalanced fixed Hunan target-domain test condition (Table 5).
Table 5. Performance of representative backbone baselines and the A7 subject-focused model on the fixed Hunan target-domain test set.
The baseline results demonstrate that different visual backbone architectures exhibit distinct advantages in image-level classification and building-level classification. Among the compared pure backbone models, Pure Swin Transformer-Tiny achieved the highest image-level test accuracy of 52.15%, surpassing Pure DeiT-Base Distilled’s 49.46% (Table 5). In terms of specific sample counts, both models correctly classified 97 and 92 test images respectively. By contrast, Pure DeiT-Base Distilled achieved a building-level accuracy of 51.35%, slightly higher than Pure Swin Transformer-Tiny’s 48.65%, with correct classifications of 19 and 18 building instances respectively (Table 5). This indicates that the two models differ by only one building instance in building-level performance, suggesting their pure backbone architectures perform similarly under the current fixed Hunan object-domain testing protocol, yet each demonstrates unique strengths in single-image discrimination and multi-view aggregation at the building-instance level (Table 5).
These results also demonstrate that the optimal validation accuracy in the source domain, image-level accuracy in the target domain, and building-level accuracy in the target domain are not entirely consistent. Pure DeiT-Base Distilled achieves higher validation accuracy and slightly better building-level accuracy, whereas Pure Swin Transformer-Tiny performs more effectively in image-level classification within the target domain. Therefore, this paper does not rank backbone networks based on any single accuracy metric nor interprets baseline results as indicating that Swin Transformer outperforms DeiT across all aspects. The purpose of representative backbone network experiments is to provide a performance benchmark for selecting appropriate backbone architectures, rather than to comprehensively evaluate the merits of different CNN and transformer architectures.
To further evaluate the difficulty of cross-regional architectural style recognition and compare the proposed subject-focused framework with external transfer-oriented baselines, additional experiments were conducted using foundation-model-based and domain-adaptation-related approaches. These comparisons were performed under the same fixed Hunan target-domain evaluation protocol, without using Hunan labels during training or adaptation (Table 6).
Table 6. Comparison with external baselines and statistical reliability metrics on the fixed Hunan target-domain test set.
The comparison indicates that external foundation-model-based approaches can achieve competitive image-level performance under the fixed Hunan target-domain condition (Table 6). Frozen DINOv2 with a source-only linear probe obtained an image-level accuracy of 64.52% and a building-level accuracy of 75.68%, exceeding the current A7 main model in both metrics. However, these results should be interpreted differently from the proposed subject-focused framework because the foundation-model baseline benefits from large-scale pretraining and does not aim to explicitly model architectural stylistic components.
In contrast, the purpose of A7 is not to achieve the highest numerical accuracy through large-scale generic visual representation, but to investigate whether architectural subject-focused mechanisms can improve cross-regional style recognition under a controlled backbone-based framework. Therefore, A7 remains the main model for analyzing subject-focused attention mechanisms, while external baselines provide an additional reference for understanding the difficulty and upper-bound potential of the target-domain recognition task.
Based on comprehensive experimental results and the feature requirements of architectural image recognition tasks, this study ultimately adopted the Swin Transformer as the unified backbone network for the A1–A8 ablation experiments. The rationale is as follows: Firstly, Pure Swin Transformer-Tiny achieved the highest image-level accuracy in the target domain among pure backbone models, indicating relatively stronger image-level discrimination capability among the tested pure backbone models under the current target-domain evaluation protocol. Secondly, the Swin Transformer generates multi-level feature representations through hierarchical window attention mechanisms and shift window techniques, enabling simultaneous processing of local components as well as global compositional elements [34]. Thirdly, candidate modules such as MSFF, GCSA, PFP/PA, and GBVS/AGL require integration with multi-stage feature architectures, making the Swin Transformer’s hierarchical feature architecture particularly suitable as a unified ablation platform. Finally, prior research on architectural heritage recognition has demonstrated the potential of the Swin Transformer and its attention-enhanced architecture for classifying historical and traditional buildings [35,36].
Under the same fixed Hunan target-domain testing protocol, A7 introduced the GBVS/AGL subject-focused mechanism into the Swin Transformer backbone. Compared with Pure Swin Transformer-Tiny, its image-level test accuracy increased from 52.15% to 56.45%, corresponding to an increase of 4.30 percentage points, while its building-level test accuracy increased from 48.65% to 64.86%, corresponding to an increase of 16.21 percentage points. To further evaluate whether the improvement was affected by category imbalance, A7 was additionally assessed using balanced accuracy and macro-F1. The obtained balanced accuracy and macro-F1 values were 65.63% and 57.51%, respectively. These complementary metrics indicate that the performance improvement should be interpreted together with category-level balance rather than overall accuracy alone. The larger difference observed at the building level suggests that, under the current experimental conditions, the subject-focused mechanism may be particularly beneficial to building-level classification after multi-view majority voting. However, because image-level and building-level accuracy are calculated using different statistical units, this difference should not be directly interpreted as a quantitative improvement in multi-view prediction consistency. Moreover, the result is specific to the current fixed target-domain testing protocol and does not imply that A7 has achieved mature high-precision six-class classification. Instead, it should be regarded as exploratory evidence of the potential value of subject-focused mechanisms in cross-regional architectural style classification under complex background conditions.
Further comparison of the A1–A8 ablation experiments show that the current models do not support a strong claim of high-precision six-class classification (Figure 6). Because the fixed Hunan target-domain test set contains only a limited number of independent building instances in Categories 03, 05, and 06, additional analysis was conducted on the relatively better-supported Categories 01, 02, and 04. The detailed results are provided in Supplementary Table S1. This analysis aims to examine whether the observed differences among A1–A8 remain consistent when the influence of extremely small categories is reduced. Nevertheless, the results provide evidence that some subject-focused and background-suppression configurations may have positive effects on cross-regional architectural style classification. All A1–A8 experiments were evaluated on the same fixed Hunan target-domain test set containing 186 images from 37 building instances. Clear differences were observed between image-level accuracy and building-level accuracy across the different module combinations (Figure 6). The complete building-level evaluation results, including balanced accuracy, macro-precision, macro-recall, macro-F1, and bootstrap-based confidence intervals for all A1–A8 configurations, are provided in Supplementary Table S6.
Figure 6. A1–A8 ablation comparison showing image-level and building-level performance.
At the image level, A8 achieved the highest image-level accuracy of 58.60%. This result suggests that Grad-CAM-related background constraints may provide some benefit for individual-image classification by encouraging the model to attend more strongly to building-subject regions in certain samples. However, the building-level accuracy of A8 was only 54.05%, which was lower than its image-level accuracy. This indicates that its advantage in image-level classification did not translate into equally strong building-level performance after multi-view majority voting (Figure 6).
At the building level, A4, A5, and A7 all achieved the highest building-level accuracy of 64.86%. These results indicate that, under the current experimental conditions, configurations involving subject priors, saliency guidance, or subject-focused mechanisms can achieve stronger building-level classification performance after multi-view majority voting. Compared with image-level classification of individual photographs, building-level classification is more closely aligned with the practical objective of architectural research, which is to determine the stylistic attribution of an entire building rather than the category of a single image (Figure 6).
Although A4, A5, and A7 achieved the same building-level accuracy, their category-balanced performance was not identical. Supplementary Table S6 shows that A7 obtained the highest balanced accuracy (65.63%) among the A1–A8 configurations, indicating relatively stronger performance consistency across the six categories under the imbalanced Hunan target-domain condition. In contrast, some models with comparable overall accuracy exhibited lower macro-level metrics, suggesting that their performance was more influenced by dominant categories. Therefore, building-level accuracy alone is insufficient to fully characterize model behavior, and balanced metrics and macro-averaged indicators are considered together in the selection of the main model.
Considering image-level accuracy, building-level accuracy, balanced metrics, bootstrap uncertainty estimation, model complexity, and performance on small-sample categories, A7 was selected as the main model for the current analysis. Bootstrap analysis further indicated that the building-level performance of A7 should be interpreted within an uncertainty range rather than as a single deterministic value. The confidence interval results are provided in Table S4. A7 introduced the GBVS/AGL subject-focused mechanism into the base Swin framework and achieved an image-level accuracy of 56.45% and a building-level accuracy of 64.86%. It also correctly classified the only Category 03 building instance in the Hunan test set. By contrast, although A8 achieved the highest image-level accuracy, its building-level performance was less stable. Therefore, A8 is retained as the best image-level result and as supplementary evidence regarding the effect of Grad-CAM-related constraints, rather than as the main model for building-level analysis (Figure 6).

4.2. Domain Shift Quantification and In-Distribution Comparison

The comparison results demonstrate a clear performance difference between geographically held-out evaluation and in-distribution evaluation. Under the geographically held-out Hunan evaluation, the building-level accuracy, balanced accuracy, and macro-F1 were 48.65%, 41.73%, and 38.49%, respectively. By contrast, when source-domain and Hunan samples were pooled under a random building-level split, the corresponding values increased to 59.74 ± 4.31%, 59.59 ± 5.48%, and 58.46 ± 4.95%, respectively (Table 7).
Table 7. Comparison between geographically held-out target-domain evaluation and in-distribution evaluation under different statistical metrics.
This difference indicates that regional separation introduces additional difficulty beyond the intrinsic complexity of six-class architectural style recognition. When Hunan samples are included in the training and testing distribution simultaneously, the model benefits from access to region-specific visual characteristics and local stylistic variations. However, under the geographically held-out protocol, the model must transfer stylistic representations learned from non-Hunan regions to buildings shaped by different historical backgrounds, construction traditions, and photographic conditions. Therefore, the lower performance under the fixed Hunan target-domain evaluation should be interpreted as evidence of domain shift rather than simply insufficient model capacity.
To further visualize the distributional difference between source-domain and target-domain architectural images, PCA visualization was conducted using building-level mean features extracted from a fixed ImageNet-pretrained ResNet-50 feature extractor. This shows the projected feature distributions of source-domain and Hunan target-domain samples (Figure 7). Although considerable overlap exists between the two domains, the target-domain samples are not completely embedded within the source-domain distribution, indicating the existence of regional feature differences (Figure 7).
Figure 7. Feature-distribution difference between source domain and Hunan target domain.
It should be noted that the PCA visualization is used only as a qualitative auxiliary analysis rather than a quantitative measurement of domain discrepancy. The actual magnitude of domain shift should be evaluated through quantitative indicators such as feature-distance measurements or domain-classifier-based metrics. Therefore, this is interpreted as visual evidence supporting the existence of source–target distribution differences rather than as a standalone statistical proof of domain shift (Figure 7).
These findings support the necessity of the cross-regional target-domain validation strategy adopted in this study. Compared with random-split evaluation, geographically isolated testing provides a more challenging but more realistic scenario for evaluating whether architectural style features can be transferred across regions. Accordingly, the subsequent analysis focuses not only on improving classification accuracy, but also on understanding how subject-focused mechanisms, background suppression, and multi-view aggregation influence model robustness under regional domain differences.

4.3. Effects of Subject-Focused Mechanisms Under Cross-Regional Target-Domain Evaluation

The comparison between A1 and A7 provides evidence regarding the potential contribution of the GBVS/AGL subject-focused mechanism under the current cross-regional target-domain evaluation protocol. Compared with the baseline configuration, A7 achieved higher image-level accuracy (56.45% vs. 53.76%) and building-level accuracy (64.86% vs. 62.16%) (Figure 6). These differences indicate that the introduction of subject-focused guidance was associated with improved classification performance under complex target-domain image conditions; however, the results should be interpreted as comparative evidence under the current experimental setting rather than as proof of the universal superiority of the mechanism. Under the fixed target-domain testing protocol used in this study, and in the presence of complex background factors such as trees, sky, roads, surrounding buildings, and partial occlusion in images of modern educational and cultural buildings in Hunan, A7 achieved higher image-level and building-level accuracy than A1. This result suggests that the GBVS/AGL subject-focused mechanism may help improve target-domain classification performance under complex background conditions.
To avoid evaluating ablation models solely through overall accuracy, Table 1 reports complementary statistical indicators, including balanced accuracy, macro precision, macro recall, macro-F1, and bootstrap-based confidence intervals. These metrics provide additional information regarding category-balanced performance and uncertainty under the imbalanced fixed Hunan target-domain test condition.
A4, A5, and A7 all achieved the highest building-level accuracy of 64.86%. However, the complementary balanced metrics show differences among these configurations. A7 obtained the highest balanced accuracy (65.63%) and macro-F1 value (57.51%) among the tested ablation models, suggesting relatively better category-balanced performance under the imbalanced target-domain condition (Table 8). These results indicate that building-level accuracy alone is insufficient to characterize model behavior, and that balanced metrics provide additional evidence for evaluating performance across categories.
Table 8. Statistical comparison of A1–A8 ablation models on the fixed Hunan target-domain test set.
Although A7 achieved a numerically higher building-level accuracy than A1, the paired bootstrap confidence interval and McNemar tests did not indicate a statistically significant difference between the two models (Table 9). Therefore, A7 should not be interpreted as statistically superior to A1 based solely on accuracy improvement. Instead, its selection as the main model is based on the combined consideration of balanced performance, model simplicity, building-level behavior, and interpretability under the current target-domain evaluation setting.
Table 9. Paired building-level statistical comparisons between principal ablation models.
However, the experiments also show that adding more modules does not necessarily lead to better performance. A3 achieved an image-level accuracy of 51.61% and a building-level accuracy of 51.35%, corresponding to a building-level accuracy change of −0.26 percentage points relative to image-level accuracy (Figure 8). This indicates that MSFF alone did not consistently improve building-level classification under the current evaluation protocol. One possible explanation is that additional multi-scale feature aggregation may increase sensitivity to local visual variations in transitional stylistic cases; however, this interpretation requires further investigation using larger datasets and additional interpretability analysis. A2 achieved a building-level accuracy of 59.46%, which was higher than its image-level accuracy; however, its improvement in distinguishing adjacent stylistic boundaries was not consistent. In particular, the confusion between Categories 02 and 04 remained pronounced.
Figure 8. Difference between building-level and image-level accuracy across A1–A8 configurations.
These differences describe the relationship between image-level prediction and building-level aggregation behavior rather than direct quantitative improvements in prediction accuracy. Because image-level and building-level accuracy are calculated using different statistical units, the observed changes should be interpreted as descriptive evidence of multi-view aggregation effects rather than as independent performance gains. Among them, A5 showed the largest positive change, at +8.95 percentage points, while A7 showed a positive change of +8.41 percentage points (Figure 8). However, because image-level accuracy and building-level accuracy are calculated using different statistical units, these differences should be interpreted descriptively and should not be directly equated with quantitative improvements in multi-view prediction consistency.
By contrast, although A8 achieved the highest image-level accuracy, its building-level accuracy was 4.55 percentage points lower than its image-level accuracy (Figure 8). This suggests that the Grad-CAM-related configuration may enhance certain single-image classification patterns, but the resulting predictions were less consistent after multi-view aggregation in this experiment. Accordingly, considering its relatively concise model structure, balanced metric performance, building-level classification behavior, bootstrap uncertainty estimation, and interpretability advantages, A7 was selected as the main model for subsequent boundary-case and attention analysis. This selection does not imply statistically significant superiority over every other configuration, but rather reflects the most balanced performance profile under the current cross-regional target-domain evaluation condition.

4.4. Boundary Case Analysis

Based on the archived results of the A1–A8 experiments and the confusion matrix of A7, the three stylistic boundary pairs that warrant the greatest attention are 02/04, 03/04, and 05/06 (Figure 9). These confusion patterns should not be interpreted simply as model failures. Instead, they provide exploratory evidence for examining stylistic transitions and ambiguous boundaries within the classification system. However, these interpretations are based on the current target-domain samples and representative cases, rather than constituting independent validation of architectural style boundaries.
Figure 9. Confusion matrix and boundary-pair analysis of A7 highlighting stylistically adjacent categories (02/04, 03/04, and 05/06).
(1)
Category 02, Western eclectic architecture, and Category 04, Chinese revival hybrid architecture, constitute the most prominent confusion boundary. In the image-level confusion matrix of A7, 20 images from Category 02 were classified as Category 04, while 14 images from Category 04 were classified as Category 02, resulting in a total of 34 bidirectional confusion cases for the 02/04 boundary. This was the most pronounced confusion relationship among the key boundary pairs. When Western Eclectic buildings incorporate Chinese-style roofs, eaves, traditional motifs, or locally hybridized components, the model may treat these visually salient traditional elements as dominant cues and consequently classify the buildings as Chinese revival hybrid architecture. Conversely, when Western classical composition, column systems, pediments, or facade order are more prominent in Category 04 buildings, the model may classify them as Category 02. Therefore, the 02/04 confusion can be interpreted as reflecting an important transitional relationship between “Western eclecticism as the dominant framework with partial incorporation of Chinese symbolic elements” and “traditional revival as the dominant framework combined with Western compositional features” within the operational classification system used in this study (Figure 9). This interpretation should be understood as a morphological and historical explanation of model confusion rather than a definitive reclassification of individual buildings.
(2)
A clear boundary also exists between Category 03, Chinese revival palace-style architecture, and Category 04, Chinese revival hybrid architecture. The A7 confusion matrix shows that five images from Category 03 were classified as Category 04, while two images from Category 04 were classified as Category 03, resulting in seven bidirectional confusion cases for the 03/04 boundary. Both categories may contain large roofs, eaves, traditional ornamentation, and axial composition, making it difficult for the model to distinguish consistently between palace-style revival and hybrid revival. Unlike the 02/04 confusion, the 03/04 confusion primarily reflects an internal subdivision within Chinese traditional revival architecture. The key question is whether traditional roofs and palace-style symbols play a dominant compositional role or appear only as partial elements within a hybrid facade expression (Figure 9). Because Category 03 contains only one building instance in the Hunan test set, the observed confusion pattern should be considered as a case-based indication of stylistic ambiguity rather than a statistically stable category-level conclusion.
(3)
Category 05, Chinese revival modern-style architecture, and Category 06, the “New Architecture” of the modern period, namely early modern architecture, form another important boundary. The A7 confusion matrix shows that no Category 05 images were classified as Category 06, whereas five Category 06 images were classified as Category 05, resulting in five confusion cases for the 05/06 boundary. This indicates that some Category 06 images showed visual similarity with Chinese revival modern-style architecture under the current feature representation, leading the model toward Category 05 predictions. One possible reason is that both categories may exhibit relatively simplified massing, more regular window organization, and reduced ornamentation, making it difficult for the model to determine whether traditional symbols still play a dominant role. From an architectural–historical perspective, this confusion should not be regarded solely as a technical error; it also reflects the continuity of the transition from traditional revival toward modern architectural expression (Figure 9).
It is worth noting that the 03/06 pair was included in this study as a control boundary rather than as a primary boundary for discussion. In the A7 confusion matrix, the number of bidirectional confusion cases between Categories 03 and 06 was zero. This suggests that the most important confusion patterns identified in this study were concentrated among stylistically adjacent or strongly transitional categories, rather than being distributed equally across all category pairs.
Representative cases further illustrate these boundary issues (Table 10). These examples were selected because they represent typical confusion patterns observed in the A1–A8 experiments. They are not intended to provide additional quantitative validation, but rather to support architectural interpretation of the confusion relationships identified by the model. The ground-truth category of both HN013 and HN027 is Category 02, yet both tended to be classified as Category 04 across multiple model configurations, indicating that their facade composition and decorative language occupy a difficult boundary between Western eclectic architecture and Chinese revival hybrid architecture (Table 10). HN030 also belongs to Category 02, but A7 classified it correctly, whereas A5 and A8 tended to classify it as Category 04. This suggests that the GBVS/AGL subject-focused mechanism helped recover the correct building-level classification for this particular building instance. HN031 is the only Category 03 building instance in the Hunan test set and was correctly classified by both A7 and A8 (Figure 10). However, because the category contains only one building instance, this result cannot be interpreted as evidence of stable class-level generalization. HN022 and HN103 illustrate the difficulty of classifying Category 06, early modern architecture. HN022 tended to be classified as Category 04 or Category 05 across different model configurations, reflecting an unstable boundary between early modern architecture, Chinese revival hybrid architecture, and Chinese revival modern-style architecture. By contrast, HN103 was correctly classified as Category 06 by A8, suggesting that Grad-CAM-related constraints may have helped the model attend to relevant modern architectural features in this specific case (Figure 10).
Table 10. Representative boundary cases selected from the Hunan target-domain test set for architectural interpretation of model confusion.
Figure 10. Representative architectural boundary cases illustrating model confusion patterns and corresponding morphological interpretations.

4.5. Interpretation of Rare-Class Results

The results for Categories 03, 05, and 06 should be interpreted with caution when using small sample sizes in the target domain. On the fixed Hunan test set, Class 03 and Class 05 each contain only one building instance, while Class 06 contains merely two. As the correct or incorrect classification of individual buildings significantly impacts the building-level performance for their respective Categories, these results are therefore not presented as evidence of stable category-level generalization capability. Instead, they are treated as representative case observations used to examine model behavior under small-sample target-domain conditions. Specifically, Class 03 is represented by HN031, Class 05 by HN100, and Class 06 consists of HN022 and HN103. The following sections further illustrate the support provided by target-domain samples for these sparse categories, as well as the differences in architectural-level classification results and attention responses across various model configurations (Figure 11).
Figure 11. Small-sample target-domain cases and model-sensitive recognition behavior for Categories 03, 05, and 06. (a) Sparse target-domain support in the fixed Hunan test set. (b) Building-level outcome matrix for the selected rare-class cases under A5, A7, and A8. (c) Model-sensitive attention evidence comparing A7 and A8.
For Category 03, both A7 and A8 correctly classified HN031, indicating that the models were able to capture certain visual characteristics of Chinese revival palace-style Architecture in this specific case (Figure 11). The model showed relatively clear responses to roof hierarchy, the central composition of the main facade, and traditional revival architectural components. However, because Category 03 contains only one building instance in the Hunan test set, this result should be regarded as an individual successful case rather than as statistical evidence of stable classification ability for the category.
For Category 05, HN100 is the sole example of a China traditional revival modern-style building in the fixed Hunan test set, comprising a total of seven test images. At the architectural level, A5, A7, and A8 all classified HN100 as Category 05; specifically, A5 and A7 achieved 5/7 correct images, while A8 achieved 6/7 correct images. These results demonstrate that, for this specific case, the model can to some extent identify the combination relationship between modern scale, simplified traditional patterns, and modern decorative expressions. Notably, when compared with Category 06 early modern architecture, HN100 retains distinct characteristics of traditional revival modernism, enabling it to achieve consistent multi-view prediction accuracy across A5, A7, and A8 (Figure 11).
Nevertheless, the correct classification of HN100 should not be overinterpreted. Because Category 05 contains only one building instance in the Hunan test set, the relatively stable results for HN100 indicate only that the models performed well in this individual case. They do not demonstrate that the models can reliably distinguish all examples of Chinese revival modern-style architecture. In other words, the Category 05 result should be understood as a positive case under small-sample target-domain conditions rather than as evidence of reliable category-level generalization. Future research should include additional Category 05 building instances from Hunan and other comparable regions to examine whether the models can consistently recognize the relationship among modern massing, the translation of traditional symbols, and decorative simplification in Chinese revival modern-style architecture.
By contrast, the results for Category 06 show greater instability. HN022 was frequently misclassified as Category 04 or Category 05 across multiple model configurations, indicating that early modern architecture with transitional visual characteristics may still be drawn by the model toward traditional revival-related categories (Figure 10). The results for HN022 in Figure 10 show that, even though some individual images were classified as Category 06 by A8, multi-view majority voting did not recover the correct building-level classification, indicating insufficient multi-view prediction consistency across different images of the same building. HN103 presents a different pattern: it was misclassified as Category 05 by A5 and A7 but was correctly classified as Category 06 by A8 (Figure 11). This suggests that Grad-CAM-related constraints may, in some individual early modern architecture cases, help the model attend to simplified facades, modern massing, or reduced ornamentation. However, this observation remains exploratory and requires validation with a larger number of independent building instances.
Overall, the small-sample results for Categories 03, 05, and 06 indicate that the current model is more suitable for auxiliary screening and boundary-case analysis than for direct high-confidence automatic classification across all six style categories. This conclusion is consistent with the statistical limitations caused by the limited number of independent building instances in Categories 03, 05, and 06. The correct classification of HN031 and HN100 shows that the model can capture certain architectural style cues in individual sparse-category cases, whereas the differing performances of HN022 and HN103 indicate that early modern architecture remains one of the less stable categories in the current classification system. The most reliable conclusion of this section is therefore not that the model has achieved mature six-class automatic classification, but that some subject-focused configurations can improve building-level classification performance to a certain extent, while the confusion matrix and representative cases reveal architecturally meaningful stylistic boundaries among 02/04, 03/04, and 05/06.

5. Discussion

5.1. Methodological Implications and Statistical Interpretation

The experimental results of this study indicate that, in the style classification of modern educational and cultural buildings, a model must learn not only “what visual style a building exhibits”, but also “which regions of the image actually belong to the building subject.” Compared with general image classification tasks, field photographs of historical buildings often contain more complex backgrounds, including trees, sky, roads, surrounding buildings, pedestrians, signage, and temporary attachments. These elements may occupy substantial portions of an image and may form spurious correlations with geographical location, photographic context, or conservation condition. If a model relies excessively on such non-building information, its classification results may deviate from the architectural features that genuinely define style.
Beyond overall accuracy, the statistical evaluation in this study indicates that model reliability should be interpreted through multiple complementary indicators. Supplementary Table S1 further provides a category-supported evaluation by excluding the extremely small-sample Categories 03, 05, and 06. The results show that the relative ranking of A1–A8 configurations may vary when evaluated under different category-support conditions, indicating that model comparison should consider not only overall accuracy but also the distribution of available building instances. The detailed A1–A8 comparison in Supplementary Table S6 further demonstrates that models with similar building-level accuracy may exhibit different levels of category-balanced performance, emphasizing the necessity of considering multiple evaluation metrics under imbalanced target-domain conditions. The differences between image-level accuracy, building-level accuracy, balanced accuracy, and macro-F1 demonstrate that a single performance value cannot fully describe model behavior under imbalanced target-domain conditions. Bootstrap confidence intervals further provide an estimation of uncertainty caused by the limited number of independent building instances. Therefore, the reported results should be understood as performance evidence under a specific cross-regional evaluation setting rather than as a definitive measurement of universal classification capability.
The representative backbone baseline results further demonstrate that the selection of Swin Transformer was not an isolated decision. EfficientNet-B0 showed relatively weak performance on the fixed Hunan target domain, suggesting that a lightweight CNN may have difficulty fully capturing complex facade composition and transitional stylistic relationships in the cross-regional fine-grained architectural style classification task considered here. DeiT-Base Distilled and Swin Transformer-Tiny achieved comparable building-level performance, indicating that transformer-based models showed relatively stable performance in this building-level classification task under the current evaluation protocol. Because Swin Transformer employs hierarchical window-based attention and can model both local architectural components and broader spatial relationships, it is well aligned with the characteristics of this task, in which roofs, eaves, windows, column systems, decorative patterns, and overall facade proportions jointly contribute to stylistic interpretation. The purpose of this study is not to demonstrate that one general-purpose visual backbone is absolutely superior to all others, but rather to examine, within a Swin Transformer framework that is compatible with architectural interpretation, whether subject-focused and background-suppression mechanisms can further improve building-level classification performance.
Accordingly, the value of subject-focused mechanisms lies not only in potential improvements in accuracy, but also in making the subject–background relationship explicit within architectural image classification. The A1–A8 ablation experiments show that A4, A5, and A7 all achieved a building-level accuracy of 64.86%, suggesting that subject priors, saliency guidance, and subject-focused mechanisms may provide a certain stabilizing effect at the building-instance level. Among these configurations, A7 introduced the GBVS/AGL subject-focused mechanism into the base Swin framework and achieved an image-level accuracy of 56.45% and a building-level accuracy of 64.86%; it was therefore selected as the main model for the subsequent qualitative analysis because it provided a relatively balanced combination of image-level performance, building-level performance, statistical reliability, and model simplicity under the current cross-regional target-domain condition. Under the present cross-regional target-domain setting, dataset scale, and experimental conditions, this relatively concise subject-focused configuration achieved a better balance among image-level performance, building-level performance, and model complexity.
To further examine whether model attention corresponds to architecturally meaningful evidence, this study established an architectural diagnostic-region annotation procedure and compared manually defined morphological regions with predicted-class Grad-CAM attention maps. The diagnostic regions were annotated according to the operational architectural criteria summarized in Supplementary Table S2, including roof morphology, facade composition, opening rhythm, veranda and column systems, decorative components, and Sino-Western compositional relationships. Representative cases covering the six style categories are presented, while the complete quantitative correspondence results are reported in Supplementary Table S5 (Figure 12). The detailed quantitative correspondence results, including diagnostic morphological regions, ROI area ratio, attention energy, activation coverage, IoU, enrichment ratio, and alignment interpretation, are reported in Supplementary Table S5. And the detailed quantitative correspondence results, including diagnostic morphological regions, ROI area ratio, attention energy, activation coverage, IoU, enrichment ratio, and alignment interpretation, are also reported in Supplementary Table S5.
Figure 12. Correspondence between architectural diagnostic regions and Grad-CAM attention maps for representative style categories.
The correspondence analysis was not intended to demonstrate that the model possesses human-like architectural cognition. Instead, it evaluates whether the spatial distribution of model attention is consistent with expert-defined architectural evidence. Four quantitative indicators, including ROI attention energy, top-20 activation coverage, intersection-over-union (IoU), and enrichment ratio, were calculated to measure the degree of spatial agreement between model attention and annotated diagnostic regions. These measurements provide a more systematic evaluation of Grad-CAM responses beyond subjective visual interpretation.
For Category 01, HN003 was correctly classified as Western veranda architecture, with a relatively high alignment between annotated diagnostic regions and model attention (Figure 12). The Grad-CAM activation was mainly concentrated on the veranda arches, column systems, corridor rhythm, and facade openings, which correspond to the diagnostic features identified by architectural annotation. The ROI energy of 0.708 indicates that a substantial proportion of discriminative activation was located within expert-defined architectural regions.
For Category 02, HN011 was incorrectly classified as Category 01, Western veranda architecture. Unlike the high-alignment cases, its Grad-CAM activation showed limited correspondence with the annotated diagnostic regions. Although some attention responses appeared around windows, openings, and facade edges, a considerable proportion of activation extended beyond the annotated regions, resulting in a low alignment level with an ROI energy of 0.191 (Figure 12). This case indicates that the model remained sensitive to general facade rhythm and opening characteristics but failed to capture the category-specific differences between Western eclectic and veranda-style architectural features.
For Category 03, HN031 achieved correct classification as Chinese revival palace-style architecture and demonstrated high alignment between architectural diagnostic regions and model attention (Figure 12). The attention distribution corresponded strongly with the Chinese-style roof profile, upturned eaves, central decorative elements, and overall roof composition. The ROI energy of 0.656 suggests that the model relied substantially on the morphological evidence manually identified by experts.
For the Category 02/04 boundary case, HN036 represents a typical misclassification between Chinese revival hybrid architecture and Western eclectic architecture. The ground-truth category was Category 04, but the model predicted Category 02. The correspondence analysis shows that the model focused on facade rhythm, arched openings, and Western facade organization, while the hybrid relationship between Chinese decorative elements and Western compositional structures was not sufficiently distinguished (Figure 12). With an ROI energy of 0.433 and moderate alignment, this case suggests that the model captured several relevant facade characteristics but still encountered difficulty in separating transitional stylistic boundaries.
For Category 05, HN100 was correctly classified as Chinese revival modern-style architecture. The Grad-CAM response showed moderate correspondence with manually annotated regions, particularly in relation to geometric facade organization, simplified decorative motifs, and modernized roof treatment (Figure 12). However, the activation also extended beyond the annotated diagnostic regions, indicating that the model may still incorporate broader facade context when distinguishing this category. The ROI energy reached 0.445, reflecting partial but not complete correspondence.
For Category 06, HN103 was correctly classified as early modern/new architecture and showed moderate alignment between attention responses and diagnostic regions. The model attention was associated with functional geometric organization, repetitive window arrangements, vertical facade elements, and reduced ornamentation (Figure 12). Nevertheless, part of the activation extended outside the annotated regions, resulting in an ROI energy of 0.455. This indicates that although subject-focused attention mechanisms improved the model’s sensitivity to early-modern architectural evidence, contextual visual information was not completely eliminated.
Overall, the ROI–Grad-CAM correspondence analysis demonstrates that subject-focused strategies can improve the spatial consistency between model attention and architectural diagnostic features, particularly when decisive morphological characteristics are visually prominent. High-alignment examples such as HN003 and HN031 indicate that the model can effectively respond to style-related architectural evidence, whereas moderate- and low-alignment cases reveal remaining challenges caused by overlapping stylistic features, transitional architectural expressions, and contextual interference.
Therefore, Grad-CAM correspondence should be interpreted as evidence of spatial consistency rather than direct proof of architectural reasoning. The combination of manually annotated diagnostic regions and quantitative correspondence measurements provides a more objective interpretation framework, but future studies should further incorporate component-level segmentation, expert multi-annotator validation, and larger category-balanced datasets to establish stronger connections between architectural knowledge and visual representation learning. Beyond attention visualization, the quantitative results of A1–A8 also provide insights into the effectiveness and limitations of the proposed optimization strategies. The combination of L-Softmax and class-balanced weighting contributes to improving feature discrimination by enlarging angular margins among adjacent style categories and reducing bias caused by category imbalance. These mechanisms are particularly relevant for difficult classification boundaries, including Categories 02/04, 03/04, and 05/06, where architectural styles share overlapping morphological characteristics.
However, the A1–A8 experiments also demonstrate that improving model complexity does not necessarily guarantee better target-domain performance. The effectiveness of these mechanisms remains jointly constrained by data quality, category ambiguity, image background interference, and limited target-domain samples. For example, A3 achieved a building-level accuracy of only 51.35%, whereas A8 obtained the highest image-level accuracy of 58.60% but a lower building-level accuracy of 54.05%. These results indicate that individual modules or stronger local attention constraints may improve certain image-level responses while not necessarily enhancing building-level consistency. This observation further highlights the importance of evaluating architectural style recognition at the building-instance level. Unlike conventional object classification tasks, architectural style interpretation is usually based on multiple visual representations of the same building, including primary facades, secondary facades, architectural details, partially occluded views, and photographs captured under different conditions. Therefore, building-level classification through multi-view majority voting provides a more comprehensive evaluation of whether the model can recognize the overall stylistic characteristics of a building instance.
A single misclassified image does not necessarily indicate failure in architectural style recognition, because different viewpoints may emphasize different morphological features and generate inconsistent model responses. In contrast, building-level prediction aggregates multiple visual observations and better reflects the practical application requirements of architectural heritage surveys, digital documentation, and auxiliary style classification. Accordingly, the combination of image-level and building-level accuracy provides a more reliable evaluation framework for deep learning-based architectural style recognition.

5.2. Architectural Interpretation of Confusion

The most architecturally meaningful result of this study is not that the model has achieved high-precision automatic classification, but that the observed confusion patterns correspond to potentially meaningful transitional relationships within the operational stylistic spectrum of modern architecture defined in this study. The Grad-CAM correspondence analysis provides additional evidence for interpreting these boundary cases. Rather than considering misclassification only as prediction failure, the alignment between model attention and manually defined diagnostic regions helps identify whether errors originate from ambiguous stylistic boundaries or from attention deviations toward non-decisive visual information. The confusion matrix of A7 indicates that 02/04, 03/04, and 05/06 are the three boundary pairs that warrant the greatest attention. These patterns not only reflect classification difficulties encountered by the model, but also reveal the historical complexity of the interaction between Chinese and Western architectural languages, the transformation of traditional revival forms, and the emergence of modernist expression in modern educational and cultural buildings in Hunan.
The 02/04 confusion represents one of the most important architectural findings of this study. Category 02, Western eclectic architecture, would generally be expected to be identified primarily through Western classical, eclectic, or imported compositional characteristics, including axial symmetry, column orders, pediments, arches, and Western decorative motifs. However, in modern architectural practice in Hunan and comparable regions, some Western eclectic buildings also incorporate Chinese-style roofs, eave treatments, traditional motifs, or localized decorative elements, thereby creating visual overlap with Category 04, Chinese revival hybrid architecture. In the A7 confusion matrix, 20 images from Category 02 were classified as Category 04, while 14 images from Category 04 were classified as Category 02. The total bidirectional confusion for the 02/04 boundary therefore reached 34 images, making it the most prominent boundary relationship in the current experiment. This result suggests that the model is sensitive to visually salient traditional components. It also indicates that manual classification should further distinguish between two different stylistic logics: Western eclecticism as the dominant framework with partial incorporation of Chinese symbolic elements, and Chinese traditional revival as the dominant framework combined with Western compositional features.
The regional context of modern educational and cultural buildings in Hunan further intensifies the complexity of this boundary. Compared with coastal treaty-port cities, the development of modern architecture in Hunan was influenced by multiple factors, including inland modernization, educational transformation, local intellectual communities, the spread of missionary education, and the exploration of national forms. Consequently, its stylistic combinations are not fully equivalent to those found in treaty-port or coastal-city samples. Some buildings adopted Western facade compositions while simultaneously incorporating traditional roofs or national-form symbols, thereby forming transitional conditions between Categories 02 and 04. The model’s repeated misclassifications should therefore not be regarded solely as technical errors, but also as reflections of the stylistic continuity that architectural–historical classification itself must address. Nevertheless, these interpretations remain dependent on the adopted annotation criteria and available target-domain samples.
The 03/04 confusion should be understood as an important boundary within traditional revival architecture. Category 03, Chinese revival palace-style architecture, and Category 04, Chinese revival hybrid architecture, may both include large roofs, eaves, traditional ornamentation, entrance axes, and national-form symbols. The key distinction lies in their compositional hierarchy. Category 03 places greater emphasis on the dominant role of palace-style roofs and monumental composition, whereas Category 04 more often reflects a hybrid relationship between traditional components and modern windows, Western compositional principles, or geometric decoration. In the A7 confusion matrix, five images from Category 03 were classified as Category 04, while two images from Category 04 were classified as Category 03, resulting in seven bidirectional confusion cases for the 03/04 boundary. This indicates that, when distinguishing subtypes within traditional revival architecture, the model still has difficulty consistently identifying the hierarchical relationships among traditional roofs, entrance composition, and overall massing organization.
Accordingly, for Category 03/04 samples, model outputs are more appropriately regarded as prompts for expert review rather than as final determinations. This is particularly important for Categories 03 and 04 because the available Hunan target-domain samples do not provide sufficient independent instances for fully stable statistical validation. Researchers should further examine factors such as roof scale, entrance treatment, column systems, eave decoration, facade axes, and the spatial organization of building groups or campuses. In representative traditional revival cases among modern educational and cultural buildings in Hunan, palace-style or hybrid revival cannot be reduced to superficial formal decoration. Such expressions are often related to Republican-era discussions of Chinese national form, campus spatial order, national symbolism, and local cultural identity. The interpretation of these confusion cases therefore requires consideration of their broader architectural–historical context.
The 05/06 confusion reflects the difficulty of distinguishing buildings during the transition from traditional revival toward early modern architectural expression. Category 05, Chinese revival modern-style architecture, retains certain traditional symbols or tendencies toward national form, although its massing, structure, and facade treatment have already become increasingly modernized. Category 06, the “New Architecture” of the modern period, namely early modern architecture, places greater emphasis on functionalism, simplified massing, reduced ornamentation, and modern structural expression. In field photographs, both categories may display reduced decoration, regular window arrangements, clearly articulated volumes, and simplified facades. As a result, the model may have difficulty determining whether traditional symbols still play a dominant role. The A7 confusion matrix shows that five images from Category 06 were classified as Category 05, while no Category 05 images were classified as Category 06. This indicates that the confusion was primarily directional, with early modern architecture in Category 06 being drawn toward Chinese revival modern-style architecture in Category 05.
The comparison between HN022 and HN103 further illustrates this issue. The ground-truth category of HN022 is Category 06; however, the visualization results show strong responses in the roof and surrounding environmental regions, and the model ultimately classified it as Category 04. This suggests that, when images of early modern architecture contain prominent roof features, environmental occlusion, or revival-related visual cues, the model may place insufficient emphasis on modern facade characteristics. By contrast, HN103 could be classified as Category 06, indicating that early modern architecture can still be correctly identified when the model focuses on key regions associated with modern massing, vertical composition, or simplified facades. Together, these cases indicate that Category 06 is not inherently unrecognizable, but that its reliable classification remains highly dependent on image quality, subject visibility, and whether the model focuses effectively on the building itself. This instability is also consistent with the cross-regional domain-shift condition identified in Section 4.2, where differences in visual context and regional expression increase the difficulty of transferring learned stylistic representations. From an architectural–historical perspective, the confusion matrix may be interpreted as a preliminary “computational stylistic spectrum”. It cannot replace architectural–historical judgment, but it can indicate which categories exhibit greater visual similarity and which building instances should be prioritized for subsequent component-level review. For persistently misclassified cases, researchers should return to the building-instance level and examine roofs, eaves, entrances, column systems, windows, decorative motifs, facade proportions, and massing relationships. These morphological observations should then be considered together with construction dates, designers, building functions, histories of renovation or extension, and local architectural contexts.
Therefore, the discussion in this study does not simply treat model misclassification as failure. Instead, it uses misclassification as an entry point for architectural–historical interpretation. The 02/04 confusion reveals the stylistic overlap between Western eclectic architecture and Chinese revival hybrid architecture. The 03/04 confusion reveals the finer boundary between palace-style and hybrid expressions within traditional revival architecture. The 05/06 confusion reveals the transitional process through which traditional symbols gradually weakened while modern architectural language progressively emerged. These findings indicate that, under the current experimental conditions, the deep learning model is more appropriately used as an auxiliary tool for architectural heritage style classification than as an automatic classification system intended to replace expert judgment.

5.3. Limitations and Future Work

Although the subject-focused mechanisms showed certain positive effects on building-level classification, this study still has several clear limitations.
(1)
Target-domain sample size, category imbalance, and statistical uncertainty The fixed Hunan target-domain test set contains a limited number of independent building instances in several categories. Specifically, Categories 03 and 05 each contain only one building instance, while Category 06 contains only two building instances. Although balanced metrics, macro-level evaluation indicators, and bootstrap confidence intervals were introduced to reduce the influence of category imbalance and estimate uncertainty, they cannot overcome the fundamental limitation caused by insufficient independent samples. Therefore, the results of these categories should be interpreted as case-based observations rather than evidence of stable class-level generalization.
Moreover, the current statistical validation remains constrained by the limited number of target-domain building instances. More rigorous paired statistical comparisons between different models require larger and more balanced multi-regional datasets. Future research should prioritize expanding independent building instances from Hunan and comparable regions, establishing more comprehensive validation sets, and incorporating formal statistical significance testing to further verify the robustness of architectural style recognition strategies.
(2)
The cross-regional target-domain evaluation adopted in this study intentionally increases the difficulty of style transfer; however, it also introduces additional uncertainty caused by regional architectural differences, photographic conditions, and historical development contexts. Future studies should include multi-regional target domains to examine whether the learned representations can generalize beyond Hunan.
(3)
The current model should not be regarded as a mature high-precision six-class classification system. A8 achieved the highest image-level accuracy of 58.60%, whereas A4, A5, and A7 achieved the highest building-level accuracy of 64.86%. The inconsistency between image-level and building-level performance indicates that improving individual image recognition does not necessarily lead to more reliable building-instance classification.
Therefore, the proposed framework is more appropriately positioned as an exploratory and auxiliary approach for digital architectural heritage analysis rather than a fully automated classification tool. Building-level evaluation remains particularly important because architectural style interpretation usually relies on multiple views of the same building, including main facades, secondary facades, architectural components, partially occluded views, and images captured under different conditions. Future studies should further improve multi-view consistency and investigate whether the framework can support practical architectural survey and documentation workflows.
(4)
Subject-focused mechanism and architectural knowledge supervision limitation. Although subject-focused mechanisms improved the spatial consistency between model attention and building regions, they cannot completely eliminate background interference or resolve all stylistic ambiguities. The Grad-CAM correspondence analysis based on manually annotated architectural diagnostic regions demonstrates that some correctly classified cases, such as HN003 and HN031, show strong alignment between model attention and architectural evidence. However, moderate- and low-alignment cases, such as HN011, HN036, HN100, and HN103, indicate that the model may still respond to broader contextual information, facade rhythm, or visually correlated but non-decisive features.
In addition, the annotation of modern Chinese architectural styles remains a complex disciplinary issue. Although this study established morphological criteria based on previous architectural research and conducted researcher-based review of diagnostic regions, transitional stylistic boundaries may still involve expert interpretation. Future research should introduce larger expert validation groups, quantitative inter-annotator agreement analysis, building-subject segmentation, object detection, and component-level annotations to establish stronger connections between architectural knowledge and visual representation learning.
(5)
Methodological extension and reproducibility limitation. All A1–A8 ablation experiments in this study were conducted using the unified Swin Transformer backbone network. Therefore, the current results primarily demonstrate the relative contribution of different subject-focused modules within the Swin framework and do not conclusively establish their universal effectiveness across other transformer or CNN architectures. Although Pure DeiT-Base Distilled achieved slightly better performance than Pure Swin Transformer-Tiny in backbone comparison experiments, the complete A1–A8 module evaluation was not replicated within the DeiT framework. Future studies should conduct cross-backbone transfer experiments to examine the general applicability of subject-focused mechanisms.
Furthermore, the original concept of multi-view contrastive learning remains a promising direction for architectural image analysis. Images of the same building captured from different viewpoints, lighting conditions, or occlusion levels should maintain consistent stylistic representations. Future research could introduce multi-view contrastive objectives to improve representation consistency and enhance building-level prediction stability. Meanwhile, more systematic organization of fixed validation sets, training procedures, and experimental records is required to improve reproducibility.
Future research can proceed in three main directions. First, the range of research samples should be expanded by supplementing Categories 03, 05, and 06, with particular emphasis on increasing the number of independent building instances. Second, building-subject segmentation and component-level annotation should be introduced to enable the model to distinguish more clearly between the building subject and background interference. Third, Grad-CAM results, attention maps, and confusion matrices could be further transformed into component-level interpretive indicators to examine the roles of roofs, column systems, entrances, windows, eaves, and decorative patterns in different style categories. In this way, deep learning results could contribute not only to classification performance, but also to the architectural–historical interpretation of stylistic boundaries.

6. Conclusions

This study addressed the six-class-style classification of modern educational and cultural buildings in Hunan by establishing a deep learning experimental framework based on cross-regional training and fixed Hunan target-domain testing, and by systematically comparing eight module combinations (A1–A8). Unlike conventional architectural image classification studies based on random data splitting, this study retained all Hunan samples as a fixed target-domain test set to examine whether a style classification model trained on non-Hunan samples could be transferred to the regional context of modern educational and cultural buildings in Hunan. The experimental results do not support a strong claim of high-precision automatic classification across all six architectural style categories. Complementary evaluation using overall accuracy, balanced accuracy, macro-F1, and bootstrap confidence intervals further indicates that the reported performance should be interpreted within the specific cross-regional target-domain setting rather than as evidence of universal classification capability. Accordingly, the proposed framework is explicitly positioned as an auxiliary tool for preliminary screening, digital documentation, stylistic analysis, and the identification of cases requiring further examination, rather than as a replacement for expert-based architectural–historical identification and professional judgment. Nevertheless, the experimental results suggest that subject-focused mechanisms can provide useful support for architectural style recognition under conditions involving complex image backgrounds, regional domain shift, and ambiguous stylistic boundaries.
(1)
Cross-regional target-domain testing provides a stricter evaluation setting for architectural style classification than conventional random image splitting when data are isolated at the building-instance level. In this study, all Hunan samples were retained as a fixed target-domain test set, and multi-view images of the same building were prevented from appearing simultaneously in the training and testing stages. This reduced the possibility of artificially inflated performance caused by the model memorizing the appearance of individual buildings. Under this setting, the model had to address domain shifts arising from regional differences, variations in photographic conditions, differences in background environments, and stylistic continuity. The results indicate that models trained on non-Hunan samples can achieve a certain degree of transfer capability in the Hunan target domain. The additional domain-shift analysis confirms that the Hunan target domain presents measurable distribution differences from the source domain, indicating that the observed performance reflects a challenging transfer scenario rather than conventional in-domain classification. However, the current performance also shows that cross-regional architectural style classification remains substantially constrained by target-domain sample size, class distribution, and differences in local architectural expression. Therefore, the cross-regional experimental design used in this study serves not only to evaluate classification performance but also as an important methodological basis for examining practical model transferability.
(2)
The best image-level and building-level results were achieved by different model configurations, demonstrating a clear distinction between image-level classification performance and the stability of building-level classification. A8 achieved the highest image-level accuracy of 58.60%, suggesting that Grad-CAM-related background constraints provided some assistance in the classification of individual images. However, its building-level accuracy was 54.05%, which was 4.55 percentage points lower than its image-level accuracy. This indicates that the advantage observed in individual-image classification did not translate consistently into building-level classification across multiple views of the same building. By contrast, A4, A5, and A7 all achieved the highest building-level accuracy of 64.86%. Among them, A7 introduced the GBVS/AGL subject-focused mechanism into the base Swin framework and achieved an image-level accuracy of 56.45% and a building-level accuracy of 64.86%. Considering its image-level performance, building-level performance, balanced evaluation indicators, model configuration, and behavior on small-sample categories, A7 was selected as the main model for the current qualitative and architectural interpretation analysis, whereas A8 was retained as the model with the best image-level result and as supplementary evidence regarding the effect of Grad-CAM-related constraints.
(3)
Subject-focused mechanisms and building-level multi-view evaluation showed complementary tendencies under the current experimental conditions, although adding more modules did not necessarily lead to better classification performance. The building-level accuracy of A7 was 8.41 percentage points higher than its image-level accuracy, while A5 showed a corresponding positive change of 8.95 percentage points. These results indicate that, for some model configurations, building-level accuracy after multi-view majority voting was higher than image-level accuracy based on individual images. However, because image-level accuracy and building-level accuracy are calculated using different statistical units, the building-level accuracy change relative to image-level accuracy should be interpreted as a descriptive result and should not be directly equated with a quantitative improvement in multi-view prediction consistency. At the same time, the results of A3 and A8 show that adding multi-scale features alone or strengthening attention at the individual-image level does not necessarily improve building-level classification. Grad-CAM and attention visualizations further indicate that, in some cases, the model attended to style-related regions such as verandas, entrance axes, window organization, facade rhythm, and building-subject outlines. Nevertheless, cases such as HN036, HN022, and HN103 still exhibited intertwined stylistic cues, responses to background context, or inconsistent predictions across different views. Therefore, the significance of subject-focused mechanisms does not lie in having completely eliminated background interference. Rather, they provide a testable technical pathway for encouraging the model to rely more on building-subject features than on incidental environmental information. Meanwhile, building-level evaluation based on multiple views is more consistent with the practical logic of architectural research, in which stylistic attribution is ultimately determined at the level of the building rather than solely from a single image.
(4)
The confusion results reveal three key stylistic boundary patterns within the operational classification system adopted in this study, indicating that some misclassifications are not merely technical noise but are related to the continuity of the modern architectural style spectrum itself. The 02/04 confusion was the most prominent and stable boundary in the current experiments, reflecting the overlap between Western Eclectic Architecture and Chinese Revival Hybrid Architecture in terms of Western compositional order, traditional revival components, and decorative language. The 03/04 confusion reflects the internal boundary between palace-style and hybrid expressions within Chinese traditional revival architecture. Its interpretation requires consideration of the dominant role of palace-style roofs, overall monumental composition, and the organizational relationship between traditional and modern components. The 05/06 confusion reflects the continuity between Chinese Revival Modern-Style Architecture and Early Modern Architecture during the gradual weakening of traditional symbols and the increasing prominence of modern architectural language. Therefore, model misclassification in this study is not only a technical error to be reduced, but also a potential clue for identifying transitional building instances, detecting stylistic boundaries, and selecting cases for subsequent expert review. For such difficult cases, final stylistic classification should return to the examination of roofs, eaves, entrances, column systems, windows, decorative patterns, facade proportions, and massing relationships, while also considering construction dates, design backgrounds, and local architectural contexts.
(5)
The results for Categories 03, 05, and 06 must be interpreted cautiously under small-sample target-domain conditions, and the current study cannot support stable class-level generalization conclusions for these categories. In the fixed Hunan target-domain test set, Categories 03 and 05 each contain only one building instance, while Category 06 contains only two building instances. Consequently, the correct or incorrect classification of a single building can substantially affect the corresponding building-level result. Although balanced metrics and bootstrap confidence intervals were introduced to better characterize uncertainty, they cannot compensate for the fundamental limitation caused by insufficient independent building instances. The correct classification of HN031 and HN100 indicates that the model can capture certain architectural style cues in individual sparse-category cases. However, these results do not demonstrate that the model can reliably classify all examples of Chinese revival palace-style architecture or Chinese revival modern-style architecture. The different performances of HN022 and HN103 across model configurations and evaluation levels further indicate that early modern architecture remains one of the less stable categories in the current classification system. Accordingly, the current results more appropriately support the following conclusion: some subject-focused configurations can improve building-level classification performance and help reveal stylistically meaningful architectural boundaries, but the model at this stage is more suitable for auxiliary screening, digital documentation, and difficult-case identification than for high-confidence automatic classification across all six architectural style categories.
In summary, the main contribution of this study lies in proposing a research pathway that integrates architectural–historical style classification, cross-regional target-domain testing, subject-focused deep learning, building-level multi-view evaluation, and stylistic boundary interpretation. The six-class style classification system provides an operational architectural–historical framework for interpreting model outputs. Subject-focused and background-suppression mechanisms offer an exploratory methodological direction for building-subject recognition in complex real-world photographic environments. Meanwhile, the comparison between image-level and building-level evaluation, multi-view aggregation results, confusion matrices, and interpretability visualizations enable model outputs to support the identification of transitional stylistic cases and subsequent expert review.
Future research should prioritize increasing the number of independent building instances in Categories 03, 05, and 06, further improving comparisons among representative backbone networks under fully unified data splits and evaluation protocols, and introducing building-subject segmentation, object detection, multi-view consistency learning, and component-level annotation to improve model reproducibility, statistical reproducibility, interpretability, and architectural–historical analytical capacity. Future studies should also conduct repeated cross-regional evaluations with multiple target domains and larger independent samples to further examine statistical robustness and generalization uncertainty. Furthermore, architectural components such as roofs, columns, entrances, windows, eaves, and decorative patterns could be transformed into quantifiable visual features, allowing deep learning research to move beyond “style classification” toward computational analysis of stylistic characteristics.

Supplementary Materials

The following supporting information can be downloaded at https://www.mdpi.com/article/10.3390/buildings16183678/s1, Table S1. Performance comparison of A1–A8 models on relatively better-supported categories (Categories 01, 02, and 04) of the fixed Hunan target-domain test set; Table S2: Operational morphological criteria for six categories.; Table S3: Post-1949 cases and inclusion evidence; Table S4: Building-level bootstrap confidence intervals for A7 on the fixed Hunan target-domain test set.; Table S5: Full architectural diagnostic-region and Grad-CAM alignment results; Table S6. Complete building-level performance comparison of A1–A8 ablation experiments on the fixed Hunan target-domain test set.; Table S7. Training convergence statistics of A1–A8 ablation experiments; Dataset S1: Supplementary Materials Module Experiment Data; Dataset S2: Supplementary Materials Statistical Analysis and Model Evaluation Data.

Author Contributions

Conceptualization, J.L. and J.Y.; methodology, J.L. and J.Y.; software, J.L.; validation, J.L., J.Y., S.S. and B.P.; formal analysis, J.L., J.Y., B.P., S.S., Y.Y. and J.G.; investigation, J.L., J.Y. and B.P.; resources, J.L., J.Y. and B.P.; data curation, J.L., J.Y., B.P. and S.L.; writing—original draft preparation, J.L., J.Y. and B.P.; writing—review and editing, J.L., J.Y. and B.P.; visualization, J.L., J.Y., B.P. and X.Z.; supervision, J.L., J.Y., B.P., S.L., S.S. and Y.Y.; project administration, J.L., J.Y. and B.P.; funding acquisition, J.Y. All authors have read and agreed to the published version of the manuscript.

Funding

This paper was supported by the Scientific Research Fund of the Hunan Provincial Education Department in China (22A0214).

Data Availability Statement

The image dataset contains ongoing research materials and some images collected from public online sources; therefore, the full dataset is not publicly released at this stage. Experimental logs, module configurations, and derived performance tables can be made available by the corresponding author upon reasonable request.

Acknowledgments

The authors extend sincere thanks to Changsha University of Science and Technology for its research support and resources. Gratitude is also expressed to Jun Yan for their guidance on this study, and to the reviewers and editors for their constructive comments.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Pan, G. History of Chinese Architecture, 7th ed.; China Architecture & Building Press: Beijing, China, 2015; pp. 406–429. [Google Scholar]
  2. Liu, Y. Introduction to Modern Chinese Architectural History; Commercial Press: Beijing, China, 2019; pp. 2–10, 25–34. [Google Scholar]
  3. Lai, D. Studies in Modern Chinese Architectural History; Tsinghua University Press: Beijing, China, 2007; pp. 181–273, 289–312. [Google Scholar]
  4. Liu, Y. From Case Accumulation to Field Expansion: Possibilities for Deepening Research on Modern Chinese Architectural History. Architect 2020, 1, 129–133. [Google Scholar]
  5. Liu, Y. Main Lines and Periodization of the Development of Modern Chinese Architecture. Archit. J. 2012, 10, 70–75. [Google Scholar]
  6. Zhang, C. “Yangshi” and “Style” in Chinese Architectural Discourse from the 1890s to the 1950s. Archit. J. 2023, 7, 96–101. [Google Scholar] [CrossRef]
  7. Zhang, F. Review and Prospect of Research on Modern Chinese Architectural History. South Archit. 1994, 2, 3–12. [Google Scholar]
  8. Zhang, F. Research on Modern Chinese Architectural History and the Conservation of Modern Architectural Heritage. J. Harbin Inst. Technol. Soc. Sci. Ed. 2008, 10, 12–26. [Google Scholar] [CrossRef]
  9. Fujimori, T.; Zhang, F. Veranda Style: The Origin of Modern Chinese Architecture. Archit. J. 1993, 5, 33–38. [Google Scholar]
  10. Lai, D. An Academic Overview of Architecture in China’s Nationalist-Controlled Areas during the Full-Scale War of Resistance, Part I. Archit. J. 2025, 12, 111–116. [Google Scholar] [CrossRef]
  11. Lai, D. An Academic Overview of Architecture in China’s Nationalist-Controlled Areas during the Full-Scale War of Resistance, Part II. Archit. J. 2026, 2, 115–126. [Google Scholar] [CrossRef]
  12. Wang, X. The Creative Transformation of Traditional Concepts in Modern Architectural Education and the Construction of National Architectural Theory. Earthq. Resist. Eng. Retrofit. 2026, 48, 195. [Google Scholar]
  13. Zhang, M. Early Progress of Modernization in Hunan; Hunan People’s Publishing House: Changsha, China, 2017; pp. 100–125, 137–145. [Google Scholar]
  14. Lin, H. Abolishing Teaching Posts in Prefectural, Subprefectural, and County Schools and the Late Qing New Policies. Soc. Sci. Front 2025, 9, 152–166. [Google Scholar]
  15. Hou, B.; Yang, H.; Sun, H.; Shi, X. Research on a Digital Conservation Model for Modern Educational Buildings: The Case of Building No. 6 on the Minglun Campus of Henan University. Urban Archit. 2026, 23, 16–19. [Google Scholar] [CrossRef]
  16. Mathias, M.; Martinovic, A.; Weissenberg, J.; Haegler, S.; Van Gool, L. Automatic architectural style recognition. Int. Arch. Photogramm. Remote Sens. Spat. Inf. Sci. 2012, XXXVIII-5/W16, 171–176. [Google Scholar] [CrossRef] [Scilit]
  17. Chu, W.-T.; Tsai, M.-H. Visual pattern discovery for architecture image classification and product image search. In Proceedings of the 2nd ACM International Conference on Multimedia Retrieval (ICMR 2012), Hong Kong, China, 5–8 June 2012; Association for Computing Machinery: New York, NY, USA, 2012; p. 27. [Google Scholar] [CrossRef] [Scilit]
  18. Yi, Y.K.; Zhang, Y.; Myung, J. House style recognition using deep convolutional neural network. Autom. Constr. 2020, 118, 103307. [Google Scholar] [CrossRef] [Scilit]
  19. Sun, M.; Zhang, F.; Duarte, F.; Ratti, C. Understanding architecture age and style through deep learning. Cities 2022, 128, 103787. [Google Scholar] [CrossRef] [Scilit]
  20. Cantemir, E.; Kandemir, O. Use of artificial neural networks in architecture: Determining the architectural style of a building with convolutional neural networks. Neural Comput. Appl. 2024, 36, 6195–6207. [Google Scholar] [CrossRef] [Scilit]
  21. Liu, J. Research on Building Style Recognition Based on Deep Learning. Master’s Thesis, Xi’an University of Architecture and Technology, Xi’an, China, 2023. [Google Scholar]
  22. Zhang, M.; Ju, W.; Liu, X. Application of Artificial Neural Networks in Architectural Style Identification: A Case Study of Modern Historical Buildings in Dalian. Huazhong Archit. 2019, 37, 43–46. [Google Scholar] [CrossRef]
  23. Han, Q.; Yin, C.; Yang, Y. A Computational Approach to Traditional Chinese Architectural Styles: Intelligent Identification and Spatial Pattern Analysis. Trop. Geogr. 2026, 46, 179–189. [Google Scholar] [CrossRef]
  24. Li, L.; Deng, K.; Chen, T. Image Classification of Traditional Residential Architecture in Eastern Fujian Based on Deep Learning. Comput. Digit. Eng. 2026, 54, 524–529. [Google Scholar]
  25. Fiorucci, M.; Khoroshiltseva, M.; Pontil, M.; Traviglia, A.; Del Bue, A.; James, S. Machine learning for cultural heritage: A survey. Pattern Recognit. Lett. 2020, 133, 102–108. [Google Scholar] [CrossRef] [Scilit]
  26. Huang, Z. Machine learning in the documentation and analysis of historical landmark buildings: A systematic review and practical demonstration. Built Herit. 2026, 10, 14. [Google Scholar] [CrossRef] [Scilit]
  27. Ottoni, A.L.C.; Ottoni, L.T.C. A deep learning approach for cultural heritage building classification using transfer learning and data augmentation. J. Cult. Herit. 2025, 74, 214–224. [Google Scholar] [CrossRef] [Scilit]
  28. Siountri, K.; Anagnostopoulos, C.-N. The Classification of Cultural Heritage Buildings in Athens Using Deep Learning Techniques. Heritage 2023, 6, 3673–3705. [Google Scholar] [CrossRef] [Scilit]
  29. Xu, H.; Sun, H.; Wang, L.; Yu, X.; Li, T. Urban Architectural Style Recognition and Dataset Construction Method under Deep Learning of Street View Images: A Case Study of Wuhan. ISPRS Int. J. Geo-Inf. 2023, 12, 264. [Google Scholar] [CrossRef] [Scilit]
  30. Zhou, J.; Xie, L.; Fricker, P.; Liu, K. ConvNeXt-L-Based Recognition of Decorative Patterns in Historical Architecture: A Case Study of Macau. Buildings 2025, 15, 3705. [Google Scholar] [CrossRef] [Scilit]
  31. Han, Q.; Yin, C.; Deng, Y.; Liu, P. Towards Classification of Architectural Styles of Chinese Traditional Settlements Using Deep Learning: A Dataset, a New Framework, and Its Interpretability. Remote Sens. 2022, 14, 5250. [Google Scholar] [CrossRef] [Scilit]
  32. Zhang, G.; Pan, Y.; Zhang, L. Deep learning for detecting building facade elements from images considering prior knowledge. Autom. Constr. 2022, 133, 104016. [Google Scholar] [CrossRef] [Scilit]
  33. Chen, C.; Li, S.; Yao, Q.; Hu, Y.; Fan, Z. Identifying and Interpreting Traditional Architectural Style Characteristics Based on Deep Learning. npj Herit. Sci. 2026. [Google Scholar] [CrossRef] [Scilit]
  34. Han, P.; Hu, S.; Xu, R. Formal Feature Identification of Vernacular Architecture Based on Deep Learning—A Case Study of Jiangsu Province, China. Sustainability 2025, 17, 1760. [Google Scholar] [CrossRef] [Scilit]
  35. Kong, Y.; Xue, P.; Xu, Y.; Li, X. An Environmental Pattern Recognition Method for Traditional Chinese Settlements Using Deep Learning. Appl. Sci. 2023, 13, 4778. [Google Scholar] [CrossRef] [Scilit]
  36. Shan, L.; Zhang, L. Application of Intelligent Technology in Facade Style Recognition of Harbin Modern Architecture. Sustainability 2022, 14, 7073. [Google Scholar] [CrossRef] [Scilit]
  37. Zhao, J.; Han, C.; Wu, Y.; Xu, C.; Huang, X.; Qi, X.; Qi, Y.; Gao, L. A Deep Learning-Based Study on Visual Quality Assessment of Commercial Renovation of Chinese Traditional Building Facades. Environ. Impact Assess. Rev. 2025, 113, 107862. [Google Scholar] [CrossRef] [Scilit]
  38. Zhong, J.; Hao, T.; Yin, J.; Li, P.; Pan, R.; Zeng, P.; Li, J.; Tai, H.; Zang, M.; Lu, S. ArchiLense: A Framework for Quantitative Analysis of Architectural Styles Based on Vision Large Language Models. In Transindividual Intelligence: Proceedings of the 7th International Conference on Computational Design and Robotic Fabrication (CDRF 2025); Liu, Y., Chai, H., Bao, D.W.N., Yuan, P.F., Eds.; Springer: Singapore, 2026; pp. 3–13. [Google Scholar] [CrossRef] [Scilit]
  39. Geirhos, R.; Jacobsen, J.-H.; Michaelis, C.; Zemel, R.; Brendel, W.; Bethge, M.; Wichmann, F.A. Shortcut learning in deep neural networks. Nat. Mach. Intell. 2020, 2, 665–673. [Google Scholar] [CrossRef] [Scilit]
  40. Zhou, K.; Liu, Z.; Qiao, Y.; Xiang, T.; Loy, C.C. Domain generalization: A survey. IEEE Trans. Pattern Anal. Mach. Intell. 2023, 45, 4396–4415. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  41. Koh, P.W.; Sagawa, S.; Marklund, H.; Xie, S.M.; Zhang, M.; Balsubramani, A.; Hu, W.; Yasunaga, M.; Phillips, R.L.; Gao, I.; et al. WILDS: A benchmark of in-the-wild distribution shifts. In Proceedings of the 38th International Conference on Machine Learning (ICML 2021); Meila, M., Zhang, T., Eds.; PMLR: Cambridge, MA, USA, 2021; Volume 139, pp. 5637–5664. [Google Scholar]
  42. Gulrajani, I.; Lopez-Paz, D. In search of lost domain generalization. In Proceedings of the International Conference on Learning Representations (ICLR 2021), Virtual, 3–7 May 2021. [Google Scholar]
  43. Selvaraju, R.R.; Cogswell, M.; Das, A.; Vedantam, R.; Parikh, D.; Batra, D. Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), Venice, Italy, 22–29 October 2017; pp. 618–626. [Google Scholar] [CrossRef] [Scilit]
  44. Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin, S.; Guo, B. Swin Transformer: Hierarchical Vision Transformer Using Shifted Windows. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada, 10–17 October 2021; pp. 10012–10022. [Google Scholar] [CrossRef] [Scilit]
  45. Wu, J.; Ying, Y.; Tan, Y.; Liu, Z. Innovative Framework for Historical Architectural Recognition in China: Integrating Swin Transformer and Global Channel–Spatial Attention Mechanism. Buildings 2025, 15, 176. [Google Scholar] [CrossRef] [Scilit]
  46. Zhang, Y.; Wang, B.; Li, J. Kangba Region of Sichuan Based on Swin Transformer Visual Model Research on the Identification of Facades of Ethnic Buildings. Sci. Rep. 2024, 14, 28742. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  47. Yuan, W.; Zhang, X.; Shi, J.; Wang, J. LiteST-Net: A Hybrid Model of Lite Swin Transformer and Convolution for Building Extraction from Remote Sensing Image. Remote Sens. 2023, 15, 1996. [Google Scholar] [CrossRef] [Scilit]
  48. Gibril, M.B.A.; Al-Ruzouq, R.; Shanableh, A.; Jena, R.; Bolcek, J.; Mohd Shafri, H.Z.; Ghorbanzadeh, O. Transformer-Based Semantic Segmentation for Large-Scale Building Footprint Extraction from Very-High Resolution Satellite Images. Adv. Space Res. 2024, 73, 4937–4954. [Google Scholar] [CrossRef] [Scilit]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Article Metrics

Citations

Article Access Statistics

Multiple requests from the same IP address are counted as one view.