1. Introduction
Autism spectrum disorder (ASD) is a neurodevelopmental disorder characterized primarily by social communication impairments, restricted interests, and repetitive stereotyped behaviors, with substantial individual variability and complex etiological mechanisms [
1]. Epidemiological studies have shown that the global prevalence and disease burden of ASD have generally increased in recent years, with a particularly pronounced burden among young children, highlighting the growing need for early screening and intervention [
2]. Traditional ASD screening methods mainly rely on caregiver-completed questionnaires or clinical observation, and they remain limited by subjectivity, delayed identification, and time-consuming procedures [
1,
3]. Therefore, developing objective, timely, and efficient auxiliary screening methods for ASD is of great significance for early identification and improved prognosis.
In recent years, with rapid advances in eye-tracking technology, computer vision, and artificial intelligence (AI), ASD auxiliary screening based on objective behavioral signals has become an increasingly active research topic. Eye-movement behavior can record learners’ fixation locations, region transitions, and temporal changes in visual scenes [
4]. However, traditional eye-tracking analysis methods mostly depend on statistical features such as fixation duration, fixation count, first fixation time, and dwell proportion within areas of interest, often combined with conventional machine learning algorithms for ASD auxiliary classification [
3,
4,
5,
6,
7]. These methods still require considerable manual processing and are limited in their ability to preserve the spatial structure and dynamic evolution of eye-movement behavior. Eye-tracking scanpaths (ETSP), which transform fixation points and saccadic paths into visual images, contain fine-grained information such as local trajectory morphology and texture distribution, while also encoding global semantic information such as overall fixation patterns, spatial layout, and cross-region dependencies. ETSPs have therefore become an important representation for visual attention modeling.
Previous studies have demonstrated that ETSP scanpath images can effectively support ASD identification and classification. Carette et al. [
8] transformed scanpaths into visual representations and verified the feasibility of using scanpath images for modeling visual attention patterns. The same research team later released an eye-tracking dataset for ASD research, providing a reproducible data basis for subsequent studies [
9]. Building on this dataset, Cilia et al. [
10] further validated the classification performance of CNNs, indicating the importance of local spatial texture features for visual attention pattern analysis. Kanhirakadavath et al. [
11] improved the representation of ASD-related visual attention differences through image augmentation and DNN-based modeling. Subsequently, Alsaidi et al. [
12] proposed a T-CNN architecture to enhance spatial feature extraction.
Benabderrahmane et al. [
13] combined scanpath images with sequential eye-movement signals and used GRU networks to model temporal dependencies. Mousli et al. [
14] used supervised contrastive learning to pretrain an encoder and then fine-tuned the classifier, improving feature stability under limited training samples and, to some extent, classification performance in cross-individual scenarios. Al-Adhaileh et al. [
15] further combined CNN and LSTM models for joint spatial–temporal modeling and achieved high classification performance on public datasets.
The broader eye-movement literature provides the behavioral context for these computational studies. Attention to socially informative regions, including the eye area, can follow different developmental trajectories in children later diagnosed with ASD [
16]. Meta-analytic evidence indicates that gaze differences between ASD and non-ASD groups are generally heterogeneous and depend on whether the displayed information is social or nonsocial [
17]. Stimulus type also materially affects measured social attention [
18], while scanpath-level analyses have reported both reduced social preference and increased persistence toward nonsocial content [
19]. Accordingly, the local branch of GLCF-Net may capture short-range trajectory density and morphology, whereas the global branch may capture broader spatial allocation and cross-region transitions; these computational features should not be interpreted as direct biological or diagnostic markers.
However, ASD is not a homogeneous condition; it involves substantial phenotypic heterogeneity and varied clinical outcomes. Eye-gaze profiles across children can be discrete and highly heterogeneous [
20,
21]. AI models that rely only on shallow features often fail to cope with cross-subject behavioral variability induced by such heterogeneity, resulting in a marked decline in generalization performance when tested on unseen participants [
22]. In the ETSP dataset, each child is typically associated with multiple scanpath images, which may share relatively stable individualized gaze habits, saccadic speed, trajectory density, and spatial preferences. If different images from the same participant appear simultaneously in the training and test sets, the model may learn participant-specific eye-movement patterns rather than transferable ASD-related visual attention patterns. Consequently, when facing unseen participants, the model may suffer from poor generalization, even though its reported performance appears falsely high.
For example, Cilia et al. [
10] further adopted participant-ID-based independent splitting, ensuring that ETSP images from the same participant were strictly assigned to the same subset. Their results showed that model accuracy decreased from 90% to 71%. In addition, the original augmentation strategy used by Kanhirakadavath et al. [
11] involved sample-level data leakage, where augmented images appeared simultaneously in the training and validation sets, leading to overly optimistic performance estimates. When the same study further evaluated a mini dataset containing one trial from each of 59 participants, the DNN model achieved an overall Accuracy of 72.88%. Moreover, although some subsequent studies did not use data augmentation and thereby avoided augmentation-induced leakage, they still did not sufficiently emphasize the principle of participant-independent splitting. Therefore, studies reporting accuracies above 90% in the existing literature may have limited practical significance if their generalization ability is weak [
12,
13,
14,
15].
In summary, insufficient model generalization is a common issue in ASD identification research. In ETSP scanpath datasets and related studies, two types of problems may lead to inflated evaluation results. The first is sample-level leakage: researchers augment the original images before performing random cross-validation on the entire augmented dataset, causing augmented copies of the same original image to be assigned to different training and validation folds. The second is the cross-subject generalization problem: even without augmentation, if the training and test sets are not divided by participant, different original images from the same participant may randomly appear in both sets. In this situation, the model has already “seen” the participant’s eye-movement pattern during training, thereby overestimating cross-individual generalization ability.
Eye-tracking scanpath images are essentially visual representations generated by mapping fixation points, saccadic paths, and their temporal changes. Local trajectory morphology, line density, color variation, and spatial neighborhood relationships between adjacent patches can reflect fine-grained differences in short-term visual exploration. In contrast, overall fixation layout, cross-region transition patterns, and spatial dependencies among different areas of interest reflect higher-level visual attention strategies.
In computer vision, CNNs are well suited for extracting local details, whereas Vision Transformers (ViTs) can complement them by capturing global context [
23,
24,
25,
26]. The complementarity between CNNs and ViTs has been widely discussed. For example, in multicellular morphology classification related to Alzheimer’s disease, researchers fused EfficientNet and ViT modules and demonstrated the effectiveness of collaborative modeling of local texture and long-range context for improving generalization on limited datasets [
27]. Li et al. [
24] proposed a parallel-branch CNN-Transformer module for change detection, using a convolutional branch to extract local features and a Transformer branch to extract global features. Through dependent and concurrent fusion, they constructed multiscale global–local representations.
Based on the above analysis, and considering that ETSP images contain both local saccadic trajectory details and global fixation distribution patterns, we extend the multiscale collaborative idea of CNN–ViT modeling to the ETSP auxiliary identification task and propose the Global–Local Collaborative Fusion Network, namely GLCF-Net. By combining local detail modeling with global relationship modeling, GLCF-Net can more comprehensively extract stable discriminative features from ETSP images, thereby improving within-dataset cross-participant generalization under participant-independent splitting.
The primary hypothesis was that global–local fusion would improve Accuracy and ROC-AUC relative to the matched ViT-B/16 configuration under participant-independent evaluation. Repeated grouped participant folds were used to characterize sensitivity to participant partitioning, while ablations assessed the contributions of the gate and CLS readout under participant-level aggregation. Comparisons with the CNN baselines were treated as reference comparisons because their trainable capacities differed. Throughout this article, “generalization” refers to unseen participants from the same dataset and acquisition conditions; independent cohorts, centers, eye trackers, stimuli, age groups, and demographic populations were not evaluated.
4. Discussion
This study focuses on ASD auxiliary screening in children and proposes a global–local collaborative fusion network, named GLCF-Net, based on eye-tracking scanpath images. The model was evaluated under a strict participant-independent split to examine its ability to distinguish visual attention patterns between children with ASD and typically developing (TD) children. Unlike conventional random image-level splitting, participant-independent splitting ensures that all scanpath images from the same child appear in only one of the training, validation, or test sets. This setting reduces the risk that the model memorizes participant-specific eye-movement patterns and provides a more cautious evaluation of its performance on unseen participants. Therefore, results obtained under this setting are more informative for assessing cross-participant auxiliary screening performance.
The repeated participant-level evaluation characterized sensitivity to participant partitioning. GLCF-Net achieved mean Accuracy 0.870 and ROC-AUC 0.936, with standard deviations of 0.049 and 0.028 across the outer folds. The corrected paired analysis did not show an Accuracy difference from any individual baseline after Holm correction; the observed fold-averaged performance differences are therefore interpreted descriptively. The separate fixed-fold analysis produced across-initialization standard deviations of 0.014 for Accuracy, 0.021 for F1-score, and 0.004 for ROC-AUC. These values quantify stochastic training variability conditional on one saved set of participant folds and do not establish stability across independent cohorts.
4.1. Classification Performance of GLCF-Net Under Participant-Independent Splitting
Under the strict participant-independent setting, GLCF-Net achieved an Accuracy of 83.52%, a Macro F1-score of 82.68%, and a ROC-AUC of 90.27% on the test set. Its overall performance was higher than that of the selected CNN-based baselines and the single-branch ViT model. Compared with random image-level splitting, participant-independent splitting avoids assigning different scanpath images from the same child to both training and test sets, thereby reducing the risk of performance overestimation caused by participant-level information overlap.
Previous ETSP-based studies have reported high classification performance under random splitting or cross-validation after data augmentation. However, such results may be affected by sample-level leakage or participant-level overlap. In contrast, Cilia et al. [
10] reported an accuracy of approximately 71% under strict participant-independent splitting, while Kanhirakadavath et al. [
11] reported an overall Accuracy of 72.88% on a mini dataset containing one trial per participant. In a similar evaluation context that emphasizes participant independence, the proposed model still achieved relatively good performance, suggesting that jointly modeling global fixation distribution and local trajectory structure is useful for ETSP-based ASD auxiliary screening.
From the class-wise results, GLCF-Net achieved a recall of 90.57% for the TD class, which was higher than the ASD recall of 73.68%. This indicates that the model identified the visual attention patterns of TD children more consistently, whereas some ASD samples were still missed, with certain ASD scanpath images being incorrectly classified as TD. From an eye-movement perspective, this may suggest that the visual attention patterns of some ASD participants partially overlapped with those of TD participants, particularly in terms of fixation distribution, scanpath density, or local trajectory organization. This observation is consistent with the substantial heterogeneity of ASD-related gaze behaviors [
20,
21] and further suggests that a single ETSP image may not fully capture participant-level variability in visual attention strategies.
4.2. Effect of Global–Local Collaborative Modeling
The baseline results show that ResNet-50 achieved relatively better Accuracy and F1-score among the CNN-based models, suggesting that residual convolutional structures can effectively capture local trajectory morphology, path density, and spatial neighborhood patterns in ETSP images. Meanwhile, the ViT model showed a certain advantage in ROC-AUC, indicating that global self-attention is helpful for modeling long-range dependencies among image patches and the overall fixation distribution [
25]. These results suggest that both local trajectory structures and global spatial layouts in ETSP images may contain discriminative information for ASD/TD auxiliary recognition. Based on this observation, GLCF-Net introduces a lightweight residual convolutional branch and combines it with a ViT-based global branch for collaborative feature modeling.
The performance of ResNet-50 indicates the potential usefulness of residual convolutional structures for ETSP local feature extraction, but does not isolate the contribution of the lightweight residual block in GLCF-Net. The ablation experiments therefore evaluated the role of the local branch. When only the ViT global branch was retained, Accuracy, Macro F1-score, and ASD recall were lower than those of the complete GLCF-Net. The lower ASD recall suggests that a single ViT branch may not sufficiently capture local scanpath morphology and fine-grained spatial structures. Because the CNN baselines trained only their final linear classifiers, whereas ViT-B/16 and GLCF-Net trained larger task-specific components, the CNN comparisons do not isolate architecture from trainable capacity. The matched ViT-B/16 comparison and the same-backbone ablations provide the more direct evidence for the added local branch, gate, and CLS readout under the present training protocol.
This finding is consistent with the characteristics of ETSP images. ETSP scanpath images contain not only global fixation distribution and cross-region transition patterns, but also local details such as trajectory density, local path morphology, short-range spatial neighborhood structure, and color variations. A single-scale or single-architecture model may be insufficient to represent these heterogeneous patterns. By using the ViT branch to model global attention relationships and the CNN branch to complement local trajectory structures, GLCF-Net can extract visual attention features from two different perspectives, thereby improving the representation of ASD–TD differences in scanpath images. Therefore, the performance improvement over single CNN or ViT baselines suggests that autism-related visual attention patterns in ETSP images may involve both fine-grained local trajectory characteristics and broader global attention allocation strategies.
This interpretation is computational rather than mechanistic. Prior eye-tracking evidence indicates that ASD-related gaze differences vary by social versus nonsocial content, stimulus type, age, and study design [
17,
18], and scanpath differences may include both reduced social preference and increased nonsocial preference [
19]. The CNN branch can be sensitive to local density, morphology, and short-range patch neighborhoods, while the ViT branch can represent overall layout and long-range spatial relations. Without stimulus-aligned areas of interest or raw fixation timing, however, neither branch can be assigned to a specific oculomotor or biological mechanism.
4.3. Interpretation of the Gated Adaptive Fusion Module
The ablation results show that the ungated model outperformed the ViT-only model in Accuracy, Macro F1-score, and ASD recall. This indicates that the introduction of the CNN local branch can improve ETSP image classification even when a relatively simple fusion strategy is used. On this basis, the complete GLCF-Net further incorporates a gated adaptive fusion module and achieves additional improvements in Accuracy, Macro Precision, Macro Recall, Macro F1-score, and ROC-AUC. These results suggest that the gated mechanism has a positive effect on integrating global and local features.
The main function of the gated adaptive fusion module is to dynamically adjust the contribution of ViT global features and CNN local features according to the responses of different samples, patches, and channels [
28]. For samples that rely more on the overall fixation layout, the model can assign greater importance to the ViT branch. For samples with more informative local trajectory morphology, path density, or spatial neighborhood structures, the model can retain more information from the CNN branch. Therefore, compared with ungated fusion, the gated module provides a more flexible feature integration mechanism and allows the model to adaptively balance local and global information according to the characteristics of different scanpath images.
However, the effect of the gated module should also be interpreted cautiously. The complete GLCF-Net and the ungated model both achieved an ASD recall of 73.68%. This indicates that although the gated module improved the overall performance, ASD precision, and ASD F1-score, it did not further improve ASD recall. In other words, the current gated fusion strategy mainly improved the overall decision boundary and prediction precision, but its effect on reducing missed ASD cases remains limited. Considering that auxiliary screening tasks require high sensitivity to potential ASD cases, future studies may incorporate class-sensitive loss functions, participant-level probability aggregation, raw temporal eye-tracking features, or multimodal behavioral indicators to further improve the sensitivity of ASD recognition.
Across the repeated participant folds, the fitted gate assigned less than half of its average weight to the ViT branch in both groups. Within the matched first repeat, the equal-fusion, no-CLS, and rotation variants each yielded lower Accuracy, F1-score, and ROC-AUC than the complete model on the same five participant folds. The separate five-run initialization analysis was conducted only for the complete GLCF-Net.
4.4. Participant-Level Prediction Stability and Misclassification Analysis
In addition to overall classification metrics and ablation results, the frame-level probability distribution at the participant level provides supplementary evidence for evaluating model stability. Since each participant usually corresponds to multiple ETSP images, the prediction of a single image may be influenced by stimulus content, image quality, scanpath sparsity, or short-term attentional state. Therefore, evaluating the model only by image-level Accuracy or AUC is not sufficient. It is also necessary to examine whether the prediction probabilities of multiple images from the same participant are relatively consistent.
At the participant level, most TD participants showed ASD prediction probabilities below the threshold of 0.5, whereas most ASD participants showed probabilities above this threshold. This indicates that the model produced prediction trends that were generally consistent with the true labels for most unseen participants. Thus, GLCF-Net not only achieved good image-level classification performance, but also showed relatively stable prediction tendencies at the participant level.
Nevertheless, some boundary cases and misclassifications were still observed. For example, TD-46 was misclassified as ASD, suggesting that this participant’s scanpath images may share certain local trajectory patterns or global fixation distributions with ASD samples. In addition, some frame-level prediction probabilities of ASD-05 and ASD-10 were close to or below the classification threshold, indicating relatively large prediction fluctuations across different images from the same ASD participant. This phenomenon suggests that a single ETSP image may not be sufficient to represent the overall visual attention characteristics of a child. The model may still face difficulties when dealing with participants whose visual attention patterns are highly heterogeneous or overlap with those of the other class. Future work may adopt participant-level probability averaging, voting-based fusion, or temporal aggregation strategies to reduce the influence of single-frame prediction fluctuations on the final auxiliary screening result.
Seven of the 54 participants were misclassified after repeat averaging, including ASD-05, ASD-16, ASD-10, ASD-22, TD-45, TD-44, and TD-49. Three errors lay within 0.05 of the decision threshold, and ASD-22 had a between-repeat probability standard deviation of 0.2879. The errors were therefore concentrated in a small subset of boundary or partition-sensitive cases, although the available image data do not identify participant-level clinical causes.
4.5. Generalization Boundary, Image Representation, and Future Work
Participant-independent evaluation prevents the same child from appearing in training and evaluation subsets, but it demonstrates only within-dataset generalization to unseen participants under the same acquisition conditions. It does not establish external or clinical generalization across independent datasets, centers, eye trackers, age groups, demographic populations, stimulus sets, or experimental paradigms. The effective sample size of 54 image-contributing participants, rather than 547 repeated images, limits precision. The five initialization runs on unchanged participant folds provided a separate estimate of run-to-run stochastic training variability, but this estimate remains conditional on one saved fold set, the present training configuration, and the available computing environment. It does not replace evaluation across independently sampled cohorts, acquisition systems, or broader training configurations.
ETSP images preserve spatial trajectory geometry, local path density, and some color-encoded progression, but rasterization cannot fully preserve ordered fixation coordinates, timestamps, durations, pupil measures, or stimulus-aligned areas of interest. Sequence-specific architectures such as Gazeformer [
29] and the Gaze Scanpath Transformer [
30] model ordered gaze information more directly. Future studies should compare image, raw-sequence, and joint representations using identical participants and stimuli, followed by preregistered external validation on independent multi-center cohorts and devices.
The measured inference latency indicates that the implementation can rapidly classify a completed ETSP image. It does not yet constitute closed-loop online eye tracking because gaze acquisition and scanpath-image construction occur before model inference. A future online system should accept streaming gaze coordinates, define a prespecified decision window, and validate end-to-end latency and calibration prospectively.
5. Conclusions
This study addressed the task of classifying children with autism spectrum disorder (ASD) and typically developing (TD) children using eye-tracking scanpath (ETSP) images, and proposed a global–local collaborative fusion network, named GLCF-Net. Based on a shared patch embedding layer, the proposed method uses a CNN local branch to extract local scanpath trajectory features and a ViT global branch to model cross-region fixation dependencies. A gated adaptive fusion module is then employed to dynamically integrate these two types of features. GLCF-Net combines local trajectory structures and global fixation distributions and was evaluated for unseen participants from the same dataset under participant-independent splitting.
The experiments were conducted on a public ETSP scanpath image dataset. A strict participant-independent splitting strategy was adopted to prevent scanpath images from the same participant from appearing simultaneously in the training and test sets. The experimental results show that GLCF-Net achieved an Accuracy of 83.52%, a Macro F1-score of 82.68%, and a ROC-AUC of 90.27%. Its overall performance was better than that of baseline models, including VGG19, ResNet-50, DenseNet, and ViT. Compared with existing ETSP-related studies that reported results under participant-independent or more conservative evaluation settings, the proposed method showed relatively good auxiliary recognition performance in terms of Accuracy and ROC-AUC. This suggests that global–local collaborative modeling is effective to some extent under strict cross-participant evaluation.
The ablation experiments further indicate that the CNN local branch can complement the limitation of ViT in modeling local trajectory structures, enabling the model to better capture local spatial details in ETSP images. The gated adaptive fusion module improves the flexibility of integrating global and local features and further enhances the overall classification performance. In addition, the participant-level prediction probability distribution shows that the model can produce prediction trends consistent with the true labels for most unseen participants. However, misclassification still occurs for boundary samples and individuals with highly heterogeneous visual attention patterns.
Overall, GLCF-Net learned ETSP-based visual attention representations under participant-independent splitting, providing auxiliary value for analyzing visual attention patterns in children with ASD. Nevertheless, this study is still limited by the small scale of the public dataset and by the fact that ETSP images compress part of the original temporal eye-tracking information. Future work may incorporate larger-scale and multi-center datasets, as well as raw eye-tracking sequences, AOI-based features, and multimodal behavioral information, to further validate the generalization ability, stability, and interpretability of the model.
After repeat averaging, participant-level Accuracy was 87.0%, F1-score was 86.3%, and ROC-AUC was 93.7%; the corresponding fold standard deviations for Accuracy, F1-score, and ROC-AUC were 0.049, 0.045, and 0.028. Independent, sequence-based, prospective, and clinical validation remains necessary to establish performance across acquisition settings and populations.