Next Article in Journal
Prestress Loss and Bi-Directional Prestress Effect of a Large-Span U-Shaped Aqueduct: Field Test and Numerical Analysis
Next Article in Special Issue
Spiking Neural Networks for the Analysis of Physiological Signals
Previous Article in Journal
Experimental Investigation of Printing Parameters in SLA 3D Printing of Plant-Based Resin Using Taguchi Method: Effects on Tensile Properties and Fracture Surface Morphology
Previous Article in Special Issue
Imaging Engineering and Artificial Intelligence in Urinary Stone Disease: Low-Dose Computed Tomography, Spectral Technologies, and Predictive Models
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Reproducible RGB Video Screening of Amyotrophic Lateral Sclerosis Using Spherical-Coordinate Landmark Correlations

by
Daniela Suárez-Hernández
,
Sulema Torres-Ramos
,
Stewart R. Santos-Arce
and
Israel Román-Godínez
*
División de Tecnologías para la Integración Ciber-Humana, CUCEI-Universidad de Guadalajara, Blvd. Marcelino García Barragán 1421, Olímpica, Jalisco 44430, Mexico
*
Author to whom correspondence should be addressed.
Eng 2026, 7(5), 238; https://doi.org/10.3390/eng7050238
Submission received: 26 March 2026 / Revised: 1 May 2026 / Accepted: 11 May 2026 / Published: 14 May 2026

Abstract

Amyotrophic lateral sclerosis (ALS) is a progressive neurodegenerative disorder for which delayed recognition may limit timely clinical management. This study investigates a reproducible computer-aided screening approach based on facial motion analysis from standard RGB video recorded during the diadochokinetic /pataka/ task. Facial landmarks were extracted using a face-mesh model and mapped into spherical coordinates to represent facial motion trajectories. Coordinated facial behavior was characterized through pairwise Pearson correlation matrices computed between landmark trajectories, yielding correlation-based descriptors of inter-region motion patterns. We compared a domain-informed Manual-24 reference configuration with data-driven feature-selection strategies (ElasticNet and mRMR) under a leakage-aware nested cross-validation design using the Toronto NeuroFace dataset. Performance was reported as mean ± standard deviation across outer folds, with sensitivity emphasized because of its relevance for screening-oriented applications. The primary configuration (mRMR, k = 3 , ϕ + kNN) achieved 61.11 ± 19.24% accuracy, 61.11 ± 9.62% sensitivity, and 61.11 ± 34.70% specificity. These results suggest that correlation-derived coordination patterns contain discriminative information for ALS/HC separation, although fold-level variability indicates that performance should be interpreted cautiously. Task-aligned comparisons with prior /pataka/-based studies highlight the influence of sensing modality, evaluation level, and uncertainty reporting on apparent performance. Overall, correlation-based facial motion descriptors combined with leakage-aware feature selection provide a transparent proof-of-concept framework for RGB video-based ALS screening, motivating validation on larger cohorts and independent datasets.

1. Introduction

Amyotrophic lateral sclerosis (ALS) is a progressive neurodegenerative disorder characterized by the degeneration of upper and lower motor neurons, leading to progressive muscle weakness, paralysis, and ultimately respiratory failure. The disease affects several motor functions, including limb mobility, speech production, swallowing, and respiratory control, significantly impairing the quality of life of affected individuals. Despite extensive research, ALS remains incurable, and current treatments are primarily palliative, aiming to slow disease progression and manage symptoms through multidisciplinary clinical care [1,2].
Early detection of ALS is clinically important because timely diagnosis allows earlier implementation of supportive interventions that may improve patient outcomes and quality of life. However, diagnosis remains challenging due to the absence of specific biomarkers and the heterogeneity of early symptoms. ALS is typically diagnosed through a combination of clinical examination, neurophysiological testing, neuroimaging, and the exclusion of other neurological conditions. Diagnostic delays and misdiagnoses are common in the early stages of ALS due to the absence of specific biomarkers and the heterogeneity of symptoms [2,3].
In recent years, computational approaches based on biomedical signals have been explored to support ALS diagnosis. Machine learning techniques have been applied to different modalities, including electromyography signals, speech analysis, and other behavioral signals, aiming to detect subtle motor impairments that may not be easily observable during standard clinical evaluations [4]. Among these modalities, facial motion analysis has recently gained increasing attention because ALS frequently affects the orofacial muscles involved in speech production, making facial movement patterns a potential source of digital biomarkers for disease detection.
Several studies have investigated the use of facial motion analysis for neurological disease detection. For example, Bandini et al. [5] proposed an automatic method for ALS detection based on video-based analysis of facial movements during speech and non-speech tasks. Their approach analyzed motion characteristics such as velocity, symmetry, and movement range, demonstrating that facial movement analysis can provide valuable information for identifying ALS-related motor impairment. Furthermore, the Toronto NeuroFace dataset introduced by Bandini et al. provides annotated video recordings of individuals with ALS, stroke patients, and healthy controls performing orofacial tasks, enabling the development of computational models for automated facial motion analysis in neurological disorders [6].
Facial analysis techniques have also been successfully applied in the study of other neurological conditions. Jin et al. demonstrated that facial expression dynamics extracted from video recordings can be used to detect Parkinson’s disease with high accuracy using machine learning techniques [7]. More recently, graph-based approaches have been proposed to model spatial relationships between facial landmarks for ALS detection, highlighting the potential of structured representations of facial motion patterns in neurological assessment [8]. These studies suggest that facial landmark trajectories may capture meaningful patterns associated with neuromotor impairment.
Building upon these advances, previous RGB-based work investigated the use of facial biometric data to detect ALS-related motor impairment during a diadochokinetic speech task [9,10]. In that setting, facial landmarks were transformed into spherical coordinates and a reduced landmark subset was obtained through correlation-based redundancy filtering before classification. Although this demonstrated the feasibility of detecting ALS-related patterns from facial motion, the selection strategy was primarily designed to reduce redundancy among landmark trajectories rather than to identify landmark-pair relationships that are maximally informative for ALS/HC discrimination.
Despite recent advances in video-based facial motion analysis for neurological assessment, several methodological gaps remain. First, some high-performing approaches rely on specialized acquisition systems, such as depth cameras or marker-based motion capture, which may limit applicability in resource-constrained clinical environments. Second, graph-based or highly engineered representations can increase algorithmic complexity and may require larger datasets to generalize reliably. Third, although spherical-coordinate landmark trajectories and correlation-based redundancy filtering have been explored previously, the use of correlation-structure descriptors combined with leakage-aware, relevance-based feature selection remains insufficiently studied for /pataka/-based ALS screening. These gaps motivate a framework that is accessible, reproducible, and explicitly designed to quantify uncertainty under small-cohort conditions.
Hence, based on previous RGB-based facial-motion analyses and to address these limitations, the present manuscript focuses on methodological extensions and a more rigorous evaluation of correlation-based descriptors for /pataka/-based ALS screening. Specifically, this work makes three main contributions. First, it compares a domain-informed Manual-24 reference configuration against automatic feature-selection strategies applied to landmark-pair correlation descriptors, emphasizing the role of relevance-aware selection through mRMR. Second, it implements feature selection strictly within the training portion of each outer cross-validation fold, reducing the risk of information leakage and improving the reproducibility of the evaluation protocol. Third, it reports fold-level mean ± SD, feature-selection behavior, and sensitivity–specificity trade-offs under task-aligned /pataka/ benchmarking, providing a more transparent assessment of variability in a small clinical cohort and enabling cautious comparison with prior studies.

2. Materials and Methods

The graphical overview shown in Figure 1 illustrates the proposed strategy for ALS detection using RGB video and spherical coordinate reference-point correlations.

2.1. Dataset and Recording Protocol

This study was conducted using the Toronto NeuroFace Dataset, an available research dataset designed for facial motion analysis in individuals with neurological disorders [6]. The dataset was collected by the Vocal Tract Visualization and Bulbar Function Laboratory at the University Health Network (Toronto, ON, Canada) and includes video recordings of both healthy control (HC) subjects and individuals diagnosed with amyotrophic lateral sclerosis (ALS) by clinical specialists.
Video recordings were acquired using an Intel® RealSense™ SR300 camera (Intel Corporation, Santa Clara, CA, USA) positioned at an approximate distance of 30–60 cm from the participant, under uniform lighting conditions and in an unconstrained clinical environment [6]. Videos were recorded at a spatial resolution of 640 × 480 pixels and a temporal resolution of approximately 50 frames per second.
For each subject, several standardized speech-related orofacial tasks were recorded, including repetitive syllable articulation sequences designed to evaluate articulatory coordination. In particular, participants performed a diadochokinetic speech task consisting of rapid repetitions of the syllable sequence /pataka/ within a single breath, a standard task for assessing speech motor control involving coordinated movements of the lips, anterior tongue, and posterior tongue [5,11]. In this study, only frontal facial video recordings corresponding to the /pataka/ task were analyzed.
A total of 18 video recordings were included in the final dataset, after excluding two recordings due to video encoding corruption. Each recording was treated as an independent sample and labeled according to the subject group defined in Table 1. Each subject contributed one /pataka/ recording to the analyzed subset; therefore, subject-wise grouping was not required for cross-validation. No additional demographic or clinical variables were incorporated into the modeling process, in order to focus exclusively on facial motion patterns derived from video data. All analyses were performed using the same set of recordings across subjects to ensure consistency.
All data used in this study were anonymized and publicly released by the dataset providers. Ethical approval and informed consent were obtained by the original data collection team; therefore, no additional institutional review board approval or participant consent was required for the present analysis [6].

2.2. Facial Landmark Extraction

Facial landmarks were extracted from each video frame using the MediaPipe® Face Mesh model, a real-time markerless approach for dense facial geometry estimation from monocular RGB images [12,13]. The model outputs a set of 468 facial landmarks per frame, providing normalized Cartesian coordinates ( x , y ) relative to the image frame, along with an estimated depth component ( z ) derived from a learned facial geometry representation.
MediaPipe® Face Mesh was selected due to its robustness to moderate variations in head pose, illumination, and facial appearance, as well as its suitability for non-invasive and markerless facial motion analysis in unconstrained recording environments [12]. This markerless paradigm is suitable for analyzing articulatory facial movements in neurological populations, avoiding the use of physical markers that may interfere with natural speech production or cause discomfort in clinical settings [14].
From the full set of detected landmarks, a subset of 54 landmarks was selected to represent anatomically and functionally relevant regions involved in speech-related facial motion, with a primary focus on the lips, jaw, chin, and lower facial regions. This subset was defined to preserve bilateral coverage of the lower face and perioral region, allowing facial symmetry and inter-region coordination to be characterized during the diadochokinetic /pataka/ task. The choice was guided by the physiological role of the lips and jaw in articulatory movement and by prior video-based studies showing that bulbar motor impairment in ALS affects speech-related orofacial motion patterns [5,6]. Landmarks from upper facial and periocular regions were excluded because they are less directly involved in the articulatory gestures required for /pataka/. This anatomically constrained subset also reduces the dimensionality of the initial correlation space before automatic feature selection, which is important in a small-cohort setting where using all 468 MediaPipe landmarks would substantially increase the number of pairwise correlation features and the risk of unstable model estimates.
Landmark coordinates were extracted for each frame of the video sequences without temporal smoothing at this stage, preserving the raw temporal dynamics of facial motion. All subsequent processing and feature construction steps were performed using these frame-level landmark trajectories.
Face-mesh detection can occasionally fail due to rapid head motion, partial occlusions, or motion blur. In such cases, frames in which the face mesh is not detected are excluded from the landmark time series. Subsequent computations (spherical mapping and correlation estimation) are performed using only the remaining valid frames for each recording.

2.2.1. Spherical Coordinate Mapping

To characterize coordinated facial motion independently of global translation and to separate motion magnitude from directional components, each landmark trajectory was mapped from Cartesian coordinates to a spherical representation. In this formulation, each landmark position is expressed relative to a fixed anatomical reference point (defined in Section 2.2.2), yielding a frame-wise displacement vector that captures local facial motion patterns. This relative mapping reduces sensitivity to camera framing and partial head translation, and it provides a compact way to analyze motion directionality through angular components.
Spherical coordinate representations have been explored in face analysis to obtain compact geometric descriptors and to support more invariant and interpretable decompositions of facial shape and expression-related geometry [15,16]. In our context, the radial component r captures the magnitude of motion relative to the reference point, while the angular components ( θ , ϕ ) describe the direction of displacement on the sphere. By analyzing the three components independently, we can investigate which aspects of facial motion carry the most discriminative signal for ALS detection in the /pataka/ task.
A practical consideration in RGB-based landmark pipelines is the reliability of the out-of-plane coordinate. Because the z coordinate is estimated rather than directly measured, the azimuthal component ϕ —which depends explicitly on Δ z —may be more sensitive to depth or pose estimation noise than r or θ . This motivates our component-wise reporting and interpretation, and it provides a plausible explanation for differences in performance across components observed in similar RGB-based settings.

2.2.2. Spherical Coordinate Transformation

Let p j ( t ) = { x j ( t ) , y j ( t ) , z j ( t ) } denote the 3D landmark coordinates of landmark j at frame t, as provided by the face landmark detector, and let p ref ( t ) = { x ref ( t ) , y ref ( t ) , z ref ( t ) } denote the reference landmark coordinates at the same frame. In our implementation, the reference landmark is fixed to the detector landmark with index 1, located near the nasal bridge in the MediaPipe® Face Mesh landmark topology. This landmark was selected as a central and relatively stable facial anchor because it is less directly affected by the large articulatory deformations of the lips and jaw during /pataka/ production. Using this point as a frame-wise reference allows each landmark trajectory to be expressed as a relative displacement within a common facial coordinate system, reducing sensitivity to global translation, camera framing, and partial head motion while preserving local motion patterns in the lower face.
For each j-th landmark and frame t, we compute the displacement vector relative to the reference landmark:
Δ p j ( t ) = p j ( t ) p ref ( t ) = { Δ x j ( t ) , Δ y j ( t ) , Δ z j ( t ) } .
The reference landmark itself was excluded from the subsequent feature construction to avoid introducing a trivial zero displacement vector.
The corresponding spherical coordinates { r j ( t ) , θ j ( t ) , ϕ j ( t ) } are then computed using the convention implemented in the experimental pipeline:
r j ( t ) = Δ x j ( t ) 2 + Δ y j ( t ) 2 + Δ z j ( t ) 2 ,
θ j ( t ) = atan 2 Δ x j ( t ) 2 + Δ z j ( t ) 2 , Δ y j ( t ) ,
ϕ j ( t ) = atan 2 Δ x j ( t ) , Δ z j ( t ) .
Angles θ j ( t ) and ϕ j ( t ) are expressed in degrees and normalized to the range [−180°, 180°) for consistency across sequences.
This relative formulation emphasizes coordinated local facial motion patterns while reducing sensitivity to global translation and partial head motion. For reproducibility, the above equations define the exact spherical coordinate convention used throughout the pipeline; no alternative ( θ , ϕ ) conventions were employed.

2.3. Correlation-Based Feature Construction and Selection

Facial motion trajectories extracted from the selected landmarks were transformed into spherical coordinates ( r , θ , ϕ ) , as described in the previous section. To capture coordinated motion patterns between facial regions, correlation matrices were computed from the temporal trajectories of the selected landmarks for each spherical coordinate component. These matrices describe pairwise relationships between landmarks and constitute the basis for feature construction in the proposed framework.

2.3.1. Feature Vector Construction

Before computing Pearson correlation matrices, landmark trajectories for each recording were standardized using z-score normalization across time (per recording and per landmark series). This step ensures that correlation estimates reflect co-variation patterns rather than differences in scale across landmarks. Then, for each video recording and for each spherical coordinate component ( r , θ , ϕ ) , a Pearson correlation matrix was computed using the temporal trajectories of the selected landmarks. The upper triangular elements of each correlation matrix, excluding the diagonal, were vectorized to form the input feature vector for subsequent analysis. This representation emphasizes inter-landmark coordination patterns rather than absolute displacement magnitudes and yields a compact, structured description of facial motion dynamics.
Recordings may differ in duration and therefore in the number of valid frames. Because Pearson correlation is computed across facial motion trajectories, the resulting correlation matrices are estimated directly from the available valid frames for each recording, without temporal resampling, padding, or sequence alignment. This design preserves each recording’s native timing while producing a fixed-size correlation representation (54 × 54 per coordinate component) for downstream modeling.
Correlation- and covariance-based representations have been widely adopted to capture dependency structures and coordinated patterns in facial and motion-related data, providing informative descriptors for classification tasks involving structured signals [17,18].

2.3.2. Feature Selection Strategies

To evaluate the impact of different selection paradigms on classification performance, three feature selection strategies were considered:
  • Manual landmark selection (baseline). A fixed subset of 24 facial landmarks previously proposed based on domain-informed anatomical considerations related to speech-related facial motion was used as a baseline configuration [9,10]. In this setting, correlation matrices and feature vectors were constructed exclusively from these 24 landmarks, and no additional automatic feature selection was applied. This strategy serves as a manually defined reference against which data-driven selection methods can be compared.
  • ElasticNet-based feature selection. Automatic feature selection was performed using ElasticNet regularization [19] applied to the correlation-based feature vectors derived from the full set of 54 landmarks. ElasticNet combines l 1 and l 2 penalties, promoting sparse yet stable solutions in the presence of highly correlated predictors, which is particularly suitable for correlation-based descriptors where strong feature dependencies are expected. This approach has been successfully applied in multiple biomedical machine learning studies, including radiomics-based cancer prognosis and signal-based diagnostic tasks such as EEG-based mental health assessment [20,21]. Feature selection was conducted using only the training data within each cross-validation fold to prevent information leakage from the test set, retaining features associated with non-zero coefficients.
  • mRMR-based feature selection. As an alternative data-driven approach, feature selection was also performed using the minimum Redundancy Maximum Relevance (mRMR) criterion [22]. This method selects features that maximize mutual information with the class labels while minimizing redundancy among the selected features. mRMR has demonstrated strong performance in biomedical applications, particularly in radiomics-based outcome prediction tasks involving highly correlated features [23]. As with ElasticNet, mRMR-based selection was applied exclusively to the training data within each cross-validation fold.
Each feature selection strategy—manual landmark selection, ElasticNet-based selection, and mRMR-based selection—was applied independently for each spherical coordinate component ( r , θ , ϕ ) , enabling a direct comparison under identical modeling and validation protocols.
To characterize the behavior of the automatic feature-selection stages, we also report the number of retained features per outer fold for ElasticNet and mRMR configurations. This analysis allows us to assess whether the automatic selectors produced compact and stable feature subsets across validation splits, in addition to their downstream classification performance.

2.4. Machine Learning Models and Validation Protocol

The correlation-based feature vectors described in Section 2.3 were used as inputs to a set of supervised machine learning classifiers. Given the limited dataset size and the structured nature of the extracted features, classical machine learning models were selected due to their robustness, interpretability, and suitability for small to medium-sized biomedical datasets where the number of features may exceed the number of observations.
The classifiers evaluated included the k-Nearest Neighbors (kNN), Decision Tree (DT), Random Forest (RF), and Multi-Layer Perceptron (MLP) models. These algorithms represent complementary learning paradigms, including instance-based learning, tree-based methods, ensemble learning, and shallow neural networks, enabling a comprehensive comparison under a unified experimental framework.
For each classifier, three feature selection configurations were evaluated: (i) a manually defined baseline using a fixed set of 24 facial landmarks, (ii) automatic feature selection using ElasticNet regularization, and (iii) automatic feature selection using the mRMR criterion. In all cases, feature selection—when applicable—was performed using training data only within each cross-validation fold, ensuring that no information from the test data was used during feature selection.
Model performance was evaluated using stratified k-fold cross-validation with k = 3 and k = 6 folds. These values were selected to balance computational efficiency and robustness of performance estimation. For each validation scheme, hyperparameter optimization was conducted using an inner stratified k-fold cross-validation procedure ( k = 3 ) applied exclusively to the training data, resulting in a nested cross-validation design [24]. This validation strategy helps obtain more reliable performance estimates in small-sample settings.
Hyperparameter search spaces for all classifiers and feature selection methods were predefined and kept fixed across experiments to ensure fair comparisons and reproducibility. Hyperparameter optimization was performed by maximizing the F1-score within the inner cross-validation loop. The complete set of hyperparameter ranges explored for each model and feature selection strategy is reported in Table S1 of the Supplementary Materials.
For example, the multilayer perceptron (MLP) was included as a compact neural-network baseline and was evaluated under the same nested validation protocol as the other classifiers. The MLP used a single hidden layer, with hidden-layer size tuned over { ( 2 ) ,   ( 50 ) ,   ( 100 ) } , activation function tuned over { tanh , relu , logistic } , solver tuned over { adam , lbfgs } , and initial learning rate tuned over { 10 4 ,   10 3 ,   10 2 } . The maximum number of iterations was fixed to 2000 and the random seed was fixed to 0.
Model performance was evaluated using standard classification metrics derived from the confusion matrix, including accuracy, sensitivity, specificity, precision, and F1-score. All metrics were computed on the held-out test folds of the outer cross-validation procedure.

3. Results

3.1. Overview of Evaluated Configurations

We evaluated three feature-selection strategies—(i) a domain-informed Manual-24 baseline, (ii) ElasticNet-based selection, and (iii) mRMR-based selection—combined with three spherical-coordinate components ( r , θ , ϕ ) and four classifiers (DT, RF, kNN, and MLP). Performance is reported as mean ± standard deviation across outer cross-validation folds, using two validation protocols ( k = 3 and k = 6 ). Unless stated otherwise, we treat the mRMR configuration evaluated under k = 3 as the primary setting for the /pataka/ task, and we use Manual-24 ( k = 3 ) as a reference baseline to contextualize the effect of automatic feature selection. Given the small number of samples and outer folds, comparisons across configurations were interpreted descriptively rather than as formal inferential tests. Therefore, differences between models are discussed in terms of sensitivity–specificity trade-offs and fold-level variability, as reflected by the reported mean ± SD values.

3.2. Results Under 3-Fold Cross-Validation (k = 3)

Table 2 summarizes the Manual-24 baseline. Under k = 3 , the highest sensitivity was obtained by kNN using the ϕ component (77.78 ± 19.24), although this was accompanied by a low specificity (33.33 ± 33.34), indicating a tendency toward false positives. The MLP classifier using ϕ showed a more balanced profile (sensitivity 61.11 ± 34.70; specificity 66.67 ± 33.34), achieving the best overall F1-score in the baseline setting (61.90 ± 26.51). Overall, baseline performance suggests that ϕ -based correlation features may contain the most discriminative signal for ALS detection, whereas r and θ yielded lower and more variable sensitivities depending on the classifier.
ElasticNet-based selection (Table 3) did not consistently increase sensitivity relative to the Manual-24 baseline. The strongest sensitivity under ElasticNet occurred for kNN with θ (66.67 ± 33.34), while ϕ -based configurations remained moderate (maximum sensitivity 50.00 ± 16.67). In general, ElasticNet tended to produce more stable performance in some configurations (e.g., identical accuracy values for DT and MLP under ϕ ), although the sensitivity gains were not consistent when compared with the Manual-24 baseline.
In contrast, mRMR-based selection (Table 4) produced several competitive configurations under k = 3 . The highest sensitivity within this setting was obtained by DT using ϕ (72.22 ± 25.46), although this configuration also showed substantial variability and lower specificity (52.78 ± 41.11). The ϕ + kNN configuration, selected as the primary setting for task-aligned comparison, provided a more balanced sensitivity–specificity profile (61.11 ± 9.62 sensitivity and 61.11 ± 34.70 specificity) with lower sensitivity dispersion. These results suggest that mutual-information-based selection can identify informative landmark-pair correlations, but performance should be interpreted in terms of the sensitivity–specificity trade-off and fold-level variability rather than as uniform superiority of a single classifier.

3.3. Results Under 6-Fold Cross-Validation (k = 6)

Under k = 6 , the Manual-24 baseline (Table 5) again showed its strongest performance for the ϕ component. The highest sensitivity in the baseline setting was obtained by kNN with ϕ (66.67 ± 40.82), achieving the highest F1-score among baseline k = 6 configurations (52.78 ± 26.70). However, this was again associated with reduced specificity (41.67 ± 37.64), reflecting a sensitivity–specificity trade-off similar to the k = 3 setting. In contrast, baseline performance using r was consistently weak for DT and RF (sensitivity 0.00 ± 0.00), suggesting that radial correlations alone do not provide robust discriminative information in this dataset.
ElasticNet results under k = 6 (Table 6) showed moderate sensitivity for ϕ using DT and RF (both 50.00, with high variability), but did not exceed the best baseline sensitivity. Moreover, several r- and θ -based configurations exhibited low sensitivity (including θ with DT at 0.00 ± 0.00), indicating that ElasticNet selection did not consistently stabilize or improve the discriminative capacity of those coordinate components under the k = 6 protocol.
Under k = 6 , mRMR-based selection (Table 7) showed its strongest sensitivity with kNN using θ (66.67 ± 40.82), together with moderate specificity (50.00 ± 44.72) and the highest F1-score among mRMR k = 6 configurations (55.56 ± 32.77). RF with θ also achieved sensitivity above 50% (58.33 ± 49.16) with comparable specificity (58.33 ± 49.16). These results suggest that, under the k = 6 protocol, θ -based correlation features may provide useful discriminative information after mRMR selection; however, the large standard deviations indicate that this pattern should be interpreted cautiously.

Feature-Selection Behavior Across Folds

The automatic feature-selection strategies showed different behaviors in terms of retained feature-set size (see Table S5 of the Supplementary Materials). For ElasticNet, the number of retained correlation features was highly variable across outer folds: median 51.5 features (range 6–1067) for k = 3 and median 159.0 features (range 4–1108) for k = 6 . By contrast, mRMR produced more controlled subset sizes because its search space was explicitly restricted to k { 25 ,   50 } , yielding a median of 25 features for both outer k = 3 and k = 6 validation, with a range of 25–50 features in both cases. This analysis characterizes the stability of the selected subset size, rather than implying that the exact selected landmark pairs were identical across folds. Overall, ElasticNet appeared more sensitive to the training partition and regularization choices, whereas mRMR provided a more constrained feature-selection mechanism under the evaluated search space. In both cases, feature selection was performed within the training portion of each outer fold.
To facilitate interpretation of the tabulated results, Figure 2 summarizes the accuracy–sensitivity–specificity profiles of representative configurations across feature-selection strategies and validation protocols. Observe that the different feature-selection strategies exhibit distinct sensitivity–specificity trade-offs. In particular, the Manual-24 baseline and mRMR configurations tend to reach the highest sensitivity values, whereas ElasticNet generally yields more moderate sensitivity profiles.
On the other hand, Figure 3 summarizes the highest sensitivity achieved for each spherical coordinate component under each feature-selection strategy and validation protocol. This component-wise view helps clarify which coordinate representation contributed the most discriminative information in each setting. This figure also reveals that the most informative spherical component is not constant across strategies. In the Manual-24 baseline, ϕ consistently provides the highest sensitivity, whereas under mRMR the θ component becomes particularly competitive under k = 6 . This supports the view that the discriminative value of the correlation-based representation depends not only on the coordinate system, but also on the feature-selection strategy used to extract the most informative landmark-pair relationships.

4. Discussion

4.1. Principal Findings

This study evaluated correlation-based facial-motion descriptors derived from spherical-coordinate landmark trajectories for ALS/HC separation during the diadochokinetic /pataka/ task. The results suggest that inter-landmark coordination patterns contain discriminative information, but performance varied across spherical components, feature-selection strategies, and validation protocols. The observed sensitivity–specificity trade-offs and fold-level variability support a cautious interpretation of the proposed approach as a proof-of-concept screening framework rather than a clinically validated diagnostic model.

4.2. Comparison with the State of the Art on the /Pataka/ Task

Table 8 provides a task-aligned context by summarizing studies that explicitly report /pataka/ performance on Toronto NeuroFace (or closely related capture settings) and marking NR when /pataka/-specific results are not reported. This alignment is essential because performance can differ substantially across NeuroFace subtasks (e.g., /pataka/ vs. spread) and task aggregation can confound comparisons.
  • Depth-based kinematics versus RGB landmark pipelines
Bandini et al. [5] reported strong /pataka/ performance using marker-less 3D/depth acquisition and a compact set of kinematic/geometry features, achieving subject-level sensitivity of 90.0% and specificity of 75.0%, with 83.3% accuracy (Table 8). Despite this robust reporting, Bandini’s results are not directly comparable to ours in sensing modality and evaluation style. Their pipeline relies on RGB-D/3D depth measurements, which provide higher geometric fidelity (notably for out-of-plane motion), whereas our approach operates on RGB-derived landmarks with estimated depth. Moreover, Bandini reports subject-level performance via majority voting under LOSO-CV, while we report fold-averaged performance as mean ± SD, explicitly quantifying variability across splits. These differences should be considered when benchmarking /pataka/-based ALS detection across modalities and protocols. In practice, depth-based performance can be interpreted as a favorable upper-bound under richer capture conditions, while our method targets broader accessibility with standard RGB video. This characteristic may facilitate the deployment of video-based screening tools in clinical environments where specialized sensing hardware is not available.
  • Graph learning on landmarks
Gomes et al. [8] reported subject-level /pataka/ performance of 66.6% accuracy, 70.0% sensitivity, and 63.6% specificity using Facial Point Graphs and a graph neural network (GNN) formulation. While their representation learning approach differs from ours (GNN-based graph learning vs. correlation-structure descriptors with embedded feature selection), both results support the premise that structured dependency information between facial regions is informative for ALS detection in /pataka/. Notably, despite relying on comparatively lightweight models and a more compact feature representation, our primary configuration achieved sensitivity and specificity within a comparable task-level range, although with fold-level variability (Table 8). This suggests that correlation-based coordination descriptors may retain clinically relevant signal without requiring the full complexity of graph representation learning, but strict performance equivalence should not be inferred.
Finally, GNN-based methods often benefit from larger training sets (or from pre-training/transfer learning) to improve generalization. In the available description, Gomes et al. do not report using pre-training or transfer learning, which may increase sensitivity to limited sample sizes and cohort heterogeneity. Differences in evaluation level (subject voting vs. fold-level reporting) and in how selection/tuning are nested within validation can also affect strict comparability; we discuss these benchmarking caveats in Section 4.3.
  • Correlation-based facial symmetry features
An extended correlation-based landmark configuration reported by Suárez-Hernandez [10] presents higher /pataka/ sensitivity (88.89%) at 66.67% accuracy and 50.0% specificity for a setting using the θ component and an MLP model (Table 8). This provides a useful sensitivity-focused reference point within the same dataset and task setting. However, strict numerical comparability depends on evaluation details (e.g., CV design, model-selection criteria, and whether tuning/selection steps are nested), which can influence sensitivity estimates; see Section 4.3 for comparability caveats.
A further methodological difference concerns how dimensionality is reduced prior to modeling. In the thesis, the reduced set of 24 landmarks is obtained through correlation-based redundancy filtering (with a fixed threshold) to remove near-duplicate landmark trajectories, yielding a deterministic index set. This strategy is reproducible given the threshold and the elimination rule, but it optimizes primarily for reducing redundancy rather than for maximizing predictive relevance to the diagnostic label. In contrast, our primary pipeline operates on correlation-derived features and applies mRMR within the training folds to select features that jointly maximize relevance to ALS/HC discrimination while minimizing redundancy. This distinction is important in correlation-feature spaces, where many candidate pairs can be mutually correlated: redundancy pruning alone may discard features that are individually redundant yet jointly informative, whereas relevance-aware selection can retain discriminative structure under a leakage-aware evaluation protocol.

4.3. Leakage-Safe Evaluation and Comparability Caveats

A central methodological aspect of our study is the leakage-aware evaluation design. Feature selection (mRMR) is performed strictly within the training portion of each outer cross-validation split, and performance is computed only on the corresponding held-out fold. This nesting is essential in high-dimensional settings, where selecting features (or tuning hyperparameters) using information beyond the training data can lead to optimistically biased estimates. In addition, we report fold-averaged metrics as mean ± SD, providing an explicit measure of stability across splits. This protocol ensures that model selection, feature selection, and performance estimation remain strictly separated, reducing the risk of information leakage between training and evaluation stages.
These design choices also clarify why strict numerical comparisons with prior work can be challenging even when the dataset and task are aligned. Several studies report subject-level voting or LOO-style protocols, often as point estimates without dispersion measures, which makes it difficult to assess variability and robustness under partition changes. Therefore, while task-aligned comparisons remain informative (Table 8), differences in evaluation level (subject voting vs. fold-level reporting), uncertainty reporting, and nesting of selection/tuning steps should be considered when interpreting relative performance.
This principle applies not only to relevance-aware selectors (e.g., mRMR) but also to redundancy-reduction steps such as correlation-threshold filtering: when computed outside the evaluation loop, such preprocessing can inadvertently incorporate information from the full dataset and lead to optimistically biased performance estimates.

4.4. Coordinate-Wise Interpretation of Discriminative Facial Motion

Across our experiments, performance differences across spherical components suggest that coordinate choice materially affects separability. Interpreting these effects in biomechanical terms, ϕ may reflect lateralized coordination patterns, while θ may capture vertical or jaw-related coordination relevant to articulatory control. These interpretations are consistent with the hypothesis that bulbar impairment affects coordinated lower-face motion during diadochokinetic speech, but should be considered associations rather than direct physiological measurements.
In addition, component-wise reliability may be affected by landmark estimation accuracy: because MediaPipe® provides an estimated out-of-plane coordinate, the azimuthal component ϕ can be more sensitive to depth/pose noise than r or θ , which may partially explain component-dependent performance differences in RGB-based pipelines.

4.5. Clinical Operating Points: Sensitivity Versus Specificity

From a clinical perspective, sensitivity is often prioritized in screening-like settings to reduce missed ALS cases, particularly when confirmatory assessment exists downstream. However, high sensitivity can incur reduced specificity, increasing false positives. Our results highlight the importance of selecting operating points according to the intended clinical context rather than relying on a single aggregate metric. In this sense, task-aligned reporting (Table 8) helps clarify trade-offs across methods under /pataka/.

4.6. Variability and Robustness

Standard deviations for several metrics are relatively large, indicating sensitivity to data partitioning and potential cohort heterogeneity. The use of outer-fold reporting quantifies this uncertainty, but also emphasizes the need for larger cohorts and external validation. Future evaluations could incorporate repeated cross-validation and confidence intervals to provide tighter uncertainty characterization. Particularly, although neural-network models can overfit in very small cohorts, the MLP was included as a compact comparative baseline under the same leakage-aware evaluation protocol used for the other classifiers. Its results should therefore be interpreted as part of a controlled algorithmic comparison rather than as evidence that higher-capacity neural models are preferable in this dataset.

4.7. Limitations

This study has several limitations. First, the analysis was conducted on a single dataset with 18 /pataka/ recordings, which limits statistical power and generalizability. Because the outer validation schemes provide only a small number of fold-level estimates, formal statistical testing between configurations would have low power and could lead to overinterpretation. Therefore, performance differences were treated as descriptive and interpreted together with their fold-level variability.
Second, the pipeline relies on RGB-derived facial landmarks, including an estimated depth coordinate, which can be affected by pose, illumination, tracking errors, and uncertainty in out-of-plane motion. This limitation is particularly relevant when interpreting differences between spherical coordinate components.
Third, the focus on the /pataka/ task improves task-level comparability with prior work but limits generalization to other speech or nonspeech facial tasks. Differences in evaluation level (subject voting vs. fold-averaged reporting), uncertainty reporting, and nesting of feature selection or hyperparameter tuning also complicate strict numerical comparisons with prior studies.
Finally, no external validation cohort was available. Larger independent datasets are required to evaluate generalization, calibration, and clinically meaningful operating thresholds for screening-oriented use.

4.8. Future Directions

Future work should prioritize external validation on independent cohorts and larger task-aligned datasets. It would also be valuable to investigate how improved depth fidelity, either through depth sensors or enhanced monocular depth estimation, affects /pataka/-based screening. Finally, interpretability analyses that map selected correlation features to anatomically meaningful landmark-pair interactions, together with calibration of decision thresholds and cost-sensitive evaluation, may improve clinical transparency and deployment relevance.

5. Conclusions

This work evaluated RGB video-based ALS screening during the diadochokinetic /pataka/ task using correlation-structure descriptors derived from spherical-coordinate facial landmark trajectories. By representing coordinated facial motion through Pearson correlation matrices and evaluating multiple learning strategies under a leakage-aware nested cross-validation design, the study provides a reproducible benchmark for correlation-based facial motion analysis on the Toronto NeuroFace dataset.
The results suggest that correlation-derived coordination patterns contain discriminative information for ALS/HC separation, although performance depends on the spherical coordinate component, feature-selection strategy, classifier, and validation protocol. The primary mRMR-based configuration achieved competitive sensitivity and specificity while providing a constrained feature-selection mechanism relative to ElasticNet. However, the observed fold-level variability underscores the need for cautious interpretation in small clinical cohorts.
Compared with prior /pataka/-based studies, depth-based acquisition and engineered kinematic features remain strong references, suggesting that geometric fidelity is an important factor for future improvement in RGB-based pipelines. Overall, correlation-based facial motion analysis combined with leakage-aware feature selection represents a transparent proof-of-concept approach for accessible video-based ALS screening. Future work should prioritize external validation, larger cohorts, improved depth estimation, and clinically oriented threshold calibration.

Supplementary Materials

The following supporting information can be downloaded at: https://www.mdpi.com/article/10.3390/eng7050238/s1, Table S1: Complete hyperparameter configuration used in the experimental pipeline; Table S2: Most frequently selected hyperparameter values across outer cross-validation folds for the manual landmark selection baseline (24 landmarks); Table S3: Most frequently selected hyperparameter values across outer cross-validation folds for ElasticNet-based feature selection; Table S4: Most frequently selected hyperparameter values across outer cross-validation folds for mRMR-based feature selection; Table S5: Number of retained correlation features across outer folds by method, validation scheme, and spherical-coordinate component.

Author Contributions

Conceptualization, D.S.-H., S.T.-R., S.R.S.-A. and I.R.-G.; methodology, S.T.-R. and I.R.-G.; software, D.S.-H., S.T.-R. and I.R.-G.; validation, S.T.-R., S.R.S.-A. and I.R.-G.; formal analysis, S.T.-R., S.R.S.-A. and I.R.-G.; investigation, D.S.-H.; resources, S.T.-R. and I.R.-G.; data curation, D.S.-H. and S.R.S.-A.; writing—original draft preparation, S.T.-R. and I.R.-G.; writing—review and editing, S.T.-R., S.R.S.-A. and I.R.-G.; visualization, D.S.-H.; supervision, S.T.-R., S.R.S.-A. and I.R.-G. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding. Daniela Suárez-Hernández received a postgraduate scholarship from the Science, Humanities, Technology, and Innovation Secretariat (SECIHTI) (CVU: 1275381).

Institutional Review Board Statement

This study used anonymized third-party data from the Toronto NeuroFace Dataset. Ethical approval for the original data collection was obtained by the dataset providers; therefore, no additional institutional review board approval was required for this secondary analysis.

Informed Consent Statement

Informed consent was obtained by the original dataset providers from all subjects involved in the study. No additional consent was required for the present secondary analysis of anonymized data.

Data Availability Statement

Restrictions apply to the availability of these data. The data were obtained from the Toronto NeuroFace dataset providers and are available from the corresponding source with the permission of the dataset owners. The full implementation (preprocessing, feature construction, and training/evaluation scripts) is available at https://github.com/lasaid-udg/als-facemesh-correlation-screening.git (accessed on 13 April 2026).

Acknowledgments

Portions of the research reported in this paper use the NeuroFace Database collected by Yana Yunusova and the Vocal Tract Visualization and Bulbar Function Laboratory teams at UHN–Toronto Rehabilitation Institute and Sunnybrook Research Institute, respectively. The data collection was financially supported by the Michael J. Fox Foundation, NIH–NIDCD, Natural Sciences and Engineering Research Council of Canada, Heart and Stroke Foundation Canadian Partnership for Stroke Recovery, and AGE-WELL NCE.

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

The following abbreviations are used in this manuscript:
ALSAmyotrophic lateral sclerosis
CVCross-validation
DTDecision tree
GNNGraph neural network
HCHealthy control(s)
kNNk-nearest neighbors
LOSOLeave-one-subject-out
MLPMultilayer perceptron
mRMRminimum Redundancy Maximum Relevance
RGBRed–green–blue
RGB-DRed–green–blue plus depth
RFRandom forest
SDStandard deviation
SoAState of the art

References

  1. Wijesekera, L.C.; Leigh, P.N. Amyotrophic lateral sclerosis. Orphanet J. Rare Dis. 2009, 4, 3. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  2. Hardiman, O.; van den Berg, L.H.; Kiernan, M.C. Amyotrophic lateral sclerosis. Nat. Rev. Dis. Prim. 2017, 3, 17071. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  3. Brown, R.H.; Al-Chalabi, A. Amyotrophic Lateral Sclerosis. N. Engl. J. Med. 2017, 377, 162–172. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  4. Grollemund, V.; Pradat, P.F.; Querin, G.; Delbot, F.; Le Chat, G.; Pradat-Peyre, J.F.; Bede, P. Machine Learning in Amyotrophic Lateral Sclerosis: Achievements, Pitfalls, and Future Directions. Front. Neurosci. 2019, 13, 135. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  5. Bandini, A.; Green, J.R.; Taati, B.; Orlandi, S.; Zinman, L.; Yunusova, Y. Automatic Detection of Amyotrophic Lateral Sclerosis from Video-Based Analysis of Facial Movements. In Proceedings of the 13th IEEE International Conference on Automatic Face and Gesture Recognition (FG); IEEE: New York, NY, USA, 2018; pp. 150–157. [Google Scholar] [CrossRef] [Scilit]
  6. Bandini, A.; Rezaei, S.; Guarin, D.; Kulkarni, M.; Lim, D.; Boulos, M.; Zinman, L.; Yunusova, Y.; Taati, B. A New Dataset for Facial Motion Analysis in Individuals with Neurological Disorders. IEEE J. Biomed. Health Inform. 2021, 25, 1111–1119. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  7. Jin, B.; Qu, Y.; Zhang, L.; Gao, Z. Diagnosing Parkinson Disease Through Facial Expression Recognition: Video Analysis. J. Med. Internet Res. 2020, 22, e18697. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  8. Gomes, N.B.; Yoshida, A.; Roder, M.; de Oliveira, G.C.; Papa, J.P. Facial Point Graphs for Amyotrophic Lateral Sclerosis Identification. arXiv 2023, arXiv:2307.12159. http://arxiv.org/abs/2307.12159.
  9. Suárez-Hernández, D.; Santos-Arce, S.R.; Torres-Ramos, S.; Salido-Ruiz, R.A.; Román-Godínez, I. Amyotrophic Lateral Sclerosis Detection Using Facial Symmetry Analysis with Machine Learning Techniques. In Proceedings of the XLVII Mexican Conference on Biomedical Engineering; Flores Cuautle, J.d.J.A., Benítez-Mata, B., Reyes-Lagos, J.J., Hernandez Acosta, H.Y., Ames Lastra, G., Zuñiga-Aguilar, E., Del Hierro-Gutierrez, E., Salido-Ruiz, R.A., Eds.; Springer: Cham, Switzerland, 2025; pp. 239–248. [Google Scholar]
  10. Suárez-Hernández, D. Detección de Esclerosis Lateral Amiotrófica Mediante Técnicas de Aprendizaje Automático en Datos Biométricos. Master’s Thesis, Universidad de Guadalajara, Centro Universitario de Ciencias Exactas e Ingenierías (CUCEI), Guadalajara, Jalisco, México, 2024. [Google Scholar]
  11. Duffy, J.R. Motor Speech Disorders: Clues to Neurologic Diagnosis. In Parkinson’s Disease and Movement Disorders; Adler, C.H., Ahlskog, J.E., Eds.; Humana Press: Totowa, NJ, USA, 2000; pp. 23–54. [Google Scholar] [CrossRef] [Scilit]
  12. Lugaresi, C.; Tang, G.; Nash, H.; McClanahan, C.; Uboweja, E.; Hays, M.; Zhang, F.; Chang, C.-L.; Yong, M.G.; Lee, J.; et al. MediaPipe: A Framework for Building Perception Pipelines. arXiv 2019, arXiv:1906.08172. [Google Scholar] [CrossRef] [Scilit]
  13. Grishchenko, I.; Ablavatski, A.; Kartynnik, Y.; Raveendran, K.; Grundmann, M. Attention Mesh: High-fidelity Face Mesh Prediction in Real-time. arXiv 2020, arXiv:2006.10962. http://arxiv.org/abs/2006.10962.
  14. Bandini, A.; Orlandi, S.; Giovannelli, F.; Felici, A.; Cincotta, M.; Clemente, D.; Vanni, P.; Zaccara, G.; Manfredi, C. Markerless Analysis of Articulatory Movements in Patients With Parkinson’s Disease. J. Voice 2016, 30, 766.e1–766.e11. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  15. Sharpe, J.; Hancock, E.R. Recognising Facial Expressions Using Spherical Harmonics. In Structural, Syntactic, and Statistical Pattern Recognition; Lecture Notes in Computer Science; Springer: Berlin/Heidelberg, Germany, 2008; Volume 5358, pp. 249–258. [Google Scholar] [CrossRef] [Scilit]
  16. Liu, P.; Wang, Y.; Huang, D.; Zhang, Z.; Chen, L. Learning the Spherical Harmonic Features for 3D Face Recognition. IEEE Trans. Image Process. 2014, 23, 914–925. [Google Scholar] [CrossRef] [Scilit]
  17. Acharya, D.; Huang, Z.; Paudel, D.P.; Van Gool, L. Covariance Pooling for Facial Expression Recognition. IEEE Trans. Image Process. 2018, 27, 5597–5610. [Google Scholar] [CrossRef] [Scilit]
  18. Kavitha, R.; Thangavel, P. Face Analysis Using Row and Correlation Based Local Directional Pattern. Multimed. Tools Appl. 2021, 80, 19711–19734. [Google Scholar] [CrossRef] [Scilit]
  19. Zou, H.; Hastie, T. Regularization and Variable Selection via the Elastic Net. J. R. Stat. Soc. Ser. 2005, 67, 301–320. [Google Scholar] [CrossRef] [Scilit]
  20. Renton, M.; Fakhriyehasl, M.; Weiss, J.; Milosevic, M.; Laframboise, S.; Rouzbahman, M.; Han, K.; Jhaveri, K. Multiparametric MRI radiomics for predicting disease-free survival and high-risk histopathological features for tumor recurrence in endometrial cancer. Front. Oncol. 2024, 14, 1406858. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  21. Hassan, M.; Kaabouch, N. Impact of Feature Selection Techniques on the Performance of Machine Learning Models for Depression Detection Using EEG Data. Appl. Sci. 2024, 14, 10532. [Google Scholar] [CrossRef] [Scilit]
  22. Peng, H.; Long, F.; Ding, C. Feature Selection Based on Mutual Information: Criteria of Max-Dependency, Max-Relevance, and Min-Redundancy. IEEE Trans. Pattern Anal. Mach. Intell. 2005, 27, 1226–1238. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  23. Wong, T.L.J.; Teng, X.; Leung, W.; Cai, J. MULTI-modal radiomics to predict early treatment response from PSA (prostate specific antigen) decline in prostate cancer patients under stereotactic body radiotherapy in MR-Linac. J. Radiat. Res. Appl. Sci. 2024, 17, 100841. [Google Scholar] [CrossRef] [Scilit]
  24. Varma, S.; Simon, R. Bias in error estimation when using cross-validation for model selection. BMC Bioinform. 2006, 7, 91. [Google Scholar] [CrossRef] [Scilit] [PubMed]
Figure 1. Graphical methodology overview for ALS detection using RGB video and spherical coordinate reference-point correlations.
Figure 1. Graphical methodology overview for ALS detection using RGB video and spherical coordinate reference-point correlations.
Eng 07 00238 g001
Figure 2. Summary of representative performance profiles across feature-selection strategies and validation protocols. The figure shows accuracy, sensitivity, and specificity (mean ± SD across outer folds) for selected configurations taken directly from Table 2, Table 3, Table 4, Table 5, Table 6 and Table 7. Sensitivity is emphasized as the primary screening-oriented metric, while accuracy and specificity are shown to illustrate the associated trade-offs.
Figure 2. Summary of representative performance profiles across feature-selection strategies and validation protocols. The figure shows accuracy, sensitivity, and specificity (mean ± SD across outer folds) for selected configurations taken directly from Table 2, Table 3, Table 4, Table 5, Table 6 and Table 7. Sensitivity is emphasized as the primary screening-oriented metric, while accuracy and specificity are shown to illustrate the associated trade-offs.
Eng 07 00238 g002
Figure 3. Component-wise summary of the highest sensitivity values obtained for each feature-selection strategy and validation protocol. Each cell reports the maximum sensitivity (%) achieved for a given spherical component (r, θ , or ϕ ), together with the classifier that attained that value, based directly on Table 2, Table 3, Table 4, Table 5, Table 6 and Table 7. This visualization highlights how the most informative spherical component varies across methods and validation settings.
Figure 3. Component-wise summary of the highest sensitivity values obtained for each feature-selection strategy and validation protocol. Each cell reports the maximum sensitivity (%) achieved for a given spherical component (r, θ , or ϕ ), together with the classifier that attained that value, based directly on Table 2, Table 3, Table 4, Table 5, Table 6 and Table 7. This visualization highlights how the most informative spherical component varies across methods and validation settings.
Eng 07 00238 g003
Table 1. Class distribution of the /pataka/ video recordings included in this study. Each recording corresponds to one subject and is treated as a single sample.
Table 1. Class distribution of the /pataka/ video recordings included in this study. Each recording corresponds to one subject and is treated as a single sample.
TaskHC (n)ALS (n)Total (n)
/pataka/10818
Table 2. Classification performance (mean ± SD across outer folds) for the Manual-24 baseline using 3-fold cross-validation. Metrics are reported in percent (%).
Table 2. Classification performance (mean ± SD across outer folds) for the Manual-24 baseline using 3-fold cross-validation. Metrics are reported in percent (%).
VariableAlgorithmAccuracySensitivitySpecificityPrecisionF1-Score
Manual-24 baseline
3-fold cross validation
rDT44.44 ± 9.6222.22 ± 19.2461.11 ± 9.6233.33 ± 28.8726.67 ± 23.09
rRF44.44 ± 9.6211.11 ± 19.2469.45 ± 4.8116.67 ± 28.8713.33 ± 23.09
rMLP44.44 ± 19.2538.89 ± 34.7052.78 ± 24.0630.56 ± 33.6833.33 ± 33.34
rkNN44.44 ± 9.6255.56 ± 50.9241.67 ± 22.0530.00 ± 26.4638.09 ± 32.99
θ DT55.56 ± 9.6238.89 ± 34.7072.22 ± 25.4650.00 ± 23.5735.56 ± 33.56
θ RF50.00 ± 0.0038.89 ± 9.6261.11 ± 9.6244.44 ± 9.6240.00 ± 0.00
θ MLP44.44 ± 9.6238.89 ± 9.6252.78 ± 24.0641.67 ± 14.4337.78 ± 3.85
θ kNN50.00 ± 16.6750.00 ± 16.6752.78 ± 24.0647.22 ± 20.9746.67 ± 17.64
ϕ DT55.55 ± 25.4638.89 ± 34.7072.22 ± 25.4644.44 ± 50.9240.00 ± 40.00
ϕ RF50.00 ± 16.6733.33 ± 57.7472.22 ± 25.4625.00 ± 35.3622.22 ± 38.49
ϕ MLP66.66 ± 28.8761.11 ± 34.7066.67 ± 33.3469.44 ± 33.6861.90 ± 26.51
ϕ kNN50.00 ± 16.6777.78 ± 19.2433.33 ± 33.3450.00 ± 16.6757.94 ± 8.36
Table 3. Classification performance (mean ± SD across outer folds) for ElasticNet using 3-fold cross-validation. Metrics are reported in percent (%).
Table 3. Classification performance (mean ± SD across outer folds) for ElasticNet using 3-fold cross-validation. Metrics are reported in percent (%).
VariableAlgorithmAccuracySensitivitySpecificityPrecisionF1-Score
ElasticNet
3-fold cross validation
rDT50.00 ± 16.6722.22 ± 19.2477.78 ± 19.2422.22 ± 50.0030.00 ± 26.46
rRF50.00 ± 16.6738.89 ± 34.7061.11 ± 9.6244.44 ± 9.6241.11 ± 8.39
rMLP55.56 ± 9.6244.45 ± 17.0066.67 ± 25.4644.44 ± 34.7041.27 ± 36.06
rkNN50.00 ± 16.6750.00 ± 16.6752.78 ± 24.0647.22 ± 20.9746.67 ± 17.64
θ DT44.44 ± 9.6233.33 ± 40.8261.11 ± 9.6227.78 ± 41.8330.16 ± 28.70
θ RF44.44 ± 9.6238.89 ± 9.6252.78 ± 24.0638.89 ± 9.6237.78 ± 3.85
θ MLP44.44 ± 9.6250.00 ± 16.6741.67 ± 22.0541.67 ± 14.4343.49 ± 12.28
θ kNN50.00 ± 0.0066.67 ± 33.3441.67 ± 22.0546.67 ± 5.7751.43 ± 9.90
ϕ DT50.00 ± 0.0038.89 ± 9.6261.11 ± 9.6244.44 ± 9.6240.00 ± 0.00
ϕ RF50.00 ± 16.6738.89 ± 34.7061.11 ± 9.6233.33 ± 33.3435.56 ± 33.56
ϕ MLP50.00 ± 0.0050.00 ± 16.6750.00 ± 0.0043.06 ± 13.9945.71 ± 9.90
ϕ kNN55.55 ± 25.4638.89 ± 34.7072.22 ± 25.4644.44 ± 50.9240.00 ± 40.00
Table 4. Classification performance (mean ± SD across outer folds) for mRMR using 3-fold cross-validation. Metrics are reported in percent (%).
Table 4. Classification performance (mean ± SD across outer folds) for mRMR using 3-fold cross-validation. Metrics are reported in percent (%).
VariableAlgorithmAccuracySensitivitySpecificityPrecisionF1-Score
mRMR
3-fold cross validation
rDT38.89 ± 34.7033.33 ± 57.7450.00 ± 50.0025.00 ± 35.3622.22 ± 38.49
rRF27.78 ± 19.2411.11 ± 19.2436.11 ± 37.588.33 ± 14.439.52 ± 16.49
rMLP27.78 ± 19.2422.22 ± 38.4930.55 ± 4.8116.67 ± 28.8719.05 ± 32.99
rkNN50.00 ± 28.8755.56 ± 50.9250.00 ± 16.6738.89 ± 34.7044.45 ± 38.49
θ DT44.45 ± 25.4638.89 ± 34.7047.22 ± 24.0633.33 ± 28.8735.71 ± 31.13
θ RF61.11 ± 9.6250.00 ± 16.6769.45 ± 4.8155.56 ± 9.6252.22 ± 13.47
θ MLP44.44 ± 9.6250.00 ± 16.6738.89 ± 9.6238.89 ± 9.6243.49 ± 12.28
θ kNN38.89 ± 9.6238.89 ± 34.7041.67 ± 22.0525.00 ± 25.0030.16 ± 28.70
ϕ DT61.11 ± 25.4672.22 ± 25.4652.78 ± 41.1161.67 ± 37.5362.78 ± 25.62
ϕ RF61.11 ± 19.2450.00 ± 16.6772.22 ± 25.4661.11 ± 34.7053.33 ± 23.09
ϕ MLP61.11 ± 34.7055.56 ± 50.9261.11 ± 34.7050.00 ± 50.0052.38 ± 50.17
ϕ kNN61.11 ± 19.2461.11 ± 9.6261.11 ± 34.7061.11 ± 34.7059.05 ± 20.07
Table 5. Classification performance (mean ± SD across outer folds) for the Manual-24 baseline using 6-fold cross-validation. Metrics are reported in percent (%).
Table 5. Classification performance (mean ± SD across outer folds) for the Manual-24 baseline using 6-fold cross-validation. Metrics are reported in percent (%).
VariableAlgorithmAccuracySensitivitySpecificityPrecisionF1-Score
Manual-24 baseline
6-fold cross validation
rDT33.33 ± 21.080.00 ± 0.0066.67 ± 40.820.00 ± 0.000.00 ± 0.00
rRF33.33 ± 21.080.00 ± 0.0058.33 ± 37.640.00 ± 0.000.00 ± 0.00
rMLP44.44 ± 17.2241.67 ± 49.1650.00 ± 31.6230.00 ± 27.3930.56 ± 34.02
rkNN33.33 ± 0.0041.67 ± 49.1633.33 ± 40.8223.33 ± 22.3625.00 ± 27.39
θ DT38.89 ± 32.7741.67 ± 49.1650.00 ± 44.7240.00 ± 41.8333.34 ± 36.52
θ RF44.44 ± 34.4350.00 ± 54.7750.00 ± 44.7236.67 ± 41.5036.11 ± 42.71
θ MLP50.00 ± 45.9550.00 ± 54.7750.00 ± 44.7241.67 ± 49.1644.45 ± 50.19
θ kNN50.00 ± 34.9658.33 ± 49.1641.67 ± 49.1646.67 ± 36.1344.45 ± 38.97
ϕ DT44.45 ± 27.2225.00 ± 41.8350.00 ± 44.7225.00 ± 28.8719.45 ± 30.58
ϕ RF50.00 ± 27.8941.67 ± 49.1658.33 ± 37.6440.00 ± 41.8333.34 ± 36.52
ϕ MLP61.11 ± 13.6150.00 ± 44.7266.67 ± 40.8262.50 ± 25.0041.67 ± 32.92
ϕ kNN50.00 ± 27.8966.67 ± 40.8241.67 ± 37.6450.00 ± 31.6252.78 ± 26.70
Table 6. Classification performance (mean ± SD across outer folds) for ElasticNet using 6-fold cross-validation. Metrics are reported in percent (%).
Table 6. Classification performance (mean ± SD across outer folds) for ElasticNet using 6-fold cross-validation. Metrics are reported in percent (%).
VariableAlgorithmAccuracySensitivitySpecificityPrecisionF1-Score
ElasticNet
6-fold cross validation
rRF44.45 ± 27.2225.00 ± 41.8366.67 ± 40.8237.50 ± 47.8722.22 ± 34.43
rMLP38.89 ± 25.0925.00 ± 41.8358.33 ± 37.6430.00 ± 44.7222.22 ± 34.43
rkNN33.33 ± 21.0825.00 ± 41.8341.67 ± 49.1620.83 ± 25.0016.67 ± 25.82
θ DT22.22 ± 17.210.00 ± 0.0041.67 ± 37.640.00 ± 0.000.00 ± 0.00
θ RF38.89 ± 32.7725.00 ± 41.8350.00 ± 44.7230.00 ± 44.7222.22 ± 34.43
θ MLP50.00 ± 40.8341.67 ± 49.1658.33 ± 49.1650.00 ± 50.0038.89 ± 44.31
θ kNN38.89 ± 32.7733.33 ± 51.6441.67 ± 49.1629.17 ± 34.3624.45 ± 38.10
ϕ DT55.55 ± 27.2250.00 ± 44.7258.33 ± 37.6450.00 ± 44.7247.22 ± 40.02
ϕ RF55.56 ± 34.4350.00 ± 54.7758.33 ± 37.6440.00 ± 41.8338.89 ± 44.31
ϕ MLP44.45 ± 27.2225.00 ± 41.8358.33 ± 37.6430.00 ± 44.7222.22 ± 34.43
ϕ kNN38.89 ± 32.7741.67 ± 49.1641.67 ± 37.6433.33 ± 40.8233.34 ± 36.52
rDT44.45 ± 27.2225.00 ± 41.8366.67 ± 40.8237.50 ± 47.8722.22 ± 34.43
Table 7. Classification performance (mean ± SD across outer folds) for mRMR using 6-fold cross-validation. Metrics are reported in percent (%).
Table 7. Classification performance (mean ± SD across outer folds) for mRMR using 6-fold cross-validation. Metrics are reported in percent (%).
VariableAlgorithmAccuracySensitivitySpecificityPrecisionF1-Score
mRMR
6-fold cross validation
rDT27.78 ± 25.0925.00 ± 41.8325.00 ± 41.8316.67 ± 23.5716.67 ± 25.82
rRF27.78 ± 25.0916.67 ± 40.8241.67 ± 49.168.33 ± 16.668.33 ± 20.41
rMLP44.44 ± 17.2241.67 ± 49.1658.33 ± 37.6436.67 ± 41.5030.56 ± 34.02
rkNN27.78 ± 25.0925.00 ± 41.8333.33 ± 40.8222.22 ± 40.3719.45 ± 30.58
θ DT50.00 ± 34.9633.33 ± 51.6475.00 ± 41.8350.00 ± 50.0027.78 ± 44.31
θ RF55.55 ± 27.2258.33 ± 49.1658.33 ± 49.1658.33 ± 28.8744.45 ± 38.97
θ MLP33.33 ± 29.8241.67 ± 49.1625.00 ± 27.3925.00 ± 27.3930.56 ± 34.02
θ kNN55.55 ± 27.2266.67 ± 40.8250.00 ± 44.7255.55 ± 38.9755.56 ± 32.77
ϕ DT44.44 ± 34.4333.33 ± 51.6458.33 ± 49.1633.33 ± 47.1425.00 ± 41.83
ϕ RF33.33 ± 29.8241.67 ± 49.1633.33 ± 40.8230.55 ± 40.0230.56 ± 34.02
ϕ MLP38.89 ± 32.7733.33 ± 51.6441.67 ± 37.6422.22 ± 40.3725.00 ± 41.83
ϕ kNN50.00 ± 34.9658.33 ± 49.1650.00 ± 44.7247.22 ± 45.2447.22 ± 40.02
Table 8. Task-aligned comparison against prior work for the diadochokinetic /pataka/ task (ALS vs. HC) in the Toronto NeuroFace ecosystem. Sensitivity is emphasized as the primary metric. Additional metrics are reported only when explicitly available for the same task and evaluation level. Prior work reports subject-level point estimates, whereas our results are reported as fold-level mean ± SD.
Table 8. Task-aligned comparison against prior work for the diadochokinetic /pataka/ task (ALS vs. HC) in the Toronto NeuroFace ecosystem. Sensitivity is emphasized as the primary metric. Additional metrics are reported only when explicitly available for the same task and evaluation level. Prior work reports subject-level point estimates, whereas our results are reported as fold-level mean ± SD.
WorkModalityValidation/ModelSensitivitySpecificityAccuracy
This work (Primary)–mRMR, k = 3 RGB (2D) ϕ + kNN; outer CV k = 3 (fold mean ± SD) 61.11   ±   9.62 61.11   ±   34.70 61.11   ±   19.24
This work (Reference)–Manual-24, k = 3 RGB (2D) ϕ + MLP; outer CV k = 3 (fold mean ± SD) 61.11   ±   34.70 66.67   ±   33.34 66.66   ±   28.87
Suárez-Hernández [10]RGB (2D)MLP (RNA), coordinate θ , CV k = 6 (best sensitivity reported) 88.89 % 50.0 % 66.67 %
Bandini et al. (2018) [5]RGB-D (3D/depth)Subject-based classification (majority vote); LOSO-CV; PATAKA 90.0 % 75.0 % 83.3 %
Gomes et al. (2023) [8]RGB (2D)Subject-based
classification
70.0 % 63.6 % 66.6 %
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Suárez-Hernández, D.; Torres-Ramos, S.; Santos-Arce, S.R.; Román-Godínez, I. Reproducible RGB Video Screening of Amyotrophic Lateral Sclerosis Using Spherical-Coordinate Landmark Correlations. Eng 2026, 7, 238. https://doi.org/10.3390/eng7050238

AMA Style

Suárez-Hernández D, Torres-Ramos S, Santos-Arce SR, Román-Godínez I. Reproducible RGB Video Screening of Amyotrophic Lateral Sclerosis Using Spherical-Coordinate Landmark Correlations. Eng. 2026; 7(5):238. https://doi.org/10.3390/eng7050238

Chicago/Turabian Style

Suárez-Hernández, Daniela, Sulema Torres-Ramos, Stewart R. Santos-Arce, and Israel Román-Godínez. 2026. "Reproducible RGB Video Screening of Amyotrophic Lateral Sclerosis Using Spherical-Coordinate Landmark Correlations" Eng 7, no. 5: 238. https://doi.org/10.3390/eng7050238

APA Style

Suárez-Hernández, D., Torres-Ramos, S., Santos-Arce, S. R., & Román-Godínez, I. (2026). Reproducible RGB Video Screening of Amyotrophic Lateral Sclerosis Using Spherical-Coordinate Landmark Correlations. Eng, 7(5), 238. https://doi.org/10.3390/eng7050238

Article Metrics

Back to TopTop