1. Introduction
Parkinson’s disease (PD) is a progressive neurodegenerative disorder that affects around 10 million people worldwide [
1]. This disease is characterized by different motor and non-motor symptoms, including rigidity, resting tremor, bradykinesia, depression, sleep disorders, and others [
2,
3].
As the disease progresses, patients exhibit a progressive deterioration of one or several symptoms, which therefore negatively impacts their quality of life. Given that motor and non-motor alterations progress differently among patients, timely, accurate, and constant monitoring is required to make the administration of appropriate/individualized treatment possible [
4]. Continuous and unobtrusive monitoring constitute current challenges for all health systems across the globe [
5]. Such challenges motivate the research community to develop automatic monitoring tools that allow the characterization of disease progression in a timely and accurate manner, considering different biomarkers. Among the existing biomarkers, speech arises as one of the most promising given the simplicity, unobtrusive nature, and low cost of its monitoring [
6].
Speech disorders affect around 89% of PD patients [
7,
8] and it is known that such alterations appear in the early stages of the disease [
9]; therefore, speech analysis is not only a promising biomarker with which to monitor the disease progression, but also to detect it early on [
10]. From a clinical perspective, dysarthric speech is characterized by producing slow, weak, and imprecise muscle movements involved in speech production. These patterns can be measured and quantified by different means. For instance, the authors in [
11] used a cohort of 63 PD patients and 35 healthy controls and evaluated their vocalic articulation by measuring the area of the vocal triangle, the vowel articulation index, and the Phi index. All these measures are based on the vocalic formants, which provide a simple and well-known method to evaluate articulation in speech. With this simple approach, the authors found that the area of the vocalic triangle is significantly smaller in dysarthric speakers than in healthy controls. With similar approaches, based on vocalic formants extracted from speech spectrograms and also with temporal-related measures, the authors in [
12] studied vowel acoustics in dysarthric speakers suffering from different disorders, including ALS, Parkinson’s, Huntington’s, and others. The authors reported that metrics designed to model vowel distinctiveness are more sensitive and specific predictors of dysarthria.
Besides studies on speech production, alterations in other physiological processes like respiration have also been studied. For instance, in [
13], the authors evaluated how abnormal respiration in PD patients also affects speech production. The study included a cohort of 55 patients assessed over 15 days. Several motor skills were evaluated, including postural stability, speech production, respiration (at rest and during physical exercises), etc. Patients were grouped according to their dysarthria severity, i.e., mild, moderate, and severe. Their findings indicate that those patients who showed more dysarthria severity also exhibited other comorbidities like abnormal respiration and posture. As a conclusion, the authors claimed that dysarthria and altered respiration are closely related in PD patients and such alterations affect other motor skills important for daily living, like postural stability. Another study [
14] involved evaluating intelligibility changes in PD patients (native Spanish speakers) who followed the Lee Silverman Voice Treatment (LSVT-LOUD) therapy. The cohort included a total of 15 patients with PD. According to their results, the authors indicate that those dysarthric patients who performed the therapy improved their intelligibility (perceptually evaluated). From an engineering perspective, most studies are focused on using hand-crafted features like jitter, shimmer, Mel spectra, etc. [
15,
16,
17]. The authors in [
18] introduced the use of Gated Recurrent Units to extract phonological posteriors. This approach was later refined in [
19] and effectively used to model how PD patients produce the phonemes in continuous speech recordings. The same pattern had been reported a couple of years before in [
20], where the authors found that phonological classes like plosives, vowels, and fricatives are the most sensitive to the motor deterioration in PD speech. A similar approach was reported in [
21], where the authors showed that phonological posteriors are more effective for detecting PD than classical approaches based on Mel spectra, Mel Frequency Cepstral Coefficients (MFCCs), or the well-known eGeMAPS [
22]. A similar finding was reported in a more recent work [
23], where the authors not only included classical features like articulation and prosody but also foundation models like Wav2vec to make comparisons w.r.t. the phonetic posteriors.
The reviewed state of the art suggests that most works have been focused on classifying PD patients vs. HC subjects; however, longitudinal evaluation remains under-explored. One of the reasons for this is due to the fact that it is hard to get access to collect speech recordings of PD patients, and such a difficulty is maximized when the patient is requested to be recorded several times within a given period of time. The work of [
24] constitutes a seminal study in this direction. The authors evaluated changes in the speech production of 80 PD patients and 60 HC subjects within a time-frame of up to 12 months. The main result indicated a significant deterioration in speech quality, especially in speech velocity and articulation. Following the previous idea, the authors in [
25] introduced the use of Gaussian Mixture Models (GMMs) to automatically model the level of dysarthria in a set with 7 patients who were recorded a total of 16 times between 2012 and 2016. The collection process was involved four recordings taken during the same day, every 2 h. The time between each recording session was around one year. The main conclusion of the study was that it is possible to accurately trace the dysarthria level of PD patients longitudinally.
After reviewing the existing literature, it is found that some researchers have addressed the topic considering standard and (to some extent) sophisticated speech modelling methodologies. However, there is one gap that still needs attention: the use of methods that allow clinical interpretability is still an open question. Building on our previous studies, where phonetic posteriors are introduced as a promising tool to model speech production in PD patients [
19,
26], in this paper, we analyse how different phonological classes change while the disease is progressing.
3. Methodology
Figure 2 illustrates the general methodology proposed in this study. Monologue and read-text recordings of the PC-GITA corpus are used to create phonological models with the Phonet module, originally introduced in [
18]. It provides values for the phonological posteriors associated with each phonological class for a given phoneme. Therefore, feature vectors are created by grouping posteriors according to their phonological class. Each feature vector is used to train individual classifiers per phonological class. Resulting models are evaluated on the PC-GITA recordings following a 10-fold cross validation strategy. Optimal meta-parameters of the classifiers are stored to be used later in the evaluation of longitudinal recordings. Notice that no further optimization is performed in this step. The final aim is to evaluate which phonological classes are more sensitive to disease progression, which ultimately would provide additional information to the expert neurologist and speech/language therapist to make appropriate decisions regarding medication and therapy updates. Details of each stage in the methodology are provided in the next subsections.
3.1. Phonological Representation
Phonological representations are created with Phonet, which is based on bi-directional Gated Recurrent Units (GRUs) that were trained with 17 h of Latin American Spanish recordings. Notice that the outcome of such a model was named in the literature as Phonemic Identifiability [
26]. Among several properties, it demonstrated a great sensitivity to model cognitive decline in PD patients. Phonet constitutes one of the modules of the Disvoice toolkit (
https://github.com/jcvasquezc/Disvoice (accessed on 22 February 2026)). The outcome includes a total of 21 phonemes which are grouped into 18 phonological classes, as described in
Table 5. The overall accuracy of Phonet in modelling the 18 phonological classes is 86.6% [
18]. Notice that silence is considered as one of the classes.
Phonological posteriors are calculated in segments of 80 ms length with an overlap of 40 ms. Therefore, each audio recording has a variable length representation. To create fixed-length vectors, six statistical functionals are computed, namely mean, standard deviation, maximum, minimum, skewness, and kurtosis. Ultimately, each audio is represented with a total of 18 6-dimensional feature vectors, one per phonological class.
3.2. Automatic Classification of PD vs. HC Subjects
After creating phonological representations for all speakers in the PC-GITA corpus, the resulting feature vectors were used to train a Support Vector Machine (SVM) classifier. Three kernels were evaluated: linear (Lin), sigmoid (Sig), and Gaussian with a radial-basis function (RBF). Optimization of the corresponding meta-parameters was performed following a speaker-independent 10-fold nested and stratified cross-validation. Notice that after this stage, there will be 18 independent models, one per phonological class. Each model is given by its corresponding optimal parameters: C for the case of linear and sigmoid kernels or C and for the Gaussian and its support vectors.
3.3. Longitudinal Analysis
Phonological posteriors are also extracted from recordings of the Longitudinal corpus. Thus, speakers in this stage are also represented as 18-dimensional feature vectors. The last stage in the methodology consists of taking those 18 optimal classification models found in the previous stage and computing their corresponding output, i.e., classification scores, for the recordings in the Longitudinal corpus. Since no further optimization was performed at this stage, the Longitudinal corpus constitutes an independent test set, which besides being recorded in realistic conditions, allows the evaluation of disease progression over three years. It is important to stress the fact that none of the speakers in this corpus were included in the previous stage (the one based on PC-GITA).
4. Experiments and Results
Two main experiments were performed in this work: (i) automatic classification of PD vs. HC subjects considering the phonological posteriors extracted with the Phonet module and within the context of PC-GITA and the (ii) automatic evaluation of dysarthria level progression, also based on the corresponding phonological posterior, within the context of the Longitudinal corpus. Two speech tasks were considered separately in both experiments: monologue and read texts.
4.1. Classification of PD vs. and HC
There was one independent classifier per phonological class. Thus, a total of 18 classification results were found per speech task. This experiment allowed us to evaluate which are the most sensitive phonological classes prior to continuing with the longitudinal evaluation. All results per phonological class are included in
Table 6. Due to space limitations, only kernels associated with the best results are shown. The upper part of the table indicates results of the monologue while the bottom part includes those of the read texts. In both cases, the average performance across phonological classes is around 65% accuracy, which despite not being very high, provides the additional advantage of being clinically interpretable. Values and corresponding phonological classes with accuracies above 70% are highlighted in bold.
Notice that both speech tasks (monologue and read text) indicate that among the most discriminative phonological classes are back and open. This indicates that the movement and control of articulators involved in the production of vowel phonemes, e.g., tongue, jaw, and lips, play a crucial role in the evaluation of speech production. Other classes that were shown to be relevant for automatic discrimination were flap and lateral, which implies that the abnormal production of the /l/ and /R/ phonemes is also characteristic of Parkinson’s speech. Notice that the production of these phonemes also requires proper control of the tongue, as in the case of the above-mentioned classes. Finally, to some extent, these observations can be summarized by the fact that the consonantal class yielded good classification results when the read texts were considered. Notice that this phonological class includes phonemes like /m/, /n/, /l/, /t/, /p/, and /k/, which require the accurate control of the tongue, velum, and lips.
Based on the aforementioned results, the next experiment considers the evaluation of disease progression considering the phonological posteriors. The main hypothesis is that there are certain phonological classes that allow disease progression to be observed more accurately, and those phonemes are such that they require the control of specific articulators.
4.2. Automatic Dysarthria Level Progression Evaluation
This experiment considers speakers of the longitudinal corpus as the test set exclusively. This means that models resulting from
Section 4.1 are taken as they are, without any further optimization of meta-parameters. One classification model is created from each phonological class. Dysarthria level progression is evaluated in two scenarios: (i) on the set of 16 patients in the longitudinal corpus and (ii) on a subset of 7 patients who showed effective dysarthria progression according to the mFDA values in the longitudinal corpus. Effective progression between session 1 and session 3 was considered as the inclusion criterion.
4.2.1. Sensitivity Analysis of the Longitudinal Recordings
This experiment intends to find which phonological classes best model PD progression, considering the three recording sessions of the longitudinal corpus.
Figure 3 shows two radar plots that include the classification scores per phonological class obtained from monologues and read texts. Only results from the 16 patients of the longitudinal corpus are considered. This is the reason why we associate this analysis with a sensitivity measurement. The result of the analysis is summarized in
Figure 3. Three different plots are included per speech task: one per recording session. Notice that both figures show minimal changes in different phonological classes across recording sessions. Notice also that, by comparing sessions 1, 2, and 3, we are implicitly considering a time span of 12 months, which may not be long enough to observe changes in the dysarthria level of PD patients. According to [
24], changes in the motor skills of PD patients are not clearly observable within a year.
In the context of SVM classifiers, classification scores are directly associated with the distance of the given sample to the separating hyperplane. Therefore, through this analysis we intend to study the relationship between those scores and the mFDA values of corresponding samples. The main assumption is that the larger the classification score, the more advanced the dysarthria level of the corresponding patient. This assumption is supported by the fact that, if the dysarthria level of a given patient increases, the classifier will be more confident regarding its belonging to a particular class, and will therefore locate that particular speaker at a more distant point in the latent space, resulting in a larger distance. Of course, the opposite can also happen, i.e., one patient can be located far away from the separating hyperplane in the first recording session, and get closer in the second or in the third recording session. This could happen due to medication effects or mood changes, but these phenomena are beyond the scope of this paper.
One possible way of objectively evaluating whether phonological posteriors are modelling disease progression across the three recording sessions is to measure the area of each radar-plot resulting from each recording session.
Table 7 shows the results of such area computations. Notice that there is a clear increase in the area of session 1 vs. session 3 for both speech tasks, monologue and read text. However, although the areas show the progression of dysarthria level, it is hard to establish which are the phonological classes that best explain such progression. For this reason, we decided to focus on modelling only those patients for which the mFDA labels evidenced effective dysarthria progression.
4.2.2. Modelling of Effective Dysarthria Level Progression
The results of the labelling process for the longitudinal corpus according to the mFDA scale are indicated in
Table 3. As mentioned above, twelve months may be a very short period of time to observe objective changes in speech production patterns. One of the main complexities of PD is the fact that it does not progress the same way for all patients. This makes its modelling a real challenge, regardless of the bio-signal, e.g., speech, gait, handwriting, facial expression, etc. In the case of this work, based on the labels of the mFDA scale, we observed that not all patients showed dysarthria progression from session 1 to session 3. Given this scenario, we decided to take a subset of patients, considering only those who showed progression between these two sessions. After performing this filtering, a total of seven patients from the original sixteen were selected. These patients are highlighted in bold in
Table 3. The classification scores distribution obtained for patients in session 3–session 1 is illustrated in
Figure 4. Monologue and read text are included on the left and right side plots, respectively. Distributions in green represent the group of patient who showed dysarthria level progression, and distributions in blue represent those patients who did not show progression. Notice that in both cases, the group of patients who showed progression are slightly shifted to the right.
Once the subset of speakers was created, we repeated the previous processes and created the radar plots to analyse the behaviour of specific phonological classes. The result is shown in
Figure 5.
Similarly to the analysis performed for the complete longitudinal cohort, dysarthria progression was objectively measured according to the areas of the radar plots per session. The results are included in
Table 8. Notice that these more specific analyses show that the area increase in the case of monologues is minimal, therefore inconclusive regarding specific phonological classes. Conversely, for read text, the area indicated an increase from 0.6695 to 0.8858 (36.8% of increase). Notice that in this experiment, we were focused on those patients who showed dysarthria progression according to a clinical criterion (mFDA values). If this had not been the basis of our consideration, we would have only considered the results in
Figure 3 and
Table 7. This would have masked the patterns that we can observe when focusing only on these seven patients. For instance, when observing the radar plots in
Figure 5, we can conclude that strident, dental, pause, back, and continuant phonological classes are the ones that explain such an increase. Based on the equivalence between phonological classes and phonemes indicated in
Table 5, we observe that what predominantly models dysarthria level progression is the control of the
lips (phoneme /f/)) and
tongue (phonemes /t/, /d/, /l/, /s/, /a/, /o/, and /u/).
5. Discussion
This study introduced a methodology that enables the measurement of differences in the impact of dysarthria on different phonological classes, therefore reflecting its impact on different processes associated with speech production, including respiration, articulation, timing, and others. According to our findings, this methodology is suitable for the monitoring of speech deterioration in PD patients over time using recordings collected every month over several years. If the monitoring period is long enough, we could show which specific phonemes are systematically mispronounced over time (showing more degradation than others). This study introduced the use of phonological posteriors to assess the dysarthria level progression in a cohort of PD patients who were recorded three times in a time frame of three years (once per year). The ground truth was established as the values of the mFDA scale assigned by expert S&L therapists. Phonological models were based on a set with 18 Gated Recurrent Units (GRUs) that estimate the posterior probability of each phoneme to be “correctly pronounced” according to a set with 18 phonological classes. A total of 18 models, one per phonological class, were created per patient after analyzing the recordings of each session (three in total). Eighteen classification models (one per phonological class) for distinguishing between PD and HC subjects based on phonological posteriors were trained with the PC-GITA corpus. These 18 models were used to test the three recording sessions per speaker in the longitudinal analysis. No further optimization was performed for these classification models, i.e., the longitudinal corpus was used as a test set, representing a realistic scenario in which a patient comes to the clinic and requests to be evaluated to know whether his/her dysarthria level is progressing or not. The deterioration of “pronunciation correctness” was measured as the distance of each sample to the separating hyperplane created with the classifier. The mFDA labelling process showed that not all patients exhibit dysarthria level progression. This was validated in the first experiment. Afterward, we focused on evaluating only those patients who showed dysarthria level progression according to the mFDA. This filtering resulted in a subset of seven patients. When analyzing the recordings of those patients in session 1 vs. session 3, it was possible to observe specific patterns that might explain the progression of dysarthria. Specifically, strident, dental, pause, back, and continuant phonological classes arise as the ones that explain such a progression, which means that specific phonemes like /t/, /d/, /l/, /s/, /a/, /o/, and /u/ can be directly associated with this phenomenon.
Other studies in the literature have addressed the topic of the automatic evaluation and monitoring of PD symptoms considering speech-related biomarkers. Spectral-based analysis, articulation, and voice quality measures are among the most common [
30], while respiration emerges as a biomarker due to its close relation with speech production [
31]. Syllable counting, duration, and pause ratio have also been considered [
32]. According to a review [
33], phonation, articulation, and prosody features are found in most of the existing works. This paper introduces a new aspect of speech that is based on measuring how “correctly” the phonemes are being pronounced. According to our literature review, this is the first time this speech aspect has been used to evaluate the progression of dysarthic speech symptoms in PD patients.
Although specific phonemes showed certain patterns in our experiments, it is important to stress that this work set out to study the progression of PD over time being objectively evidenced in speech patterns. Such patterns allowed us to observe dysarthria progression from a broad perspective, rather than detecting specific patterns in specific phonemes with an association to PD symptoms. We believe that these findings provide objective evidence about possible abnormalities in broader speech-related processes like respiration, therefore contributing to a better understanding of the relationship between speech production patterns and other speech-related processes affected when suffering from PD.
Limitations: The small sample size considered in the longitudinal evaluations constitutes a clear limitation for this study. In fact, one-half of the cohort showed “improvement” in their mFDA score, and this phenomenon was not addressed in the paper. This could be a result of the fact that not every PD patient shows dysarthria, but it could also be because the scale is not sensitive enough. Further investigation is necessary to clarify this aspect. Additionally, this reduced sample size also limited our exploration of possible differentiating patterns between male and female patients. Therefore, increasing the number of participants would make the observations reported in this paper stronger, resulting in models that are more generalizable and suitable to the establishment of further discussion regarding specific gender-related disease behaviour. Additionally, this work has not addressed the problem of changing acoustic conditions over time, which constitutes a challenge in speech processing, especially when recordings are collected in at-home conditions.