Abstract
Machine-learning gait studies often report high accuracy for separating neurodegenerative diseases, but rarely test whether this reflects pathology or the demographic differences that accompany disease in convenience cohorts. We audited the PhysioNet Gait in Neurodegenerative Disease Database (63 subjects: controls, Parkinson’s disease [PD], Huntington’s disease [HD], amyotrophic lateral sclerosis [ALS]), classifying 37 stride-rhythm features under subject-disjoint nested leave-one-subject-out cross-validation. A random forest reached a macro-averaged one-vs-rest AUROC of 0.832 (95% CI 0.759–0.904; permutation p = 0.001). Four complementary audits—demographic-only baselines, confound adjustment, feature-family ablation, and age matching—show that this overstates gait-specific discrimination: five anthropometric variables alone reached an AUROC of 0.812, statistically indistinguishable from the gait model (p = 0.701), and the de-confounded gait model (0.755) fell further to 0.585 once age was matched across groups. HD was the exception, essentially unaffected by adjustment (0.812 → 0.804) and carried by variability and long-range-correlation measures. In an independent, age-matched cohort recorded with a different sensor (n = 110, PD versus control only), the same vulnerability reappeared through a different variable: the gait model (0.828) showed no statistically significant difference from walking speed alone (0.746, p = 0.11), while age and anthropometry fell to chance; this cohort contained no HD subjects, so the HD finding remains externally unvalidated. The shortcut vulnerability generalises across cohorts; the demographic variable that drives it does not. Four-class differential diagnosis is not demonstrated on this benchmark; HD gait arrhythmicity is the one signal that withstands internal sensitivity analysis. We recommend that accuracy claims from small, demographically imbalanced gait cohorts be accompanied by this kind of audit.