Background: Thyroid function tests are frequently requested, but manufacturer reference intervals often lack age- or sex-based stratification. This study established indirect reference intervals for thyroid stimulating hormone (TSH), free thyroxine (fT4) and free triiodothyronine (fT3) in a large adult population. We evaluated unsupervised algorithms and conventional partitioning to identify subgroups, compared multiple limit estimation methods, and assessed the diagnostic performance of the derived intervals against an independent, clinically defined external validation cohort.
Methods: Data from 37,255 adults (age ≥ 18) collected between 2022 and 2024 were analyzed; a reference population of 5870 individuals was established following standardized exclusion criteria. Age-based subgroups were identified through hierarchical clustering with the Elbow method, and variable importance was assessed using Random Forest analysis. Reference intervals were calculated using non-parametric, Bhattacharya, refineR, and reflimR algorithms, applied to three population frameworks: (1) the total population without stratification, (2) subgroups derived from hierarchical clustering and (3) subgroups defined by the conventional Harris–Boyd partitioning method. Diagnostic performance was subsequently evaluated in an independent external cohort (National Health and Nutrition Examination Survey [NHANES];
N = 2297) for all estimated reference intervals. Three classification scenarios were assessed: TSH-only, fT4-only, and combined TSH + fT4, with sensitivity, specificity, Youden index, and decision curve analysis performed for each.
Results: Random Forest analysis identified age as the dominant variable influencing TSH, fT4 and fT3 distributions (mean decrease in accuracy: TSH 41.62, fT4 44.18, fT3 43.7), while sex showed the lowest impact. Clustering yielded six age-based subgroups for analytes. Harris–Boyd partitioning yielded six age-based subgroups for TSH, two sex-based subgroups for fT4, and six combined age-and-sex subgroups for fT3. TSH limits were broadly concordant across all three approaches (six-subgroup partitioning: 0.36–0.68 to 4.75–5.67 mIU/L). For fT4, conventional (sex-based) and clustering (age-based) partitioning produced similar ranges (11.33–20.08 pmol/L), except reflimR’s notably lower limit (10.90 pmol/L). For fT3, conventional (age + sex) and clustering (age-only) partitioning showed comparable ranges (3.36–7.03 pmol/L), with clustering revealing a clearer age-related decline in the oldest group. Diagnostic performance varied markedly by analyte. TSH-only classification achieved positive discrimination across all 13 methods (Youden index: 0.173–0.239). In contrast, fT4-only and combined TSH + fT4 classifications performed at or below chance for most methods, with 75% of the cohort falling outside fT4 reference intervals, indicating an inter-platform harmonization issue rather than a partitioning failure. Decision curve analysis confirmed TSH-only classification’s superiority, exceeding universal testing from pt ≈ 0.20 onward across all methods.
Conclusions: Age-stratified reference intervals combined with limit estimation showed potential diagnostic advantages over manufacturer and non-stratified intervals; however, an independent external validation using a clinically defined outcome indicated that this advantage was not consistently reproduced and was dependent on the clinical decision threshold considered. These results suggest that age stratification and algorithm choice merit further clinically adjudicated validation before broad clinical adoption, and that unsupervised clustering offers a practical, objective alternative to manual subgrouping for laboratories pursuing this approach.
Full article