2.1. Data Sources
This was a pre-specified, descriptive, and analytical cross-sectional secondary analysis of two public adult audiometric datasets. The pooled German HÖRSTAT (Oldenburg/Emden, 2010–2012) and Aalen (2008–2009) sample comprised 3105 adults aged 18–97 years with manual air-conduction audiometry at 250, 500, 1000, 2000, 3000, 4000, 6000, and 8000 Hz in both ears, obtained using the HDA200 circumaural headphones and an ascending procedure as described by Holube and von Gablenz [
8]. The US National Health and Nutrition Examination Survey (NHANES) is a continuous, nationally representative probability sample of the US civilian non-institutionalised population conducted by the National Center for Health Statistics. Adult air-conduction audiometry at 500, 1000, 2000, 3000, 4000, 6000, and 8000 Hz was available in cycles 2011–2012 (G), 2015–2016 (I), and 2017–March 2020 pre-pandemic (P). NHANES 2013–14 (cycle H) did not include the audiometric examination module and was therefore excluded from the present analysis. Cycles G and I tested ages 20–69 years; cycle P tested ages 70+ years among adults. Combined, these cycles contributed 10,575 adult participants with at least one audiometric examination attempted. Both datasets are publicly available; no additional ethical approval was required for this secondary analysis.
2.2. Variables and Harmonisation
We harmonised the two datasets to a common schema of bilateral air-conduction thresholds at 500–8000 Hz (the HÖRSTAT 250 Hz measurement was excluded because NHANES does not measure it). HÖRSTAT missing-value codes 111 (not measured) and 222 (not valid) were recoded to missing. NHANES “no response” (666) and “could not obtain” (888) were likewise recoded to missing in primary analyses. The 1 kHz measurement used AUXU1K1 (first measurement); the AUXU1K2 retest was used only for reliability auditing. NHANES additionally releases retest thresholds at 6 and 8 kHz (AUXR6K, AUXR8K); these were examined for the purpose of a high-frequency boundary check and are reported in
Section 3.11. Sex was coded as male/female and age as years. Regular hearing-aid use was harmonised across NHANES cycles as follows: cycles G and I (AUQ146 = 1 “ever worn hearing aid” AND AUQ152 ∈ {1, 2, 3} “worn almost every day/most days/about half the time in past year”); cycle P (AUQ147 = 1 “now use hearing aid/amplifier/implant” AND AUQ153 ∈ {1, 2, 3}). The HÖRSTAT HearAids variable was treated as regular use by design (coded 1 = aided, 2 = unaided). Cycle-specific skip patterns meant HA questions were asked only to adults 20–69 in cycle G, to all adults in cycle I, and to adults 70+ in cycle P; hearing-aid analyses respected these age bands.
Self-rated hearing status in NHANES was captured by AUQ054 on a five-level Likert scale (1 = Excellent, 2 = Good, 3 = A little trouble, 4 = Moderate trouble, 5 = A lot of trouble; a sixth level, “Deaf”, was recoded to 5 for modelling because n = 22 across cycles). AUQ054 was administered to all respondents in all three cycles with near-complete response (non-missingness > 99.9%). A primary binary indicator “any hearing trouble” was defined as AUQ054 ≥ 3.
2.3. Derived Audiometric Metrics
For each participant and ear, we computed PTA4 as the mean threshold at 500, 1000, 2000, and 4000 Hz and PTA68 as the mean threshold at 6000 and 8000 Hz, using complete-case logic (all constituent frequencies present for that ear). Bilateral validity required both ears to have a computable metric. Better-ear values were the minimum across the two ears. Interaural differences were |ear-R − ear-L| for each metric. PTA68 was chosen as the primary high-frequency summary, rather than the more common HF4 (3, 4, 6, 8 kHz), to eliminate conceptual overlap with the 4 kHz term already included in PTA4; HF4 and a PTA346 summary (3, 4, 6 kHz) aligning with Hoffman et al. were examined as secondary metrics.
2.5. Statistical Analysis
Prevalence estimates were reported with 95% Wilson confidence intervals [
12] for unweighted HÖRSTAT estimates and logit-transformed, design-adjusted confidence intervals for survey-weighted NHANES estimates (Taylor-linearized variance estimator as implemented in samplics.TaylorEstimator). The NHANES design used masked stratum (SDMVSTRA) and primary sampling unit (SDMVPSU) variables. Combined sampling weights across cycles were constructed from cycle-specific NHANES MEC weights [
13]: WTMEC2YR/2 for cycles G and I and WTMECPRP for cycle P [
14]. Because the adult audiometric examination was administered to ages 20–69 in cycles G and I and to ages 70+ only in cycle P, the adult age modules are disjoint by cycle. No additional period rescaling of the cycle P weight was applied, because cycle P is the sole source of the 70+ domain; the pooled weighted estimate for all adults is consequently a synthetic combination of a 2011–2016 working-age frame and a 2017–March 2020 frame for ages 70+, and we refer to it as the pooled NHANES analytic sample rather than as a single contemporaneous US adult cross-section. Strata were verified to be non-overlapping across cycles, so SDMVSTRA and SDMVPSU were used directly in the pooled design. As a sensitivity analysis, rescaling all cycles to a common 7.2-year reference period (a procedure appropriate only when cycles contribute overlapping populations) under-weights the cycle-P-sourced 70+ domain and shifts the overall weighted estimate only modestly, from 24.9% to 24.5%; the small magnitude reflects the small population share of the 70+ domain and confirms that this overall-population companion estimate is not sensitive to this choice. PTA4-discordant HF prevalence was stratified by sex and by ten-year age groups (20–29, 30–39, …, 70–79, 80+).
Logistic regression models estimated the association of age (per 10 years) and sex with (1) PTA4-discordant HF restricted to PTA4-normal participants, (2) PTA4 asymmetry, and (3) PTA68 asymmetry. NHANES models were fitted as design-based weighted logistic regressions with first-order Taylor-linearized variance estimation incorporating the survey weights, SDMVSTRA strata, and SDMVPSU primary sampling units (design degrees of freedom = number of PSUs minus number of strata); HÖRSTAT models used standard binomial GLMs. An age × sex interaction was tested for model (1) in both cohorts. Cohen’s κ for PTA4 > 25 vs. PTA68 > 25 binary agreement was computed per cohort [
15], with design-weighted computation for NHANES.
In NHANES, a further logistic model predicted “any hearing trouble” (AUQ054 ≥ 3) from audiometric phenotype (four-level categorical), age, and sex, with design-adjusted SEs as above. A secondary continuous dose–response model regressed any-trouble on PTA68 per 10 dB, age per decade, and sex, restricted to PTA4-normal adults.
Seven sensitivity analyses (S1–S7) varied: the HF metric (HF4, S1; PTA346, S2); the classification thresholds (a matched 20 dB HL threshold applied to both the PTA4-normal criterion, PTA4 ≤ 20 dB HL, and the high-frequency cutoff, PTA68 > 20 dB HL, S3; and a 40 dB HL moderate-severe HF cutoff, S4); the handling of NHANES “could not obtain” codes (S5; recoded to 100 dB HL); the handling of HÖRSTAT thresholds exceeding audiometer maximum output (S6; recoded to max rather than max + 5 dB); and the handling of the NHANES no-response code at 6 or 8 kHz (S7; recoded to the audiometer ceiling). Alternative asymmetry cutoffs (10 and 20 dB) were examined separately and are reported in
Supplementary Table S6. A forest plot of PTA4-discordant HF prevalence under each specification was produced to visualise robustness.
High-frequency measurement reliability was examined using the NHANES 6 and 8 kHz retest thresholds (AUXR6K, AUXR8K), comparing each retested ear’s initial and repeat value. Because those retests are contributed almost entirely by severely impaired ears, boundary sensitivity was additionally assessed by simulation on the NHANES analytic sample. Independent Gaussian error was added to each constituent frequency at per-frequency standard deviations of 3, 5, and 7.6 dB (the last being the single-measurement standard deviation implied by the observed 6–8 kHz retest difference standard deviation of 10.8 dB under an assumption of independent, equal-variance repeat measurements, 10.8/√2 = 7.6), and ear-specific and better-ear PTA68 were recomputed within each replicate before reapplying the survey weights, with the better-ear designation held at its observed value so that error entered through the two constituent frequencies of that ear; 2000 replicates were drawn per setting under a fixed random seed, and errors were assumed independent across frequencies and ears. A simulation–extrapolation (SIMEX) analysis assuming a per-frequency standard deviation of 5 dB used λ = 0, 0.5, 1.0, 1.5, and 2.0 with 400 replicates per λ and quadratic extrapolation to λ = −1. These are sensitivity scenarios rather than survey estimates and are reported without confidence intervals. Analyses used Python 3.12 (Python Software Foundation, Wilmington, DE, USA) with the pandas, numpy, scipy, statsmodels, samplics, matplotlib and pyreadstat packages (all open-source, obtained from the Python Package Index).
The primary endpoint (PTA4-discordant HF), the secondary asymmetry and unilateral-like endpoints, and the seven sensitivity specifications were finalised before data extraction. The AUQ054 self-rated hearing analyses, the Cohen’s κ agreement analysis, and the age × sex interaction test were added as planned extensions after the primary analyses were completed and cross-cohort reproducibility was observed. None of these analyses were registered on a public platform. Reporting follows the STROBE guideline for cross-sectional studies [
16] (the checklist can be found in the
Supplementary Material).