1. Introduction
Halitosis (oral malodor) is a frequent complaint and remains clinically relevant because it can impair social interactions, self-confidence, and overall well-being. In population-level evidence syntheses, halitosis shows a substantial pooled prevalence across settings, indicating it is not a niche issue restricted to specialty clinics [
1]. Beyond prevalence, halitosis is increasingly recognized as a patient-centered condition with measurable psychosocial consequences: systematic review evidence supports a significant association between halitosis and worse oral health-related quality of life (OHRQoL), suggesting that malodor can translate into a meaningful functional and emotional burden for affected individuals [
2,
3].
Most cases of halitosis are intra-oral and arise from microbial degradation of proteins and peptides with production of volatile compounds, especially volatile sulfur compounds (VSCs). Classic and contemporary literature emphasizes the importance of local ecological drivers such as plaque accumulation, gingival inflammation, and tongue coating in initiating and sustaining malodor [
4,
5,
6]. In support of this biologic framework, quantitative evidence indicates that periodontitis is associated with halitosis, consistent with deeper anaerobic niches and increased VSC production in inflammatory periodontal environments [
7,
8,
9]. Studies have also detected key malodor-related compounds, including hydrogen sulfide and methyl mercaptan, in periodontal pockets, reinforcing the plausibility of periodontal contributions to malodor phenotypes [
10,
11]. In addition, oral malodor is often described as multifactorial, with salivary hypofunction and xerostomia potentially amplifying odor intensity through reduced clearance, altered buffering, and shifts in microbial homeostasis [
8]. For this reason, standardized saliva collection and flow assessment methods are important when investigating halitosis correlates and potential modifiers [
10].
A further challenge in halitosis research is measurement heterogeneity. Organoleptic assessment is widely used clinically and captures perceived odor intensity, while device-based approaches (e.g., halitometry) estimate VSC burden; however, meta-analytic findings suggest these methods do not always correlate strongly, implying that reliance on a single metric may incompletely characterize clinical and patient-relevant severity [
9]. Accordingly, combining clinician-rated measures (e.g., organoleptic scoring) with patient-reported outcomes can provide a more comprehensive understanding of malodor burden. The Halitosis Associated Life-Quality Test (HALT), including validated language adaptations such as the Romanian version, enables quantification of halitosis-related quality-of-life impact in young adult populations [
12].
Orthodontic therapy represents a clinically important context for malodor research because appliances can alter plaque retention, gingival inflammation, and oral hygiene behaviors. Fixed appliances introduce plaque-retentive structures and may increase gingival inflammation, which can plausibly worsen malodor; clinical orthodontic studies have therefore evaluated halitosis alongside periodontal indices during treatment [
13]. In contrast, clear aligners are removable, potentially facilitating tooth cleaning and interdental hygiene. Evidence syntheses indicate that aligner therapy is generally associated with better periodontal parameters than fixed appliances, supporting the hypothesis that aligners could be linked to lower malodor burden through improved plaque and gingival control [
14,
15].
Despite this rationale, comparative studies that simultaneously include untreated controls, integrate both clinician-rated malodor and patient-reported halitosis burden, and evaluate key modifiers such as plaque, tongue coating, and salivary flow in young adults remain limited, particularly in adolescent and young adult populations. Because cross-sectional comparisons across orthodontic treatment modalities may also reflect differences in baseline oral hygiene motivation, treatment selection, and care-seeking behavior, any observed group contrasts should be interpreted as associations rather than causal effects of appliance type alone. Therefore, the present study compares halitosis burden (HALT), organoleptic malodor severity, and OHRQoL (OHIP-14 [
15]) across clear aligner users, fixed-braces patients, and controls, and examines how oral indices and salivary measures relate to clinically meaningful malodor phenotypes.
2. Materials and Methods
2.1. Study Design and Setting
This cross-sectional observational study was conducted to compare oral malodor burden and oral health-related quality of life among clear aligner users, conventional fixed-brace patients, and non-orthodontic controls, with a planned focus on salivary flow as a clinically relevant modifier. The study was implemented in two university-affiliated dental clinics (orthodontics and preventive dentistry) between March 2024 and November 2025. All procedures followed the principles of the Declaration of Helsinki, and participation was voluntary.
Ethics approval was obtained from the institutional review board prior to recruitment. All participants received written and verbal explanations of the study aims, procedures, and confidentiality safeguards. Written informed consent was collected from each adult participant. For participants under 18 years, written consent was obtained from a parent/legal guardian, together with written assent from the participant. Personal identifiers were removed at entry into the study database, and data were stored on a password-protected workstation accessible only to the study team.
2.2. Participants, Recruitment, and Grouping
Because orthodontic participants and controls were recruited from different clinical streams, the control group may have differed in care-seeking profile, oral hygiene motivation, and baseline oral health status in ways that could influence plaque, tongue coating, and malodor independently of appliance category. This concern was addressed analytically through adjustment for measured oral hygiene behaviors and clinical oral indices, but residual selection bias and incomplete group comparability cannot be excluded. Participants were recruited consecutively from routine orthodontic follow-up visits (aligner and fixed-brace groups) and from preventive-care appointments (controls). Eligibility criteria were: age 15–35 years, stable general health, and willingness to complete questionnaires and undergo breath and oral assessments. To reduce major confounding, we excluded individuals with: (1) systemic conditions known to markedly affect breath (uncontrolled diabetes, advanced liver/kidney disease), (2) upper respiratory infection in the prior 10 days, (3) antibiotic use in the prior 4 weeks, (4) current daily smoking, and (5) active untreated periodontal disease requiring specialist intervention (screened clinically at the visit). Participants were also excluded if they had used strong antiseptic mouthwash (chlorhexidine) within 24 h, because this could artificially lower malodor measures. To reduce major confounding from advanced periodontal disease, participants were excluded if screening examination showed findings suggestive of active periodontal disease requiring specialist management, including generalized bleeding with visible inflammation, periodontal pocketing judged clinically incompatible with routine preventive or orthodontic follow-up, or other signs prompting referral for periodontal evaluation. Because this was not a dedicated periodontal study, full-mouth periodontal charting was not performed; this is acknowledged as a limitation.
The age range of 15–35 years was selected to capture the adolescent and young adult population most commonly encountered in contemporary orthodontic practice in the participating clinics, while limiting inclusion of older adults with potentially greater age-related periodontal and salivary variability. Age was also retained as a covariate in multivariable analyses to reduce residual confounding.
The study classified participants by current treatment modality (clear aligners versus conventional labial fixed appliances) rather than by individual commercial system or bracket prescription. Detailed information such as aligner brand, attachment design, use of elastics, bracket slot system, ligature type, or prophylaxis schedule was not uniformly available across all participants and was therefore not analyzed. The final analytic sample included 184 participants, distributed into three groups: aligners (n = 62), fixed braces (n = 64), and controls (n = 58). Group definitions were set a priori: (1) aligner users currently wearing removable clear aligners ≥ 3 months, (2) fixed-brace patients currently wearing conventional labial fixed appliances ≥ 3 months, and (3) controls without current orthodontic treatment and without orthodontic treatment in the prior 12 months. Because behavior strongly influences oral malodor, we also recorded pre-specified modifiers for subgroup analyses: frequent snacking (≥3 snack episodes/day vs. fewer), xerostomia symptoms, and salivary flow status (low vs. normal). For participants undergoing orthodontic treatment, treatment duration was defined as the time from treatment initiation to the study enrollment/assessment visit.
A priori sample size planning targeted adequate power to detect a medium difference in HALT scores across three groups (effect size f ≈ 0.28, α = 0.05, power = 0.80), yielding a minimum sample near 168 participants; enrollment was continued to 184 to reduce instability in subgroup analyses (especially low-flow strata).
2.3. Measures and Data Collection
All assessments were performed in a standardized morning window (09:00–11:30). Participants were instructed to avoid eating or drinking (water allowed) for 2 h beforehand and to avoid pungent foods/alcohol for 12 h. They were also asked not to brush, floss, chew gum, or use mouthwash during the 2 h pre-assessment window. At the visit, the team recorded demographic data, orthodontic status, and oral hygiene behaviors (brushing frequency, interdental cleaning, tongue cleaning, and mouthwash use). Dietary behavior was simplified into a clinically practical classification: frequent snacking (≥3 episodes/day) vs. infrequent snacking, assessed using a brief standardized interview and a one-page checklist.
Primary outcomes were:
- •
HALT (Halitosis Associated Life-Quality Test) total score (0–100), assessing perceived halitosis impact.
- •
Organoleptic scoring was analyzed both as an ordinal clinical severity measure (0–5) and using two pre-specified thresholds with distinct purposes. First, a score of ≥2 was used descriptively to indicate clinically relevant malodor prevalence at the group level. Second, a stricter threshold of ≥3 was used in secondary predictive analyses and figures to represent a moderate-to-severe malodor phenotype. This distinction was adopted to separate broad clinical prevalence estimates from modeling of more pronounced malodor severity. To maintain terminology consistency throughout the manuscript, organoleptic ≥2 is referred to as clinically relevant malodor, whereas organoleptic ≥3 is referred to as a moderate-to-severe malodor phenotype.
- •
OHIP-14 total score (0–56), reflecting oral health-related quality of life.
To provide context beyond questionnaires alone, we assessed oral conditions that plausibly influence malodor:
- •
Plaque index using a simplified, standardized 0–3 scale averaged across index tooth sites. Gingival inflammation was also recorded using a simplified gingival index on a 0–3 scale, where higher scores indicated greater visible inflammatory burden during routine clinical examination. This variable was included as an ancillary descriptive oral health measure rather than as a primary predictor in the main multivariable models.
- •
Tongue coating score on a 0–3 scale (0 = none, 3 = heavy coating), recorded under consistent lighting with the tongue gently protruded.
- •
Unstimulated whole salivary flow (mL/min) using a practical passive drool technique: participants sat upright and allowed saliva to pool and drip into a pre-weighed container for 5 min. Flow rate was calculated as volume/time. Low flow was defined as <0.25 mL/min, consistent with common clinical thresholds for reduced resting salivary output.
Two examiners completed pre-study calibration sessions to harmonize organoleptic scoring, plaque grading, tongue-coating assessment, and salivary-flow collection instructions. In a subset of participants, duplicate organoleptic scoring was performed for quality control; however, formal inter- and intra-examiner reliability coefficients were not retained for final reporting and are therefore acknowledged as a limitation that may have contributed measurement variability in this subjective outcome.
Complete examiner blinding to treatment group was not feasible because fixed appliances and aligners were clinically visible during assessment. To reduce expectancy bias, organoleptic scoring followed the same pre-specified rubric, assessment sequence, and standardized morning assessment conditions for all groups; however, because masking was not feasible, some observer expectancy cannot be excluded.
Compliance with the pre-assessment instructions was verified on arrival using a brief standardized checklist and verbal confirmation before breath and oral measurements were performed. When participants reported clear protocol deviations that could materially influence breath assessment, the assessment was deferred and rescheduled within the routine clinic workflow.
Height and weight were recorded at the study visit and body mass index (BMI, kg/m2) was calculated as weight divided by height squared. BMI was included as a baseline descriptive characteristic because it may relate indirectly to oral health behaviors and dryness-related symptoms.
2.4. Statistical Analysis
All analyses were conducted using R statistical software (R Foundation for Statistical Computing, Vienna, Austria; version 4.2) with a two-sided significance threshold of p < 0.05. Continuous variables were inspected visually (histograms/Q–Q plots) and formally (Shapiro–Wilk tests) to guide parametric vs. non-parametric testing. For comparisons across the three orthodontic groups, we used one-way ANOVA for normal outcomes (total HALT and OHIP-14) with Tukey post hoc tests for pairwise contrasts. For outcomes with non-normal distributions or clear ordinal behavior (organoleptic scores), we used Kruskal–Wallis tests with appropriate post hoc procedures. Categorical variables (prevalence of organoleptic ≥3, frequent snacking) were compared using chi-square tests, switching to Fisher’s exact test when expected cell counts were small.
To explore patterning beyond basic group comparisons, we pre-specified two additional analytic layers. First, we evaluated whether salivary flow modified appliance-related effects using interaction terms and subgroup analyses (low-flow vs. normal-flow strata). Second, to examine whether a broader clinical feature set improved discrimination of a moderate-to-severe malodor phenotype (organoleptic ≥ 3), we compared a base model with an expanded model including plaque, tongue coating, and salivary flow. These analyses were intended as exploratory, hypothesis-generating internal comparisons rather than as validated clinical prediction tools. Model performance was summarized using the area under the ROC curve (AUC), and incremental performance (ΔAUC) was quantified via bootstrap resampling. Decision-curve analysis (DCA) was used only as an exploratory internal visualization of model behavior across clinically reasonable risk thresholds (0.10–0.50), and not as evidence of clinical readiness, transportability, or external validity.
Finally, associations among key continuous measures (HALT, organoleptic score, OHIP-14, plaque, tongue coating, salivary flow) were examined with Spearman correlation coefficients to avoid over-assuming linearity.
3. Results
Across the three groups (aligners, fixed braces, controls), participants were similar in age (22.7–23.1 years;
p = 0.706), sex distribution (female: 56.2–66.1%;
p = 0.519), and BMI (22.6–22.9 kg/m
2;
p = 0.677), indicating demographic similarity, but not full clinical comparability. As expected, orthodontic treatment duration was longer in the fixed-braces group than in aligners (10.9 ± 2.8 vs. 8.6 ± 2.7 months;
p < 0.001). Oral hygiene behaviors differed meaningfully: brushing ≥2/day was more frequent in orthodontic patients than controls (88.7% aligners, 85.9% fixed braces vs. 63.8% controls;
p = 0.001), while daily flossing was highest among aligner users (69.4%) compared with fixed braces (43.8%) and controls (36.2%) (
p < 0.001). Tongue cleaning was also most common in aligners (62.9% vs. 32.8% fixed braces and 48.3% controls;
p = 0.003). Mouthwash use was highest in fixed braces (67.2%) compared with aligners (56.5%) and controls (36.2%) (
p = 0.002), whereas snacking frequency and sugary-snack self-report were similar across groups (
p = 0.414 and
p = 0.944). Taken together, the behavioral differences indicate that between-group contrasts should not be interpreted as arising from appliance status alone (
Table 1).
Fixed braces showed the least favorable clinical profile, with higher plaque (1.7 ± 0.3) and gingival inflammation (1.7 ± 0.2) compared with aligners (plaque 1.2 ± 0.3; gingival 1.3 ± 0.2) and controls (plaque 1.6 ± 0.6; gingival 1.6 ± 0.3), all with strong between-group differences (
p < 0.001 for plaque and gingival indices). Tongue coating was also higher in fixed braces (1.4 ± 0.3) than aligners (1.1 ± 0.3) and controls (1.1 ± 0.3) (
p < 0.001). Unstimulated salivary flow was lowest in fixed braces (0.3 ± 0.1 mL/min) versus 0.4 ± 0.1 mL/min in both aligners and controls (
p < 0.001). Although the prevalence of objectively “low salivary flow” (<0.25 mL/min) numerically favored aligners (9.7%) over fixed braces (23.4%) and controls (13.8%), this categorical comparison did not reach conventional significance (
p = 0.094). Xerostomia symptoms, however, were significantly more frequent in fixed braces (42.2%) than aligners (17.7%) and controls (25.9%) (
p = 0.009), indicating a clinically important subjective dryness burden in the fixed-braces group, as seen in
Table 2.
The fixed-braces group had the highest burden across all primary patient-centered outcomes, including HALT total score (53.7 ± 6.2) compared with controls (46.3 ± 6.4) and aligners (41.7 ± 7.4) (
p < 0.001). Clinician-rated malodor followed the same gradient (organoleptic 2.9 ± 0.4 fixed braces vs. 2.4 ± 0.6 controls vs. 2.2 ± 0.6 aligners;
p < 0.001), alongside worse oral health-related quality of life (OHIP-14: 18.6 ± 4.7 fixed braces vs. 15.9 ± 4.3 controls vs. 13.8 ± 4.8 aligners;
p < 0.001). Using the descriptive threshold of organoleptic ≥2, clinically relevant malodor prevalence was 96.9% in fixed braces, 79.3% in controls, and 66.1% in aligners (
p < 0.001), as seen in
Table 3. Because this low threshold may capture mild or borderline odor findings and is lower than cutoffs used in some prior reports, these percentages should be interpreted as broad descriptive indicators of detectable malodor burden within this clinic-based cohort rather than as diagnostic prevalence estimates of established halitosis [
16,
17]. Post hoc testing confirmed a consistent gradient across outcomes, but these between-group differences should be interpreted with caution because recruitment structure and oral hygiene differences may also have contributed.
Within both salivary-flow strata, fixed braces consistently showed worse outcomes than aligners and controls. Among participants with low flow, fixed braces had higher HALT (57.1 ± 5.9) than aligners (45.6 ± 6.7) and controls (48.4 ± 3.6) (
p < 0.001), higher organoleptic scores (3.1 ± 0.3 vs. 2.4 ± 0.4 vs. 2.7 ± 0.4;
p = 0.003), and higher OHIP (20.4 ± 4.3 vs. 14.7 ± 1.9 vs. 16.9 ± 4.6;
p = 0.018). The same pattern held in normal-flow participants (HALT: 52.6 ± 5.9 fixed braces vs. 41.3 ± 7.4 aligners vs. 45.9 ± 6.7 controls;
p < 0.001; organoleptic: 2.9 ± 0.4 vs. 2.2 ± 0.6 vs. 2.4 ± 0.6;
p < 0.001; OHIP: 18.1 ± 4.7 vs. 13.7 ± 5.1 vs. 15.8 ± 4.2;
p < 0.001). Importantly, the group×flow interaction was not significant for HALT (0.81), organoleptic (0.952), or OHIP (0.816), indicating that salivary flow did not meaningfully modify the relative differences between orthodontic groups, even though low flow tended to coincide with worse absolute scores (
Table 4).
Frequent snacking was associated with higher absolute symptom burden, but the orthodontic-group pattern remained consistent regardless of snacking status. In frequent snackers, fixed braces had the highest HALT (55.4 ± 5.4) compared with controls (47.6 ± 6.7) and aligners (44.8 ± 7.2) (
p < 0.001), as well as higher organoleptic scores (3.1 ± 0.3 vs. 2.4 ± 0.6 vs. 2.3 ± 0.6;
p < 0.001) and OHIP (18.9 ± 3.8 vs. 16.1 ± 4.8 vs. 14.4 ± 4.4;
p < 0.001). In infrequent snackers, the same hierarchy persisted (HALT: 52.1 ± 6.4 fixed braces vs. 45.4 ± 6.1 controls vs. 38.8 ± 6.3 aligners;
p < 0.001; organoleptic: 2.8 ± 0.4 vs. 2.4 ± 0.6 vs. 2.1 ± 0.6;
p < 0.001; OHIP: 18.4 ± 5.4 vs. 15.9 ± 3.9 vs. 13.3 ± 5.2;
p < 0.001). Plaque index differences mirrored clinical risk (frequent snackers: 1.8 ± 0.3 fixed braces vs. 1.6 ± 0.6 controls vs. 1.3 ± 0.3 aligners;
p < 0.001), and interaction terms were non-significant (group×snacking: HALT 0.237; organoleptic 0.181; OHIP 0.848; plaque 0.514), suggesting snacking does not substantially change the between-group contrasts (
Table 5).
Bivariate associations showed that patient-reported halitosis burden (HALT) aligned most strongly with clinician-rated malodor (ρ = 0.6,
p < 0.001), indicating substantial concordance between subjective impact and clinical assessment. HALT also correlated moderately with OHIP (ρ = 0.4,
p < 0.001), plaque (ρ = 0.4,
p < 0.001), and tongue coating (ρ = 0.4,
p < 0.001), supporting a coherent pattern linking malodor-related quality-of-life impairment to oral biofilm and tongue findings. Salivary flow correlated inversely with HALT (ρ = −0.3,
p < 0.001), organoleptic (ρ = −0.3,
p < 0.001), and OHIP (ρ = −0.3,
p < 0.001), consistent with dryness as a contributor to worse breath and QoL. Snacking frequency had smaller positive correlations with HALT (ρ = 0.2,
p < 0.01), organoleptic (ρ = 0.2,
p < 0.05), and tongue coating (ρ = 0.2,
p < 0.05), suggesting a weaker but detectable behavioral component (
Table 6).
After multivariable adjustment, the strongest independent predictor of HALT was the organoleptic score: each +1 increase was associated with a +6.2-point higher HALT (95% CI 4.3 to 8.0;
p < 0.001; partial R
2 20.2%). Tongue coating was also a major driver (+5.1 per +1; 95% CI 2.4 to 7.8;
p < 0.001; partial R
2 7.6%), as was OHIP-14 (+0.3 per point; 95% CI 0.1 to 0.5;
p < 0.001; partial R
2 7.2%). Xerostomia symptoms independently added a clinically meaningful increment (+2.4; 95% CI 0.6 to 4.1;
p = 0.008; partial R
2 4.0%). In contrast, being in the fixed-braces group versus aligners was not statistically significant once these clinical and behavioral factors were included (β = 2.2;
p = 0.101), and the fixed braces×low-flow interaction was also non-significant (β = 0.9;
p = 0.663), consistent with the stratified analyses. Overall fit was high (R
2 = 0.8; adjusted R
2 = 0.7; global
p < 0.001), and multicollinearity appeared acceptable (VIF range 1.1–2.6). This pattern suggests that the higher self-reported burden observed in fixed-brace patients was largely accounted for by concurrent clinical severity and symptom correlates rather than appliance category alone (
Table 7).
Greater plaque and tongue coating were the strongest predictors of worse organoleptic severity: each +1 increase in plaque index carried an OR of 4.3 (95% CI 2.1–9.1;
p < 0.001), while each +1 increase in tongue coating index carried an OR of 9.9 (95% CI 3.7–26.3;
p < 0.001), indicating a very steep escalation in the odds of being in a higher malodor category with tongue biofilm. Low salivary flow also increased severity (OR 2.5; 95% CI 1.1–5.9;
p = 0.036). Fixed braces remained an independent risk factor versus aligners even after adjustment (OR 4.9; 95% CI 1.9–12.5;
p < 0.001), whereas controls did not differ from aligners (OR 1.1;
p = 0.764). Female sex showed a modest but significant association with higher severity (OR 1.9; 95% CI 1.0–3.5;
p = 0.034), while snacking, flossing, tongue cleaning, and age were not significant in this model. Thus, appliance category appears more closely linked to clinician-rated malodor severity than to self-reported halitosis burden once other concurrent clinical correlates are considered (
Table 8).
Among patients with fixed braces, the heatmap shows that the chance of more noticeable bad breath (organoleptic ≥ 3) rises sharply when both plaque and tongue coating are high. The highest-risk pattern is the high plaque + high tongue coating cell, where 93.5% of patients met the ≥3 threshold (n = 8). In contrast, when tongue coating is low and plaque is low-to-mid, the proportion with organoleptic ≥3 stays much lower (for example, 36.6% (n = 7) in the low tongue + low plaque cell). When the same cells are split by salivary flow, the pattern is generally similar, but some “high plaque/high tongue” combinations remain high even when flow is normal, suggesting that the combination of plaque + tongue coating is a strong driver of worse breath in fixed-brace patients (
Figure 1).
In exploratory internal analyses, the expanded model showed higher apparent discrimination than the simpler model (AUC 0.902 vs. 0.968). However, because model development, bootstrap estimation, and decision-curve analysis were all performed within the same modest dataset, with some small subgroup strata, these estimates are likely optimism-prone and should be regarded only as exploratory internal performance summaries rather than as evidence of transportable or clinically ready prediction models (
Figure 2).