Next Article in Journal
RETRACTED: Lucas et al. Health Impact and Psychosocial Perceptions Among French Medical Residents During the SARS-CoV-2 Outbreak: A Cross-Sectional Survey. Int. J. Environ. Res. Public Health 2021, 18, 8413
Previous Article in Journal
Rational Risk or Cultural Resistance: Co-Produced Findings on Structural Barriers to Drug and Alcohol Recovery Among South Asian Muslim Communities in England
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Using Generative AI to Improve Readability, Actionability, and Health Literacy of COPD Caregiver Materials

Department of Health Science, University of Alabama, Tuscaloosa, AL 35487, USA
*
Author to whom correspondence should be addressed.
Int. J. Environ. Res. Public Health 2026, 23(9), 1167; https://doi.org/10.3390/ijerph23091167
Submission received: 13 July 2026 / Revised: 22 August 2026 / Accepted: 31 August 2026 / Published: 7 September 2026

Highlights

Public health relevance—How does this work relate to a public health issue?
  • Many COPD caregivers rely on educational materials that may be difficult to understand because of high reading levels and complex medical language.
  • Improving the accessibility of caregiver education is important for reaching people with varying levels of health literacy and supporting their role in COPD care.
Public health significance—Why is this work of significance to public health?
  • A three-phase workflow combining generative AI with human-guided health literacy review improved the readability and clarity of COPD caregiver materials while preserving key content and actionability.
  • The approach addresses both the linguistic complexity of health information and the potential for experts to overestimate what is clear and understandable to intended audiences.
Public health implications—What are the key implications or messages for practitioners, policy makers and/or researchers in public health?
  • Public health researchers and practitioners can use this workflow to adapt educational materials more efficiently while retaining human review for accuracy and usability.
  • Similar approaches could be applied to patient education materials across other populations and settings where clear, accessible information is needed.

Abstract

Informal caregiving for individuals with chronic obstructive pulmonary disease (COPD) places substantial emotional and physical demands on caregivers. However, many educational resources remain difficult to use due to high reading levels, dense medical language, and limited accessibility. Given that a large proportion of U.S. adults experience challenges with health literacy, existing caregiver toolkits may not provide the clear, actionable guidance needed to support caregiver self-efficacy or improve patient outcomes. This study introduces a reproducible, three-phase workflow that uses generative artificial intelligence (AI) to simplify and refine caregiver education materials, with application to the COPD Caregiver’s Toolkit. Four AI tools were used to simplify text, followed by human-guided refinement using the Sydney Health Literacy Lab (SHeLL) health literacy editor to ensure conceptual accuracy and usability. Readability indices and blinded assessments were used to evaluate changes in understandability and actionability across versions. AI-assisted simplification improved readability and clarity while maintaining high levels of actionability and preserving key instructional content. This scalable, adaptable approach may help public health researchers and practitioners improve the accessibility of educational resources across settings. Multi-phase, human-in-the-loop AI workflows may reduce linguistic complexity and improve the dissemination of essential health information to caregiving populations.

1. Introduction

Chronic Obstructive Pulmonary Disease (COPD) is a progressive respiratory illness that severely impacts both those diagnosed and the family members or friends who provide daily care [1]. Informal caregivers, who typically have little or no professional training in disease management, often face challenges navigating complex healthcare systems, managing medications, and responding to sudden exacerbations. These demands create persistent uncertainty and emotional, physical, and financial strain [2,3]. As the disease progresses, caregiving demands intensify, further increasing the burden experienced [4]. Limited access to structured guidance or formal support networks compounds this strain, leaving many caregivers feeling isolated and unprepared [5].
The strain of COPD caregiving extends across social, psychological, and economic domains [6]. Caregivers commonly report elevated levels of anxiety and depression, and prolonged social isolation can deepen emotional distress [7]. Sleep disruption and continuous physical effort often lead to fatigue and chronic stress [5]. Financial hardship is also prevalent, as caregiving responsibilities can interfere with employment and reduce household income [2]. Over time, the combined pressures contribute to declining physical and mental health among caregivers, increasing the risk of burnout and reducing their capacity to sustain effective care [8].
Low self-efficacy in caregiving abilities leads to higher hospitalization rates among COPD patients, underscoring the critical need for caregiver education and support [9]. Without adequate resources, caregivers may develop ineffective coping strategies that negatively affect their well-being and patients’ health outcomes [9]. Family members often engage in overprotective behaviors that, while well-intentioned, restrict patients’ daily autonomy and contribute to a decline in their physical activity [4]. Developing targeted, accessible, and evidence-based resources to enhance caregivers’ knowledge, confidence, and well-being is necessary to improve health outcomes for both caregivers and patients.
Existing interventions designed to support caregivers, such as self-management counseling and the Carer Support Needs Assessment Tool (CSNAT) [10], demonstrated some effectiveness in addressing caregiver burden. However, these resources may not be accessible or well-adapted to caregivers’ evolving needs due to complex language use and inadequate responsiveness to disease progression or caregiving constraints [2,11]. Caregivers report that self-management counseling provides only partial support, leaving them feeling unprepared to handle the complexities of COPD care [3]. Additionally, while tools like the CSNAT help identify gaps in caregiver support, their implementation is often hindered by healthcare providers’ time constraints and the lack of standardized caregiver assessments [7]. These limitations highlight the need for more comprehensive and accessible educational resources that support caregivers throughout the disease trajectory.
The COPD Caregiver’s Toolkit [12], developed by the National Heart, Lung, and Blood Institute (NHLBI) and the Respiratory Health Association (RHA), is a valuable resource for caregivers. However, a systematic evaluation of the Toolkit’s readability and usability from a health-literacy perspective is currently lacking. Given that nearly 90% of U.S. adults struggle with health literacy [13], ensuring that educational resources are accessible and easy to understand is essential for adequate caregiver support. By leveraging AI-driven resource analysis, this research aims to improve the accessibility, usability, and overall effectiveness of this widely disseminated caregiver education tool. Using the Toolkit, we sought to (1) develop a reproducible process that combines generative AI-assisted text simplification with human-guided health literacy refinement and (2) evaluate changes in readability, understandability, and actionability across versions of the materials. The study was designed as an initial evaluation of the proposed approach and resulting materials.

2. Methods

We conducted a methodological development and comparative evaluation study using a three-phase process to adapt and evaluate existing COPD caregiver education materials. The methodological component involved developing a standardized process in which content from the COPD Caregiver’s Toolkit was simplified using four generative AI tools and then refined through human review using the Sydney Health Literacy Lab (SHeLL) health literacy editor. The evaluation component compared versions of the materials using established readability measures and blinded assessments of understandability and actionability. The study focused on changes in the characteristics of the materials.

2.1. Phase I: Evaluation of Generative AI for Enhancing Toolkit Readability

Design. An experimental text-comparison design evaluated four generative AI models—ChatGPT® Plus (OpenAI), Gemini® Advanced (Google), Copilot® Pro (Microsoft), and Perplexity® Pro (Perplexity AI)—to inform model selection within a broader AI-assisted material development workflow. The goal was to determine how effectively AI revisions improved text accessibility to a fifth-grade reading level and to compare each model’s processing efficiency in February 2025. These AI tools were selected because they were among the most widely recognized and publicly accessible generative AI tools available to health consumers at the time of the study.
Procedure. The Toolkit text was entered into each model using the standardized prompt, “Simplify the following text to a 5th-grade reading level.” Each model was prompted four times, consistent with similar studies using generative AI [14,15], generating independent outputs for analysis. To ensure consistency, each AI tool received the same standardized prompt and source text using identical procedures. Models were accessed through their standard user interfaces; system prompts, temperature, and other unavailable model parameters were not controlled. Prompts, source materials, procedures, and generated outputs were standardized and documented to support methodological reproducibility, although exact replication of probabilistic AI-generated outputs could not be guaranteed. Readability scores were computed for the original and AI-modified versions, and the percentage change for each metric was calculated. Processing time (in seconds) was also recorded to assess efficiency. Following comparison of readability scores across the four AI-generated versions, the version demonstrating the strongest overall readability performance for each module or section was retained for subsequent refinement.
Measures. Readability was evaluated using four validated indices: Flesch Reading Ease (FRE), Flesch–Kincaid Grade Level (FKGL), Gunning Fog Index (GFI), and Simple Measure of Gobbledygook Index (SMOGI), consistent with AMA and NIH recommendations [16,17]. FRE scores range from 0 to 100, with higher scores indicating greater readability [18], while FKGL, GFI, and SMOGI estimate the years of education required for comprehension [19,20]. Scores were calculated using the WordCalc Readability Formulas tool (https://www.wordcalc.com/readability/, accessed on 13 February 2025). The primary outcome was the percent change in readability scores between the original and AI-enhanced versions. Secondary outcomes included model processing time. Readability improvements were evaluated against health-literacy benchmarks: FKGL ≤ 6 and FRE ≥ 70 for general comprehension and ≥80 for optimal accessibility [21]; GFI ≤ 7 [22]; and SMOGI between 6 and 8 [23].
Data Analysis. Descriptive statistics summarized readability scores and processing times. Percent change (%Δ) calculations quantified each AI model’s effectiveness across six Toolkit components (Modules 1–5 and Checklists and Forms). Descriptive analyses were used because the study compared versions of the same source materials rather than observations sampled from a population. Pre- and post-AI readability scores were compared to assess overall improvement and potential to enhance patient education materials (PEMs).

2.2. Phase II: Refining AI Enhancements

Design. Phase II operationalized human-in-the-loop refinement by applying the SHeLL Health Literacy Editor [24] to AI-generated outputs to ensure alignment with established health communication standards. Different AI models could be selected for different modules, depending on which version demonstrated the best alignment with health literacy.
Procedure. The selected AI-based simplified text for each module/section was compared with the original Toolkit to identify changes or omissions that could alter the meaning, intent, or essential content of the source material. Information judged necessary to preserve the original meaning and instructional purpose was manually restored or revised before further evaluation. This process was used to maintain semantic fidelity to the source material; however, no formal objective assessment of semantic equivalence was conducted. The full AI-modified Toolkit was then entered into the SHeLL Editor, which evaluates linguistic complexity, structural clarity, and person-centered communication. Metrics include total characters, words, unique words, sentences, and paragraphs, as well as proportions of long words and sentences, complex lists, passive constructions, and potentially stigmatizing language. The tool also measures lexical density, type-token ratio, and standardized type-token ratio, indicating vocabulary variety and information richness [25].
Data Analysis. SHeLL outputs were recorded and summarized. Content accuracy was evaluated by comparing the AI-modified Toolkit to the original to confirm conceptual clarity and preservation of key information. Refinements were made based on SHeLL feedback using Microsoft Word’s Track Changes feature, and a summary report described the modifications.

2.3. Phase III: Health-Literacy Demands Assessment

Design and Procedure. Phase III evaluated whether the AI-assisted development process resulted in materials that met established thresholds for understandability and actionability. A blinded comparative assessment evaluated the understandability and actionability of the final AI-enhanced Toolkit using a validated PEM evaluation tool. Two graduate research assistants, trained in public health, independently assessed the original and AI-modified versions, which were formatted identically to reduce bias.
Measures. A modified version of the Patient Education Materials Assessment Tool for Printed Materials (PEMAT-P) [26] was used to assess understandability and actionability. Items irrelevant to the text-only format were marked “Not Applicable.” The Understandability domain evaluated clarity, organization, and word choice, while the Actionability domain assessed the clarity of recommended actions and supporting guidance.
Data Analysis. PEMAT-P item-level scores were analyzed to examine inter-rater reliability and descriptive differences. Cohen’s Kappa (κ) assessed inter-rater reliability for binary items [27], interpreted using established benchmarks [28]. Domain-level percentages for Understandability and Actionability were calculated for each version to summarize performance.

2.4. Ethical Considerations

A committee on the protection of human participants reviewed the study protocol. The research was deemed exempt from Institutional Review Board (IRB) oversight, as it involved only secondary analysis of AI-generated text and no human participants or identifiable data. The IRB determined that the proposed activity is not research involving human subjects as defined by federal regulations.

3. Results

3.1. Phase I: Evaluation of Generative AI for Enhancing Toolkit Readability

The original version of the COPD Caregiver Toolkit demonstrated limited readability across its six components (Modules 1–5 and Checklists and Forms) as measured by standard readability metrics (Table 1). FRE scores ranged from 57.4 to 65.7, suggesting the content may be challenging for individuals with lower health literacy. FKGL scores ranged from 6.3 to 9.2, with most components exceeding the recommended 5th- to 6th-grade reading level for PEMs. Similarly, GFI scores ranged from 9.3 to 11.7, and SMOGI scores ranged from 9.6 to 12.0, further indicating the complexity of the original material.
Figure 1 and Figure 2 present the mean ± SD readability outcomes for the COPD Caregiver’s Guide and its Generative AI–revised versions, with Figure 1 comparing FKGL, GFI, and SMOG indices across guide components and Figure 2 displaying corresponding Flesch Reading Ease (FRE) scores. The AI-generated version with the strongest overall readability performance for each module was selected for further review and refinement. Based on these comparisons, Perplexity-generated text was retained for Modules 1, 2, 4, and 5, with FRE scores ranging from 82.6 to 95.4, FKGLs from 2.3 to 4.5, and SMOGIs from 7.8 to 8.2. Gemini-generated text was selected for Module 3 (FRE = 87.9; FKGL = 3.2), while Copilot-generated text was selected for the Checklists and Forms (FRE = 84.6; FKGL = 4.1). Thus, the version carried forward for human review combined content generated by multiple AI models rather than relying on the output of a single model.
The heat map in Figure 3 shows how readability changes varied across modules and AI models. Greener cells indicate greater improvements in readability relative to the original Toolkit, whereas redder cells indicate declines. No single AI model produced the strongest results across all modules/sections. Therefore, the output showing the strongest overall pattern of positive readability change for each module/section, as indicated by the predominance and magnitude of green-tinted cells in the heat map, was selected for subsequent human review and refinement. This resulted in a combined version consisting of Perplexity-generated text for Modules 1, 2, 4, and 5, Gemini-generated text for Module 3, and Copilot-generated text for the Checklists and Forms.

3.2. Phase II: Refining AI-Enhancements

To assemble a complete AI-modified version of the COPD Caregivers Toolkit, the outputs from different AI models were selectively combined based on module-specific performance, demonstrating a modular approach to AI-assisted material development. Overall, the assembled version demonstrated improved readability across all sections, with most modules written at or below a fifth-grade reading level.
Specifically, the AI-modified Toolkit consisted of 16,670 characters, 2624 words (416 long words), 819 unique words, and 284 paragraphs. While only three long sentences and one sentence containing a potentially overwhelming list were flagged, the editor recommended converting long lists into dot points to improve readability, which is now reflected in the modified version. The text complexity score was 25.6%, suggesting approximately one-quarter of the document may present challenges for individuals with limited literacy due to advanced vocabulary or complex sentence structure; as a result, 39 long words were shortened, five sentence structures were adjusted, and 416 words were altered with thesaurus suggestions from the SHeLL editor.
The SHeLL Editor flagged 174 terms for which simpler alternatives were available in the built-in thesaurus, many of which were related to medical terminology. For example, the acronym “COPD” was defined early in the document and reiterated throughout, indicating a limitation of the Editor’s contextual sensitivity. Nonetheless, simpler terms were used when they did not compromise content accuracy (e.g., medication adherence was replaced with “sticking to taking medications”). The Editor flagged numerous uncommon terms and acronyms, prompting selective simplification or definition where clarity could be improved without compromising accuracy. Several acronyms, such as HVAC (Heating, Ventilation, and Air Conditioning), GERD (Gastroesophageal Reflux Disease), and POLST (Physician Orders for Life-Sustaining Treatment), were defined upon first mention in accordance with SHeLL recommendations. Passive voice was used sparingly, with only four instances identified, all of which were revised to the active voice.
No paragraphs were found to exceed eight sentences or 150 words, indicating well-controlled paragraph lengths. Importantly, the Editor identified no non-person-centered or stigmatizing language, suggesting that the tone of the AI-modified version remained respectful and aligned with recovery-oriented communication principles. Lexical density was measured at 3.0, indicating a relatively information-rich text. The type-token ratio was 0.31, and the standardized type-token ratio was 0.61, reflecting a moderate level of vocabulary diversity. These values suggested that while some repetition supported clarity, there was still enough variety to maintain engagement. No changes were required for paragraph structure or person-centered language based on these findings.
The Editor’s feedback guided module-specific revisions. In Module 1, several terms (e.g., emphysema, bronchitis, and expulsion) were simplified to reduce complexity. Definitions or rewording were also added for GERD (gastroesophageal reflux), pulmonary (lung), and pulmonary rehabilitation (lung wellness exercises). Three sentences were shortened, and 65 flagged words were replaced with simpler alternatives. Module 2 featured 195 uncommon words and thesaurus-flagged alternatives; terms like adherence (sticking to medications) and HVAC (air conditioning units) were clarified, and repetitive acronym flags for COPD were acknowledged but not revised further, as the term was already defined. Module 3 contained one flagged long sentence and one instance of passive voice, both of which were revised. Although COVID-19 was flagged, it was deemed universally familiar and not modified further.
In Module 4, thesaurus-based recommendations were incorporated to reduce linguistic complexity. While Power of Attorney was already defined, the acronym POLST was not defined in this module and was subsequently clarified. Module 5 contained three instances of passive voice and two unnecessary intensifiers, such as “super,” both of which were removed. Although respite care and POLST were defined earlier in the toolkit, they were reintroduced in Module 4 for continuity. The “How to Stay Healthy” section was revised to improve flow, and content was added to provide actionable strategies for stress reduction, including suggestions such as getting enough sleep, exercising, and making time for oneself. A more robust discussion of support networks was also added, including faith communities, social work departments, and adult day care centers.
Across all modules, the Editor highlighted complex and uncommon terms, many of which had already been defined. However, the AI-modified version occasionally omitted essential content from the original Toolkit. In Module 1, the statement that COPD has no cure but can be managed was missing and was reintroduced to reassure caregivers. The section “Is it a symptom of COPD—or is it something else?” was also missing and was restored. Module 4 initially failed to emphasize the value of early symptom recognition; language was added to reinstate this message, along with additional details about how medications and antibiotics can help shorten flare-up duration. In Module 5, new content was added to elaborate on methods for stress reduction and signs of depression, and the caregiver support section was expanded.

3.3. Phase III: Health-Literacy Demands Assessment

Understandability. The AI-enhanced version of the Toolkit (Version A) received higher overall understandability scores than the original version (Version B). Both raters assigned Version A an understandability score of 85%, compared to lower scores for the original version (46% and 69%). Inter-rater reliability for PEMAT-P understandability items was moderate for Version A and substantial for Version B. Specifically, Cohen’s kappa κ for understandability items was κ = 0.625 for Version A and κ = 0.67 for Version B, indicating moderate agreement between raters [28]. These levels of agreement support consistency in ratings across domains for content clarity, word choice, numerical clarity, and organizational structure.
Descriptive differences in item-level ratings revealed several areas where Version A outperformed Version B across key understandability domains. In the content domain, Version A was rated more favorably on Item 1, whether the material makes its purpose completely evident, receiving perfect scores from both raters. In the domain of word choice and style, Version A was rated higher on Items 3 and 4, indicating greater use of everyday language and a more consistent definition of medical terms. Version A also received a score of 5 or higher on Item 5, which evaluates the use of the active voice.
In the use of numbers domain, both versions scored similarly on Items 6 and 7, with both raters noting that numbers were clear and the materials did not require user calculations. For organization, Version A received higher scores on Item 8, reflecting more effective chunking of information. Item 9, which evaluates the use of informative section headers, was scored equally across both versions. In the layout and design domain, Version A outperformed Version B on Item 12, which assesses the use of visual cues to emphasize key information. Overall, Version A demonstrated stronger performance across most understandability domains, particularly in content clarity, word usage, and organizational features.
Actionability. Both Toolkit versions performed similarly in terms of actionability. Student 1 assigned Version A a perfect score of 100%, while Student 2 rated it at 80%. Version B received identical actionability scores of 80% from both raters. Inter-rater reliability for actionability items showed strong agreement (κ = 0.99) for both the original and AI-modified versions of the toolkit. Item-level differences indicated few distinctions between Toolkit versions. Both versions scored equally on Items 20 through 25, suggesting they clearly identified recommended actions with manageable steps and, when appropriate, tangible tools such as checklists or planners.

4. Discussion

AI-assisted modifications improved the readability of the COPD caregiver materials while maintaining high levels of understandability and actionability. Improvements varied across models and modules/sections, with Perplexity producing the strongest readability gains for most modules, Gemini performing best for Module 3, and Copilot for the Checklists and Forms. ChatGPT also improved readability but required longer processing times. These findings indicate that no single AI model consistently performed best across the Toolkit and support selecting AI-generated content based on the characteristics of individual modules/sections rather than relying on a single model.
The SHeLL-guided review provided an important additional step beyond AI-based simplification. Although the selected AI-generated versions improved readability, subsequent review identified opportunities to further simplify language and improve clarity while also restoring essential information that had been omitted or altered during AI processing. These findings reinforce the value of combining generative AI with structured human review rather than relying on AI-generated revisions alone.
From a health communication and caregiver education standpoint, these findings support a potential theoretical link between AI-driven simplification and caregiver cognition. Readability may influence how easily information can be processed, particularly because caregivers often process information under stress, time scarcity, and emotional burden, which can constrain working memory. Simplifying sentence structure and terminology may reduce the linguistic demands associated with interpreting caregiver education materials and allow readers to focus more readily on care-related information.
This potential cognitive pathway helps explain why improvements in readability could ultimately support caregiving behavior, particularly given that more than half of U.S. adults read at or below a sixth-grade level [29]. When instructional materials are misaligned with the reader’s literacy, caregivers may disengage, misinterpret guidance, or rely on partial understanding. Clearer language and explicit definitions have the potential to support comprehension, confidence, and perceived self-efficacy, which are established predictors of behavioral follow-through [21]. However, comprehension, self-efficacy, behavioral change, and clinical outcomes were not assessed in the present study. Accordingly, the observed improvements in readability and understandability should be interpreted as improvements in characteristics of the educational materials rather than evidence of improved caregiver or patient outcomes.
Notably, the AI-modified version performed as well as, or better than, the original on measures of understandability and actionability, including the use of step-by-step instructions, checklists, and planners. These structural features may support procedural learning and reduce ambiguity, which is especially important for caregivers managing complex or unfamiliar tasks [30]. The AI-modified text showed the most apparent advantage in elements relevant to comprehension and information processing, such as the active voice, more precise explanations of medical terminology, and overall clarity. These features may help minimize interpretive errors and facilitate accurate execution of care instructions [31].
At the same time, these findings clarify the conditions under which AI-driven simplification is likely to be beneficial and where it may be insufficient. AI tools appear well-suited to addressing linguistic barriers and implicit author bias that can arise when materials are written by experts for non-expert audiences [32,33]. In this study, AI-assisted revisions partially mitigated those gaps without requiring a complete redesign of content. However, simplification alone cannot compensate for the absence of contextual support, individualized clinical guidance, or the emotional and social dimensions of caregiving. Caregivers may still struggle when recommendations conflict with lived constraints, cultural expectations, or access limitations. In such cases, AI-enhanced materials should be viewed as complementary supports rather than standalone solutions.
This research contributes to health communication and behavioral science by offering an early, structured demonstration of how AI-driven simplification can reshape caregiver-facing materials in ways that may reduce linguistic barriers to information processing. The findings are consistent with the theoretical proposition that readability may influence caregivers’ understanding and perceived capability, which could, in turn, shape behavior; however, testing this pathway requires prospective evaluation with caregivers. Future implementation studies should extend this study by examining real-world caregiver use, comprehension under realistic conditions, and the extent to which AI-assisted simplification must be paired with clinical reinforcement or human support to produce sustained behavioral change.

Limitations and Strengths

Several limitations should be considered when interpreting these findings. First, this study did not include direct user testing with caregivers. Traditional readability metrics rely largely on word and sentence characteristics and therefore provide only indirect estimates of how well readers understand the content. Without usability testing or assessments of comprehension, confidence, decision-making, or emotional response, it remains unclear whether the AI-modified materials improve real-world caregiving behaviors or outcomes. Although the observed gains in readability and understandability are promising, they should be viewed as preliminary indicators rather than evidence of clinical or behavioral impact.
Second, preservation of essential content was assessed by comparing the revised materials with the original Toolkit rather than through a formal measure of semantic equivalence or independent clinical expert review. Although omitted or altered content was restored when identified, future studies should incorporate objective measures of semantic equivalence and clinical expert review to verify that AI-based simplification preserves the accuracy and meaning of the source material.
Third, although the PEMAT-P assessments demonstrated acceptable inter-rater agreement, the assessment was conducted by two trained graduate research assistants and did not include caregivers or clinical experts. The PEMAT-P findings therefore provide a standardized assessment of material characteristics but should not be interpreted as a substitute for end-user usability testing or clinical review.
Fourth, generative AI is rapidly evolving, and the performance of commercial models may change as developers update model architectures, training, interfaces, and system behavior. The models in this study were evaluated at a specific point in time; consequently, the relative performance observed here may not persist across later model versions. This limitation supports interpreting the primary contribution of the study as the reproducible workflow rather than evidence favoring a particular commercial AI model.
Finally, comparisons of processing time across AI models should be interpreted cautiously. External factors, such as internet connectivity and system performance, may have influenced the observed differences. Still, processing time offers a practical indicator of feasibility for real-world implementation.
Despite these limitations, the study has several strengths. The use of a reproducible, three-phase workflow enabled systematic comparison across four generative AI models, providing insight into how different tools simplify health materials while preserving key content. The inclusion of human-guided refinement using the Sydney Health Literacy Lab editor helped ensure alignment with established health communication standards. The evaluation approach also strengthens the findings. In addition to readability indices, the study incorporated blinded assessments of understandability and actionability using two trained raters and the PEMAT-P tool, demonstrating acceptable inter-rater reliability. Our approach offers a practical framework for improving the accessibility of health education materials.

5. Conclusions

We applied a structured workflow to a COPD caregiver Toolkit and compared outputs across multiple generative AI models. In doing so, we brought together two domains often treated separately: health literacy performance beyond grade level (i.e., understandability and actionability) and dissemination feasibility. Findings suggest that AI-assisted simplification should be conceptualized less as a technical exercise and more as an implementation strategy with the potential to facilitate information use under stress, time constraints, and uneven literacy. From a policy standpoint, the results indicate the need for clearer standards governing the development and review of patient education materials. Current guidance often emphasizes readability thresholds but does not consistently address whether materials are understandable and actionable. Incorporating these domains should strengthen the usability of publicly disseminated caregiver education resources.
In addition, AI-assisted workflows raise questions about transparency, quality control, and oversight. Policies that encourage integration of human review with AI-generated content may help ensure that materials remain accurate and aligned with established health literacy principles. In practice, generative AI offers a feasible approach for improving the clarity and accessibility of caregiver education materials. Whether these improvements subsequently help caregivers manage responsibilities or navigate high-demand situations requires evaluation with end users.
Although this workflow was evaluated using COPD caregiver materials, the underlying process is not inherently disease-specific. The same staged approach—AI-assisted simplification, human-guided health literacy refinement, and standardized assessment—could potentially be adapted to educational materials addressing other chronic diseases, caregiving populations, and health contexts. Application to languages other than English may also be feasible but would require additional validation. Readability measures, linguistic conventions, cultural context, and AI model performance may differ across languages; therefore, adaptation should incorporate language-appropriate assessment methods and review by individuals with relevant linguistic and cultural expertise. Future studies should evaluate the workflow across disease areas, populations, and languages before broader generalizability is assumed.
Future prospective studies involving caregivers should also determine whether improvements in readability, understandability, and actionability translate into better comprehension and accurate interpretation of health information, more informed decision-making, greater caregiver self-efficacy, changes in caregiving behaviors, and, ultimately, improved patient outcomes. Such studies would help establish whether improvements in the characteristics of educational materials produce meaningful benefits for the people who use them.

Author Contributions

Conceptualization, M.L.S.; methodology, M.L.S.; software, M.L.S., A.B., O.H. and R.S.; validation, M.L.S., G.A., S.F.; formal analysis, M.L.S., G.A., S.F., A.B. and R.S.; investigation, M.L.S., G.A., S.F., A.B. and R.S.; resources, M.L.S. and S.W.; data curation, M.L.S., A.B., O.H. and R.S.; writing—original draft preparation, M.L.S., G.A., A.B. and S.F.; writing—review and editing, M.L.S., G.A., R.S. and S.W.; visualization, M.L.S. and R.S.; supervision, M.L.S.; project administration, M.L.S.; funding acquisition, M.L.S. All authors have read and agreed to the published version of the manuscript.

Funding

This research was supported by The University of Alabama through the Championing Resources for Excellence in Arts/Humanities Transformation Efforts (CREATE) Program and the Division for Research (Award No. A25-0097-001).

Institutional Review Board Statement

Ethical review and approval were waived for this study because it involved only secondary analysis of AI-generated text and did not involve human participants or identifiable data. The Institutional Review Board at The University of Alabama determined that the activity did not constitute human subjects research as defined by federal regulations (IRB ID: 25-01-8333).

Informed Consent Statement

Informed consent was waived because the study did not involve human participants or the collection of identifiable private information and was limited to the analysis and AI-assisted revision of publicly available educational materials.

Data Availability Statement

Data supporting the findings of this study may be available from the corresponding author upon reasonable request, subject to applicable copyright and other restrictions.

Acknowledgments

During the preparation of this manuscript, the authors used ChatGPT Plus (OpenAI) and Grammarly Pro to assist with language refinement, clarity, grammar, and organization. ChatGPT® Plus (OpenAI), Gemini® Advanced (Google), Copilot® Pro (Microsoft), and Perplexity® Pro (Perplexity AI) were also used as part of the study methodology to simplify caregiver education materials, as described in the Methods. All AI-assisted content was reviewed, verified, and edited by the authors, who take full responsibility for the content of this publication.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Suresh, M.; Young, J.; Fan, V.; Simons, C.; Battaglia, C.; Simpson, T.L.; Fortney, J.C.; Locke, E.R.; Trivedi, R. Caregiver experiences and roles in care seeking during COPD exacerbations: A qualitative study. Ann. Behav. Med. 2021, 56, 257–269. [Google Scholar] [CrossRef] [Scilit]
  2. Cruz, J.; Marques, A.; Machado, A.; O’Hoski, S.; Goldstein, R.; Brooks, D. Informal caregiving in COPD: A systematic review of instruments and their measurement properties. Respir. Med. 2017, 128, 13–27. [Google Scholar] [CrossRef] [Scilit]
  3. Marques, A.; Cruz, J.; Brooks, D. Interventions to support informal caregivers of people with chronic obstructive pulmonary disease: A systematic literature review. Respiration 2021, 100, 1230–1242. [Google Scholar] [CrossRef] [Scilit]
  4. Siltanen, H.; Jylhä, V. Supporting the supporter: A focus on families of patients with chronic obstructive pulmonary disease. JBI Database Syst. Rev. Implement. Rep. 2019, 17, 2212–2213. [Google Scholar] [CrossRef] [Scilit]
  5. Noonan, M.C.; Wingham, J.; Taylor, R.S. “Who cares?” The experiences of caregivers of adults living with heart failure, chronic obstructive pulmonary disease and coronary artery disease: A mixed methods systematic review. BMJ Open 2018, 8, e020927. [Google Scholar] [CrossRef] [Scilit]
  6. Yi, M.; Jiang, D.; Jia, Y.; Xu, W.; Wang, H.; Li, Y.; Zhang, Z.; Wang, J.; Chen, O. Impact of caregiving burden on quality of life of caregivers of COPD patients: The chain mediating role of social support and negative coping styles. Int. J. Chronic Obstr. Pulm. Dis. 2021, 16, 2245–2255. [Google Scholar] [CrossRef] [Scilit]
  7. Lippiett, K.A.; Richardson, A.; Myall, M.; Cummings, A.; May, C.R. Patients and informal caregivers’ experiences of burden of treatment in lung cancer and chronic obstructive pulmonary disease (COPD): A systematic review and synthesis of qualitative research. BMJ Open 2019, 9, e020515. [Google Scholar] [CrossRef] [Scilit]
  8. Zhang, N.; Tian, Z.; Liu, X.; Yu, X.; Wang, L. Burden, coping and resilience among caregivers for patients with chronic obstructive pulmonary disease: An integrative review. J. Clin. Nurs. 2023, 33, 1346–1361. [Google Scholar] [CrossRef] [Scilit]
  9. Reinhard, S.C.; Given, B.; Petlick, N.H.; Bemis, A. Supporting family caregivers in providing care. In Patient Safety and Quality: An Evidence-Based Handbook for Nurses; Hughes, R.G., Ed.; Agency for Healthcare Research and Quality: Rockville, MD, USA, 2008. [Google Scholar]
  10. Ewing, G.; Grande, G. Development of a carer support needs assessment tool (CSNAT) for end-of-life care practice at home: A qualitative study. Palliat. Med. 2012, 27, 244–256. [Google Scholar] [CrossRef] [Scilit]
  11. Bryant, J.; Mansfield, E.; Boyes, A.; Waller, A.; Sanson-Fisher, R.; Regan, T. Involvement of informal caregivers in supporting patients with COPD: A review of intervention studies. Int. J. Chronic Obstr. Pulm. Dis. 2016, 11, 1587–1596. [Google Scholar] [CrossRef] [Scilit]
  12. U.S. Department of Health and Human Services. The COPD Caregiver’s Toolkit; National Heart, Lung, and Blood Institute: Bethesda, MD, USA, 2024. Available online: https://www.nhlbi.nih.gov/education/copd-learn-more-breathe-better/copd-caregivers-toolkit (accessed on 7 February 2025).
  13. Kutner, M.; Greenberg, E.; Jin, Y.; Paulsen, C. The Health Literacy of America’s Adults: Results from the 2003 National Assessment of Adult Literacy (NCES 2006–483); U.S. Department of Education: Washington, DC, USA, 2006. [Google Scholar]
  14. Kirchner, G.J.; Kim, R.Y.; Weddle, J.B.; Bible, J.E. Can artificial intelligence improve the readability of patient education materials? Clin. Orthop. Relat. Res. 2023, 481, 2260–2267. [Google Scholar] [CrossRef] [Scilit]
  15. Rouhi, A.D.; Ghanem, Y.K.; Yolchieva, L.; Saleh, Z.; Joshi, H.; Moccia, M.C.; Suarez-Pierre, A.; Han, J.J. Can artificial intelligence improve the readability of patient education materials on aortic stenosis? A pilot study. Cardiol. Ther. 2024, 13, 137–147. [Google Scholar] [CrossRef] [Scilit]
  16. Daraz, L.; Morrow, A.S.; Ponce, O.J.; Farah, W.; Katabi, A.; Majzoub, A.; Seisa, M.O.; Benkhadra, R.; Alsawas, M.; Larry, P.; et al. Readability of online health information: A meta-narrative systematic review. Am. J. Med. Qual. 2018, 33, 487–492. [Google Scholar] [CrossRef] [Scilit]
  17. Leonard Grabeel, K.G.; Russomanno, J.; Oelschlegel, S.; Tester, E.; Heidel, R.E. Computerized versus hand-scored health literacy tools: A comparison of SMOG and Flesch-Kincaid in printed patient education materials. J. Med. Libr. Assoc. 2018, 106, 38–45. [Google Scholar] [CrossRef] [Scilit]
  18. Flesch, R. A new readability yardstick. J. Appl. Psychol. 1948, 32, 221–233. [Google Scholar] [CrossRef] [Scilit]
  19. Gunning, R. The Technique of Clear Writing; McGraw-Hill: New York, NY, USA, 1968. [Google Scholar]
  20. McLaughlin, G.H. SMOG grading: A new readability formula. J. Read. 1969, 12, 639–646. [Google Scholar]
  21. Badarudeen, S.; Sabharwal, S. Assessing readability of patient education materials: Current role in orthopedics. Clin. Orthop. Relat. Res. 2010, 468, 2572–2580. [Google Scholar] [CrossRef] [Scilit]
  22. Friedman, D.B.; Hoffman-Goetz, L. A systematic review of readability and comprehension instruments used for print and web-based cancer information. Health Educ. Behav. 2006, 33, 352–373. [Google Scholar] [CrossRef] [Scilit]
  23. Doak, C.C.; Doak, L.G.; Root, J.H. Teaching Patients with Low Literacy Skills, 2nd ed.; J.B. Lippincott: Philadelphia, PA, USA, 1985. [Google Scholar]
  24. The SHeLL EDITOR. Available online: https://www.healthliteracysolutions.com.au/ (accessed on 14 February 2025).
  25. Ayre, J.; Bonner, C.; Muscat, D.M.; Dunn, A.G.; Harrison, E.; Dalmazzo, J.; Mouwad, D.; Aslani, P.; Shepherd, H.L.; McCaffery, K.J. Multiple automated health literacy assessments of written health information: Development of the SHeLL (Sydney Health Literacy Lab) Health Literacy Editor v1. JMIR Form. Res. 2023, 7, e40645. [Google Scholar] [CrossRef] [Scilit]
  26. Shoemaker, S.J.; Wolf, M.S.; Brach, C. Development of the Patient Education Materials Assessment Tool (PEMAT): A new measure of understandability and actionability for print and audiovisual patient information. Patient Educ. Couns. 2014, 96, 395–403. [Google Scholar] [CrossRef] [Scilit]
  27. McHugh, M.L. Interrater reliability: The kappa statistic. Biochem. Med. 2012, 22, 276–282. Available online: https://pubmed.ncbi.nlm.nih.gov/23092060/ (accessed on 1 February 2025). [CrossRef] [Scilit]
  28. Landis, J.R.; Koch, G.G. The measurement of observer agreement for categorical data. Biometrics 1977, 33, 159–174. [Google Scholar] [CrossRef] [Scilit]
  29. Mamedova, S.; Pawlowski, E. Adult literacy in the United States; National Center for Education Statistics: Washington, DC, USA, 2019. Available online: https://nces.ed.gov/pubs2019/2019179/index.asp (accessed on 1 February 2025).
  30. Mamun, T.I. Beyond algorithms: Human-centered explainability for clinicians, patients, and caregivers. Proc. Int. Symp. Hum. Factors Ergon. Health Care 2026, 15, 126–130. [Google Scholar] [CrossRef] [Scilit]
  31. Tucker, C.A. Promoting personal health literacy through readability, understandability, and actionability of online patient education materials. J. Am. Heart Assoc. 2024, 13, e033916. [Google Scholar] [CrossRef] [Scilit]
  32. Gopal, D.P.; Chetty, U.; O’Donnell, P.; Gajria, C.; Blackadder-Weinstein, J. Implicit bias in healthcare: Clinical practice, research and decision making. Future Healthc. J. 2021, 8, 40–48. [Google Scholar] [CrossRef] [Scilit]
  33. Rooney, M.K.; Santiago, G.; Perni, S.; Horowitz, D.P.; McCall, A.R.; Einstein, A.J.; Jagsi, R.; Golden, D.W. Readability of patient education materials from high-impact medical journals: A 20-year analysis. J. Patient Exp. 2021, 8, 2374373521998847. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Comparison of readability indices (mean ± SD) for the COPD Caregiver’s Guide and generative AI–revised versions across guide components.
Figure 1. Comparison of readability indices (mean ± SD) for the COPD Caregiver’s Guide and generative AI–revised versions across guide components.
Ijerph 23 01167 g001
Figure 2. Mean ± SD of Flesch Reading Ease (FRE) scores for original and AI-revised COPD Caregiver’s Guide components.
Figure 2. Mean ± SD of Flesch Reading Ease (FRE) scores for original and AI-revised COPD Caregiver’s Guide components.
Ijerph 23 01167 g002
Figure 3. Percent change in readability metrics across modules and AI models for COPD Caregiver’s Toolkit. Note. FRE = Flesch Reading Ease; FKGL = Flesch–Kincaid Grade Level; GFI = Gunning Fog Index; SMOG = Simple Measure of Gobbledygook.
Figure 3. Percent change in readability metrics across modules and AI models for COPD Caregiver’s Toolkit. Note. FRE = Flesch Reading Ease; FKGL = Flesch–Kincaid Grade Level; GFI = Gunning Fog Index; SMOG = Simple Measure of Gobbledygook.
Ijerph 23 01167 g003
Table 1. Readability scores of the original version of The COPD Caregiver’s Toolkit.
Table 1. Readability scores of the original version of The COPD Caregiver’s Toolkit.
Guide ComponentFREFKGLGFISMOG
Module 163.47.59.910.6
Module 26281111.3
Module 365.77.610.010.8
Module 462.87.610.411.0
Module 557.49.211.712.0
Checklist/Forms62.26.39.39.6
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Stellefson, M.L.; Avery, G.; Flora, S.; Boles, A.; Henry, O.; Sivalingam, R.; Williams, S. Using Generative AI to Improve Readability, Actionability, and Health Literacy of COPD Caregiver Materials. Int. J. Environ. Res. Public Health 2026, 23, 1167. https://doi.org/10.3390/ijerph23091167

AMA Style

Stellefson ML, Avery G, Flora S, Boles A, Henry O, Sivalingam R, Williams S. Using Generative AI to Improve Readability, Actionability, and Health Literacy of COPD Caregiver Materials. International Journal of Environmental Research and Public Health. 2026; 23(9):1167. https://doi.org/10.3390/ijerph23091167

Chicago/Turabian Style

Stellefson, Michael L., Gracie Avery, Sarah Flora, Ansley Boles, Olivia Henry, Rakshan Sivalingam, and Savanna Williams. 2026. "Using Generative AI to Improve Readability, Actionability, and Health Literacy of COPD Caregiver Materials" International Journal of Environmental Research and Public Health 23, no. 9: 1167. https://doi.org/10.3390/ijerph23091167

APA Style

Stellefson, M. L., Avery, G., Flora, S., Boles, A., Henry, O., Sivalingam, R., & Williams, S. (2026). Using Generative AI to Improve Readability, Actionability, and Health Literacy of COPD Caregiver Materials. International Journal of Environmental Research and Public Health, 23(9), 1167. https://doi.org/10.3390/ijerph23091167

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop