1. Introduction
The non-invasive diagnosis of idiopathic pulmonary fibrosis (IPF), and its distinction from the other 200-plus interstitial lung diseases (ILDs), is important for prognosis and treatment yet remains subjective and dependent on subspecialty expertise. Fibresolve (IMVARIA Inc., Berkeley, CA, USA) is the first machine-learning system authorized by the U.S. Food and Drug Administration (FDA) to aid this diagnosis, having received De Novo marketing authorization (DEN220040) as an adjunct in the diagnosis of IPF prior to invasive testing [
1]. Analyzing the complete chest computed tomography (CT) volume in three dimensions, it learns a broad set of imaging correlates of IPF and can indicate the diagnosis even in cases that lack a classic usual interstitial pneumonia (UIP) appearance [
2]. In the pivotal PUFAIR study [
3], Fibresolve version 1 (v1) met both of its pre-specified co-primary endpoints, contributing to the technology’s FDA authorization in 2024.
A central challenge for AI-enabled diagnostics is that their algorithms are expected to improve iteratively over time, yet significant modifications have traditionally required separate regulatory submissions, lengthening timelines for making such improvements available for clinical use. The FDA’s Predetermined Change Control Plan (PCCP) framework was finalized in December 2024 to address this challenge: the framework permits pre-specified modifications to an authorized device to be implemented and validated against the originally authorized version without a new marketing submission, provided each change remains within an agreed scope [
4]. The present study reviews the application of this framework to the Fibresolve technology, and assesses the resulting performance changes. We advance the software from version 1 to version 2 (v2), which augments the original CT-only classifier with a vision-transformer model incorporating age, sex, and forced vital capacity (FVC) in an ensemble, while retaining the v1 CT-only model as a fallback when these ancillary variables are unavailable. In alignment with the PCCP framework, version 2 is then validated head-to-head against the originally authorized version 1.
To briefly review the state of the disease: IPF is the most common idiopathic interstitial pneumonia, and although historically a progressive disease with a poor prognosis, it now follows a more favorable course with antifibrotic therapy [
5,
6,
7]. Effective management depends on distinguishing IPF from the other ILDs [
8,
9]. This distinction can be made non-invasively when a UIP pattern on CT is accompanied by compatible clinical findings, but in practice it is subjective and only moderately reproducible even among experts (weighted
0.48–0.65) [
10,
11]. When the imaging is inconclusive, a surgical lung biopsy may be required but carries appreciable morbidity and mortality [
12], and access to the subspecialty expertise that these judgments demand is often delayed [
13,
14]. These limitations motivated the development of objective, automated tools, and machine-learning analysis of chest CT in particular has been shown to identify fibrosis-related features beyond standard imaging criteria and to correlate with histopathology and long-term outcomes in fibrotic lung disease [
15,
16,
17,
18,
19]. These clinical challenges drove the development of Fibresolve v1 and contributed to methodologies for improvement for v2.
Two developments in machine learning motivated the v2 update. First, transformer architectures, originally introduced for sequence modeling [
20] and subsequently adapted to images as the Vision Transformer (ViT) [
21], have become a leading approach in medical image analysis, where global self-attention (i.e., the ability to gather more holistic image context) can capture the diffuse, spatially distributed patterns that characterize fibrotic lung disease [
22]. Second, multimodal models that combine imaging with structured clinical data may be anticipated to outperform imaging-only models for disease classification and prognosis [
23]. FVC in particular is a validated and clinically meaningful measure in IPF and is therefore a natural candidate input [
24,
25]. Reflecting this, subsequent work has demonstrated that a multimodal Fibresolve classifier also predicts mortality across ILDs [
26], complementing CT-based prognostic findings reported in independent registry cohorts [
27,
28]. The ongoing evolution of datasets, machine learning architecture, and addition of multimodal inputs all contribute to potential algorithm performance improvements.
To validate this update under the PCCP with maximal internal validity, we re-analyzed the identical 300-patient PUFAIR cohort with both versions and compared their performance using paired statistical methods. The purpose of this study is: (1) to describe the version 2 architecture and its fallback design; (2) to compare the diagnostic performance of version 2 with version 1 across the full dataset and the key thin-slice diagnostic subgroup, following the analysis structure of the original PUFAIR study; and (3) to formally assess the non-inferiority of version 2 relative to version 1.
3. Results
As both Fibresolve versions were evaluated on the same fixed cohort assembled for PUFAIR, its baseline demographic, clinical, and technical characteristics are identical to those reported previously [
3]. The 300 patients had a median age of 62 years and an approximately even sex distribution (49.7% female); the cohort was predominantly White (85.7%) and non-Hispanic (88.7%), and 69.0% were ever-smokers. Percent-predicted forced vital capacity (ppFVC) was recorded in 271 patients and ranged widely, from preserved (above 75% in 34.7%) to substantially reduced (below 50% in 19.7%). Imaging originated from scanners spanning four manufacturers and 23 distinct models, with slice thickness ranging from 1 to 5 mm, and surgical pathology was available in 94.7% of cases (284/300). Under the reference standard, IPF accounted for 27.7% (83/300) of final diagnoses; the most common non-IPF diagnoses were unclassifiable ILD (18.0%), chronic hypersensitivity pneumonitis (16.7%), and nonspecific interstitial pneumonia (10.7%), with the remainder distributed across numerous less common ILD subtypes.
The analysis of greatest clinical relevance is the thin-slice diagnostic CT subset (N = 137; 40 IPF, 97 non-IPF), which reflects the diagnostic-quality imaging on which Fibresolve is intended to be used. In this subgroup, Fibresolve v2 achieved a sensitivity of 57.5% (95% CI 42.2–71.5) and a specificity of 84.5% (76.0–90.4), compared with 55.0% (39.8–69.3) and 82.5% (73.7–88.8) for v1 (
Table 1,
Figure 2A). PPV rose from 56.4% to 60.5% and the diagnostic odds ratio from 5.75 (95% CI 2.55–12.98) to 7.40 (95% CI 3.21–17.03), with substantially overlapping confidence intervals indicating this difference should not be interpreted as statistically distinguishable. Agreement between versions was high (94.9%,
= 0.87; McNemar
p = 1.00). The clinical value is sharpest in the indeterminate subset of these cases, those not confirmed as IPF non-invasively and otherwise considered for biopsy (N = 124), where v2 delivered a non-invasive diagnostic yield of 56.2% (39.3–71.8) at 87.0% specificity (
Table 2,
Figure 3), exceeding the v1 yield of 53.1% and comparing favorably with the 36.1% (95% CI 33.4–38.9%) yield reported for transbronchial biopsy [
8].
Across the full dataset (N = 300), which additionally includes non-dedicated and thick-slice acquisitions and therefore represents a harder test, v2 retained its advantage: sensitivity was 44.6% (34.4–55.3) versus 41.0% (31.0–51.7) and specificity 88.0% (83.0–91.7) versus 86.6% (81.5–90.5), with PPV improving from 54.0% to 58.7%, NPV from 79.3% to 80.6%, and the diagnostic odds ratio from 4.50 to 5.91 (
Table 3,
Figure 2B). The two versions agreed on 96.7% of cases (
= 0.90) and the McNemar test was not significant (exact
p = 1.00). Confidence intervals for the likelihood ratios and diagnostic odds ratios overlapped substantially between versions in all three analyses, consistent with the non-significant paired comparisons.
In pre-specified non-inferiority testing against a margin, v2 was non-inferior to v1 for both endpoints. The paired difference in sensitivity (v2 − v1) was percentage points (95% CI to ) and the paired difference in specificity was points (95% CI to ); the non-inferiority null hypothesis was rejected for both (p < 0.001). Thus, v2 met the non-inferiority criterion while showing full-dataset point estimates that were all equal to or better than those of v1.
Because the ancillary clinical inputs required by v2 were not available for every patient, performance was additionally examined separately for the two pathways. FVC was available in 271 of 300 patients (90.3%), who were therefore evaluated by the v2 ensemble; the remaining 29 (9.7%) lacked FVC and were routed to the v1 CT-only fallback. Among the 271 v2-pathway cases (75 IPF, 196 non-IPF), v2 achieved a sensitivity of 42.7% (32/75; 95% CI 32.1–53.9) and a specificity of 87.8% (172/196; 95% CI 82.4–91.6). Among the 29 fallback cases (8 IPF, 21 non-IPF), the CT-only pathway achieved a sensitivity of 62.5% (5/8; 95% CI 30.6–86.3) and a specificity of 90.5% (19/21; 95% CI 71.1–97.3). In the thin-slice subset (N = 137), FVC was available in 121 patients (88.3%); the v2 pathway (36 IPF, 85 non-IPF) achieved 55.6% sensitivity (20/36; 95% CI 39.6–70.5) and 84.7% specificity (72/85; 95% CI 75.6–90.8), and the 16 fallback cases (4 IPF, 12 non-IPF) yielded 75.0% sensitivity (3/4; 95% CI 30.1–95.4) and 83.3% specificity (10/12; 95% CI 55.2–95.3). The fallback strata contained only eight and four IPF-positive cases, respectively; their estimates are correspondingly imprecise and are reported to characterize pathway usage rather than to support a formal pathway comparison. Within these limits, the fallback pathway showed no evidence of degraded performance relative to the v2 pathway.
Performance was examined across demographic, clinical, and technical subgroups (
Figure 4). Subgroup-level differences in both sensitivity and specificity were generally small relative to their confidence intervals and are not adjusted for multiple comparisons. These analyses are exploratory and intended only to assess for gross inconsistency across acquisition and demographic strata, not to support subgroup-specific claims. Performance was otherwise broadly consistent across CT manufacturers, clinical sites, and slice-thickness groups, supporting generalizability across acquisition settings. As anticipated, sensitivity was lowest in patients with preserved lung function (ppFVC > 75%), where it was 11.8% for both versions, indicating that the multimodal inputs did not improve detection of the milder or earlier disease that is most difficult to identify from imaging; both versions performed best in cases with reduced FVC and on thinner-slice acquisitions.
Since prior ILD literature commonly excludes certain diagnoses, performance was re-estimated under sequential exclusion of UILD, CTD-ILD, and CHP cases (
Table 4). Sensitivity was unchanged by these exclusions for both versions (44.6% v2, 41.0% v1 in the full dataset), as expected, since the IPF-positive cases are unaffected. Specificity was stable or improved, most notably when UILD cases were excluded; differences between the versions were small and in both directions across the exclusion scenarios, with version 2 modestly higher in most and marginally lower in the most stringent (e.g., 88.0% vs 89.0% when all three categories were excluded from the full dataset). Results in the thin-slice subgroup were consistent.
4. Discussion
This study evaluated Fibresolve v2, an algorithmic update to Fibresolve v1, the first FDA-authorized AI-enabled tool for supporting the diagnosis of IPF [
1]. The update was implemented under the device’s Predetermined Change Control Plan (PCCP), the regulatory framework that permits pre-specified modifications to an authorized device to be implemented and validated head-to-head against the originally authorized version, provided each change remains within an agreed scope and preserves safety and effectiveness [
4]. Rather than retraining the model freely, we made a single pre-specified modification, a multimodal vision-transformer branch incorporating age, sex, and FVC, and validated it against the deployed version on the locked PUFAIR cohort [
3].
In this paired re-analysis, Fibresolve v2 matched or exceeded v1 across the full dataset and both diagnostic subsets and was non-inferior to v1 for both sensitivity and specificity against the pre-specified margin, while preserving high agreement ( = 0.87–0.90; case-level concordance 94.9–96.7%). The update maintained the established behavior of the deployed model while improving classification in a small number of borderline cases. Although v2 consistently matched or exceeded v1 across the overall cohort and most subgroups, the magnitude of improvement was modest and the paired comparisons were not statistically significant.
The principal change in v2 was the addition of a vision-transformer branch that integrates age, sex, and FVC alongside CT imaging. On the full dataset, v2 identified additional true IPF cases while simultaneously reducing false-positive classifications, resulting in improvements across several diagnostic metrics. Although the effect is modest, these findings are consistent with prior related studies demonstrating the value of integrating imaging and structured clinical variables in diagnostic models [
23] and support the concept that complementary clinical information can improve adjudication of diagnostically challenging cases without substantially altering overall model behavior. Although age, sex, and FVC are expected to correlate to some degree with CT-derived disease severity, the observed improvement of v2 over v1 across sensitivity, specificity, and diagnostic odds ratio indicates that these variables contribute incremental information beyond what CT imaging alone captures, rather than simply duplicating it; isolating the marginal contribution of each individual variable would require constructing additional model configurations outside the two versions specified under the current PCCP scope.
Performance was stable across CT manufacturers, acquisition protocols, slice thicknesses, and clinical sites for both versions, and the v2 update preserved that stability. Limited generalizability across imaging environments has been a recurring barrier to clinical deployment of machine-learning systems, so this consistency supports broader applicability. The fallback design reinforces this robustness: when age, sex, or FVC are unavailable, the system reverts to the validated CT-only v1 model, ensuring uninterrupted workflow and preserving case coverage. From a clinical implementation perspective, this fallback approach provides a practical and safe degradation pathway while enabling incremental adoption of multimodal inputs.
The clinical utility of Fibresolve v1 established in PUFAIR [
3] was maintained, and modestly strengthened, in v2. The model has shown potential impact on indeterminate fibrotic-ILD cases that would otherwise undergo invasive sampling. In this “digital biopsy” subgroup, v2 achieved a non-invasive diagnostic yield of 56.2% and maintained high specificity (87.0%), comparable to or exceeding the 36.1% (95% CI 33.4–38.9%) yield reported for transbronchial biopsy [
8]. Given that the median interval from initial assessment to final diagnosis in this cohort was 213 days [
3], an earlier indication could shorten the diagnostic timeline and accelerate access to antifibrotic therapy, which slows lung-function decline in IPF [
30,
31] and across progressive fibrosing ILD more broadly [
32].
The processes established under the PCCP framework enabled efficient incremental improvement of this FDA-regulated technology while preserving the safeguards appropriate to such a tool. By pre-specifying the scope of anticipated modifications, the methodology for implementing them, and the criteria for evaluating their impact, the framework allows model updates to proceed without a new regulatory review for each iteration—time-consuming and costly for developers and regulatory agencies alike—while constraining those updates to a defined envelope of permissible modification. The result is a structure oriented toward steady, controlled improvement: changes are advanced only when performance is maintained or improved, with continuous monitoring for performance drift and explicit attention to downside risk. Within this framework, head-to-head comparison of the modified model against the prior authorized version on a locked, adjudicated reference set of known cases provides a far more efficient validation pathway for algorithm improvements than commissioning a new clinical study for each revision. This matters because model performance continues to advance, as demonstrated here. Under prior regulatory paradigms, incorporating such improvements into an authorized clinical product could take years; the PCCP framework compresses that timeline substantially, allowing validated performance gains to reach clinical use far sooner—which, for a diagnostic tool, translates directly into earlier access to improved diagnostic accuracy for patients. For Fibresolve, this translates to upwards movements in sensitivity and specificity, translating to better diagnostic assessments for patients.
This study has several limitations. First, the analysis re-used the PUFAIR cohort and therefore represents a paired validation of the v2 update rather than a new distinct independent external validation study. Second, several subgroup analyses were limited by small sample sizes, which constrained the interpretation of stratum-level differences. Third, sensitivity remained limited among patients with normal lung function, suggesting that early or very mild disease remains a challenging diagnostic scenario. Lastly, ancillary clinical variables were not available for all patients; although the fallback architecture mitigates this limitation by reverting to the CT-only v1 model, some cases were evaluated using the original CT-only pathway.