Validation of Automated Bacterial Suspension Preparation by Colibri® and Plate Streaking by WASP® for Antibiotic Disk Diffusion Susceptibility Testing
Round 1
Reviewer 1 Report
Comments and Suggestions for AuthorsDear Authors,
The manuscript addresses an important and practical bottleneck in AST workflows and is well written and timely. I support the rationale for standardizing inoculum preparation, as manual adjustment to a 0.5 McFarland standard is both error-prone and time-consuming. The validation on a broad strain set and the high categorical agreement are clear strengths, and the study is generally well aligned with current laboratory automation trends.
That said, several presentation and standards-alignment issues should be corrected to improve clarity and reproducibility:
-
Please state explicitly—early in Methods and again where appropriate—that disk diffusion testing requires a 0.5 McFarland inoculum per CLSI, and add a formal citation to CLSI alongside EUCAST throughout the text. While the manuscript indicates preparation of a 0.5 McFarland suspension in both manual and automated arms, the standards cross-referencing should be made explicit so readers understand the requirement applies regardless of the interpretive guideline used.
-
In the Introduction/Discussion, briefly acknowledge that a nontrivial proportion of published AST reports either do not adjust to 0.5 McFarland or fail to report how inoculum density was standardized; framing this issue and citing representative examples will underscore the value of the standardization focus of your study.
-
The automated suspension/pre-analytical system should be documented with sufficient technical detail for reproducibility. Please describe the mechanism by which the device targets 0.5 McFarland, how it detects and corrects under- or over-density, what acceptance/rejection criteria are used, how the device signals out-of-range measurements, the sterilization/cleaning procedures between runs, and any consumables or media specifics. Include photographs or schematic figures of the workflow and user interface. This level of detail is essential in a validation paper.
-
Report the actual inoculum density ranges achieved by each approach. Even if the automated system targets exactly 0.5, please provide the observed distribution (e.g., mean, range, SD) for both manual and automated suspensions. If you applied acceptance windows (for example, 0.50–0.59 McFarland), state them explicitly. These ranges should be summarized in a dedicated table.
-
Ensure all example plate images conform to EUCAST/CLSI reading criteria. Figure panels should avoid overlapping/merging inhibition zones and excessive confluent growth that complicates caliper measurements; if necessary, replace with plates carrying fewer disks so that zones are discrete. Also verify that manual readings did not include bacterial growth within the inhibition zone boundary. The manuscript itself notes more confluent growth with the novel method; select images that still meet reading criteria.
-
Move table titles above the tables and add a comparative table that (i) lists the inoculum ranges for manual vs automated preparations, (ii) contrasts zone-diameter distributions for key antibiotics, and (iii) reports whether differences are statistically significant and whether they alter categorical calls (R/S/I). This will help readers judge practical equivalence, not just overall CA.
-
Include operational metrics—hands-on time, total turnaround time, labor requirements, and approximate per-test consumable costs—for the manual vs automated workflows. Since a central claim is efficiency and standardization, quantifying these gains will strengthen the conclusions.
-
Expand the Discussion regarding species/colony morphologies for which the automated approach is most suitable or potentially problematic (e.g., mucoid producers, swarming Proteus, very adherent/“sticky” colonies, tiny or friable colonies). Mapping system performance to these phenotypes will guide adoption in diverse labs.
-
Quality control is essential. Please document inclusion of CLSI/EUCAST-recommended QC organisms with acceptable zone ranges during the study period (and the frequency of QC). If QC strains were not run alongside test batches, this is a major omission that needs to be corrected or clearly acknowledged as a limitation.
-
You appropriately acknowledge that broth microdilution—the reference method—was not performed. Keep this limitation visible and, if possible, clarify whether any discordances were further adjudicated by an orthogonal method.
Author Response
Reviewer 1.
The manuscript addresses an important and practical bottleneck in AST workflows and is well written and timely. I support the rationale for standardizing inoculum preparation, as manual adjustment to a 0.5 McFarland standard is both error-prone and time-consuming. The validation on a broad strain set and the high categorical agreement are clear strengths, and the study is generally well aligned with current laboratory automation trends.
We thank this reviewer for reviewing our manuscript.
That said, several presentation and standards-alignment issues should be corrected to improve clarity and reproducibility:
1. Please state explicitly—early in Methods and again where appropriate—that disk diffusion testing requires a 0.5 McFarland inoculum per CLSI, and add a formal citation to CLSI alongside EUCAST throughout the text. While the manuscript indicates preparation of a 0.5 McFarland suspension in both manual and automated arms, the standards cross-referencing should be made explicit so readers understand the requirement applies regardless of the interpretive guideline used.
Thank you for this helpful suggestion. We have now explicitly stated early in the Methods section (lines 255-257) that disk diffusion testing was performed using a 0.5 McFarland inoculum in accordance with CLSI guidelines, and we have added the appropriate CLSI reference.
2. In the Introduction/Discussion, briefly acknowledge that a nontrivial proportion of published AST reports either do not adjust to 0.5 McFarland or fail to report how inoculum density was standardized; framing this issue and citing representative examples will underscore the value of the standardization focus of your study.
We agree with this suggestion. We added a brief section in the introduction (lines 47-50) acknowledging this issue and citing two relevant articles.
3. The automated suspension/pre-analytical system should be documented with sufficient technical detail for reproducibility. Please describe the mechanism by which the device targets 0.5 McFarland, how it detects and corrects under- or over-density, what acceptance/rejection criteria are used, how the device signals out-of-range measurements, the sterilization/cleaning procedures between runs, and any consumables or media specifics. Include photographs or schematic figures of the workflow and user interface. This level of detail is essential in a validation paper.
Thank you for this remark, we agree with this suggestion and added some additional information regarding the technical details on lines 253-254, 256-268, 270-279, 284-286, lines 302-303 to increase reproducibility. We also added a schematic figure of the workflow (line 269, figure 3). However, some details (such as the specific tolerance range for the 0.5McFarland suspension when using Colibri) are not publicly available.
4. Report the actual inoculum density ranges achieved by each approach. Even if the automated system targets exactly 0.5, please provide the observed distribution (e.g., mean, range, SD) for both manual and automated suspensions. If you applied acceptance windows (for example, 0.50–0.59 McFarland), state them explicitly. These ranges should be summarized in a dedicated table.
Thank you, this is a very interesting remark. Colibrí is a commercially available automated system, registered as an In Vitro Diagnostic Medical Device (IVD) with CE marking and validated by Copan. After contact with the application specialist, we noticed a mistake: the information regarding the Colibrí suspension preparation was not completely correct, since Colibrí does not prepare a 0.5 McFarland suspension. We corrected this mistake and clarified this issue on lines 259-274 with an extra schematic figure. We apologize for any confusion this may have caused.
5. Ensure all example plate images conform to EUCAST/CLSI reading criteria. Figure panels should avoid overlapping/merging inhibition zones and excessive confluent growth that complicates caliper measurements; if necessary, replace with plates carrying fewer disks so that zones are discrete. Also verify that manual readings did not include bacterial growth within the inhibition zone boundary. The manuscript itself notes more confluent growth with the novel method; select images that still meet reading criteria.
Thank you for your comment regarding the plate images. We have carefully reviewed the example plate images and confirm that they conform to EUCAST/CLSI reading criteria. While some inhibition zones may slightly overlap, this does not affect accurate measurements. We ensured that manual readings did not include confluent bacterial growth within the inhibition zone boundaries, and all images were chosen to clearly demonstrate discrete, measurable zones.
Furthermore, several previously published studies have successfully used six antibiotic disks on a single round Mueller-Hinton agar plate, supporting the validity of our approach. For example: Herroelen, P.H., Heestermans, R., Emmerechts, K. et al. Validation of Rapid Antimicrobial Susceptibility Testing directly from blood cultures using WASPLab®, including Colibrí™ and Radian® in-Line Carousel. Eur J Clin Microbiol Infect Dis 41, 733–739 (2022).
6. Move table titles above the tables and add a comparative table that (i) lists the inoculum ranges for manual vs automated preparations, (ii) contrasts zone-diameter distributions for key antibiotics, and (iii) reports whether differences are statistically significant and whether they alter categorical calls (R/S/I). This will help readers judge practical equivalence, not just overall CA.
We thank the reviewer for this suggestion and appreciate the intention to provide a more detailed comparison between manual and automated inoculation and measurement. However, we believe that including an additional comparative table with raw zone diameters and statistical analyses would not be practical or methodologically meaningful in this context.
First, our dataset comprises more than 2000 individual inhibition zones, which would result in a very extensive and unwieldy table. Moreover, the raw individual zone values were not systematically collected.
Second, we consider statistical testing on raw zone differences to be of limited relevance, as clinical interpretation (R/S/I categorization) does not always correlate linearly with small variations in zone diameter. For example, a non-significant difference of 1 mm may still result in a major or very major error, while a statistically significant difference elsewhere might have no practical consequence. We therefore believe that reporting categorical agreement provides a more direct and clinically meaningful measure of comparability between both approaches.
We hope the reviewer understands our rationale for keeping the comparison at this level and believe that the current presentation adequately demonstrates the practical equivalence between the methods.
7. Include operational metrics—hands-on time, total turnaround time, labor requirements, and approximate per-test consumable costs—for the manual vs automated workflows. Since a central claim is efficiency and standardization, quantifying these gains will strengthen the conclusions.
Thank you for this remark. We acknowledge that our study does not provide precise quantitative data on the difference in hands-on time and total turnaround time (TAT) between the automated and manual workflows. The exact time savings are very difficult to determine. We are aware that this is a limitation of our study, and future work could aim to systematically measure and compare hands-on time and TAT across different workflows. Therefore, we added a statement clarifying this issue on lines 195-202.
However, it is important to note that a faster TAT is not necessarily the primary goal of automation in this study; rather, the main objectives are to reduce hands-on time, increase standardization, and improve traceability within the workflow.
8. Expand the Discussion regarding species/colony morphologies for which the automated approach is most suitable or potentially problematic (e.g., mucoid producers, swarming Proteus, very adherent/“sticky” colonies, tiny or friable colonies). Mapping system performance to these phenotypes will guide adoption in diverse labs.
Thank you for this suggestion. In our experience, we did not encounter any issues with specific species or colony morphologies. All tested phenotypes were processed successfully using the automated workflow. We added a brief statement regarding species/colony morphology in the discussion on lines 179-181.
9. Quality control is essential. Please document inclusion of CLSI/EUCAST-recommended QC organisms with acceptable zone ranges during the study period (and the frequency of QC). If QC strains were not run alongside test batches, this is a major omission that needs to be corrected or clearly acknowledged as a limitation.
We confirm that EUCAST-recommended QC-organisms (ATCC and NCTC-strains) with acceptable zone ranges were included in this study. These strains, with date of inclusion, are listed in the supplementary data under “EUCAST strains”. We added some information regarding this issue on lines 311-315 under “2.5 quality control”.
10. You appropriately acknowledge that broth microdilution—the reference method—was not performed. Keep this limitation visible and, if possible, clarify whether any discordances were further adjudicated by an orthogonal method.
Thank you for pointing this out. Due to limited resources and personnel availability, it was not feasible to perform broth microdilution testing for each isolate alongside the disk diffusion validation. The study was therefore designed to focus on comparative performance between the automated and manual disk diffusion workflows. We acknowledge that this is a limitation of our study and mentioned this on lines 187-193.
Reviewer 2 Report
Comments and Suggestions for Authors
The validation study appears highly relevant to current clinical microbiology practices. However, a few points merit further clarification:
- Rationale for Blood Agar Use: What was the justification for selecting blood agar as the medium? Was it intended to support the growth of specific organisms or as per the automated standards?
- Inclusion of Fastidious Organisms: Why were fastidious organisms, such as Streptococcus species, excluded from the validation? Including such organisms could enhance the robustness and applicability of the findings.
- Test Duplication for Accuracy: Should susceptibility tests be routinely duplicated to ensure consistency and reliability of results provided to clinicians?
- Use of MIC as a Control: Why was Minimum Inhibitory Concentration (MIC) testing not incorporated as a control or reference method to benchmark automated outcomes?
- Ethical Approval Considerations: Was ethical approval obtained from the relevant institutional review board, and if so, could the approval details be included to reinforce the study’s compliance with research standards?
- Supplementary table: Please add full form of antibiotics used.
Author Response
The validation study appears highly relevant to current clinical microbiology practices. However, a few points merit further clarification:
We thank this reviewer for reviewing our manuscript.
1. Rationale for Blood Agar Use: What was the justification for selecting blood agar as the medium? Was it intended to support the growth of specific organisms or as per the automated standards?
Blood agar was used as this represents the standard primary culture medium in our laboratory for routine bacterial isolation and AST preparation. Its use ensures optimal growth of a broad range of clinically relevant organisms, including fastidious species, and aligns with our validated in-house procedures. To clarify this, we added this information to lines 253-254.
2. Inclusion of Fastidious Organisms: Why were fastidious organisms, such as Streptococcus species, excluded from the validation? Including such organisms could enhance the robustness and applicability of the findings.
Thank you for this remark. However, we did include two Streptococcus species in this study: Streptococcus agalactiae (N=8) and Streptococcus pyogenes (N=4) (mentioned on line 213). You can find a detailed list of included species and strains in the Supplementary data.
3. Test Duplication for Accuracy: Should susceptibility tests be routinely duplicated to ensure consistency and reliability of results provided to clinicians?
We thank the reviewer for this thoughtful comment regarding test duplication. In our study, the overall categorical agreement between the automated and manual methods was 96.3%, even without retesting discrepant results. Only 0.4% very major errors were observed, all of which resolved upon repeat testing.
Given these findings, we consider that routine duplication of susceptibility tests is not required for ensuring consistency and reliability in routine use. Occasional repeat testing of unexpected or discrepant results remains good laboratory practice, but our data indicate that the automated approach provides reproducible and clinically reliable results comparable to the manual method without systematic duplication.
4. Use of MIC as a Control: Why was Minimum Inhibitory Concentration (MIC) testing not incorporated as a control or reference method to benchmark automated outcomes?
We appreciate the reviewer’s suggestion to include MIC testing as an additional reference method. However, due to limited resources and personnel availability, it was not feasible to perform broth microdilution testing for each isolate alongside the disk diffusion validation. The study was therefore designed to focus on comparative performance between the automated and manual disk diffusion workflows.
We acknowledge that this is a limitation of our study and mentioned this on lines 187-193.
5. Ethical Approval Considerations: Was ethical approval obtained from the relevant institutional review board, and if so, could the approval details be included to reinforce the study’s compliance with research standards?
This study was conducted exclusively using bacterial strains, without any involvement of human or animal material, clinical samples, or identifiable patient data. Therefore, formal approval from an institutional ethics committee was not required. Nevertheless, all experimental procedures were performed in accordance with institutional and international standards for good laboratory practice and research integrity. We added this statement on lines 361-364.
- Supplementary table: Please add full form of antibiotics used.
Thank you for this suggestion. We added a full form of the antibiotics used to the supplementary data under ‘Antibiotics list’.
Reviewer 3 Report
Comments and Suggestions for AuthorsThe manuscript presents a well-conceived study evaluating automated antimicrobial susceptibility testing (AST) using Colibri® and WASP® systems compared with a manual disk diffusion workflow. This topic is highly relevant as clinical microbiology laboratories continue to adopt total laboratory automation to improve standardization and throughput. The study includes a diverse set of isolates with clinically important resistance mechanisms, and the analysis of categorical agreement is clearly presented. However, there are several points that require clarification to strengthen the interpretation of the results and ensure alignment with established AST validation standards.
Areas for Improvement:
-
Reference Method Selection and ISO 20776-2 Interpretation
The manuscript states that categorical agreement and error classification were performed according to ISO 20776-2, while manual disk diffusion is used as the reference comparator method. However, ISO 20776-2 specifies broth microdilution as the required reference method for AST performance evaluations. Disk diffusion is clinically accepted and widely used, but it is not the reference standard within the ISO validation framework.
To improve clarity, the manuscript would benefit from explicitly stating that:-
Disk diffusion was chosen as the comparator because it reflects the current routine workflow in the authors’ laboratory.
-
Therefore, the resulting major and very major error rates should be interpreted as method-to-method agreement, rather than ISO-standard accuracy estimates.
Additionally, if feasible, resolving the discrepant isolates using broth microdilution would strengthen the validity of the conclusions.
-
-
Clarification of Colony Selection and Inoculum Preparation
Since the Colibri® selects colonies automatically while manual testing relies on technologist choice, it would be helpful to describe whether the same colony type and growth morphology were ensured across methods. This clarification would help rule out strain heterogeneity or colony selection bias as a source of method disagreement. -
Handling of Discrepant Results
The authors retested major and very major discrepancies, which helped refine agreement rates. It may be helpful to elaborate on whether retesting was performed using freshly subcultured isolates and whether any heteroresistance or mixed colony morphotypes were observed upon review. A brief explanation would provide additional context on why initial disagreement occurred and why the automated system may produce more confluent growth. -
Interpretation of Error Patterns by Antibiotic Class
Several discrepancies occurred in antibiotic classes known to be sensitive to inoculum effects (e.g., fluoroquinolones, carbapenems, glycopeptides). A brief discussion acknowledging these intrinsic challenges in disk diffusion interpretation would help contextualize the performance findings. -
Minor Points
-
A short statement quantifying hands-on time reduction would support the conclusion regarding workflow improvement.
-
Consider adding a brief note on whether future multi-center validation is planned, as reproducibility across laboratory environments will be important for broader adoption.
-
Author Response
The manuscript presents a well-conceived study evaluating automated antimicrobial susceptibility testing (AST) using Colibri® and WASP® systems compared with a manual disk diffusion workflow. This topic is highly relevant as clinical microbiology laboratories continue to adopt total laboratory automation to improve standardization and throughput. The study includes a diverse set of isolates with clinically important resistance mechanisms, and the analysis of categorical agreement is clearly presented. However, there are several points that require clarification to strengthen the interpretation of the results and ensure alignment with established AST validation standards.
We thank this reviewer for reviewing our manuscript.
1. Reference Method Selection and ISO 20776-2 Interpretation
The manuscript states that categorical agreement and error classification were performed according to ISO 20776-2, while manual disk diffusion is used as the reference comparator method. However, ISO 20776-2 specifies broth microdilution as the required reference method for AST performance evaluations. Disk diffusion is clinically accepted and widely used, but it is not the reference standard within the ISO validation framework.
To improve clarity, the manuscript would benefit from explicitly stating that:
-
- Disk diffusion was chosen as the comparator because it reflects the current routine workflow in the authors’ laboratory.
- Therefore, the resulting major and very major error rates should be interpreted as method-to-method agreement, rather than ISO-standard accuracy estimates.
Additionally, if feasible, resolving the discrepant isolates using broth microdilution would strengthen the validity of the conclusions.
Thank you for this remark. We included the two suggested sentences explicitly. Indeed, adding broth microdilution would have strengthen the validity of the conclusions, but is not feasible at this point. We fully acknowledge this as a limitation and highlighted this limitation in lines 187-194.
2. Clarification of Colony Selection and Inoculum Preparation
Since the Colibri® selects colonies automatically while manual testing relies on technologist choice, it would be helpful to describe whether the same colony type and growth morphology were ensured across methods. This clarification would help rule out strain heterogeneity or colony selection bias as a source of method disagreement.
Thank you for this suggestion. Both the automated and manual methods were initiated from the same frozen bacterial stock, ensuring that identical strains were used across all tests. During colony selection—whether performed manually or by the Colibri® system—colonies were visually inspected for morphological uniformity. In cases where mixed or morphologically distinct colonies were observed, the corresponding culture plate, suggesting potential strain heterogeneity, was excluded from the validation. This approach minimized the risk of colony selection bias as a source of methodological discrepancy. We added this information on lines 169-172.
3. Handling of Discrepant Results
The authors retested major and very major discrepancies, which helped refine agreement rates. It may be helpful to elaborate on whether retesting was performed using freshly subcultured isolates and whether any heteroresistance or mixed colony morphotypes were observed upon review. A brief explanation would provide additional context on why initial disagreement occurred and why the automated system may produce more confluent growth.
Thank you for this helpful suggestion. Retesting of major and very major discrepancies was performed using freshly subcultured isolates. Upon review, no mixed or morphologically distinct colonies were observed. Nevertheless, potential heteroresistance remains an important consideration when interpreting discordant results. Additional information and a relevant reference addressing this aspect have been incorporated into the revised manuscript (lines 169-174).
4. Interpretation of Error Patterns by Antibiotic Class
Several discrepancies occurred in antibiotic classes known to be sensitive to inoculum effects (e.g., fluoroquinolones, carbapenems, glycopeptides). A brief discussion acknowledging these intrinsic challenges in disk diffusion interpretation would help contextualize the performance findings.
Thank you for this valuable suggestion. We have briefly addressed this point in the Discussion section (lines 172-174) and included an appropriate reference to acknowledge the potential influence of inoculum effects on certain antibiotic classes.
- Minor Points
- A short statement quantifying hands-on time reduction would support the conclusion regarding workflow improvement.
Thank you for this remark. We acknowledge that our study does not provide precise quantitative data on the difference in hands-on time and total turnaround time (TAT) between the automated and manual workflows. The exact time savings are very difficult to determine. We are aware that this is a limitation of our study, and future work could aim to systematically measure and compare hands-on time and TAT across different workflows. Therefore, we added a statement clarifying this issue on lines 195-202.
However, it is important to note that a faster TAT is not necessarily the primary goal of automation in this study; rather, the main objectives are to reduce hands-on time, increase standardization, and improve traceability within the workflow.
- Consider adding a brief note on whether future multi-center validation is planned, as reproducibility across laboratory environments will be important for broader adoption.
Thank you for this helpful suggestion. We have added a brief note addressing this point on lines 349-350.
Round 2
Reviewer 1 Report
Comments and Suggestions for AuthorsDear Authors,
I appreciate seeing your revisions on the manuscript file. I think that this manuscript will help researchers who are working on antimicrobial susceptibility analysis when they are preparing the inoculum concentration. I would only suggest that you should add only CLSI-related citations in the revised file, and also you need to add them in a correct format
Author Response
I appreciate seeing your revisions on the manuscript file. I think that this manuscript will help researchers who are working on antimicrobial susceptibility analysis when they are preparing the inoculum concentration. I would only suggest that you should add only CLSI-related citations in the revised file, and also you need to add them in a correct format.
-We thank the reviewer for reviewing our manuscript once again. We would, however, appreciate further clarification regarding the request to cite only CLSI. As described in the manuscript, our laboratory follows EUCAST guidelines and breakpoints. In this context, excluding all EUCAST references in favour of CLSI alone would, in our view, not accurately reflect the scientific framework within which the study was conducted. We therefore consider it essential to retain the EUCAST citations alongside the CLSI ciations.

