Next Article in Journal
Development of High-Throughput Serum Bactericidal Assays for Bordetella pertussis to Evaluate BPZE1
Previous Article in Journal
Community-Led Defaulter Tracking for Catch-Up Vaccination: Implementation Experience in Uganda, 2022 and 2024
Previous Article in Special Issue
Disease and Economic Burden Averted by Hib Vaccination in 160 Countries: A Machine-Learning Analysis
 
 
Article
Peer-Review Record

The Mismatch Between Professionally Produced Vaccine Content and Audience Demand on Chinese Short-Form Video Platforms: A Cross-Platform Content Analysis

Vaccines 2026, 14(6), 491; https://doi.org/10.3390/vaccines14060491
by Yuqi Fu 1, Yuan Dang 2, Yuming Liu 2 and Yangmu Huang 2,*
Vaccines 2026, 14(6), 491; https://doi.org/10.3390/vaccines14060491
Submission received: 12 April 2026 / Revised: 26 May 2026 / Accepted: 27 May 2026 / Published: 30 May 2026

Round 1

Reviewer 1 Report

Comments and Suggestions for Authors

Abstract:
• Clearly indicate the study design in the methods section of the abstract.
• Specify the data collection period.

Introduction:
• Line 41: Provide a brief overview of vaccine hesitancy and refusal rates in China.

Methods:
• Lines 107-109: Present vaccine-related terms in both Chinese and English.
• To ensure transparency, add a flow diagram that details the number of videos collected from each platform, the number and reasons for exclusions, and the final number included in the analysis.
• Indicate the software used for video extraction and duplicate removal.
• Clarify the data extraction process and describe the measures taken to ensure data accuracy.
• Line 116: Specify whether a data extraction form was utilized and explain how it was validated.
• Explain the process for including covariates in the multivariate Bayesian negative binomial regression models. State whether univariate analysis was performed and under what conditions.
• Clearly justify the use of MCMC in the methods section.

Results:
• Lines 216-228: Instead of presenting findings without numerical data, include the actual numbers in the text. Report the median (IQR) engagement across different producer categories.
• When stating, "Medical professional accounts achieved the highest median engagement across all four metrics (190 likes, 63 shares, 51 favorites, and 21 comments),” also provide the IQR for each metric.
• Similarly, for “medical official institution accounts showed the lowest medians (25 likes, 5 shares, 4 favorites, and 2 comments),” include the IQRs.
• Add box plots to the section titled “Who provides vaccine information and who gains engagement.”
• In the subsection “How content structures translate into engagement outcomes,” include the posterior mean and 95% confidence intervals as shown in Figure 2.
• Label Figure 2 as 2a, 2b, 2c, and 2d.

Discussion:
• How do you relate the findings of this study to vaccination indicators like vaccination coverage in China?

•Discuss the advantages of employing multivariate Bayesian negative binomial regression models in your study.

 

Comments on the Quality of English Language

The quality of the English language must be improved.

Author Response

We would like to express our sincere gratitude to Reviewer 1 for the detailed and valuable comments that helped us improve the quality of our manuscript. We have addressed all comments as follows: 

Comments 1: Clearly indicate the study design in the methods section of the abstract.

Response 1: We thank the reviewer for this helpful suggestion. In response, we revised the Methods section of the Abstract (Page 1, Lines 13) to explicitly describe the study as a cross-sectional quantitative content analysis.

 

Comments 2: Specify the data collection period.
Response 2: We thank the reviewer for this suggestion. In response, we clarified the video retrieval period in the Abstract (Page 1, Lines15). During revision, we also recognized that the original wording in the Methods section could potentially be misinterpreted as referring to the publication period of the videos rather than the retrieval period. Therefore, we revised the corresponding description in the Methods section (Page 3, Lines 116–118) to more clearly indicate that videos were retrieved during the specified data retrieval window regardless of their original publication date.

 

Comments 3: Line 41: Provide a brief overview of vaccine hesitancy and refusal rates in China.

Response 3: We thank the reviewer for this helpful suggestion. In response, we added a brief overview of vaccine hesitancy, refusal, and vaccination uptake in China in the Introduction (Page 1, Lines 39–42). Specifically, we noted that several voluntary vaccines in China continue to exhibit substantial hesitancy and refusal rates, while vaccination coverage for multiple vaccines remains below optimal levels across different populations.

 

Comments 4: To ensure transparency, add a flow diagram that details the number of videos collected from each platform, the number and reasons for exclusions, and the final number included in the analysis.

Response 4: We thank the reviewer for this valuable suggestion regarding transparency in the screening process. During revision, we re-examined the preprocessing workflow and recognized that data retrieval, deduplication, eligibility assessment, and exclusion of records with incomplete analytical variables were conducted iteratively across multiple stages of dataset development, rather than through a strictly sequential screening pipeline.

As a result, precise intermediate counts at each processing stage could not be retrospectively reconstructed with sufficient consistency to support a fully accurate flow diagram. To avoid potentially misleading or inconsistent reporting, we therefore decided not to include a formal flow diagram in the revised manuscript.

Instead, we revised the Methods section to provide a clearer and more accurate narrative description of the data retrieval and preprocessing procedures, including explicit inclusion and exclusion criteria, duplicate handling procedures, and the final analytic sample size (n = 3,752), which remained stable throughout all analyses.

 

Comments 5: Indicate the software used for video extraction and duplicate removal.

Response 5: We thank the reviewer for this valuable suggestion. In response, we clarified the software and workflow used for video retrieval, preprocessing, and duplicate removal in the Methods section (Page 3, Lines 123–132). Specifically, we added details indicating that raw video metadata were collected using Python-based automated scripts through platform APIs, while subsequent preprocessing procedures, including deduplication, manual verification, eligibility screening, and exclusion of incomplete records, were conducted in Microsoft Excel (Office LTSC).

 

Comments 6: Clarify the data extraction process and describe the measures taken to ensure data accuracy.

Response 6: We thank the reviewer for this suggestion. In response, we further clarified the data preprocessing and variable extraction procedures in the Methods section (Page 3, Lines 132–138), including the extracted variables and the use of manual verification and predefined screening criteria to improve data consistency and accuracy prior to analysis.

 

Comments 7: Line 116: Specify whether a data extraction form was utilized and explain how it was validated.

Response 7: We thank the reviewer for this suggestion. In response, we clarified in the Methods section (Page 3, Lines 125–131) that a structured data extraction spreadsheet was developed in Microsoft Excel to standardize variable collection across videos. We also clarified that variables were determined through discussion within the research team prior to formal extraction and included publicly accessible video- and account-level information directly obtainable during the extraction process.

 

Comments 8: Explain the process for including covariates in the multivariate Bayesian negative binomial regression models. State whether univariate analysis was performed and under what conditions.

Response 8: We thank the reviewer for this important suggestion. In response, we clarified the covariate selection process in the Methods section (Page 5, Lines 203–209). Specifically, we added details indicating that separate univariate negative binomial regression analyses were first conducted for each engagement outcome, and covariates showing significant associations were subsequently included in the corresponding multivariable Bayesian negative binomial regression models. We also clarified that final adjusted models were simplified to improve model parsimony while retaining substantively relevant covariates.

 

Comments 9: Clearly justify the use of MCMC in the methods section.

Response 9: We thank the reviewer for this valuable suggestion. In response, we further clarified the rationale for using Markov chain Monte Carlo (MCMC) estimation in the Methods section (Page 5, Lines 200–203). Specifically, we added a statement indicating that the Bayesian framework with MCMC estimation was used to enable stable parameter estimation and uncertainty quantification, particularly for subgroup analyses involving overdispersed engagement count outcomes.

 

Comments 10: Lines 216-228: Instead of presenting findings without numerical data, include the actual numbers in the text. Report the median (IQR) engagement across different producer categories.

Response 10: We sincerely thank the reviewer for this suggestion. We have added Table 1 in the revised manuscript, which reports the median (IQR) of all four user engagement metrics across all samples and account types. We have also supplemented the IQR values for the highest and lowest engagement groups on Page 6, Lines 263–264, with references to the full table for additional details. The data presentation has been significantly improved as requested.

 

Comments 11: When stating, "Medical professional accounts achieved the highest median engagement across all four metrics (190 likes, 63 shares, 51 favorites, and 21 comments),” also provide the IQR for each metric.

Response 11: We thank the reviewer for this suggestion. As noted in Response 10, we have added median (IQR) values for all engagement metrics across account types in Table 1 and the revised text (Page 6, Lines 250–252). The corresponding IQRs for medical professional accounts are now fully reported in the updated table.

 

Comments 12: Similarly, for “medical institution official media showed the lowest medians (25 likes, 5 shares, 4 favorites, and 2 comments),” include the IQRs.

Response 12: We thank the reviewer for this comment. This issue has been addressed together with Comment 10 and Comment 11 by adding median (IQR) values for all account types in Table 1 and the revised text (Page 6, Lines 255–257), including medical institution official media.

 

Comments 13: Add box plots to the section titled “Who provides vaccine information and who gains engagement.”

Response 13: We thank the reviewer for this valuable suggestion. In response, we have added box plots (Figure 1) to the section “Who provides vaccine information and who gains engagement” to visually present the distribution of user engagement metrics across account types. This figure complements Table 1 by illustrating distributional patterns and group differences in engagement across producer categories. The corresponding changes can be found in the revised manuscript (Page 7, Lines 265–272).

 

Comments 14: In the subsection “How content structures translate into engagement outcomes,” include the posterior mean and 95% confidence intervals as shown in Figure 2.

Response 14: We sincerely thank the reviewer for this helpful suggestion. We agree that reporting posterior means together with 95% credible intervals (the appropriate metric within our Bayesian framework) improves the transparency of Bayesian results and have carefully considered implementing this throughout the manuscript.

However, given the large number of outcomes and subgroup-specific models, fully reporting credible intervals in the main text would introduce substantial redundancy and reduce readability. We therefore did not fully itemize all subgroup-specific credible intervals in the main text in order to preserve readability and avoid substantial redundancy.

Instead, we retain key posterior estimates in a consistent format in the main text, while the complete posterior estimates, including 95% credible intervals, are provided in the Supplementary Materials. This approach ensures both interpretability and reporting consistency while maintaining full statistical transparency.

 

Comments 15: Label Figure 2 as 2a, 2b, 2c, and 2d.

Response 15: We thank the reviewer for this helpful suggestion. In response, we have updated the figure labeling to (a), (b), (c), and (d) in the revised manuscript. Additionally, due to the insertion of the new box plot figure (now Figure 1) requested in Comment 13, the original Figure 1 and Figure 2 have been renumbered as Figure 2 and Figure 3, respectively. All corresponding figure references in the text have been updated accordingly.

 

Comments 16: How do you relate the findings of this study to vaccination indicators like vaccination coverage in China?

Response 16: We thank the reviewer for this important suggestion. In response, we have further clarified the potential public health relevance of our findings in the Discussion section. Specifically, we added a statement noting that prior research has shown that appropriately framed, professionally grounded vaccine communication may improve vaccine-related knowledge and risk perception, thereby potentially supporting vaccination intention and uptake (Page 12, Lines 461–465). This addition helps contextualize our engagement findings in relation to broader vaccination-related outcomes while avoiding overinterpretation beyond the scope of the present study.

 

 

Comments 17: Discuss the advantages of employing multivariate Bayesian negative binomial regression models in your study.

Response 17: We thank the reviewer for this valuable suggestion. In response, we added discussion on the advantages of employing the multivariate Bayesian negative binomial regression models in the present study. Specifically, we clarified that this modeling approach enabled simultaneous examination of multiple engagement dimensions across professional subgroups, thereby providing a more integrated understanding of thematic engagement heterogeneity in short-form vaccine communication. This discussion has been added in the Discussion section (Page 12, Lines 456–460).

Reviewer 2 Report

Comments and Suggestions for Authors

The paper presents an intersection between novel forms of media dissemination and healthcare. It has a quite restricted audience given cultural-sensitive aspect. In my view, several areas need to be improved.

  • Page 1. The abstract is overly lengthy and does not use points estimates or effect sizes to convey information. It also fails to explain why these video platforms were selected. An impact statement including daily users might help.
  • Page 2. Some language errors could have occurred. For example, in lines 46-48, please, clarify the grammatical inconsistency in “evidences shows…”. Do you refer to one body of evidence or multiple studies?
  • Page 3, lines 95–96, I feel you need to justify the use of engagement metrics as proxies for “audience demand”. Also, authors should clarify how this construct is conceptually distinguished from more general attention or platform-driven media.
  • Page 4, lines 168–171. Could you report on formal intercoder reliability statistics (e.g., Cohen’s κ or Krippendorff’s α) to substantiate coding consistency? This would enhance the paper greatly, instead of relying solely on consensus.
  • Page 5, lines 184–188. Subgroup analyses were pre-specified? Also, how multiple comparisons were addressed to avoid inflated type I error in the models?
  • Page 6, lines 223–226, there’s a rather award phrase (i.e., “suggest that visibility and engagement are not concentrated among the most institutionally authoritative sources”). In my view, this appears to be an interpretive statement, hence requiring empirical support.
  • Page 7, lines 262–266. I always ask authors to be kind to readers. Therefore, how should they interpret the reported log-scale coefficients? Interpretability for the audience is key nowadays.
  • Page 8, lines 291–292. Correct the grammatical inconsistency in “effectiveness content were negatively associated” and clarify whether this refers to the same operationalized theme as defined in Section 2.3.
  • Page 8, lines 306–312, please, provide a clearer explanation or example of how the “overall mismatch index” is calculated and how its scale should be interpreted substantively. For instance, what constitutes a large vs. small mismatch.
  • Page 11, lines 422–426, please, reconcile the statement that formal intercoder reliability statistics were not reported with the earlier description of coding validation (Page 4, lines 168–171), and clarify whether reliability was assessed but omitted or not calculated.
  • Page 11, lines 432–435, please, correct the typographical issue in “user demographics.Such work” and ensure consistent spacing and formatting in this section.
Comments on the Quality of English Language

Typos and grammar issues are present. Proofreading by a fluent English speaker is needed.

Author Response

We greatly appreciate Reviewer 2 for the thoughtful and constructive comments that have helped us refine our study and presentation. Our point-by-point responses are listed below:

Comments 1: Page 1. The abstract is overly lengthy and does not use points estimates or effect sizes to convey information. It also fails to explain why these video platforms were selected. An impact statement including daily users might help.

Response 1: Thank you for this valuable suggestion. In response, we substantially revised and streamlined the abstract to improve clarity and conciseness. We also incorporated quantitative point estimates and effect-size information into the Results section, including the proportion of different content producers, representative engagement patterns, regression coefficient ranges, and mismatch index values, so that the main findings are conveyed more directly and quantitatively.

In addition, we clarified the rationale for selecting Douyin, Xiaohongshu, and Kuaishou as the study platforms. Specifically, we added platform user-scale information based on QuestMobile industry statistics, noting that these platforms have approximately 850 million, 350 million, and 410 million daily active users, respectively (Page 2, Lines 54–56). This revision was added to better demonstrate the representativeness and public communication influence of these major short-form video platforms in China.

 

Comments 2: Page 2. Some language errors could have occurred. For example, in lines 46-48, please, clarify the grammatical inconsistency in "evidences shows…". Do you refer to one body of evidence or multiple studies?

Response 2: We sincerely thank the reviewer for this helpful comment. We have corrected the grammatical inconsistency in "evidences shows" by revising the wording to appropriately refer to evidence from multiple studies. In response to this suggestion, we further carefully reviewed the manuscript and revised similar linguistic issues to ensure consistency and improve overall language accuracy throughout the paper.

 

Comments 3: Page 3, lines 95–96, I feel you need to justify the use of engagement metrics as proxies for "audience demand". Also, authors should clarify how this construct is conceptually distinguished from more general attention or platform-driven media.

Response 3: Thank you for this insightful comment. We revised the manuscript to further clarify the conceptual scope of "audience demand" and justify the use of engagement metrics in this study (Page 3, Lines 101–107). Specifically, we clarified that the study focuses on engagement-based audience responsiveness reflected through observable interaction behaviors on short-form video platforms, rather than normative health information needs or actual vaccination intentions. We also explicitly acknowledged the potential influence of platform recommendation mechanisms and broader attention dynamics on engagement metrics. In addition, related clarification has been added to the limitations section (Page 13, Lines 480–484).

 

Comments 4: Page 4, lines 168–171. Could you report on formal intercoder reliability statistics (e.g., Cohen's κ or Krippendorff's α) to substantiate coding consistency? This would enhance the paper greatly, instead of relying solely on consensus.

Response 4: Thank you for this valuable suggestion. We revised the manuscript to provide a clearer description of the coding validation procedure. Specifically, during the initial coding stage, two researchers independently reviewed a random 10% subsample using the predefined codebook, with agreement exceeding 95% across reviewed coding items. Formal intercoder reliability statistics such as Cohen's κ or Krippendorff's α were not calculated in the present study (Page 5, Lines 189–194). This was partly due to the complexity of the coding structure, which involved sentence-level multi-label thematic coding and variable-length video transcripts, making the application of standardized intercoder reliability coefficients less straightforward. We have acknowledged this as a methodological limitation in the revised manuscript.

 

Comments 5: Page 5, lines 184–188. Subgroup analyses were pre-specified? Also, how multiple comparisons were addressed to avoid inflated type I error in the models?

Response 5: Thank you for this important comment. We revised the Methods section to clarify that subgroup analyses across all account types were conducted as part of the overall study design to examine heterogeneity in theme–engagement associations and to support the subsequent supply–demand mismatch analyses (Page 5, Lines 209–215). Although subgroup analyses were performed for all account types, the present study primarily reports results for medical professionals and medical institution official media because of their central relevance to professional vaccine communication, rather than selective result reporting.

Regarding the potential inflation of type I error, inference in the Bayesian regression framework was based on posterior distributions and 95% credible intervals rather than dichotomous null-hypothesis significance testing. Accordingly, we interpreted results based on overall posterior effect patterns and uncertainty estimates rather than isolated significance decisions, and formal multiple-comparison correction procedures were therefore not applied. Relevant clarification has been added to the Methods section (Page 5, Lines 215–217).

 

Comments 6: Page 6, lines 223–226, there's a rather award phrase (i.e., "suggest that visibility and engagement are not concentrated among the most institutionally authoritative sources"). In my view, this appears to be an interpretive statement, hence requiring empirical support.

Response 6: Thank you for this helpful comment. We apologize for the overly assertive and interpretive wording in the original Results section. In response, we revised the relevant sentence to reduce the conclusiveness of the statement and align it more closely with the empirical findings presented in this study (Page 6, Lines 257–259).

 

Comments 7: Page 7, lines 262–266. I always ask authors to be kind to readers. Therefore, how should they interpret the reported log-scale coefficients? Interpretability for the audience is key nowadays.

Response 7: Thank you for this helpful suggestion. We agree that interpretation of the reported coefficients should be made clearer for readers. In response, we revised the Results section and figure legend to provide a more reader-friendly explanation of the log-scale coefficients and their relative interpretation in terms of expected engagement (Page 8, Lines 311–312).

 

Comments 8: Page 8, lines 291–292. Correct the grammatical inconsistency in "effectiveness content were negatively associated" and clarify whether this refers to the same operationalized theme as defined in Section 2.3.

Response 8: Thank you for pointing this out. We corrected the grammatical inconsistency in the revised manuscript and clarified that the term used here refers to the same operationalized thematic category defined in Section 2.3. We also appreciate your reminder regarding the inconsistent terminology used across sections, and have revised the wording accordingly to improve conceptual consistency throughout the manuscript (Page 10, Lines 341).

 

Comments 9: Page 8, lines 306–312, please, provide a clearer explanation or example of how the "overall mismatch index" is calculated and how its scale should be interpreted substantively. For instance, what constitutes a large vs. small mismatch.

Response 9: Thank you for this helpful comment. We revised the Methods and Results sections to further clarify the construction and interpretation of the overall mismatch index (Page 10, Lines 252–255). Specifically, we now explicitly state that the mismatch indices were constructed as relative comparative measures within the study sample rather than standardized scales with predefined thresholds. Accordingly, larger values indicate comparatively greater divergence between thematic content supply and engagement-based audience demand across account types.

Detailed calculation procedures remain provided in Supplementary Method S2. Because the index was intended for relative comparison across account types rather than absolute classification, we did not introduce additional numerical examples or cutoff-based interpretations in the main text, as doing so could overextend the interpretation of the scale beyond its intended use.

 

Comments 10: Page 11, lines 422–426, please, reconcile the statement that formal intercoder reliability statistics were not reported with the earlier description of coding validation (Page 4, lines 168–171), and clarify whether reliability was assessed but omitted or not calculated.

Response 10: Thank you for pointing out this inconsistency. We revised the relevant sections to clarify the distinction between preliminary coding consistency assessment and formal intercoder reliability testing, and to clarify that formal intercoder reliability statistics were not calculated in the present study. We also ensured consistent wording throughout the manuscript regarding the coding validation process.

 

Comments 11: Page 11, lines 432–435, please, correct the typographical issue in "user demographics.Such work" and ensure consistent spacing and formatting in this section.

Response 11: Thank you for pointing this out. We corrected the typographical issue in the revised manuscript and carefully reviewed the surrounding section to ensure consistent spacing and formatting throughout.

Round 2

Reviewer 1 Report

Comments and Suggestions for Authors

I do not have any further comment.

Comments on the Quality of English Language

The quality of the English language must be improved.

Author Response

Comment:
“I do not have any further comment.”

Response:

We sincerely thank the reviewer for the careful evaluation of our manuscript throughout the review process and for acknowledging the improvements made in the revised version. We greatly appreciate the reviewer’s valuable feedback, which has helped strengthen the manuscript.

In this revision, we further carefully proofread the manuscript and made additional minor language and editorial improvements throughout the paper to enhance readability and clarity. All revisions have been highlighted in the revised manuscript.

Reviewer 2 Report

Comments and Suggestions for Authors

I am convinced that the current revision covered my methodological concerns and, albeit context-limited, the paper adds substantial information to the field.

Author Response

Comment:
“I am convinced that the current revision covered my methodological concerns and, albeit context-limited, the paper adds substantial information to the field.”

Response:
We sincerely thank the reviewer for the positive evaluation of our revised manuscript and for recognizing the contribution of our study. We appreciate the reviewer’s constructive comments throughout the review process, which greatly helped improve the methodological rigor and overall clarity of the manuscript.

In addition, following the reviewer’s and editor’s suggestions, we have further carefully proofread the manuscript and made additional minor language and editorial revisions to improve readability and English expression throughout the paper. All changes have been highlighted in the revised manuscript.

Back to TopTop