Abstract
AI-generated news is increasingly presented through combinations of text and visual content, making complete user-facing stimuli an important topic for credibility evaluation. This exploratory empirical evaluation examined perceived message credibility ratings across six implemented AI-generated news stimuli, organized by two selected stimulus domains—design industry and quantum computing—and three implemented presentation configurations: text only, an image with researcher-designated higher correspondence, and an image with researcher-designated lower correspondence. A total of 211 students from design-related disciplines evaluated all six stimuli in a fixed order, yielding 1266 ratings. Credibility ratings differed across the stimuli corresponding to the three implemented presentation configurations, F(1.95, 406.54) = 9.16, p < 0.001, partial η2 = 0.042, and the pattern of ratings across the implemented configurations differed between the selected stimulus domains, F(1.94, 404.48) = 14.94, p < 0.001, partial η2 = 0.067. Within the design-industry stimuli, the text-only stimulus was rated higher than the stimulus with a researcher-designated higher-correspondence image. Within the quantum-computing stimuli, both image-present stimuli were rated higher than the text-only stimulus, and the researcher-designated lower-correspondence image stimulus received the highest rating. The six implemented stimuli showed different credibility-rating patterns, underscoring the importance of considering visual attributes together with the content and presentation context in which they occur. Because story identity, presentation configuration, and serial position were not independently crossed, the findings describe the six implemented stimuli rather than isolated causal effects of visual presentation.
1. Introduction
Generative artificial intelligence is increasingly used to produce information for multimodal digital interfaces. In AI-generated news and related content systems, users encounter combinations of text, images, formatting, and authorship information rather than text alone. Credibility evaluation is therefore also an interface-design problem: visible presentation choices can shape how users interpret and judge generated content [1,2,3,4].
This study examines perceived message credibility, defined as a user’s judgment of the reliability and believability of an individual report. This construct is narrower than trust in journalism, media credibility, or confidence in AI as a technology [5,6]. The distinction matters because the present data consist of ratings of six specific reports; they do not measure broader institutional trust.
Existing research has examined machine authorship, disclosure labels, source attribution, and text quality, but visual configurations in AI-generated news have received less systematic attention. The practical question is not simply whether an image is present. It is whether different text-and-image configurations are associated with different credibility judgments, and whether those patterns remain similar across communication contexts. The present study addresses this question with two selected stimulus domains and three implemented presentation configurations.
The present study addresses this gap through an exploratory empirical evaluation of perceived message credibility across six implemented AI-generated news stimuli. The stimuli were organized by two selected domains—design industry and quantum computing—and three implemented presentation configurations within each domain: text only, an image with researcher-designated higher correspondence, and an image with researcher-designated lower correspondence. This structure permits statistical comparison of rating patterns across the implemented configurations within and between the selected stimulus domains. At the same time, each configuration was paired with a different fictional story and appeared at a fixed serial position. The resulting comparisons therefore concern the six implemented stimuli and cannot identify an independent causal effect of presentation configuration, image presence, or researcher-designated correspondence.
The study uses the heuristic–systematic model and the MAIN model as interpretive perspectives for considering credibility-related cues in AI-mediated information environments. It does not directly test familiarity, prior knowledge, cue-processing mechanisms, or participant-perceived image–text correspondence. Accordingly, the analysis focuses on observed rating patterns and statistical comparisons across the implemented stimulus configurations, rather than on directional hypotheses or psychological mechanisms.
2. Literature Review
2.1. Credibility Evaluation of AI-Generated Information
Credibility has been studied at the levels of source, message, and medium. For an individual news report, perceived message credibility concerns whether the presented account appears believable, reliable, and plausible [5,6]. This judgment may draw on content coherence, apparent evidence, production quality, and available source information. Studies of online and AI-generated information likewise show that audiences do not rely on a single cue; they combine message features with assumptions about who or what produced the content [1,2,3,4].
Research on automated and AI-assisted journalism has concentrated mainly on human-versus-machine authorship, disclosure labels, source attribution, and textual quality [7,8,9,10,11,12,13,14,15]. The reported effects are not uniform. AI attribution can alter credibility judgments under some conditions, whereas other studies find small, conditional, or absent differences. Recent work also documents changing transparency practices and heterogeneous audience responses to AI-labeled news [7,8,9,10,11,12,13,14,15,16,17]. These findings support treating authorship disclosure as one cue within a broader evaluation environment rather than as a universally decisive signal. The role of complete multimodal presentation configurations remains comparatively underexplored.
2.2. Visual Information and Multimodal Credibility Cues
Images can increase vividness, provide apparent evidence, and contribute to a professional interface. Visual-framing and audience research indicates that images can shape emotional responses, perceived credibility, and attitudes toward news [18,19,20,21,22]. Complementary computational research treats image–text consistency as an important component of multimodal credibility assessment [23,24]. Visual appearance can influence users’ credibility perceptions of online news [25]. More broadly, credibility evaluation also varies with prior knowledge, message plausibility, and source credibility [26,27,28,29]. Visual enrichment should therefore not be assumed to create a general credibility premium.
Image–text correspondence is especially relevant to multimodal AI systems because a visually plausible image may appear evidential even when its relationship to the report is weak. Higher- and lower-correspondence configurations can therefore be treated as distinct researcher-designated combinations of textual and visual information. Taken together, prior work suggests that the credibility value of an image depends on how it functions within the complete message configuration rather than on image presence alone.
2.3. Human-Centered AI and Credibility-Aware Interface Design
The heuristic–systematic model distinguishes relatively effortful evaluation of message content from reliance on accessible judgment cues, while allowing both modes of processing to coexist [30,31]. The MAIN model similarly explains how technological affordances can activate heuristics that inform credibility judgments [32]. A recent meta-analysis found that heuristic credibility cues often matter in digital environments, although their effects vary across cues and contexts [33]. These frameworks are used here as interpretive perspectives rather than as directly tested process models.
The models do not imply that a cue has a fixed direction of influence. The evidentiary value of an image may depend on the report being evaluated and on what users can infer from the available text. A visual element may support evaluation in one context but appear decorative, distracting, or inconsistent in another. The present study therefore treats stimulus domain as a contextual condition and asks whether the observed configuration pattern differs across two selected domains; it does not treat domain as a validated measure of familiarity or prior knowledge. Overall, these perspectives support evaluating visual interface cues as context-sensitive inputs to credibility judgment, while leaving the underlying processing mechanisms for direct testing in future research.
2.4. Research Gap and Research Questions
Prior research has examined AI authorship, disclosure, source attribution, textual quality, and visual features in credibility evaluation, but less attention has been given to how users rate complete AI-generated news stimuli that combine text and images in different implemented configurations. The present study addresses this gap through an exploratory empirical evaluation of six AI-generated news stimuli. These stimuli represent three implemented presentation configurations within each of two selected stimulus domains: design industry and quantum computing.
The primary objective was to compare perceived message credibility ratings across the implemented configurations and examine whether this pattern differed between the two selected stimulus domains. The presentation-configuration term and the domain × configuration interaction provide a statistical structure for comparing the six implemented stimuli. However, because presentation configuration was linked to different fictional stories and fixed serial positions, these statistical comparisons do not identify independent causal effects of configuration, image presence, or researcher-designated correspondence.
RQ1: What patterns of perceived message credibility ratings are observed across the three implemented presentation configurations—text only, researcher-designated higher-correspondence image, and researcher-designated lower-correspondence image—within the two selected AI-generated news stimulus domains?
RQ2: How does the observed pattern of ratings across the implemented presentation configurations differ between the selected design-industry and quantum-computing stimulus domains?
The study and analysis plan were not preregistered. Familiarity, prior knowledge, and cue-processing mechanisms were not directly measured, and researcher-designated correspondence was not validated through a participant manipulation check. Nondirectional research questions were therefore retained rather than introducing retrospective directional hypotheses. The classroom-level disclosure comparison and participant-characteristic analyses were secondary and exploratory; they do not support population-level causal inference regarding disclosure.
3. Materials and Methods
3.1. Study Design
The primary research structure concerned the 2 × 3 within-participant comparison of two selected stimulus domains and three implemented presentation configurations, evaluated as single-item perceived message credibility across six implemented stimuli. The primary analysis retained a 2 × 3 × 2 mixed analytical structure, with classroom-level AIGC authorship disclosure included as a secondary exploratory between-participants term. These repeated-measures terms represented the six stimuli rather than independently crossed manipulations.
The analysis treated selected stimulus domain as a within-participant term, comprising design industry and quantum computing. These domains were selected to provide contrasting communication contexts for the study materials. Each domain was represented by one topic family; therefore, the observed comparisons concern the selected stimulus domains and cannot isolate topic content, familiarity, prior knowledge, or disciplinary relevance as independent explanatory factors.
The analysis treated implemented presentation configuration as a within-participant term, with three levels: text only, an image with researcher-designated higher correspondence, and an image with researcher-designated lower correspondence. These labels represent the research team’s assessment of relative image–text correspondence and were not validated through a participant manipulation check. For readability, subsequent references retain these labels while preserving their researcher-designated status.
AIGC authorship disclosure was treated as a between-participants term. Four intact classes participated: two classes completed the oral-and-written AIGC authorship-disclosure condition, and two different classes completed the no-disclosure condition. Disclosure was therefore assigned at the intact-class and session level rather than through individual randomization.
Every participant evaluated the same six stimuli in the same fixed order. Within each selected stimulus domain, each implemented presentation configuration was paired with a different fictional story. Consequently, configuration, story identity, and serial position cannot be separated. The statistical comparisons are therefore interpreted as observed rating patterns across the six implemented stimuli, rather than as isolated causal effects of presentation configuration, image presence, or researcher-designated correspondence. Disclosure-related comparisons are secondary and exploratory because classroom and session differences cannot be independently separated from disclosure.
3.2. Participants and Sample Cleaning
The primary analytic sample retained all 211 students from design-related disciplines. Each participant evaluated six news items, yielding 1266 repeated-measures observations. Because identical ratings across six stimuli can reflect either low engagement or a genuine absence of perceived differences, invariant responding was not used as a primary exclusion criterion.
Twenty-three participants (10.9%) assigned exactly the same credibility rating to all six stimuli. A prespecified exclusion rule for such responses had not been preregistered. These participants were therefore retained in the primary analysis, and an additional sensitivity analysis excluded them, producing a non-invariant-response sample of 188 participants and 1128 observations.
The primary sample included 106 participants in the oral-and-written disclosure classes and 105 in the no-disclosure classes. Mean age was 20.06 years (SD = 1.62; range = 18–25). The sample comprised 161 women (76.3%) and 50 men (23.7%), including 79 lower-division undergraduates (37.4%), 109 upper-division undergraduates (51.7%), and 23 master’s students (10.9%). Complete participant characteristics are reported in Supplementary Table S1.
The sample size was determined by the number of eligible students available in the four participating classes rather than by an a priori power analysis. Confidence intervals and multiple sensitivity analyses were used to evaluate the precision and robustness of the reported associations.
Data-quality checks identified no missing analysis variables, duplicate participant identifiers, or ratings outside the 1–7 response scale. Median questionnaire-completion time was 188 s (IQR = 159; range = 51–1121). Five participants completed the questionnaire in less than 60 s, 23 in less than 90 s, and 51 in less than 120 s. Because no completion-time threshold had been specified in advance, all participants were retained in the primary analysis; post hoc sensitivity analyses applied these three thresholds.
3.3. Study Materials
Tencent Yuanbao was used to generate the study news texts and accompanying images from standardized prompts. The platform was selected because it was accessible to the research team during stimulus preparation and provided text- and image-generation functions within a single user-facing interface; the platform itself was not the object of evaluation. The prompts controlled the news format, approximate length, paragraph structure, byline, dateline, intended image content, and 16:9 aspect ratio. The resulting materials were presented as multimodal information stimuli in a controlled mobile-questionnaire evaluation scenario. The study examined user judgments of presented AI-generated content; it did not assess navigation, interactive usability, or the technical performance of the generation system. The research team reviewed candidate materials for basic readability and format consistency, but no independent pretest quantified text plausibility, linguistic similarity, image realism, esthetic quality, or perceived professionalism.
The study materials comprised six AI-generated news stimuli representing the two selected stimulus domains and three implemented presentation configurations: text only, an image with researcher-designated higher correspondence, and an image with researcher-designated lower correspondence. Within each topic family, the three configurations used standardized formatting, but each configuration was paired with a distinct fictional story. The three configurations therefore organized the statistical comparisons, while the observed ratings remained specific to the six implemented stimuli. All six stimuli were presented to every participant in the same fixed order.
The six selected study stimuli and the retained standardized prompt templates are provided in the Supplementary Materials to support material-level reproducibility. The Supplementary Materials also includes the available image-selection criteria and relevant supplementary analyses. The stimuli were selected to instantiate the planned domain-by-configuration structure and contrasting communication contexts; they do not provide exhaustive coverage of news topics, visual designs, or AI-generated news formats.
The two questionnaire versions used the same item wording, news materials, fixed stimulus order, response scales, and evaluation procedure. Their non-item instructions differed where required by the disclosure condition: Group A received oral-and-written AIGC authorship disclosure before opening the questionnaire and in the written instructions, whereas Group B received no AIGC disclosure before or during the questionnaire and was debriefed afterward.
3.3.1. News Topics
The design-industry stimulus domain concerned the impact of AIGC on the design industry. It was selected as a communication context involving design processes, creative production, and professional practice.
The quantum-computing stimulus domain concerned a breakthrough in quantum-computing applications. It was selected as a contrasting communication context containing specialized technical information.
The two selected stimulus domains were used to organize comparisons across contrasting topic contexts. Because one topic family represented each domain, stimulus content and other unmeasured topic-related characteristics cannot be separated. Accordingly, comparisons are interpreted as differences between the two selected stimulus domains rather than as effects of familiarity, prior knowledge, disciplinary relevance, or topic content considered independently.
3.3.2. News Text Generation
All news texts were generated through Tencent Yuanbao from standardized Chinese-language prompts. The prompts specified a consistent news format, approximate length, paragraph structure, byline, and dateline. Each item was approximately 300 Chinese characters long and consisted of a single paragraph. Multiple candidate outputs were reviewed, and the final texts were selected through discussion among the supervising faculty member and research-team members.
Within each topic family, the three presentation configurations followed the same standardized news template, approximate length, single-paragraph structure, byline, and dateline format. However, each configuration was paired with a different fictional event. This standardization reduced format differences while retaining the story-identity confounding described in Section 3.1.
3.3.3. Image Generation and Researcher-Designated Correspondence
The images were generated using the image-generation function available through the Tencent Yuanbao web interface between 30 and 31 October 2025. The interface did not provide a verifiable build-level model identifier; no precise image-model version is therefore reported. Standardized Chinese-language prompts specified a 16:9 aspect ratio and intended image content. No reference images were used.
Multiple candidate images were reviewed by the supervising faculty member and research-team members using intended topic correspondence, basic readability, visual clarity, and consistency with the standardized 16:9 presentation as discussion criteria. No independent external ratings or preregistered numerical selection thresholds were collected. The selected images were proportionally cropped to standardize presentation and remove the platform-generated “AI-generated” label in the lower-right corner. The label was removed consistently from all selected images to avoid introducing an unplanned authorship-disclosure cue; no other image content was edited.
The text-only stimulus contained no news image. In both image-present configurations, the image appeared below the corresponding news text in the same position. Although broad generation settings and aspect ratio were standardized, the implemented images could still differ in salience, complexity, esthetics, and other uncontrolled visual characteristics.
Researcher-designated higher correspondence and researcher-designated lower correspondence reflected the research team’s assessment of the image–text combinations, based on standardized prompts and discussion. No participant manipulation check was administered, so participants’ perceptions of this distinction remain unverified.
3.3.4. Use of Generative AI
Tencent Yuanbao was used to generate candidate study news texts and images. The image-generation function was accessed between 30 and 31 October 2025; the interface did not expose a verifiable build-level model identifier. OpenAI Codex was used to assist with R code development and checking, English translation, language editing, figure preparation, and manuscript formatting. The authors selected all materials, made all analytical and interpretive decisions, verified statistical results against the source data, reviewed all AI-assisted outputs, and took full responsibility for the manuscript.
3.4. Study Procedure
Data were collected during four classroom-based administrations conducted at different times in mid-February 2026. Each disclosure condition comprised two intact classes. Participants completed the online questionnaire individually on their mobile phones while a researcher was present. They were instructed not to consult browsers, social media, or other external sources during the evaluation. Completion time was recorded, and the questionnaire generally took no more than 15 min.
The questionnaire’s first page described the study purpose, voluntary participation, expected completion time, and confidential handling of responses. Participants indicated consent by voluntarily proceeding after reading this information; no separate consent checkbox was used. The original questionnaire requested names solely to reconcile records during data organization. Names were removed before the final analytic dataset was created, and only participant numbers were retained. Immediately before Group A opened the questionnaire, the researcher delivered a standardized oral instruction explaining that the news items had been generated using AI and asking participants to assess credibility using their own judgment and experience. The wording closely paralleled the written disclosure shown before the news task. Group B received no AI-authorship information before or during questionnaire completion.
The evaluation phase comprised the six implemented news stimuli, presented to all participants in the same fixed order. After reading each stimulus, participants rated its single-item perceived message credibility on a 1–7 scale. Because each serial position corresponded to one specific stimulus, order, fatigue, comparison, and carryover effects remain possible alternative explanations for differences among the implemented stimuli.
After evaluating the six news stimuli, participants completed 10 credibility-evaluation items. These items described their assessments of content-quality cues, source- and presentation-related cues, and AIGC-related perception cues when judging news credibility.
After the six news evaluations and the credibility-cue items, participants reached the end of the questionnaire. Participants in the no-disclosure condition were then informed that the news materials had been generated using AI and that the reported events and organizations were fictional. No participant discontinued the questionnaire during administration.
The relationship among the four intact classes, classroom-level disclosure comparison, fixed six-stimulus sequence, user-evaluation measures, and analysis strategy is summarized in Figure 1.
Figure 1.
Study design and human-centered evaluation overview. Four intact design-related classes completed the study in classroom sessions in mid-February 2026. Two classes received oral-and-written AIGC authorship disclosure, whereas two classes received the no-disclosure condition; disclosure was not individually randomized. All participants evaluated the same six news stimuli in a fixed order. Each presentation configuration was tied to one fictional story within each selected stimulus-domain condition. The primary analysis retained all 211 participants. Sensitivity analyses examined invariant responding, completion-time thresholds, Gaussian linear mixed-effects models, and ordinal cumulative-link mixed models. Correspondence categories were researcher-designated and were not independently validated by a participant manipulation check. The schematic was created by the authors. Abbreviation: n.s., not statistically significant; the disclosure comparison is secondary and exploratory.
3.5. Measures
3.5.1. Primary Outcome
The primary outcome was single-item perceived message credibility for an individual news report. The study intentionally operationalized credibility as a focused global evaluation of the presented message rather than as a multidimensional trust or credibility construct. After each stimulus, participants answered the Chinese item corresponding to “How credible do you consider the following news report to be?” Responses ranged from 1 (“completely not credible”) to 7 (“completely credible”).
The identical single item was administered after every report to preserve within-participant comparability while limiting repeated-measurement burden across the six stimuli. The outcome is a report-level global judgment, not a validated multidimensional credibility scale.
Each participant provided six single-item perceived message credibility ratings, one for each implemented stimulus. These ratings served as the dependent variable in the mixed analysis of variance, linear mixed-effects model, and ordinal mixed-model sensitivity analysis.
3.5.2. Credibility-Evaluation Cues
The post-task credibility-evaluation items comprised three researcher-defined conceptual domains and 10 single-item indicators concerning content quality, source and presentation cues, and AIGC-related perceptions. They were collected for exploratory description and were neither combined into a total score nor treated as validated multidimensional credibility dimensions. Because perceived message credibility was measured with one global item, internal-consistency reliability (e.g., Cronbach’s alpha) was not applicable to the primary outcome; the 10 cue indicators were likewise not combined into a composite scale.
The 10 indicators were collected once at the participant level after all six news evaluations rather than separately for each stimulus. They therefore cannot be linked to a specific topic, story, or presentation configuration and cannot be interpreted as mediators or stimulus-specific process measures.
3.5.3. Participant Characteristics
Participant characteristics included age, gender, academic level, frequency of AI use, and frequency of news consumption.
Gender comprised the categories female and male. Academic level was classified according to the survey as lower-division undergraduate, upper-division undergraduate, or master’s student. Frequency of AI use comprised four levels: never, occasionally (several times per month), frequently (several times per week), and daily. Daily online-news consumption comprised four levels: never actively reading news, 1–30 min, 30 min–1 h, and more than 1 h.
These variables were used to describe the sample and assess observed differences between the two classroom-level disclosure groups. Separate exploratory models then examined their associations with perceived message credibility and whether they moderated differences across presentation configurations. Because disclosure was not individually randomized, these comparisons are descriptive or exploratory rather than baseline-equivalence tests for a randomized trial.
Age, frequency of AI use, and frequency of news consumption were standardized before inclusion in the exploratory models. Because age and academic level were strongly related both conceptually and empirically, participant characteristics were entered into separate adjusted models rather than simultaneously in a single model, thereby reducing the influence of multicollinearity on interpretation.
3.6. Data-Analysis Strategy
All analyses were conducted in R 4.5.0. Data management and visualization primarily used tidyverse; the statistical analyses used afex, emmeans, lme4, car, effectsize, ordinal, and related supporting packages.
The primary analysis retained all 211 participants. Quality screening examined missing values, duplicate identifiers, out-of-range ratings, invariant responding, and completion time. Because no invariant-response or completion-time exclusion rule had been preregistered, these indicators were not used to define the primary sample. Sensitivity analyses excluded the 23 invariant responders and separately retained participants with completion times of at least 60, 90, or 120 s.
Sample size, mean, SD, median, IQR, range, and 95% confidence interval were calculated for each of the 12 classroom-level disclosure × domain × presentation cells. Figure 2a presents raw single-item perceived message credibility distributions collapsed across disclosure classes, with the pattern across the three implemented configurations shown separately by selected stimulus domain. Because each presentation configuration was linked to a specific story and fixed serial position, stimulus position was perfectly confounded with stimulus identity and could not be estimated as an independent order predictor. The descriptive summaries therefore characterize the six implemented stimuli.
Figure 2.
Perceived credibility across the selected stimulus domains and implemented presentation configurations. (a) Observed distributions in the full sample. Violins represent score distributions, embedded boxplots show medians and interquartile ranges, and diamonds indicate arithmetic means. (b) Estimated marginal means and 95% confidence intervals from the primary mixed ANOVA, averaged across classroom-level disclosure groups. N = 211 participants and 1266 repeated observations. Story identity and fixed serial position were confounded with presentation configuration; the graph therefore describes the six implemented stimuli rather than story-independent visual-format effects. The plots were created by the authors from the study data. Note. Higher- and lower-correspondence categories were designated by the research team and were not validated through an independent participant manipulation check.
The principal model retained a 2 × 3 × 2 mixed ANOVA. Its primary research structure concerned the repeated-measures comparison of selected stimulus domain and implemented presentation configuration; classroom-level AIGC authorship disclosure was included as a secondary exploratory between-participants term. Type III sums of squares were used. Greenhouse–Geisser-corrected degrees of freedom and epsilon were reported for model terms involving presentation configuration. Each model term was accompanied by F, p, partial eta squared, and its 95% confidence interval. The presentation-configuration term and the domain × configuration interaction compare stimulus ratings within the causal limits described in Section 3.1. The mixed ANOVA treated participants as the observational units, with stimulus domain and implemented presentation configuration specified as repeated-measures terms and disclosure as a between-participants term. Because disclosure was allocated across only four intact classes, disclosure-related tests may be affected by unmodelled classroom clustering and are treated as secondary and exploratory.
For statistically significant omnibus model terms, simple-effects analyses and pairwise comparisons were conducted using estimated marginal means. Holm adjustment was applied to multiple comparisons, with a two-sided significance threshold of α = 0.05. These comparisons describe within-domain rating differences among the implemented stimuli. The primary within-participant omnibus results, estimated marginal means, and rating patterns are presented in Figure 2 and Table 1 and Table 2; the complete 2 × 3 × 2 omnibus results are reported in Supplementary Table S14.
Table 1.
Primary within-participant ANOVA terms for single-item perceived message credibility in the full sample.
Table 2.
Pairwise comparisons within each selected stimulus-domain condition in the full sample.
Robustness analyses included a random-intercept linear mixed-effects model with sum-to-zero contrasts and a cumulative-link mixed model treating the 1–7 outcome as ordinal. Both models included a participant random intercept. The fixed sequence precluded an independent order-effect analysis. These analyses assessed statistical stability under alternative specifications while retaining the design’s identification limits. Because disclosure was assigned across only four intact classes, these models do not provide a reliable estimate of a population-level classroom-disclosure effect; disclosure findings are therefore reported as exploratory.
Analyses of participant characteristics were explicitly exploratory. Standardized differences described observed differences between the classroom-level disclosure groups; separate adjusted linear mixed-effects models tested associations with single-item perceived message credibility and statistical variation in ratings across the implemented configurations. Holm adjustment was applied across the relevant test families. These models did not remove possible classroom- or session-level confounding. Detailed results are reported in the Supplementary Materials.
A further exploratory participant-level analysis related each post-task cue indicator to mean perceived message credibility ratings across the six reports using Spearman rank correlations. Holm adjustment was applied across the 10 correlations. Because each cue was measured once after the complete stimulus sequence, these associations cannot identify stimulus-specific cue use, temporal direction, or mediation of the domain × configuration rating pattern.
Finally, the disclosure and no-disclosure groups were compared on the 10 credibility-evaluation cues using Mann–Whitney U tests. Holm-adjusted p values and rank-biserial correlations are reported in the Supplementary Materials.
4. Results
4.1. Sample and Data Quality
The primary analytic sample comprised 211 participants, contributing 1266 repeated credibility ratings. No missing values, duplicate participant identifiers, or ratings outside the 1–7 response scale were identified.
Twenty-three participants (10.9%) gave the same credibility rating to all six stimuli. Because no exclusion criterion for invariant responding had been preregistered, these cases were retained in the primary analysis and examined separately in the sensitivity analyses.
The oral-and-written disclosure classes included 106 participants, and the no-disclosure classes included 105. Participants were 18–25 years old (M = 20.06, SD = 1.62); 161 were women (76.3%) and 50 were men (23.7%). Complete participant characteristics and standardized between-class differences are reported in Supplementary Table S1 and Figure S3; completion-time checks are provided in Supplementary Table S11 and Figure S1.
4.2. Credibility Ratings Across the Six Implemented Stimuli
Single-item perceived message credibility ratings showed different patterns across the six implemented stimuli (Figure 2). The full set of cell means, standard deviations, medians, interquartile ranges, ranges, and confidence intervals is reported in Supplementary Table S2.
Within the design-industry stimuli, the text-only stimulus had the highest estimated credibility rating, whereas the stimulus containing the researcher-designated higher-correspondence image had the lowest estimated credibility rating. The stimulus containing the researcher-designated lower-correspondence image fell between them.
Within the quantum-computing stimuli, both image-present stimuli were rated above the text-only stimulus, and the stimulus containing the researcher-designated lower-correspondence image received the highest estimated credibility rating. Figure 2a shows the raw score distributions, and Figure 2b shows the model-estimated pattern.
4.3. Statistical Comparisons Across the Implemented Configurations
The stimulus-domain × presentation-configuration interaction was significant, F(1.94, 404.48) = 14.94, p < 0.001, partial η2 = 0.067, 95% CI [0.026, 0.116]. The rating pattern across the three implemented configurations differed between the two selected stimulus domains. This statistical interaction concerns the implemented stimuli, not an isolated causal effect of configuration.
Ratings also differed across the stimuli corresponding to the three implemented presentation configurations, F(1.95, 406.54) = 9.16, p < 0.001, partial η2 = 0.042, 95% CI [0.010, 0.084].
By contrast, there was no reliable overall difference between the two selected stimulus domains, F(1, 209) = 2.18, p = 0.142. Disclosure-related terms from the complete 2 × 3 × 2 model are reported in Supplementary Table S14. Because disclosure was allocated across only four intact classes rather than individually randomized, these results are treated as secondary exploratory evidence and are not interpreted as population-level disclosure effects.
Pairwise Comparisons Within Selected Stimulus Domains
Within the design-industry stimuli, the text-only stimulus received a higher credibility rating than the stimulus containing the researcher-designated higher-correspondence image (text only minus higher-correspondence image: MD = 0.242, 95% CI [0.048, 0.436], Holm-adjusted p = 0.044). The remaining pairwise comparisons were not statistically reliable.
Within the quantum-computing stimuli, both image-present stimuli received higher credibility ratings than the text-only stimulus. The stimulus containing the researcher-designated higher-correspondence image differed from the text-only stimulus by MD = 0.246 (95% CI [0.051, 0.441], Holm-adjusted p = 0.014), and the stimulus containing the researcher-designated lower-correspondence image differed from the text-only stimulus by MD = 0.606 (95% CI [0.421, 0.791], Holm-adjusted p < 0.001). The researcher-designated lower-correspondence image stimulus also received a higher credibility rating than the researcher-designated higher-correspondence image stimulus (lower- minus higher-correspondence image: MD = 0.360, 95% CI [0.184, 0.536], Holm-adjusted p < 0.001). These pairwise comparisons describe the ordering of the implemented quantum-computing stimuli.
4.4. Robustness Analyses
The stimulus-domain × presentation-configuration interaction remained statistically significant after excluding the 23 invariant responders, F(1.94, 360.70) = 14.79, p < 0.001, partial η2 = 0.074, and in samples restricted to completion times of at least 60, 90, and 120 s (all p < 0.001). The same statistical interaction was also observed in a random-intercept Gaussian mixed model, χ2(2) = 31.30, p < 0.001, and an ordinal cumulative-link mixed model, likelihood-ratio χ2(2) = 31.11, p < 0.001. These analyses assessed stability under alternative sample restrictions and model specifications; the design’s causal-identification limits remain unchanged. Disclosure-related interaction estimates varied across model specifications and were not interpreted substantively.
4.5. Exploratory Analyses
Exploratory participant-characteristic analyses yielded no robust associations with overall single-item perceived message credibility after multiplicity adjustment (all adjusted p-values were ≥0.070). Exploratory moderation tests indicated variation in the presentation-configuration pattern by academic level, χ2(4) = 26.39, Holm-adjusted p < 0.001, whereas no other participant-characteristic moderation test survived multiplicity adjustment. Given the exploratory nature of these analyses and the potential influence of classroom- or session-level confounding, this result was not interpreted substantively. Among the 10 post-task cue ratings, perceived objectivity was positively associated with mean credibility across the six reports, ρ = 0.245, Holm-adjusted p = 0.003; all other cue associations failed to survive correction. Because the cue items were collected once after the complete stimulus sequence, these analyses cannot explain the pattern of ratings across the implemented configurations. Full results appear in Supplementary Tables S7–S9 and S13.
5. Discussion
5.1. Principal Findings
This study examined single-item perceived message credibility across six AI-generated news stimuli, organized by three implemented presentation configurations within two selected stimulus domains. Within the design-industry stimuli, the text-only stimulus received the highest estimated credibility rating, and the researcher-designated higher-correspondence image stimulus received the lowest. Within the quantum-computing stimuli, both image-present stimuli were rated above text only, with the researcher-designated lower-correspondence image stimulus receiving the highest estimated rating.
The significant stimulus-domain × presentation-configuration interaction captures this difference in rating patterns between the two domains. The findings concern the complete stimuli; their design precludes attributing the differences independently to visual presentation, image presence, or researcher-designated correspondence.
The secondary disclosure comparison detected no reliable main or two-way interaction terms in this sample. Allocation across four intact classes limits causal inference and prevents treating this finding as evidence of no disclosure effect. Participant-characteristic analyses likewise identified no robust adjusted associations with overall single-item perceived message credibility.
5.2. Interpreting Context-Sensitive Credibility-Rating Patterns
The differing rating patterns are consistent with prior research considering visual and textual features together in judgments of news and online information credibility [18,19,20,21,22,23,24,25]. Here, content context and presentation configuration varied together, placing the empirical observation at the level of complete AI-generated news stimuli. The contribution of any individual feature or process remains unresolved.
HSM and the MAIN model offer interpretive perspectives on these patterns [30,31,32]. They frame visual and textual features as potential credibility-related cues whose relevance depends on the complete stimulus. MAIN directs attention to visual and source-related features, whereas HSM accommodates judgments drawing on multiple available features. These are theoretical interpretations, not tested mechanisms: cue-processing mode, cue diagnosticity, processing effort, and related cognitive mechanisms were not directly measured.
The highest estimated rating for the quantum-computing stimulus with a researcher-designated lower-correspondence image highlights the distinction between researcher classification and participant perception. Without a participant manipulation check, this ordering provides no independent evidence of a correspondence effect. Potential explanations include the particular story, image, serial position, and unmeasured visual properties: realism, esthetic appeal, complexity, novelty, emotional salience, and perceived professionalism. The design leaves their contributions unresolved.
The study’s theoretical contribution is an empirical basis for examining variation in credibility ratings across complete AI-generated news stimuli and content contexts. It motivates future Human–AI and AI-generated news research that considers textual and topical context while independently manipulating image presence and participant-perceived image–text correspondence.
5.3. Practical Implications for AI-Generated News Presentation
The observed differences suggest that evaluating AI-generated news presentations using image presence alone may be insufficient. For these six stimuli, ratings varied across combinations of story content, domain, and presentation configuration. This pattern supports context-specific evaluation of visual assets rather than assuming a uniform credibility response to relevant or less relevant images.
Pre-deployment user testing could examine ratings of specific image–text combinations across content domains and assess whether users interpret images as evidence, decoration, or misleading cues. Separate measures of perceived image–text correspondence and other visual properties would help evaluate generalizability across stories, topics, and interface settings.
Visual presentation also warrants evaluation alongside editorial verification and transparency practices. These address distinct questions: how a message is presented, how its information is verified, and whether users are aware of AI involvement. Their interaction remains untested here, supporting separate examination in future AI-generated news evaluations.
5.4. Limitations and Future Research
First, story identity, implemented presentation configuration, and serial position were not independently crossed. Each configuration accompanied a different fictional story, and all participants viewed the same six stimuli in a fixed order. Sensitivity analyses assessed the stability of the statistical pattern, not independent causal effects of configuration, image presence, or researcher-designated correspondence.
Second, researcher-designated higher correspondence and researcher-designated lower correspondence reflected the team’s assessment without a participant manipulation check. Image realism, esthetics, complexity, salience, professionalism, and other visual properties were not independently assessed. Future research should use pretests of participants’ perceived image–text correspondence, independently manipulate image presence and image–text correspondence, and include multiple stories within each content domain.
Third, only two selected stimulus domains were included, each represented by one topic family. Familiarity, prior knowledge, and disciplinary relevance were unmeasured, so domain comparisons provide no independent evidence about these constructs. Future studies should include multiple topics and domains and measure relevant participant characteristics directly before examining their associations with ratings.
Fourth, single-item perceived message credibility supported within-participant comparison while limiting measurement burden. Its scope excludes validated multidimensional measurement of trust, source credibility, or information quality. The post-task cue indicators were not stimulus-specific and provided no test of processing mechanisms or mediation. Future research could use validated multidimensional credibility measures and directly test the processes proposed by HSM and MAIN.
Fifth, the study and analysis plan were not preregistered; future work should preregister hypotheses, planned comparisons, and analysis procedures. Disclosure was assigned across four intact classes and sessions, warranting individual-level randomization or a larger clustered design in future studies. The Supplementary Materials supportsmaterial-level reproducibility through the six stimuli, available prompts, image-selection criteria, and supplementary analyses. Exact build-level image-model identifiers and unavailable generation parameters could not be reconstructed retrospectively.
6. Conclusions
This exploratory empirical evaluation examined single-item perceived message credibility across six implemented AI-generated news stimuli. The stimulus-domain × presentation-configuration interaction was significant, F(1.94, 404.48) = 14.94, p < 0.001, partial η2 = 0.067, indicating that the observed rating pattern differed between the two selected stimulus domains. These findings suggest that credibility evaluation of AI-generated news should consider visual presentation together with the content context in which it appears.
Because configuration, story identity, and serial position were linked, the findings concern these stimuli rather than isolated causal effects of visual presentation or researcher-designated correspondence. They provide an empirical basis for controlled research on image presence, participant-perceived image–text correspondence, story content, and presentation order.
Supplementary Materials
The following supporting information can be downloaded at: https://www.mdpi.com/article/10.3390/app16199637/s1: Study Materials and Supplementary Analyses (DOCX), comprising study conditions, the available questionnaire and disclosure instructions, the six news texts and images, retained generation prompts and image-selection criteria; Table S1: Participant characteristics by classroom-level AIGC authorship disclosure condition (primary sample, N = 211); Table S2: Descriptive statistics for the 12 study cells (primary sample, N = 211); Table S3: Estimated marginal means for selected stimulus-domain condition × implemented presentation configuration × disclosure (primary sample, N = 211); Table S4: Implemented presentation-configuration comparisons within selected stimulus-domain condition and disclosure classes (primary sample, N = 211); Table S5: Non-invariant-response sensitivity analysis (N = 188); Table S6: Random-intercept linear mixed-effects model robustness test (primary sample, N = 211); Table S7: Exploratory associations between participant characteristics and overall perceived credibility (primary sample, N = 211); Table S8: Exploratory moderation of implemented presentation configuration by participant characteristics (primary sample, N = 211); Table S9: Credibility-evaluation cues: descriptive statistics and disclosure-class comparisons (primary sample, N = 211); Table S10: Residual diagnostic summary for the full-sample linear mixed-effects model; Table S11: Completion-time sensitivity analyses for key effects; Table S12: Ordinal mixed-model sensitivity tests (primary sample, N = 211); Table S13: Exploratory participant-level associations between post-task credibility-evaluation cues and mean perceived message credibility (primary sample, N = 211); Table S14: Complete 2 × 3 × 2 mixed ANOVA of single-item perceived message credibility in the full sample; Figure S1: Participant flow and data-quality screening; Figure S2: Full-sample distributions of perceived message credibility across selected stimulus-domain condition, implemented presentation configuration, and classroom-level AIGC authorship disclosure condition; Figure S3: Exploratory participant-characteristic analyses in the primary sample; Figure S4: Diagnostics for the full-sample random-intercept linear mixed-effects model. Candidate-image counts, independent external ratings, and preregistered numerical selection thresholds were not recorded and are reported as limitations rather than reconstructed retrospectively.
Author Contributions
Conceptualization, C.C. and Y.L.; methodology, C.C., G.L. and Y.L.; validation, C.C.; formal analysis, C.C. and G.L.; investigation, S.Y.; data curation, G.L. and S.Y.; software, G.L.; visualization, S.Y. and Y.L.; writing—original draft preparation, G.L.; writing—review and editing, C.C., G.L. and S.Y.; supervision, C.C. All authors have read and agreed to the published version of the manuscript.
Funding
This research received no external funding.
Institutional Review Board Statement
The study was conducted in accordance with the Declaration of Helsinki and approved by the ethics review committee of Nanjing Forestry University (protocol code 2026041; approval date: 10 February 2026).
Informed Consent Statement
Informed consent was obtained from all participants before participation. Participants indicated consent by voluntarily proceeding after reading the study information on the questionnaire homepage; no separate consent checkbox was used. Participants in the no-disclosure condition were debriefed after questionnaire completion.
Data Availability Statement
The de-identified participant-level data supporting the reported findings are available from the corresponding author upon reasonable request for verification purposes. The data are not publicly deposited because the informed-consent materials did not authorize unrestricted public sharing. Access requests will be considered subject to the approved ethics protocol and applicable privacy restrictions. The analysis code is not publicly available.
Acknowledgments
The authors thank the participants for their time and the members of the research team for their assistance in reviewing and selecting the study materials. During preparation of this work, the authors used Tencent Yuanbao to generate candidate study news texts and images. The image-generation function was accessed between 30 and 31 October 2025; the interface did not provide a verifiable build-level model identifier. Final materials were selected through discussion among the supervising faculty member and research-team members. Selected images were proportionally cropped to remove the platform-generated label; no other content-level image editing was performed. OpenAI Codex assisted with R code development and checking, English translation, language editing, figure preparation, and manuscript formatting. The authors independently verified the statistical results against the source data, made all analytical and interpretive decisions, reviewed and edited all AI-assisted outputs, and take full responsibility for the publication.
Conflicts of Interest
The authors declare no conflicts of interest.
References
- Bartleman, M.; Schapals, A.K.; Dubois, E. Generative AI and the New Landscape of Automated Journalism: A Systematized Review of 185 Studies (2012–2024). Journal. Media 2026, 7, 39. [Google Scholar] [CrossRef] [Scilit]
- Kreps, S.; McCain, R.M.; Brundage, M. All the News That’s Fit to Fabricate: AI-Generated Text as a Tool of Media Misinformation. J. Exp. Political Sci. 2022, 9, 104–117. [Google Scholar] [CrossRef] [Scilit]
- Lee, S.; Nah, S.; Chung, D.S.; Kim, J. Predicting AI News Credibility: Communicative or Social Capital or Both? Commun. Stud. 2020, 71, 428–447. [Google Scholar] [CrossRef] [Scilit]
- Shi, Y.; Sun, L. How Generative AI Is Transforming Journalism: Development, Application and Ethics. Journal. Media 2024, 5, 582–594. [Google Scholar] [CrossRef] [Scilit]
- Hanimann, A.; Heimann, A.; Hellmueller, L.; Trilling, D. Believing in credibility measures: Reviewing credibility measures in media research from 1951 to 2018. Int. J. Commun. 2023, 17, 214–235. [Google Scholar]
- Choi, W.; Bak, H.; An, J.; Zhang, Y.; Stvilia, B. College students’ credibility assessments of GenAI-generated information for academic tasks: An interview study. J. Assoc. Inf. Sci. Technol. 2025, 76, 867–883. [Google Scholar] [CrossRef] [Scilit]
- Hwang, Y.; Jeong, S.H. Generative Artificial Intelligence and Misinformation Acceptance: An Experimental Test of the Effect of Forewarning About Artificial Intelligence Hallucination. Cyberpsychol. Behav. Soc. Netw. 2025, 28, 284–289. [Google Scholar] [CrossRef] [Scilit]
- Lermann Henestrosa, A.; Kimmerle, J. The Effects of Assumed AI vs. Human Authorship on the Perception of a GPT-Generated Text. Journal. Media 2024, 5, 1085–1097. [Google Scholar] [CrossRef] [Scilit]
- Leuppert, R.; Weinmann, C.; Eiden, J. AI-reporters as the future of journalism? Investigating recipients’ credibility evaluations of AI-authored news. Journalism 2025, 27, 2673–2692. [Google Scholar] [CrossRef] [Scilit]
- Ma, H.; Huang, W.; Dennis, A.R. Unintended Consequences of Disclosing Recommendations by Artificial Intelligence versus Humans on True and Fake News Believability and Engagement. J. Manag. Inf. Syst. 2024, 41, 616–644. [Google Scholar] [CrossRef] [Scilit]
- Otis, A. The effects of transparency cues on news source credibility online: An investigation of ‘opinion labels’. Journalism 2024, 25, 198–217. [Google Scholar] [CrossRef] [Scilit]
- Park, C.S.; Molla, M.A.M. Beyond uniform machine heuristics: Multidimensional audience evaluations of AI-labeled news. Journal. Media 2026, 7, 115. [Google Scholar] [CrossRef] [Scilit]
- Tandoc, E.C.; Yao, L.J.; Wu, S. Man vs. Machine? The Impact of Algorithm Authorship on News Credibility. Digit. Journal. 2020, 8, 548–562. [Google Scholar] [CrossRef] [Scilit]
- Waddell, T.F. The Effects of AI Attribution, Source Priming, and Story Topic Polarization on News Credibility. Digit. Journal. 2025, 14, 1–18. [Google Scholar] [CrossRef] [Scilit]
- Wang, S.; Huang, G. The Impact of Machine Authorship on News Audience Perceptions: A Meta-Analysis of Experimental Studies. Commun. Res. 2024, 51, 815–842. [Google Scholar] [CrossRef] [Scilit]
- Mera-Fernández, M.; Moreno-Gil, V.; Morata-Santos, M. AI-Generated Content in Spanish Media: Transparency, New Uses, and Defined Strategies. Journal. Media 2026, 7, 113. [Google Scholar] [CrossRef] [Scilit]
- Yeste Piquer, E.; Suau Martínez, J.; Sintes-Olivella, M.; Xicoy Comas, E. What If I Prefer Robot Journalists? Trust and Objectivity in the AI News Ecosystem. Journal. Media 2025, 6, 51. [Google Scholar] [CrossRef] [Scilit]
- Brantner, C.; Lobinger, K.; Wetzstein, I. Effects of Visual Framing on Emotional Responses and Evaluations of News Stories about the Gaza Conflict 2009. Journal. Mass Commun. Q. 2011, 88, 523–540. [Google Scholar] [CrossRef] [Scilit]
- Greussing, E. Powered by Immersion? Examining Effects of 360-Degree Photography on Knowledge Acquisition and Perceived Message Credibility of Climate Change News. Environ. Commun. 2020, 14, 316–331. [Google Scholar] [CrossRef] [Scilit]
- Shen, C.; Kasra, M.; Pan, W.; Bassett, G.A.; Malloch, Y.; O’Brien, J.F. Fake images: The effects of source, intermediary, and digital media literacy on contextual assessment of image credibility online. New Media Soc. 2019, 21, 438–463. [Google Scholar] [CrossRef] [Scilit]
- Vultee, F.; Burgess, G.S.; Frazier, D.; Mesmer, K. Here’s What to Know About Clickbait: Effects of Image, Headline and Editing on Audience Attitudes. Journal. Pract. 2022, 16, 1–18. [Google Scholar] [CrossRef] [Scilit]
- Weikmann, T.; Egelhofer, J.L.; Lecheler, S. Beyond Credibility: The Effects of Different Forms of Visual Disinformation. Journal. Mass Commun. Q. 2025, 102, 1020–1043. [Google Scholar] [CrossRef] [Scilit]
- Lago, F.; Phan, Q.T.; Boato, G. Visual and Textual Analysis for Image Trustworthiness Assessment within Online News. Secur. Commun. Netw. 2019, 2019, 9236910. [Google Scholar] [CrossRef] [Scilit]
- Choudhary, M.; Chouhan, S.S.; Rathore, S.S. Beyond Text: Multimodal Credibility Assessment Approaches for Online User-Generated Content. ACM Trans. Intell. Syst. Technol. 2024, 15, 1–33. [Google Scholar] [CrossRef] [Scilit]
- Wobbrock, J.O.; Hattatoglu, L.; Hsu, A.K.; Burger, M.A.; Magee, M.J. The Goldilocks zone: Young adults’ credibility perceptions of online news articles based on visual appearance. New Rev. Hypermedia Multimed. 2021, 27, 51–96. [Google Scholar] [CrossRef] [Scilit]
- Kwasniewicz, L.; Wojcik, G.M.; Schneider, P.; Kawiak, A.; Wierzbicki, A. What to Believe? Impact of Knowledge and Message Length on Neural Activity in Message Credibility Evaluation. Front. Hum. Neurosci. 2021, 15, 659243. [Google Scholar] [CrossRef] [Scilit]
- Murphy, K.M. Fake News and the Web of Plausibility. Soc. Media Soc. 2023, 9, 20563051231170606. [Google Scholar] [CrossRef] [Scilit]
- Ou, M.; Ho, S.S. Does knowledge make a difference? Understanding how the lay public and experts assess the credibility of information on novel foods. Public Underst. Sci. 2024, 33, 241–259. [Google Scholar] [CrossRef] [Scilit]
- Wertgen, A.G.; Richter, T. Source credibility modulates the validation of implausible information. Mem. Cogn. 2020, 48, 1359–1375. [Google Scholar] [CrossRef] [Scilit]
- Chaiken, S. Heuristic versus systematic information processing and the use of source versus message cues in persuasion. J. Personal. Soc. Psychol. 1980, 39, 752–766. [Google Scholar] [CrossRef]
- Chen, S.; Chaiken, S. The heuristic–systematic model in its broader context. In Dual-Process Theories in Social Psychology; Chaiken, S., Trope, Y., Eds.; Guilford Press: New York, NY, USA, 1999; pp. 73–96. [Google Scholar]
- Sundar, S.S. The MAIN model: A heuristic approach to understanding technology effects on credibility. In Digital Media, Youth, and Credibility; Metzger, M.J., Flanagin, A.J., Eds.; MIT Press: Cambridge, MA, USA, 2008; pp. 73–100. [Google Scholar]
- Cao, R.; Hashim, N.B.; Abdul Rahman, S.N. How heuristic credibility cues shape perceived credibility on social media: A meta-analysis of experimental research. Behav. Sci. 2026, 16, 1184. [Google Scholar] [CrossRef] [Scilit]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.

