Skip to Content
  • Article
  • Open Access

27 July 2026

Are Scientists Starting to Sound like ChatGPT? Stylistic Diffusion in Biomedical Abstracts in the LLM Era

and
1
School of Computing, Eastern Institute of Technology, Hawke’s Bay, Napier 4110, New Zealand
2
School of Information and Communication, Nankai University, Tianjin 300071, China
*
Author to whom correspondence should be addressed.
This article belongs to the Section Learning

Abstract

Large language models are part of everyday scholarly writing, but their influence may not always appear as obvious AI authorship. More often, it may appear through gradual changes in academic style. This study examines whether biomedical abstracts show measurable signs of LLM-associated “same-ification” after the public release of ChatGPT. Instead of classifying individual abstracts as AI-written, our study analyses population-level changes across 28,415 PubMed abstracts published between 2018 and 2025. A strict dictionary of LLM-associated stylistic markers was used alongside a control vocabulary, lexical diversity, semantic similarity, interrupted time-series modelling, journal-level heterogeneity analysis, and a study-specific validation framework called SAME-ML. The findings show a clear post-2022 stylistic shift. Mean strict marker use increased from 4.995 to 11.658 markers per 1000 words, representing a 133.4% rise, while the control vocabulary slightly declined. This increase was strongest in 2024 and 2025 and was mainly driven by markers linked to value framing, contribution signalling, connective phrasing, improvement rhetoric, and problem-solution positioning. SAME-ML further showed that the post-ChatGPT period contained a learnable stylistic signal that remained visible after topic-masked and matched-sample validation. Overall, the findings show a corpus-level rhetorical shift and provide a reproducible approach for tracking stylistic change in scholarly communication over time.

1. Introduction

When the world changes, the words people use change with it. Such changes often leave detectable fingerprints in the frequency distributions of written language [1]. Few changes have affected word usage as abruptly as the public release of ChatGPT in late 2022, and few have been so quietly pervasive. Within a single year, keyword-based estimates suggested that more than 60,000 scholarly articles, more than one percent of the literature, already bore the lexical signature of large language models (LLMs) [2], and in biomedicine the figure was far higher: at least 13.5% of 2024 PubMed abstracts showed evidence of LLM processing, an impact exceeding even that of COVID-19 on scientific vocabulary [1]. The intuitive response has been to police this influence at the level of the individual author, asking whether a given text was machine-written or human-written. Yet detection is an unstable foundation. State-of-the-art classifiers can be evaded by simple paraphrasing and, more troublingly, systematically misclassify the prose of non-native English writers as AI-generated, penalising precisely those whom the tools could most assist [3]. This suggests a more interesting possibility than machine-versus-human authorship. If the telltale words of LLMs are diffusing through the literature even where no model touched the keyboard, then the bigger change may not be that some humans are letting machines write for them but that humans are beginning to write like machines. The convergence is occurring not in what science says but in the increasingly uniform rhetoric of how science is presented.
LLMs have moved quickly from experimental natural language processing (NLP) systems into everyday writing infrastructure, supported by the scaling of GPT-3, the public launch of ChatGPT, and the later development of GPT-4 [4,5,6]. Since late 2022, these tools have become visible across scientific work because they can draft, revise, summarise, translate, and polish prose with a level of fluency that makes them attractive to researchers and students [5,7,8]. Academic writing is not only a container for results, it is also a rhetorical system through which authors signal novelty, caution, contribution, and disciplinary belonging [7,8]. Biomedical abstracts are especially important in this debate because they are short, highly searchable, widely indexed, and often used by readers and automated systems to make rapid judgements about a study [1,9].
Existing work shows that generative AI (GenAI) is already influencing science, biomedicine, and scholarly publishing, although its benefits and risks vary across tasks and contexts [7,8,9,10,11]. Editorial and publishing organisations have responded by emphasising transparency, responsibility, and the need to prevent AI tools from being treated as authors because such systems cannot take responsibility for the integrity of a manuscript [11,12,13,14]. However, disclosure policies mainly address declared use, authorship responsibility, and editorial governance rather than the quieter possibility that LLMs are reshaping the style of academic language itself [10,11,12,13,14]. At the same time, AI-text detection remains fragile, with evidence showing that detectors can fail under paraphrasing, vary across tools, and unfairly misclassify non-native English writing as AI-generated [3,15,16,17].
The strongest corpus-level evidence to date comes from studies that track LLM-associated vocabulary rather than individual authorship [1,2]. Kobak et al. showed that PubMed abstracts experienced an abrupt post-ChatGPT increase in style-heavy words and estimated that at least 13.5% of 2024 biomedical abstracts showed evidence of LLM processing [1]. Gray similarly used indicative keywords to estimate the prevalence of LLM-assisted writing across scholarly literature, showing that such signals can be measured at scale even when individual disclosure is unavailable [2]. These studies are valuable, but they leave a further question unresolved: whether the observed vocabulary changes are only traces of possible LLM use or whether they reflect a broader diffusion of rhetorical habits across academic writing [1,2].
Our study addresses that gap by reframing post-ChatGPT language change as LLM-associated same-ification, understood as a corpus-level movement toward repeated rhetorical and stylistic patterns [1,2,18]. Same-ification is important because LLM influence may appear not only through fully generated text but also through repeated patterns of polishing, connective phrasing, value framing, contribution signalling, and improvement-oriented rhetoric [1,7,8].
Accordingly, the primary aim of this study was to determine whether PubMed biomedical abstracts published after the public release of ChatGPT show a measurable population-level increase in LLM-associated stylistic markers. We compared the pre-ChatGPT (2018–2022) and post-ChatGPT (2023–2025) periods, tested whether changes exceeded those in a biomedical control vocabulary, and evaluated the robustness of the signal using distributional, temporal, journal-level, topic-controlled, matched-sample, and explainable machine-learning analyses.
To address this aim, the study examined the following three research questions:
  • RQ1. Did strict LLM-associated stylistic markers increase in PubMed biomedical abstracts after the public release of ChatGPT?
  • RQ2. Is any observed increase stronger than changes in general biomedical and academic control vocabulary, suggesting a stylistic rather than purely topical shift?
  • RQ3. Does the same-ification signal remain visible across distributional, temporal, topic-controlled, matched-sample, and explainable machine-learning validation analyses?
By focusing on corpus-level rhetorical convergence, this study reframes how LLM influence in scholarly writing can be examined. Existing work has mainly treated LLM-associated vocabulary as a signal for estimating possible AI use, but this paper examines whether such vocabulary reflects a broader diffusion of academic style. This study advances this argument through a reproducible measurement pipeline that combines strict stylistic markers, control vocabulary, interrupted time-series modelling, lexical diversity, semantic similarity, and journal-level heterogeneity analysis. It also extends corpus-based evidence with SAME-ML, a style-aware, matched, and explainable machine-learning validation framework that tests whether the observed signal is learnable, robust to topic composition, and interpretable at the feature level.

3. Methods

3.1. Methodological Positioning and Novelty

This study investigates whether public biomedical academic writing has begun to show measurable signs of LLM-associated stylistic diffusion after the public release of ChatGPT. Existing studies have mostly focused on detecting whether academic texts were generated or processed by large language models, often by identifying “excess vocabulary” or estimating possible LLM usage in published writing. For example, existing excess-vocabulary work on PubMed abstracts showed that many words increased abruptly after the emergence of ChatGPT-like LLMs, with the strongest post-2023 changes appearing in stylistic rather than purely content-related vocabulary [1].
The present study takes a different direction. It does not attempt to classify individual abstracts as AI-written or human-written. Instead, it asks whether biomedical academic writing, as a public corpus, is showing signs of same-ification, understood here as a population-level movement toward repeated lexical and rhetorical patterns associated with LLM-assisted writing. This distinction is important because the influence of LLMs may not only appear through fully generated text. It may also appear more subtly, through the gradual normalisation of particular stylistic habits, such as more frequent use of polished, evaluative, connective, and emphasis-building expressions.
The novelty of our study, therefore, lies in reframing post-ChatGPT vocabulary change as a broader phenomenon of stylistic diffusion. Rather than treating LLM-associated vocabulary only as a detection signal, this study uses it as evidence of possible language change in academic writing. In doing so, the study contributes a reproducible method for measuring how public academic language may be shifting in the LLM era.

3.2. Study Design

The study adopted a retrospective computational corpus design. PubMed abstracts published between 2018 and 2025 were used to examine changes in biomedical academic writing before and after the public release of ChatGPT in November 2022. The period from 2018 to 2022 was treated as the pre-ChatGPT baseline, while 2023 to 2025 was treated as the post-ChatGPT period. This coding is used consistently throughout the analyses: pre-ChatGPT observations satisfy y i ≤ 2022, and post-ChatGPT observations satisfy y i ≥ 2023.
This design was chosen because ChatGPT’s public release provides a defined historical boundary for examining population-level language change. The design follows the logic of interrupted time-series research, where a clearly identifiable event is used to evaluate whether an outcome departs from its earlier baseline through a level or slope change [19]. It is used as a historically meaningful exposure point after which LLM-based writing support became widely accessible through authoring, editing, and revision workflows. The unit of analysis was the individual abstract, with document-level metadata, including publication year, journal name, title, language information, word count, marker frequencies, control vocabulary frequencies, lexical diversity, and similarity-related features. The methodological workflow is summarised in Figure 1.
Figure 1. Methodological workflow for corpus construction, text processing, comparative analysis, and output generation.

3.3. Data Source and Analytical Dataset Construction

The dataset was constructed from PubMed-indexed biomedical abstracts. PubMed was selected because it provides a large, structured, publicly accessible, and time-stamped corpus of biomedical academic writing. Records were retrieved using the National Center for Biotechnology Information’s Entrez Programming Utilities, which provide programmatic access to PubMed and related NCBI databases [25]. For each year from 2018 to 2025, PubMed records were retrieved using publication-date filters. The extracted fields included PubMed ID, publication year, journal title, article title, language metadata, publication type, and abstract text. The abstract was selected as the main textual unit because abstracts are highly visible, relatively standardised, and available across a wide range of biomedical journals. This makes them suitable for comparing language patterns over time.
Records were excluded if they did not contain an abstract, were not in English, had fewer than 80 words after preprocessing, or appeared to represent non-standard publication formats. Titles containing terms such as “erratum”, “corrigendum”, “correction”, “retraction”, “withdrawn”, or “expression of concern” were removed to reduce contamination from correction notices and editorial records. After filtering, the final analytic corpus contained 28,415 abstracts, with 20,236 in the pre-ChatGPT period and 8179 in the post-ChatGPT period.
The corpus should be read as a structured analytic sample rather than as a complete census of PubMed. PubMed retrieval was capped at 5000 records per publication year, and records were retained only when the XML extraction procedure returned an eligible English abstract. The 2023 sample was smaller than the surrounding years because the extraction step yielded fewer usable abstracts for that year. For this reason, the manuscript reports 2023 explicitly, interprets it as a transitional low-yield year, and places greater evidential weight on the sustained 2024–2025 pattern rather than on a single-year discontinuity.

3.4. Text Preprocessing

All text was processed using a reproducible script-based workflow. Abstracts were converted to lowercase, non-alphabetic characters were removed, repeated whitespace was standardised, and each abstract was tokenised into word units. Word counts were calculated for every abstract and used to normalise all marker frequencies. Normalisation was necessary because abstracts differ in length. Without normalisation, longer abstracts would be more likely to contain stylistic markers simply because they contain more words. Therefore, the primary outcome was expressed as the number of LLM-associated stylistic markers per 1000 abstract words. Word-level summaries were also reported as occurrences per million words to support interpretation of individual marker trajectories.

3.5. LLM-Associated Stylistic Marker Dictionary

The main analytical measure was a strict dictionary of LLM-associated stylistic markers. These markers were defined as words that operate primarily as rhetorical, evaluative, connective, or emphasis-building language rather than as biomedical content terms. Examples include words such as delve, intricate, meticulous, pivotal, underscore, comprehensive, crucial, enhance, notably, nuanced, seamless, and transformative. The candidate marker pool was derived from the publicly released excess-vocabulary list accompanying Kobak et al. [1]. Rather than adopting the full list directly, this study refined it into a strict style-focused dictionary by retaining words associated with rhetorical, evaluative, connective, and emphasis-building language while excluding biomedical content terms and generic academic vocabulary. This was an important methodological step because broad excess-vocabulary lists may include content-heavy or generic academic words. Such terms may reflect topic shifts, journal composition, or changes in biomedical publishing rather than stylistic diffusion. The measurement logic for applying this dictionary is formalised in Section 3.7.
To improve construct validity, the strict marker set excluded generic and content-adjacent words, such as “this”, “these”, “based”, “research”, “including”, “clinical”, “treatment”, “patient”, “disease”, “analysis”, “model”, “data”, and “results”. The retained markers were selected because they more clearly reflected style, framing, emphasis, or academic polish. This strict-style version was used as the main analysis. Broader vocabulary patterns were treated only as exploratory or sensitivity evidence. The dictionary was, therefore, designed as a high-specificity measure of rhetorical style rather than a high-recall measure of all possible LLM-associated vocabulary. No marker is unique to ChatGPT; inference, therefore, rests on aggregate frequency changes across the corpus. Hence, our study examines whether the collective frequency of a restricted stylistic marker set changed after 2022 in a way that is stronger than ordinary vocabulary drift.

3.6. Control Vocabulary

A control vocabulary was included to test whether the observed pattern was specific to LLM-associated stylistic markers or simply reflected a general change in biomedical academic language. The control vocabulary contained common biomedical and academic terms, such as “patient”, “clinical”, “treatment”, “disease”, “study”, “data”, “analysis”, “method”, “results”, “risk”, and “association”. For each abstract, both the stylistic marker count and the control-word count were calculated and normalised per 1000 words. This allowed the study to compare two trends, one representing LLM-associated stylistic markers and one representing general biomedical academic vocabulary. If the stylistic-marker rate increased sharply after 2022 while the control vocabulary did not follow the same trajectory, this would strengthen the interpretation that the change reflects stylistic diffusion rather than general vocabulary growth.

3.7. Measuring LLM-Associated Same-Ification

To make the measurement procedure transparent and reproducible, same-ification was operationalised as a document-level and year-level scoring process. This process quantified whether the analytical dataset showed increased use of LLM-associated stylistic markers after the public release of ChatGPT. The algorithm, therefore, treats each abstract as an observation and estimates the density of stylistic markers relative to the document length and to a control vocabulary. To clarify how the different analytical components contribute to the measurement of LLM-associated same-ification, the study used a multi-level evidence framework, shown in Figure 2. The framework integrates the following three complementary evidence lanes: document-level evidence based on marker rate, lexical diversity, and semantic similarity; marker-level evidence based on stylistic marker growth; and journal-level evidence based on venue-specific diffusion intensity and heterogeneity.
Figure 2. Multi-level framework for detecting LLM-associated same-ification.
Let D denote the analytical dataset of PubMed abstracts, where each document, di, has a publication year, yi, a token sequence, Ti, and a word count, Ni = |Ti|. Let Ls denote the strict LLM-associated stylistic marker dictionary, and Lc denote the control vocabulary. For each abstract, the stylistic marker count, Mi, and control vocabulary count, Ci, were calculated as follows:
M i = w T i 1 w L s
C i = w T i 1 w L c
where 1(·) is an indicator function that equals 1 when the token appears in the relevant dictionary and 0 otherwise. To control for variation in abstract length, both counts were normalised per 1000 words, as follows:
R i s = M i N i × 1000               R i c = C i N i × 1000
where R i s is the LLM-associated stylistic-marker rate, and R i c is the control-vocabulary rate for abstract i. Each abstract was then assigned to either the pre-ChatGPT period or the post-ChatGPT period using the publication year, as follows:
P i = 0 , y i 2022 1 , y i 2023
At the word level, each marker, w, was evaluated by comparing its pre-ChatGPT and post-ChatGPT frequencies per million words. For each period, p, the frequency of marker w was calculated as follows:
F w , p = count w , p i p N i × 10 6
The absolute post-ChatGPT change and relative post/pre ratio were then calculated as follows:
Δ w = F w , post F w , pre
ρ w = F w , post + ϵ F w , pre + ϵ
where ε = 0.01 was added to avoid division by zero for markers that were absent or extremely rare in the pre-ChatGPT period. These word-level measures were used to identify which stylistic markers contributed most strongly to the observed post-2022 shift.
An exploratory same-ification score was also calculated to combine multiple indicators of convergence. Let Z(·) denote standardisation using the analytical dataset’s mean and standard deviation. The exploratory score for abstract i was defined as follows:
S i = Z R i s + Z Q i Z L D i
where R i s is the stylistic-marker rate, Q i is the semantic-similarity feature, and L D i is the lexical-diversity score. Higher values of S i indicate stronger evidence of same-ification only in the narrow operational sense of higher stylistic marker density, greater document-level similarity, and lower lexical diversity. The composite index was not treated as the main outcome because its value depends on how similarity and lexical diversity are operationalised. The primary inference, therefore, rests on the strict marker trajectory and its contrast with the control vocabulary; Q i and L D i are used as supplementary diagnostics for RQ3. Algorithm 1 summarises the complete analytical procedure used to measure LLM-associated stylistic diffusion and same-ification.
Algorithm 1. LLM-Associated Stylistic Diffusion and Same-ification Measurement
Input: PubMed abstract dataset, D; strict stylistic marker dictionary, Ls; control vocabulary, Lc; ChatGPT boundary year, Y0 = 2022; minimum abstract length, Nmin = 80.
Output: Document-level marker table M; yearly trend table Ty; marker-level change table W; exploratory same-ification score table S.
1Initialise M ← ∅
2for each abstract di in D do
3   Extract PMID, publication year yi, journal name, and abstract text
4   textiLOWERCASE(di.abstract)
5   textiREMOVE_NON_ALPHABETIC_CHARACTERS(texti)
6   TiTOKENISE(texti)
7   NiCOUNT(Ti)
8   if Ni < Nmin then
9   Exclude di from further analysis
10   continue
11   end if
12   marker_countiCOUNT(wTi such that wLs)
13   control_countiCOUNT(wTi such that wLc)
14   marker_ratei ← (marker_counti/Ni) × 1000
15   control_ratei ← (control_counti/Ni) × 1000
16   if yi > Y0 then
17   periodi ← Post-ChatGPT
18   else
19   periodi ← Pre-ChatGPT
20   end if
21lexical_diversityiMTLDi
22   for each direction d ∈ {→, ←} do
23   Initialise a new token segment
24    for each token position k in the current segment do
25    TTRid(k) ← COUNT_UNIQUE_TYPES(current segment after k tokens)/k
26            if TTRid(k) ≤ τ, where τ = 0.72, then
27            Record one complete factor
28                Begin a new token segmen
29            end if
30    end for
31    Fid ← fid + (1 − TTRi,finald)/(1 − τ)
32   end for
33MTLDi ← ½[(Ni/Fi→) + (Ni/Fi←)]
34semantic_similarityi ← Qγi
35    for each publication year y do
36            Aγ ← RANDOM_SAMPLE(abstracts published in year y,
37                 size = min(250, nγ), seed = 42)
38            for each abstract j ∈ Aγ do
39           ej ← SBERT(textj, model = “all-MiniLM-L6-v2”)
40            end for
41            Qγ ← 2/[mγ(mγ − 1)] × Σa<β, a,b ∈ Aγ [(eaTeβ)/(‖ea2‖eβ2)]
42         end for
43        semantic_similarityi ← Qγi
44   Append to M: {pmidi, yi, journali, Ni, marker_counti, control_counti, marker_ratei, control_ratei, lexical_diversityi, semantic_similarityi, periodi}
45end for
46Initialise Ty ← ∅
47for each publication year y represented in M do
48   marker_meanyMEAN(marker_ratei where yi = y)
49   control_meanyMEAN(control_ratei where yi = y)
50   marker_CIy ← COMPUTE_95CI(marker_ratei where yi = y)
51   control_CIy ← COMPUTE_95CI(control_ratei where yi = y)
52   Append to Ty: {y, marker_meany, control_meany, marker_CIy, control_CIy}
53end for
54Initialise W ← ∅
55for each stylistic marker w in Ls do
56   pre_ratewFREQUENCY_PER_MILLION(w in Pre-ChatGPT abstracts)
57   post_ratewFREQUENCY_PER_MILLION(w in Post-ChatGPT abstracts)
58   absolute_changewpost_ratewpre_ratew
59   relative_ratiow ← (post_ratew + 0.01)/(pre_ratew + 0.01)
60   Append to W: {w, pre_ratew, post_ratew, absolute_changew, relative_ratiow}
61end for
62Initialise S ← ∅
63for each retained abstract di in M do
64   z_markeriZSCORE(marker_ratei across all retained abstracts)
65   z_semanticiZSCORE(semantic_similarityi across all retained abstracts)
66   z_diversityiZSCORE(lexical_diversityi across all retained abstracts)
67   sameificationiz_markeri + z_semanticiz_diversityi
68   Append to S: {pmidi, sameificationi, z_markeri, z_semantici, z_diversityi}
69end for
70return M, Ty, W, S
Note. The same-ification score is used as an exploratory composite indicator, not as a direct detector of AI authorship. Higher values indicate abstracts with a stronger LLM-associated stylistic marker concentration, greater semantic similarity to the wider retained corpus, and lower lexical diversity.

3.8. Interpretable SAME-ML Validation Framework

This study proposes Style-Aware, Matched, and Explainable Machine Learning (SAME-ML) as a study-specific validation framework. SAME-ML is not a new classification algorithm. It is an integrated analytical pipeline designed to test whether a corpus-level stylistic shift is learnable, robust to alternative explanations, and interpretable. “Style-aware” refers to the use of document-level lexical, rhetorical, marker, control-vocabulary, and length features rather than individual AI-authorship labels. “Matched” refers to the comparison of post-ChatGPT abstracts with comparable pre-ChatGPT abstracts within text-derived biomedical topics. “Explainable” refers to the use of feature importance and SHAP values to identify the features contributing to period classification. The framework combines style-only classification, topic-adjusted validation, matched-sample comparison, and marker-ablation analysis.
The target was the publication period, coded as pre-ChatGPT for abstracts published from 2018 to 2022 and post-ChatGPT for abstracts published from 2023 to 2025. The predictors were derived from the same document-level feature table used in the computational corpus analysis. The style-feature set included the operational marker rate per 1000 words, control-vocabulary rate per 1000 words, marker-control ratio, marker-minus-control contrast, MTLD lexical diversity, log-transformed word count, and rhetorical-group indicators. Year-level semantic similarity and the composite same-ification index were excluded because they could encode temporal information directly. The rhetorical indicators represented value and possibility framing, scope and connective framing, evidence-and-claim signalling, improvement and integration rhetoric, problem-solution orientation, and overall rhetorical density. The strict-style corpus results remained the primary descriptive evidence. SAME-ML was used only as a validation layer to test whether the broad stylistic shift was learnable after topic, length, control-vocabulary, and matching adjustments.
The composite same-ification index was retained as an exploratory descriptive measure, as presented below in the Section 4, because it summarises the temporal patterns of marker density, semantic similarity, and lexical diversity. However, it was excluded from the SAME-ML predictors, matching variables, and matched outcomes because its year-level semantic-similarity component could encode publication period and thereby compromise the independence of the validation analyses.
The first validation experiment (E1) used style-only supervised learning. LASSO (glmnet version 5.0) logistic regression, random forest, and XGBoost (version 3.2.1.1.) classifiers were trained using only style and same-ification features, with no full-text topical vocabulary supplied to the models. This experiment asked whether the post-ChatGPT period could be ranked or separated using rhetorical features alone. The analysis, therefore, treated supervised learning as a diagnostic validation tool rather than as a detector of AI-generated writing.
The second experiment (E2) introduced topic-masked validation. Abstracts were converted into TF-IDF representations and reduced using latent semantic analysis before being grouped into ten topic clusters. Model performance was then compared under the following three conditions: a topic-only text model, a style-plus-topic model, and a topic-residualised style model. In the residualised condition, each style feature was adjusted for topic cluster and log-transformed word count before classification. This design was used to test whether the stylistic signal persisted after reducing the possibility that the post-ChatGPT pattern was mainly a consequence of changing biomedical subject matter.
The third experiment (E3) used a matched pre/post comparison. Post-ChatGPT abstracts were matched one-to-one, without replacement, with comparable pre-ChatGPT abstracts using nearest-neighbour propensity-score matching. The propensity model included log-transformed word count, control-vocabulary rate, lexical diversity, and journal grouping, while matching was constrained to occur within the same text-derived biomedical topic. Covariate balance was assessed using absolute standardised mean differences before and after matching. The matched pairs were then compared in terms of stylistic marker density, marker-control ratio, overall rhetorical density, and rhetorical-function rates. Year-level semantic similarity and the composite same-ification index were excluded from the matching design. This experiment was interpreted as a design-based robustness check rather than as a causal estimate of ChatGPT exposure.
The fourth experiment (E4) used marker-ablation analysis. Text-based XGBoost models were trained under the following four conditions: full abstract text, abstract text with stylistic markers removed, stylistic markers only, and control vocabulary only. This experiment tested where the period signal lived in the text. If performance dropped after removing stylistic markers, this would suggest that the LLM-associated marker layer contributed directly to the post-ChatGPT signal. If performance remained strong, it would suggest that the period signal was distributed more broadly across the lexical and rhetorical surface of the abstracts.
Model performance was evaluated on held-out test data using ROC AUC, PR AUC, accuracy, sensitivity, specificity, and F1 score. Because the post-ChatGPT class was smaller than the pre-ChatGPT class, ROC AUC and PR AUC were treated as the main ranking-based indicators, while threshold-dependent measures were interpreted cautiously. Complete or near-complete separation in any held-out experiment was interpreted as evidence of separability within this corpus and feature design, not as evidence of deployment-ready classification or individual AI authorship. Model interpretability was added through XGBoost gain-based feature importance and tree SHAP values [24]. These interpretability outputs helped identify whether the period signal was driven mainly by biomedical content, general vocabulary, semantic similarity, or rhetorical same-ification features.

4. Results

The results are organised around the three research questions rather than around separate analytical layers. RQ1 is answered using aggregate, yearly, distributional, and interrupted time-series evidence. RQ2 is answered by comparing strict stylistic markers with the control vocabulary and by examining the rhetorical functions of the fastest-growing markers. RQ3 is answered through secondary same-ification indicators, journal-level heterogeneity, topic-masked modelling, matched-sample comparison, marker-ablation, and interpretable SAME-ML analyses. This structure keeps the primary claim tied to marker/control evidence while treating semantic similarity, composite scores, and machine-learning outputs as validation rather than as standalone proof of LLM authorship.

4.1. RQ1: Post-ChatGPT Increase in Strict Stylistic Markers

RQ1 asked whether strict LLM-associated stylistic markers increased in PubMed biomedical abstracts after the public release of ChatGPT. The aggregate evidence answers this question affirmatively. Figure 3 shows that the strict marker rate was relatively stable before the ChatGPT boundary, remaining close to five markers per 1000 abstract words between 2018 and 2022. After the boundary, the trajectory changed substantially: The mean strict marker rate increased from 5.965 in 2023 to 10.595 in 2024 and 13.263 in 2025. This increase is interpreted as a population-level shift in rhetorical style, not as evidence that any individual abstract was written by an LLM.
Figure 3. Yearly trend in strict LLM-associated stylistic-marker rates compared with the control vocabulary. The dashed vertical line marks the ChatGPT boundary, and the ribbon shows the 95% confidence interval.
This divergence is central to the interpretation of the results. If both the strict marker vocabulary and the control vocabulary had increased in parallel, the post-2022 pattern could be explained as a broad change in biomedical abstract language, abstract length, or publication composition. Instead, the sharp growth is concentrated in the style-only marker set. The comparison, therefore, strengthens the claim that the post-2022 shift is stylistic rather than simply a by-product of general vocabulary growth.
The indexed comparison in Figure 4 makes the imbalance clearer. When 2022 is treated as the baseline, LLM-associated stylistic markers rise to more than 250% of their 2022 value by 2025. The control vocabulary remains near the baseline and falls below it in 2025. Lexical diversity also increases but at a much slower rate than the stylistic marker index. This pattern suggests that post-ChatGPT abstracts are not simply becoming less varied or more lexically impoverished. Rather, they appear to be accumulating a recognisable stylistic layer while retaining, and in some respects increasing, lexical variety.
Figure 4. Relative post-2022 change in strict stylistic markers, control vocabulary, and lexical diversity. Values are indexed to 2022 = 100 to enable comparison across differently scaled measures.
Table 1 shows that the strongest between-period difference occurred in strict stylistic marker use, with a large standardised effect (g = 0.80). Lexical diversity also increased moderately (g = 0.54), while the increase in abstract length was small (g = 0.20). Although the control-vocabulary difference was statistically significant, its effect size was negligible (g = −0.04). This contrast supports the interpretation that the post-2022 shift was concentrated in stylistic markers rather than representing uniform growth across general biomedical vocabulary.
Table 1. Descriptive and inferential comparison of pre- and post-ChatGPT abstracts.
The yearly estimates in Table 2 show that the shift was not a gradual continuation of the pre-ChatGPT baseline. Between 2018 and 2022, the strict marker rate fluctuated within a narrow band, from 4.575 to 5.383 per 1000 words. The substantial increase appears after the boundary, especially in 2024 and 2025. The 2025 mean of 13.263 is approximately 2.54 times the 2022 value. This time pattern is consistent with diffusion rather than immediate replacement: ChatGPT’s release marks the historical boundary, but the larger linguistic movement becomes more visible after a period of adoption, normalisation, and possible integration into authoring, editing, and revision workflows. The stronger evidence for RQ1, therefore, comes from the stable pre-2022 baseline and the sustained 2024–2025 increase, not from the 2023 point alone.
Table 2. Yearly strict marker and control-vocabulary rates, 2018–2025.
The distributional analysis extends the aggregate findings by showing that the post-ChatGPT shift is visible across the spread of abstract-level marker rates. As shown in Figure 5, the post-ChatGPT distribution has a higher centre and a longer upper tail, indicating that the post-2022 period contains more abstracts with comparatively high concentrations of strict stylistic markers. Strict stylistic-marker rates were substantially higher in the post-ChatGPT period (M = 11.658, SD = 11.785) than in the pre-ChatGPT period (M = 4.995, SD = 6.375). The mean difference was 6.663 markers per 1000 words, 95% CI [6.393, 6.933], Welch’s t(10,169) = 48.35, p < 0.001, Hedges’ g = 0.80, indicating a large standardised difference.
Figure 5. Distributional shift in strict LLM-associated stylistic marker use before and after the ChatGPT boundary. Violin plots show the distribution, boxplots show the median and interquartile range, and white-centred points show group means.
The shape of the post-ChatGPT distribution is important because it points to uneven diffusion. Some post-ChatGPT abstracts remain low in stylistic marker density, while others show much higher saturation. This makes the finding more realistic than a claim of universal change. The evidence suggests that LLM-associated phrasing is spreading unevenly across biomedical writing, with particular documents, venues, or subfields adopting the style more strongly than others.
The interrupted time-series model tested whether the post-2022 pattern differed from the pre-existing trend. Table 3 summarises the magnitude of the shift, while Table 4 estimates level and slope changes. The strict marker rate increased by 133.4%, whereas the control-vocabulary rate decreased slightly. The pre-ChatGPT time trend was small and not statistically significant (estimate = 0.082, p = 0.454), indicating no clear year-on-year increase in marker density before the boundary. The post-ChatGPT level shift was positive and statistically significant (estimate = 3.689, p = 0.005), and the post-ChatGPT slope change was stronger (estimate = 5.064, p < 0.001). These results directly support RQ1 by showing both a higher post-boundary level and a steeper post-boundary trajectory.
Table 3. Descriptive post-ChatGPT shift in strict stylistic marker use and comparator measures.
Table 4. Segmented regression model predicting strict LLM-associated stylistic-marker rate per 1000 words.
The model, therefore, supports an acceleration interpretation. The post-2022 period was not only higher than the pre-ChatGPT baseline, but the rate of increase also became steeper. This distinction matters for this paper’s contribution. A level shift alone would suggest a one-off vocabulary jump after the release of ChatGPT. A slope change suggests a diffusion process, where stylistic markers become increasingly normalised over time. The positive association with the control vocabulary is small, while word count has a negative coefficient, indicating that the modelled post-ChatGPT increase cannot be reduced to longer abstracts or a broad increase in common biomedical vocabulary.

4.2. RQ2: Specificity of the Shift Relative to Control Vocabulary and Rhetorical Function

RQ2 asked whether the increase was stronger than changes in general biomedical and academic control vocabulary. The marker-level results support that interpretation by showing that the largest contributors to the post-ChatGPT rise are not primarily biomedical content terms. Figure 6 and Table 5 show that the largest absolute increases occurred for markers such as “potential”, “across”, “while”, “strategies”, “challenges”, “particularly”, “demonstrated”, “enhance”, “comprehensive”, and “exhibited”. These words operate as framing, connective, evidential, or emphasis-building language. They, therefore, help explain why the aggregate pattern is better read as a stylistic shift than as a change in biomedical subject matter.
Figure 6. Strict stylistic markers with the largest post-ChatGPT absolute increases. The x-axis and horizontal line length represent the absolute increase in occurrences per million words, calculated as the post-ChatGPT rate minus the pre-ChatGPT rate. Circle size and colour both represent the post/pre frequency ratio: larger and darker-purple circles indicate stronger proportional growth from the pre-ChatGPT baseline, whereas smaller and yellow circles indicate lower proportional growth.
Table 5. Strict stylistic markers with the largest post-ChatGPT increases.
The largest absolute growth is not limited to one rhetorical family. “Potential”, “promising”, “crucial”, and “comprehensive” intensify evaluative or possibility-oriented framing. “Across”, “while”, “particularly”, and “additionally” help link claims and extend argumentative scope. “Demonstrated”, “exhibited”, “insights”, “underscore”, and “underscores” signal contribution and evidential force. “Enhance”, “enhancing”, “enhances”, and “integrating” frame research as improvement-oriented and integrative. “Strategies” and “challenges” create a problem-solution narrative. The pattern, therefore, supports RQ2 because the post-ChatGPT increase is concentrated in how abstracts frame significance, contribution, and coherence, not merely in what biomedical topics they describe.
Figure 7 separates absolute prominence from proportional growth. Some markers, such as “potential”, “across”, and “while”, are frequent in the post-ChatGPT period and contribute substantially to the aggregate increase. Other markers have lower post-ChatGPT frequency but stronger post/pre ratios, suggesting rapid emergence from a smaller baseline. Reading both dimensions together avoids overemphasising already common words while overlooking lower-frequency markers that may be more sensitive indicators of stylistic change.
Figure 7. Relationship between post-ChatGPT marker frequency and post/pre ratio among strict stylistic markers. Labels highlight selected high-growth markers.
The yearly trajectories in Figure 8 show that individual markers do not all follow the same path. Some, such as “potential” and “across”, rise steadily after the boundary. Others, such as “enhance”, “challenges”, and “strategies”, show sharper increases in 2024 and 2025. This supports a diffusion account because the markers move in a broadly convergent post-2022 direction while entering that trajectory at different speeds. The pattern is, therefore, more consistent with gradual stylistic uptake than with a single mechanical vocabulary substitution.
Figure 8. Yearly trajectories for selected high-increase strict stylistic markers. The dashed line marks the ChatGPT boundary.
Table 6 reframes the word-level results as rhetorical functions. The strongest grouped contributions come from value and possibility framing and from scope and connective framing. Evidence and claim-signalling markers also increase, indicating a stronger tendency for abstracts to guide the reader toward significance and contribution. This is where the stylistic nature of the finding becomes most visible. The language that grows after 2022 is not mainly terminology about diseases, interventions, or methods. It is language that shapes how research is positioned, polished, and rhetorically presented.
Table 6. Functional grouping of the fastest-growing strict LLM-associated stylistic markers.
The secondary indicators qualify the interpretation of RQ2. Figure 9 shows that lexical diversity increased after 2022, especially in 2025. This result is important because it rules against a crude version of same-ification in which abstracts simply become lexically poorer or semantically interchangeable. The post-ChatGPT period can contain broader vocabulary while also carrying a more concentrated layer of stylistic markers.
Figure 9. Secondary same-ification indicators, including lexical diversity, semantic similarity, and the exploratory composite same-ification index.
Semantic similarity and the composite same-ification index show a less stable pattern. Both appear to peak around 2023 and then decline in later years. This weakens any claim that biomedical abstracts became semantically identical after ChatGPT. The more defensible interpretation is narrower and more researchable: a shared rhetorical texture appears to have intensified across otherwise diverse biomedical abstracts. For this reason, the composite same-ification score is treated as supplementary evidence, while the strict marker trajectory and the marker-control contrast remain the primary empirical basis for answering RQ2.

4.3. RQ3: Robustness Across Venue, Topic, Matching, and SAME-ML Validation

RQ3 asked whether the same-ification signal remains visible across distributional, temporal, topic-controlled, matched-sample, and explainable machine-learning validation analyses. The venue-level analysis is the first robustness check. Figure 10 suggests that post-minus-pre stylistic marker shifts are not evenly distributed across publication venues. This heterogeneity is consistent with the possibility that diffusion is shaped by venue-specific writing norms, disciplinary communities, editorial practices, and article composition.
Figure 10. Exploratory journal-level heterogeneity in post-minus-pre stylistic-marker rates. The results should be interpreted cautiously because journal composition and baseline availability can influence apparent changes.
These venue-level results should be interpreted cautiously. Several journals in the exploratory table have limited or absent pre-ChatGPT observations, which can inflate apparent post-minus-pre differences. The journal analysis, therefore, supports heterogeneity rather than journal attribution. It should not be read as evidence that specific journals adopted LLM-assisted writing more heavily. A stronger extension would estimate journal or publisher random effects with post-period interactions and apply leave-one-journal or leave-one-publisher sensitivity checks. Because those models require a more balanced venue panel than the present analytic sample provides, the current manuscript treats Figure 10 as exploratory evidence for uneven diffusion.
The SAME-ML experiment extends the corpus-level evidence by testing whether the post-ChatGPT period can be distinguished using document-level stylistic features. Year-level aggregate measures were excluded so that classification relied exclusively on features measured at the individual-document level. The models were not intended to infer AI authorship for individual abstracts. Instead, they tested whether the post-2022 period contains a recurring stylistic signal that can be recovered using different machine-learning procedures.
The style-only experiment provided the first validation signal for RQ3. As shown in Figure 11 and Table 7, all three models achieved moderate above-chance discrimination on the held-out data. The sROC AUC was 0.724 for the LASSO logistic regression, 0.712 for random forest, and 0.723 for XGBoost, while the PR AUC ranged from 0.567 to 0.580. At the prespecified classification threshold of 0.50, the accuracy ranged from 0.758 to 0.762 and specificity from 0.928 to 0.944, whereas sensitivity remained comparatively low, ranging from 0.312 to 0.337. Thus, the models ranked post-ChatGPT abstracts better than chance but missed many positive-class observations at the default threshold.
Figure 11. Held-out ROC curves for the style-only models using document-level stylistic and rhetorical features. All three models demonstrated moderate above-chance discrimination between pre- and post-ChatGPT abstracts. Year-level semantic similarity and the composite same-ification index were excluded from the predictor set.
Table 7. Combined SAME-ML held-out model performance across style-only, topic-masked, and marker-ablation experiments.
The feature-importance results in Figure 12 indicate that the period-related signal was distributed across several dimensions of abstract style. Overall rhetorical density was the most influential predictor in the XGBoost model, followed by MTLD lexical diversity, stylistic-marker rate, and log-transformed abstract length. Evidence-and-claim framing and improvement-and-integration language also contributed to prediction. The ranking, therefore, suggests that the model learned a combination of rhetorical organisation, lexical diversity, marker density, and document length rather than relying on a single feature or a year-level temporal proxy.
Figure 12. XGBoost gain-based importance for the style-only model. Overall rhetorical density contributed the greatest predictive gain, followed by MTLD lexical diversity, stylistic-marker rate, and log-transformed abstract length. Importance represents relative predictive contribution and does not indicate causal influence.
The SHAP analysis in Figure 13 provides a complementary document-level interpretation. Abstract length, overall rhetorical density, lexical diversity, and stylistic-marker rate produced the largest absolute SHAP contributions. The distribution of contributions across multiple features is consistent with a multidimensional stylistic shift. SHAP values indicate how particular feature values influenced the fitted model’s predictions; they do not establish that these features caused the post-2022 change.
Figure 13. SHAP summary plot for the style-only XGBoost model. Features are ordered by mean absolute SHAP value. Positive SHAP values move predictions toward the post-ChatGPT class, whereas negative values move predictions toward the pre-ChatGPT class. Colour represents the relative feature value within the held-out sample.
The topic-masked experiment addresses an important concern: whether the post-2022 increase was caused by a change in biomedical topics rather than by a change in style. Figure 14 shows that the increase in stylistic-marker intensity was not confined to a single biomedical topic. Mean marker rates were higher in the post-ChatGPT period across all ten text-derived topic groups. For example, the mean rate increased from 46.1 to 70.7 in the expression/cell topic, from 54.6 to 76.1 in the compounds/molecular/human topic, from 62.7 to 78.9 in the systematic-review/meta-analysis topic, and from 44.7 to 58.2 in the artery/aortic/endovascular topic. Although the magnitude of the increase varied across topics, its broad direction supports the interpretation that the observed stylistic shift extends across multiple areas of biomedical research.
Figure 14. Topic-masked stylistic intensity across ten topic clusters. Each cell reports the mean operational stylistic-marker rate and sample size for pre-ChatGPT and post-ChatGPT abstracts within the topic cluster.
Figure 15 compares the topic-only, combined style-and-topic, and topic-residualised models. The combined style-and-topic model achieved the strongest held-out performance, with an ROC AUC of 0.799 and a PR AUC of 0.670. The topic-only text model achieved an ROC AUC of 0.762 and a PR AUC of 0.584, indicating that subject composition itself contained some period-related information. After the stylistic variables were residualised for topic and document-level covariates, the resulting model retained moderate discrimination, with an ROC AUC of 0.744 and a PR AUC of 0.599. Thus, topic contributed to classification, but it did not fully account for the stylistic signal.
Figure 15. Topic-masked validation of the SAME-ML signal. The combined style-and-topic model achieved the strongest performance. The topic-only and topic-residualised models both retained moderate above-chance discrimination, indicating that topic composition explains part, but not all, of the observed period-related signal.
The matched comparison provides a more conservative robustness check by comparing post-ChatGPT abstracts with more similar pre-ChatGPT abstracts. Figure 16 presents the covariate balance before and after matching. Before matching, the largest imbalance occurred in the MTLD lexical diversity, followed by log-transformed word count, whereas the control-vocabulary rate showed relatively little imbalance. Matching substantially reduced the absolute standardised mean differences for all three displayed continuous covariates, bringing each below the conventional 0.10 balance threshold. Matching was also constrained by topic group and incorporated journal grouping in the propensity-score specification. The matched analysis should, nevertheless, be interpreted as a design-based robustness check rather than as a causal estimate of ChatGPT exposure.
Figure 16. Absolute standardised mean differences before and after propensity-score matching. Matching substantially improved balance in MTLD lexical diversity and log-transformed word count while retaining the already small imbalance in control-vocabulary rate. All displayed post-matching differences were below 0.10.
The matched comparison produced a rightward shift in the post-ChatGPT distribution of stylistic marker density (Figure 17), indicating that the period difference remained visible among abstracts with comparable measured characteristics. Figure 18 shows that the largest matched differences occurred in overall rhetorical density and stylistic-marker rate. Evidence-and-claim framing also showed a substantial positive difference, while value-and-possibility language, improvement-and-integration language, scope-and-connective language, and problem-and-solution framing showed smaller positive effects. In contrast, the matched difference in the marker-control ratio was close to zero. The results, therefore, indicate broader increases in rhetorical and marker density rather than a large change in the relative marker-to-control-vocabulary ratio.
Figure 17. Distribution of stylistic marker density in the matched pre- and post-ChatGPT samples. The post-period distribution is shifted toward higher marker density after balancing the measured matching variables.
Figure 18. Matched post-minus-pre stylistic differences. Positive values indicate higher values in matched post-ChatGPT abstracts.
In the marker-ablation experiment, the full-text model achieved an ROC AUC of 0.812 and a PR AUC of 0.704 (Figure 19). Removing the prespecified stylistic-marker vocabulary produced only a modest decline to an ROC AUC of 0.803 and a PR AUC of 0.681. A model restricted to the stylistic markers achieved an ROC AUC of 0.703 and a PR AUC of 0.563, whereas the control-vocabulary-only model performed more weaker, with an ROC AUC of 0.625 and a PR AUC of 0.395. The small ROC AUC reduction of 0.009 after marker removal indicates that the prespecified marker vocabulary captures part of the period-related signal but that substantial information also remains in the wider text.
Figure 19. Marker-ablation performance across four text conditions. Removal of the prespecified stylistic markers produced a small reduction in performance, indicating that the broader textual signal is not limited to the marker dictionary.
The feature rankings in Figure 20 help identify the vocabulary associated with each ablation condition. In the full-text model, terms such as “findings”, “enhance”, “across”, “insights”, “potential”, and “comprehensive” ranked among the most influential features. After removal of the prespecified markers, terms including “findings”, “key”, “remains”, “including”, “offering”, and “purpose” continued to contribute to classification. The marker-only model emphasised terms such as “potential”, “significant”, “across”, “enhance”, “strategies”, and “comprehensive”, whereas the control-vocabulary model relied more heavily on conventional biomedical terms such as “results”, “models”, “patients”, “study”, “clinical”, and “outcomes”. These rankings indicate predictive contribution rather than the direction of association with either period.
Figure 20. Top predictive terms by XGBoost gain under each marker-ablation condition. The comparison separates general textual-period effects from stylistic-marker effects.
Figure 21 summarises the four main validation results using their experiment-specific effect measures. The style-only model achieved a held-out ROC AUC of 0.723, while the topic-residualised model achieved an ROC AUC of 0.744. In the matched analysis, the standardised post-minus-pre difference in the stylistic-marker rate was 0.533. Removing the prespecified stylistic-marker vocabulary reduced the full-text ROC AUC by only 0.010. Taken together, these results provide converging evidence of a moderate and broadly distributed post-ChatGPT stylistic shift: the signal is learnable from document-level style, remains visible after topic adjustment and matching, and is not explained exclusively by the prespecified marker vocabulary.
Figure 21. Summary of the principal SAME-ML validation results. The panels report experiment-specific metrics: style-only held-out ROC AUC, topic-residualised held-out ROC AUC, the matched standardised difference in marker rate, and the change in ROC AUC after marker removal. Because the panels use different metrics, their numerical magnitudes should not be interpreted as directly comparable effect sizes.

4.4. Synthesis of Findings

RQ1 was answered affirmatively: strict LLM-associated stylistic markers increased after the ChatGPT boundary, with the clearest growth occurring in 2024 and 2025 rather than as a one-year discontinuity in 2023. The mean marker rate rose from 4.995 per 1000 words in the pre-ChatGPT period to 11.658 in the post-ChatGPT period, a 133.4% increase. Distributional and time-series evidence show that the change is visible beyond the mean and that the post-boundary trajectory is steeper than the pre-boundary baseline.
RQ2 was also answered affirmatively, with qualification. The control vocabulary did not show the same sustained increase, and the fastest-growing words were concentrated in rhetorical functions rather than biomedical content. The shift is, therefore, best interpreted as stylistic specificity: Post-ChatGPT abstracts more frequently use language that frames value, connects claims, signals evidence, and presents research as integrative or consequential. The concurrent increase in MTLD shows that this is not lexical impoverishment. It is a narrower form of rhetorical convergence within otherwise diverse biomedical writing.
RQ3 received a cautiously affirmative answer. The signal remains visible across several validation layers, as follows: distributional comparison, interrupted time-series modelling, topic-masked analysis, matched-sample comparison, marker-ablation, and interpretable SAME-ML. The strongest claim is not that the models can detect AI-authored abstracts. The stronger and more defensible claim is that the post-ChatGPT period contains a learnable and interpretable rhetorical signal that survives several controls for topic, length, control vocabulary, and sample imbalance.
Based on our findings, the three RQs point to a specific form of same-ification. The corpus does not become semantically identical, and the evidence does not support individual authorship attribution. Instead, biomedical abstracts appear to be accumulating a shared layer of polished, contribution-signalling language. Same-ification in this study, therefore, means convergence in rhetorical presentation: diverse studies increasingly rely on similar linguistic moves to package novelty, relevance, and significance.
The evidence also defines the boundary of the claim. The 2023 sample is smaller than the surrounding years, the journal-level analysis is exploratory, marker-level rankings are descriptive unless corrected for multiple comparisons, and SAME-ML requires external replication before it can be read as generalisable classification. These limits do not remove the observed post-2022 signal, but they narrow the interpretation to corpus-level stylistic diffusion rather than causal attribution to ChatGPT or definitive detection of LLM use.

5. Discussion

This study frames post-ChatGPT change in biomedical abstracts as corpus-level rhetorical diffusion. In relation to RQ1, strict LLM-associated stylistic markers increased sharply after the ChatGPT boundary, while the pre-2022 baseline remained comparatively stable. Figure 3 and Figure 4 and Table 1 and Table 2 show that the increase became most visible in 2024 and 2025, supporting a diffusion interpretation rather than a one-time reaction to the release of ChatGPT. This matters because diffusion implies normalisation through writing, editing, reviewing, and institutional authoring practices.
RQ2 asked whether the shift was stylistic rather than merely a general change in biomedical vocabulary. The answer is cautiously affirmative. The control vocabulary did not show the same sustained post-2022 growth, and the segmented regression in Table 4 showed both a significant level shift and a stronger post-ChatGPT slope change. Figure 6, Figure 7 and Figure 8 and Table 5 and Table 6 show that the strongest markers operate as value framing, connective framing, evidence signalling, improvement rhetoric, and problem-solution language. This extends excess-vocabulary work by moving from prevalence estimation toward the rhetorical functions through which LLM-associated style becomes visible in published abstracts.
The findings also complicate a simple meaning of same-ification. Figure 9 shows that lexical diversity increased after 2022, while semantic similarity and the composite same-ification index were less stable. This means that same-ification should not be understood as semantic sameness or lexical impoverishment. A more precise interpretation is rhetorical convergence: abstracts may continue to cover diverse biomedical topics and use varied vocabulary while increasingly drawing on a shared layer of polished, contribution-signalling expressions. This distinction is important because it avoids the exaggerated claim that science is becoming textually identical. The stronger claim is that academic style may be becoming more recognisably LLM-like in how it packages importance and contribution.
RQ3 concerned whether the post-ChatGPT shift was sufficiently systematic to be recovered through machine-learning validation. The SAME-ML experiments provide converging, but appropriately moderate, evidence of period-related stylistic separability. Style-only models achieved above-chance held-out discrimination, and performance improved when topic information was combined with stylistic features. Importantly, the topic-residualised model remained predictive, indicating that topic composition did not fully explain the stylistic signal. The matched analysis also identified positive differences in rhetorical density and stylistic-marker rate after balancing the measured covariates. Finally, marker ablation produced only a small reduction in full-text performance, showing that the shift extends beyond the prespecified marker vocabulary. These findings support a corpus-level change in biomedical abstract style, but they do not establish individual AI authorship or external diagnostic validity.
This study contributes to the literature by shifting the focus from AI-authorship detection to language change. Detection studies remain valuable, but they are vulnerable to fairness, robustness, and interpretive concerns, particularly for multilingual writers [3,15,16]. The present approach does not ask whether an individual author used ChatGPT. It asks whether a public scholarly corpus shows measurable signs of stylistic diffusion. This provides a more responsible and publishable framing for studying GenAI in academic writing, especially in biomedical contexts where authorship, accountability, and trust are closely tied to publication ethics.
The practical implication is that journals and editors may need to look beyond disclosure policies and plagiarism-style detection. If LLM use gradually reshapes the rhetorical norms of abstracts, then editorial attention should also consider the quality, specificity, and evidential modesty of academic prose. A polished abstract is not necessarily a better abstract if rhetorical intensity masks uncertainty, exaggerates contribution, or makes studies sound more similar than they are. Researchers can use the proposed framework to monitor stylistic change over time, compare disciplines, and evaluate whether LLM-assisted writing is improving clarity or producing a narrower style of scholarly expression.
The analysis is based on PubMed abstracts, not full articles, peer reviews, grant applications, or author drafts. The ChatGPT boundary is historically meaningful, but it cannot prove that every post-2022 change was caused by ChatGPT. The 2023 sample was smaller than other years and was interpreted cautiously. The marker dictionary is intentionally strict, but no word is unique to LLM use, and marker-level rankings should not be treated as inferential tests without multiple-comparison correction. Journal-level heterogeneity may reflect editorial pipelines, venue composition, author populations, or publication-type changes rather than LLM uptake alone. The topic controls and SAME-ML analyses reduce, but do not eliminate, these threats to interpretation. Future work should test the framework across larger and more balanced PubMed samples, disciplines, languages, full-text corpora, and author-disclosed LLM-use settings.

6. Conclusions

The findings provide a qualified but clear answer to the title question. Biomedical abstracts are beginning to show a more ChatGPT-like rhetorical profile, not through identical content but through a shared style of value framing, connective flow, evidence signalling, improvement rhetoric, and problem-solution language. LLM-associated same-ification is therefore best understood as rhetorical convergence: the field remains diverse in what it studies but becomes more similar in how it packages significance.

Author Contributions

Conceptualisation, N.H.S.A. and B.J.J.; Methodology, N.H.S.A.; Software, N.H.S.A.; Validation, N.H.S.A. and B.J.J.; Formal analysis, N.H.S.A.; Investigation, N.H.S.A.; Writing—original draft preparation, N.H.S.A.; Writing—review and editing, N.H.S.A. and B.J.J.; Visualisation, N.H.S.A.; Supervision, B.J.J. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Data Availability Statement

Data are available upon reasonable request from the corresponding author.

Acknowledgments

During manuscript preparation, Grammarly Premium (version 1.2.279.1925) was used to support language refinement, structural organisation, and clarity checking. Canva AI-assisted (version 2.0) (www.canva.com) design features were also used to support the visual organisation of Figure 2. The figures were reviewed, edited, and finalised by the authors. These tools were not used to determine the study design, dataset, statistical analysis, interpretation of findings, or final scholarly claims. The authors take full responsibility for the accuracy, integrity, originality, references, analysis, figures, and conclusions of the manuscript.

Conflicts of Interest

The authors declare no conflict of interest.

Abbreviations

The following abbreviations are used in this manuscript:
TF-IDFTerm Frequency-Inverse Document Frequency
SAME-MLStyle-Aware, Matched, and Explainable Machine Learning
SHAPSHapley Additive exPlanations
LSALatent Semantic Analysis
MTLDMeasure of Textual Lexical Diversity
PMIDPubMed Identifier
NCBINational Center for Biotechnology Information

References

  1. Kobak, D.; González-Márquez, R.; Horvát, E.-Á.; Lause, J. Delving into LLM-Assisted Writing in Biomedical Publications through Excess Vocabulary. Sci. Adv. 2025, 11, eadt3813. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  2. Gray, A. ChatGPT Contamination: Estimating the Prevalence of LLMs in the Scholarly Literature. arXiv 2024, arXiv:2403.16887. [Google Scholar]
  3. Liang, W.; Yuksekgonul, M.; Mao, Y.; Wu, E.; Zou, J. GPT Detectors Are Biased against Non-Native English Writers. Patterns 2023, 4, 100779. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  4. Brown, T.B.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. Language Models Are Few-Shot Learners. In Advances in Neural Information Processing Systems; Curran Associates: Red Hook, NY, USA, 2020; Volume 33, pp. 1877–1901. [Google Scholar]
  5. OpenAI. Introducing ChatGPT. Available online: https://openai.com/index/chatgpt/ (accessed on 7 June 2026).
  6. OpenAI; Achiam, J.; Adler, S.; Agarwal, S.; Ahmad, L.; Akkaya, I.; Aleman, F.L.; Almeida, D.; Altenschmidt, J.; Altman, S.; et al. GPT-4 Technical Report. arXiv 2023, arXiv:2303.08774. [Google Scholar]
  7. Stokel-Walker, C.; van Noorden, R. What ChatGPT and Generative AI Mean for Science. Nature 2023, 614, 214–216. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  8. van Dis, E.A.M.; Bollen, J.; Zuidema, W.; van Rooij, R.; Bockting, C.L. ChatGPT: Five Priorities for Research. Nature 2023, 614, 224–226. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  9. Tian, S.; Jin, Q.; Yeganova, L.; Lai, P.-T.; Zhu, Q.; Chen, X.; Yang, Y.; Chen, Q.; Kim, W.; Comeau, D.C.; et al. Opportunities and Challenges for ChatGPT and Large Language Models in Biomedicine and Health. Brief. Bioinform. 2024, 25, bbad493. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  10. Wang, C.; Li, M.; He, J.; Wang, Z.; Darzi, E.; Chen, Z.; Ye, J.; Li, T.; Su, Y.; Ke, J.; et al. A Survey for Large Language Models in Biomedicine. arXiv 2024, arXiv:2409.00133. [Google Scholar]
  11. Ganjavi, C.; Eppler, M.B.; Pekcan, A.; Biedermann, B.; Abreu, A.; Collins, G.S.; Gill, I.S.; Cacciamani, G.E. Publishers’ and Journals’ Instructions to Authors on Use of Generative Artificial Intelligence in Academic and Scientific Publishing: Bibliometric Analysis. BMJ 2024, 384, e077192. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  12. Committee on Publication Ethics. Authorship and AI Tools. Available online: https://publicationethics.org/guidance/cope-position/authorship-and-ai-tools (accessed on 7 June 2026).
  13. World Association of Medical Editors. Chatbots, ChatGPT, and Scholarly Manuscripts. Available online: https://wame.org/page3.php?id=110 (accessed on 7 June 2026).
  14. Nature Portfolio. Artificial Intelligence (AI) Editorial Policies. Available online: https://www.nature.com/nature-portfolio/editorial-policies/ai (accessed on 7 June 2026).
  15. Sadasivan, V.S.; Kumar, A.; Balasubramanian, S.; Wang, W.; Feizi, S. Can AI-Generated Text Be Reliably Detected? arXiv 2023, arXiv:2303.11156. [Google Scholar]
  16. Weber-Wulff, D.; Anohina-Naumeca, A.; Bjelobaba, S.; Foltynek, T.; Guerrero-Dib, J.; Popoola, O.; Sigut, P.; Waddington, L. Testing of Detection Tools for AI-Generated Text. Int. J. Educ. Integr. 2023, 19, 26. [Google Scholar] [CrossRef] [Scilit]
  17. Mitchell, E.; Lee, Y.; Khazatsky, A.; Manning, C.D.; Finn, C. DetectGPT: Zero-Shot Machine-Generated Text Detection Using Probability Curvature. In Proceedings of the 40th International Conference on Machine Learning; PMLR: Honolulu, HI, USA, 2023. [Google Scholar]
  18. Bender, E.M.; Gebru, T.; McMillan-Major, A.; Shmitchell, S. On the Dangers of Stochastic Parrots: Can Language Models Be Too Big? In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency; ACM: New York, NY, USA, 2021; pp. 610–623. [Google Scholar] [CrossRef] [Scilit]
  19. Lopez Bernal, J.; Cummins, S.; Gasparrini, A. Interrupted Time Series Regression for the Evaluation of Public Health Interventions: A Tutorial. Int. J. Epidemiol. 2017, 46, 348–355. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  20. McCarthy, P.M.; Jarvis, S. MTLD, vocd-D, and HD-D: A Validation Study of Sophisticated Approaches to Lexical Diversity Assessment. Behav. Res. Methods 2010, 42, 381–392. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  21. Benoit, K.; Watanabe, K.; Wang, H.; Nulty, P.; Obeng, A.; Muller, S.; Matsuo, A. quanteda: An R Package for the Quantitative Analysis of Textual Data. J. Open Source Softw. 2018, 3, 774. [Google Scholar] [CrossRef] [Scilit]
  22. Reimers, N.; Gurevych, I. Sentence-BERT: Sentence Embeddings Using Siamese BERT-Networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing; Association for Computational Linguistics: Hong Kong, China, 2019; pp. 3982–3992. [Google Scholar] [CrossRef] [Scilit]
  23. Roberts, M.E.; Stewart, B.M.; Tingley, D. stm: An R Package for Structural Topic Models. J. Stat. Softw. 2019, 91, 1–40. [Google Scholar] [CrossRef] [Scilit]
  24. Lundberg, S.M.; Lee, S.-I. A Unified Approach to Interpreting Model Predictions. In Advances in Neural Information Processing Systems; Curran Associates: Red Hook, NY, USA, 2017; Volume 30. [Google Scholar]
  25. National Center for Biotechnology Information. APIs: Develop. U.S. National Library of Medicine. Available online: https://www.ncbi.nlm.nih.gov/home/develop/api/ (accessed on 7 June 2026).
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Article Metrics

Citations

Article Access Statistics

Multiple requests from the same IP address are counted as one view.