1. Introduction
Recent progress in neural text-to-speech and voice cloning has made synthetic speech increasingly realistic, raising concerns about trust, attribution, and misinformation [
1,
2]. The challenge is no longer limited to detecting artificial-sounding voices, but also includes cases where speaker identity is preserved while the spoken message is altered.
This threat is not limited to voice imitation. An attacker may retain a believable cloned voice while changing the content of the utterance, including short targeted edits that replace only a few words or segments [
3]. Such manipulation can attribute false statements to real speakers while preserving their identity cues [
1]. In these cases, detection cannot rely only on vocal realism. The transcript may still reveal disruptions in linguistic organization and internal consistency [
2].
A psycholinguistic perspective supports this transcript-first view. Research on deception has shown that production-related cues may reflect increased demands on planning, monitoring, and consistency maintenance, although such cues must be interpreted cautiously [
4,
5]. Markers such as pauses or filled pauses are not unique to deception and may also arise in normal speech planning [
5]. Some findings are even counterintuitive. For example, Arciuli et al. found that “um” occurred less often and was shorter during lying [
6]. This suggests that production-related traces may be informative, but not as simple one-to-one indicators. In this study, these psycholinguistic concepts are used as cautious background motivation rather than as direct explanations of cloned speech.
Motivated by this gap, we study fake speech from a transcript-first, psycholinguistically informed perspective using FakeSpeech+, a paired real–fake dataset designed for identity-preserving semantic manipulation. We examine two groups of interpretable transcript-level cues: (i) linguistic content organization and discourse dynamics, and (ii) compact psycholinguistic proxy cues related to production traces, including hesitation and disfluency markers. To reduce trivial length effects, we evaluate these cues under transcript-length control through residualization. Our results show that manipulated transcripts still exhibit measurable differences in discourse dynamics and production-trace variability.
This work makes two contributions. First, it introduces FakeSpeech+, a paired real–fake dataset for identity-preserving semantic manipulation in speech. Second, it presents a transcript-first analysis of linguistic and psycholinguistic features under length-controlled conditions, supported by baseline experiments.
2. Related Works
In this section, we review related work in three strands. The first focuses on linguistic and psycholinguistic features in NLP. The second summarizes feature-based studies comparing human- and machine-generated text. The third reviews existing manipulation datasets and situates the present study within this benchmark landscape.
2.1. Linguistic and Psycholinguistic Features in NLP
First, prior work has modeled linguistic content organization through discourse relations, syntactic structure, and coherence signals.
Liu and Strube in [
7] modeled content organization through discourse relations. They extracted PDTB-style relations between adjacent sentences and represented each document as a relation sequence. They showed that relation transitions correlate with coherence labels. A BiLSTM trained only on relation sequences performed close to raw-text models. Shuffling relation order reduced performance. A fusion model that combined text with relation signals improved over text-only baselines, especially for longer texts and in cross-domain settings.
Naismith et al. [
8] used GPT-4 to rate coherence with rubric-based prompting. They compared GPT-4 ratings to expert judgments and to traditional coherence metrics. GPT-4 showed much higher agreement with experts. Providing rationales helped slightly, while differences between prompt variants were small.
Second, psycholinguistic work in NLP has used interpretable cues related to affect, framing, and cognitively motivated language patterns.
Monteiro et al. [
9] introduced PsyMatrix. Their goal was to support pretrained model selection for text classification. They described each dataset using psycholinguistic, topic, and language signals. They summarized these signals across documents. They then compressed them into low-dimensional dataset embeddings. Finally, they used the embeddings to predict and rank model performance on unseen datasets.
Adkins et al. [
10] proposed a psycholinguistic NLP framework for forensic text analysis. Their goal was to assist investigations. They relied on language cues to narrow down persons of interest. They combined emotion and subjectivity signals with lexical patterns. They also tracked changes across an interview timeline. Their results showed the framework helped identify relevant suspects in their evaluation setting.
Arif et al. [
11] studied hope speech using psycholinguistic and emotional perspectives. Their goal was to represent hope-related text with interpretable features. They used LIWC-style categories, emotion lexicons, and sentiment signals. They trained machine learning models for classification. They reported strong performance, especially with boosting-based models. The study showed these features were useful for modeling affective social media text.
2.2. Feature-Based Analyses of Human- and AI-Generated Text
Muñoz-Ortiz et al. [
12] studied linguistic differences between human-written news and LLM-generated news. Their goal was to quantify how human and generated news diverge across multiple feature families. They reported consistent distributional differences in sentence-length dispersion, vocabulary variety, and syntactic structure, and they also observed systematic shifts in affective and psychometric signals.
Seals and Shalin [
13] compared student-written long-form analogies with ChatGPT-generated analogies. Their goal was to test whether psycholinguistic properties in human writing are preserved in LLM output. Using Coh-Metrix features, they achieved high classification performance and showed that multiple cohesion/readability indicators differed between the two sources.
Zaitsu and Jin [
14] analyzed human vs. ChatGPT writing in Japanese using stylometric features. Their goal was to identify measurable stylistic fingerprints that separate human texts from LLM texts in a non-English setting. They reported clear differences using function-word rate, POS bigrams, particle bigrams, and punctuation-related patterns, and showed strong separability between human and generated texts.
Nkhobo and Chaka [
15] compared student discursive essays with ChatGPT-generated essays using Coh-Metrix. Their goal was to test whether lexical diversity, syntactic complexity, and referential cohesion differ between human and generated writing. They compared student discursive essays with ChatGPT-generated essays using Coh-Metrix to examine lexical diversity, syntactic complexity, and referential cohesion. Their raw mean patterns suggested stronger lexical and referential-cohesion properties in student essays and higher syntactic complexity in ChatGPT essays. However, the overall mean differences across these indices were not statistically significant under their sample size.
Reviriego et al. [
16] focused specifically on vocabulary use and lexical diversity. Their goal was to measure whether ChatGPT reduces or expands lexical richness relative to humans across tasks. They reported that ChatGPT-3.5 generally used fewer distinct words and lower diversity than humans, while ChatGPT-4 showed similar diversity and sometimes higher values, depending on the setting.
Shaitarova et al. [
17] examined “linguistic footprints” of ChatGPT across tasks, domains, personas, and two languages (English and German). Their goal was to test how human–AI feature differences vary across settings. They found that prompting affected English output more strongly than German, and that several basic linguistic and readability features consistently distinguished human and generated text across conditions.
Prior studies have examined human-written and AI-generated text using feature-based analyses across different genres and linguistic dimensions, as summarized in
Table 1.
2.3. Fake-Speech Benchmarks and Transcript-Level Gaps
Existing manipulation datasets span text, audio, visual, and multimodal settings. Early benchmarks mainly focused on single-modality manipulations, which helped isolate modality-specific effects. More recent datasets have expanded toward multimodal settings that combine audio, visual, and textual alterations.
Existing manipulation datasets cover text, speech, visual, and multimodal settings. In text, Tesconi et al. [
18] introduced TweepFake, which includes human-written and machine-generated tweets for synthetic text detection. Pu et al. [
19] proposed the Deepfake Text Detection in the Wild dataset, which contains articles and online posts generated by several text-generation systems together with human-written texts. In Arabic, Mustafa et al. [
20] introduced the GPT-2 Arabic Deepfake Tweets dataset, highlighting the difficulty of deepfake detection in underrepresented languages.
In audio, Wu et al. [
21] developed the ASVspoof challenge series for evaluating spoofing countermeasures in automatic speaker verification systems. These datasets include both text-to-speech and voice conversion attacks.
Reimao and Tzerpos [
22] introduced the Fake or Real (FoR) dataset, which contains more than 198,000 utterances, including more than 87,000 synthetic utterances and more than 111,000 real utterances. Ali et al. [
23] also proposed a speech deepfake dataset for famous figures, focusing on targeted voice spoofing for political identities.
In the visual domain, Dolhansky et al. [
24] introduced the DFDC dataset, which contains more than 100,000 videos generated using face-swapping methods. Kwon et al. [
25] proposed KoDF, which expands demographic diversity by focusing on Korean subjects. Li et al. [
26] introduced Celeb-DF, which provides higher-quality fake videos with fewer visible artifacts than earlier benchmarks. Shao et al. [
27] extended this direction through DGM4, which includes paired image-text manipulations across both modalities.
More recent datasets moved toward multimodal manipulations. Khalid et al. [
28] introduced FakeAVCeleb, which combines face manipulation, altered audio, and lip-synchronized outputs. Cai et al. [
29] proposed AV-Deepfake1M, a large-scale benchmark covering video reenactment, text-to-speech, and voice cloning. Yi et al. [
30] introduced the Half-Truth Dataset (HAD), which focuses on partial speech manipulation, where only selected words or phrases are fake while the rest of the utterance remains real. In addition, Alsaeedi et al. [
31] introduced FakeSpeech, an audio–visual dataset designed for semantic-level speech manipulation while preserving speaker identity, vocal style, and original visual frames.
Table 2 presents the details of existing datasets.
2.4. Research Gap and Contributions
Previous studies have shown measurable differences between human and AI-generated texts, including sentence-length patterns, lexical diversity, grammatical and stylistic cues, and cohesion-related indices. However, most of this evidence comes from general human–AI text comparisons and does not fully reflect the fake-speech setting considered here, where speaker identity may remain convincing while content is altered through targeted semantic manipulation.
The motivation for this work began with the analysis problem itself. We sought to examine whether fake-speech transcripts contain reliable and interpretable cues under identity-preserving semantic manipulation, but found that existing datasets were not suitable for this purpose. Available benchmarks largely target spoofing, voice cloning, or broad multimodal manipulation, rather than controlled semantic alteration with preserved speaker identity and transcript-level analysis as a primary objective. Therefore, the gap was not only analytical but also dataset-related: the required benchmark did not exist for the kind of transcript-first investigation pursued here.
To address this problem, we adopt a transcript-first perspective and introduce FakeSpeech+, a paired real–fake dataset designed for identity-preserving semantic manipulation. We then examine interpretable transcript-level cues from two groups: linguistic content organization and discourse dynamics, and compact psycholinguistic proxy cues, including interpretive scaffolding, affect anchoring, embodied grounding, experiential and directional language, and hesitation or disfluency markers. Under length-controlled analysis, we evaluate which cues remain informative and provide an interpretable account of their statistical, predictive, and linguistic differences.
This work makes two main contributions. First, it introduces FakeSpeech+, a paired real–fake dataset designed for identity-preserving semantic manipulation in speech. Second, it presents a transcript-first statistical analysis of linguistic and psycholinguistic features under length-controlled conditions, supported by baseline experiments to assess practical utility.
3. Dataset
This section describes the dataset used in this study and explains its construction and organization.
3.1. Dataset Collection
We used the second version of our previously developed dataset, called FakeSpeech+. The earlier version was an audiovisual dataset based on phantom reading, with emphasis on speech manipulation and audiovisual plausibility rather than explicit semantic control. In contrast, FakeSpeech+ is an audio-only dataset developed specifically to investigate meaning-level manipulation in spoken content, where the central objective is to alter the intended meaning while preserving naturalness and speaker identity. The main objective of this version is to create realistic fake speech in which the intended meaning changes, thereby enabling the extraction of psycholinguistic and discourse-level features for comparing human and generated speech patterns.
To reduce confounding effects, we preserved the natural diversity of speakers and recording conditions instead of artificially standardizing them. We began with a subset of 500 real clips selected from VoxCeleb [
32]. The selection prioritized samples with clear and intelligible speech, since this was necessary to ensure reliable transcription and consistent feature extraction.
3.2. Dataset Generation Pipeline
For each real clip, we generated a corresponding fake sample by prompting ChatGPT with GPT-5.2 Instant to produce a new utterance with an opposite or contradictory intended meaning relative to the original sentence. The rewritten content was designed to remain natural and coherent. In this way, the manipulation targeted the meaning of speech rather than only its surface form.
The rewritten texts were then converted into speech using the ElevenLabs [
33] voice-cloning platform. Each voice was cloned directly from the original speaker to preserve timbre, prosody, and overall vocal identity, making the manipulated samples more realistic. The synthesized audio was exported in .wav format. A subset of generated samples was also manually reviewed to verify naturalness and speaker consistency.
After generation, all samples were transcribed, and transcript-based features were computed as the main input to our analysis. This design allowed the experiments to be conducted on the transcribed text, grounding the analysis in the underlying linguistic content of speech rather than in raw acoustic variability. Through this controlled design, the resulting fake samples preserved speaker-related vocal characteristics while allowing the analysis to focus on semantic inconsistencies rather than obvious synthesis artifacts.
3.3. Dataset Statistics
During cleaning, we removed samples that could not be processed consistently because of voice-synthesis restrictions. After filtering, the final dataset contained 980 samples, divided evenly into 490 real and 490 fake clips. In the fake subset, the spoken content was semantically manipulated while preserving speaker-related vocal characteristics, enabling a controlled comparison for feature-based analysis.
To better reflect real-world variability, FakeSpeech+ includes English speech from male and female speakers and covers a range of accents, including native English, Indian-accented English, and Arabic-accented English. The recordings also span clean, moderately noisy, and noisy conditions. This diversity helps avoid an overly narrow experimental setup and makes the dataset more representative of real speech in practice.
4. Methodology
In this section, we describe the methodology used in this study. Our analysis follows a transcript-first setting, where spoken content is first transcribed from the audio and then used as the input for all subsequent processing and analysis. We then describe the feature design, feature extraction procedure, length control through residualization, the statistical tests used in the study, and the experimental setup.
4.1. Theory-Driven Feature Design
Our feature design is motivated by a psycholinguistic view of fabrication as a process that can influence planning, monitoring, and consistency management. Rather than relying on opaque correlations, we focus on interpretable transcript-level cues that may reflect how manipulated speech departs from natural human production.
We use two feature groups. The first captures linguistic content organization and discourse dynamics, including syntactic structure, lexical distribution, and sentence-to-sentence semantic continuity. The second captures compact psycholinguistic proxy cues, including framing, affect anchoring, embodied grounding, and hesitation or disfluency markers. These cues are treated as indirect proxies rather than direct measures of mental state. This theoretical framing is used to motivate feature selection and interpretation, rather than to treat the extracted cues as direct measures of internal cognitive state.
To reduce trivial confounds, we control transcript length through residualization. This supports a more interpretable comparison between authentic and manipulated transcripts.
4.2. Feature Extraction
In this subsection, we will explore the complete operationalization setup for features extraction.
4.2.1. Transcript-Only Setting and Minimal Preprocessing
All features are computed from transcripts only, with no acoustic or visual cues. To preserve realism under automatic transcription and noisy, real-world conditions, we avoid manual transcript cleaning. The pipeline applies only minimal, computation-driven normalization: (i) sentence segmentation using spaCy (en_core_web_sm) for all sentence-based and syntactic analyses, and (ii) a lightweight regex tokenization for lexical and lexicon-count features. For lexicon-based counts, words are extracted and lowercased to ensure consistent matching across speakers and sources. No domain-specific filtering (e.g., removing fillers, punctuation normalization beyond simple counting, or transcript rewriting) is performed.
4.2.2. Linguistic Structure and Coherence Features
We operationalize transcript structure and discourse coherence using dependency parsing and sentence embedding similarity, producing 11 numeric features.
Syntactic organization and fragmentation are derived from spaCy dependency parses. We compute sentence-level dependency-tree depth using a depth-first traversal over dependency edges restricted to tokens within each sentence span and summarize this as dependency_tree_depth_mean and dependency_tree_depth_max across sentences. Clause-linking tendencies are represented by dependency-relation counts normalized per sentence: subordination_rate counts tokens whose dependency labels fall in SUBORD_DEPS = {advcl, ccomp, xcomp, acl, relcl, mark} and divides by the number of sentences, while coordination_rate counts tokens with labels in COORD_DEPS = {conj, parataxis} and divides by the number of sentences. Fragmentation is captured by a syntactic well-formed heuristic: a sentence is flagged as fragment-like if it contains no VERB or AUX; fragment_ratio is the proportion of such sentences among all transcript sentences.
Lexical diversity is captured by two length-robust measures computed on regex-tokenized words. mattr is computed using a moving-average type–token ratio with a fixed window size of 50 tokens (fallback to global TTR when transcript length ≤50). hapax_ratio is computed as the number of word types occurring exactly once divided by the total token count, measuring lexical rarity/dispersion.
Discourse continuity is modeled using semantic similarity between adjacent sentences. Sentences are embedded using SentenceTransformer all-MiniLM-L6-v2 Sentence Transformer model through the SentenceTransformers Python library version 5.2.3 in the Kaggle Notebook environment (Kaggle, San Francisco, CA, USA) with normalized embeddings, so cosine similarity reduces to a dot product. Adjacent sentence similarities are summarized as sentence_similarity_mean and sentence_similarity_std. Local incoherence is defined as the proportion of adjacent sentence pairs whose similarity falls below a fixed threshold (thr = 0.35), yielding local_incoherence_rate. To avoid redundancy with adjacent similarities, global drift is computed relative to the first sentence: topic_drift_score is the mean of (1 − cosine_similarity(first_sentence, sentence_i)) for all subsequent sentences, capturing cumulative semantic displacement across the transcript.
4.2.3. Psycholinguistic-Style Proxy Features
In addition to structural/coherence signals, we compute eight transcript-derived proxy features designed to reflect psychologically motivated language patterns (planning difficulty, experiential framing, intent/direction, interpretive scaffolding, embodied grounding, and affect anchoring). These features are lexicon- and pattern-based, computed as rates per 100 words unless otherwise stated, using the same regex tokenization and lowercase matching.
Disfluency and hesitation are modeled using multiple surface indicators. disfluency is computed as a per-100-words rate combining (i) filler counts from a configurable filler set (e.g., “uh”, “um”, “like”) plus multiword fillers (“you know”, “i mean”), (ii) immediate repetition counts (adjacent repeated tokens, e.g., “I I”). Hesitation is computed as a per-100-words rate combining filler counts and punctuation-based hesitation markers (e.g., “…”, “—”, “--”). These definitions intentionally remain transcript-only and do not depend on timing or acoustic cues.
Experience and direction are modeled via tense/temporal and intent/imperative proxies. experience is computed as a per-100-words rate combining counts of past-tense verb tags (VBD/VBN) from spaCy and a small set of past-time markers (e.g., “yesterday”, “ago”, “previously”). direction is computed as a per-100-words rate combining counts of future/intent markers (e.g., “will”, “going to”, “want to”, “plan to”, “need to”, “must”) and an imperative heuristic that flags sentences starting with a base verb (VB) from a small imperative verb list (e.g., “do”, “go”, “make”, “tell”).
Interpretive scaffolding and embodied grounding are represented by two ratios. interpretive_ratio is a per-100-words rate of interpretive and explanatory language, computed as the count of hedge phrases (e.g., “i think”, “maybe”, “it seems”) plus causal/explanatory markers (e.g., “because”, “therefore”, “as a result”). embodied_ratio is a per-100-words rate of sensory/body-related terms (e.g., “pain”, “hurt”, “breath”, “see”, “hear”, “stomach”), using a configurable lexicon intended to capture embodied grounding cues.
Emotion anchoring is operationalized as a per-100-words composite that captures explicit affect and anchoring forms. emotion_anchor combines (i) emotion word hits from a basic emotion lexicon, (ii) occurrences of anchor phrases that explicitly attach affect to the speaker (e.g., “i feel”, “it made me feel”, “i’m so”), and (iii) an intensifier–emotion bigram proxy (e.g., “very happy”, “really angry”) to reflect anchored emphasis.
Finally, cognitive_effort is computed as a transparent composite proxy reflecting production difficulty and scaffolding. It is a weighted sum of the above rates: 0.45 × disfluency + 0.35 × hesitation + 0.20 × interpretive_ratio.
4.2.4. Controlling for Length Effects (Residualization)
Many linguistically motivated features vary with transcript length. Longer transcripts simply provide more chances for structural, lexical, and discourse patterns to appear. To avoid treating these length effects as meaningful signals, we control for transcript length before analysis.
For each feature , we fit a simple linear model where word count is the predictor. We then keep the residuals as a length-controlled version of the feature. In other words, we remove the part of the feature that can be explained by length, and keep what remains. This step is a statistical control only. It is not used to infer internal mental states. Unless stated otherwise, all analyses and figures in the Findings section use residualized features.
4.2.5. Statistical Verification of Feature-Level Separation
To quantitatively assess whether the residualized features still distinguished authentic from manipulated transcripts after controlling for transcript length, we applied two statistical tests to each feature. We used the Kolmogorov–Smirnov (KS) test to compare the overall distributions of the two classes, and a permutation test on the median difference to assess whether class differences in central tendency were larger than expected by chance. To quantify the strength of the observed differences, we additionally report Cliff’s delta as an effect-size estimate. Because multiple features were tested, false discovery rate (FDR) correction was applied to adjust the resulting p-values.
These statistical analyses served as the primary basis for assessing feature-level separation after residualization. In addition, a baseline classifier was used separately to evaluate the practical predictive utility of the feature set in combination. Together, these analyses provide complementary rather than redundant evidence.
4.3. Baseline Experimental Setup
To assess the practical discriminative utility of the extracted transcript-level cues, we used a linear Support Vector Machine (SVM) as a lightweight baseline classifier. The experiments were conducted on the residualized feature set. Missing values were handled using median imputation, and all features were standardized before training. The classifier was configured with a (linear kernel, C = 1.0, class_weight = balanced, and probability = True) to support probability-based evaluation and AUC computation.
The dataset was divided using a stratified train–test split, with 30% reserved for testing and a (fixed random seed of 42) for reproducibility. Performance was evaluated using Accuracy, Balanced Accuracy, F1-score, Precision, Recall, and Area Under the ROC Curve (AUC). Where relevant, the baseline was examined under different feature configurations, including linguistic features, psycholinguistic features, and the combined feature set. Since the baseline classifier evaluates the features jointly in a multivariate setting, its results are interpreted as evidence of practical predictive utility rather than as a direct replication of the feature-by-feature statistical ranking.
5. Results and Discussion
In this section, we present the main findings of the study from two complementary perspectives. First, we ask which residualized features show statistically reliable separation between authentic and manipulated transcripts. Second, we ask whether the feature set provides practical discriminative utility in a baseline classification setting. Because these analyses answer different questions, they are treated as complementary rather than expected to produce identical rankings.
5.1. Statistical Feature-Level Results
To assess whether the residualized features distinguished authentic from manipulated transcripts, we applied the Kolmogorov–Smirnov (KS) test, a permutation test on the median difference, Cliff’s delta, and false discovery rate (FDR) correction. The results revealed substantial variation across features in the strength and robustness of their statistical separation.
Table 3 and
Table 4 summarize the statistical results for the residualized linguistic and psycholinguistic features, respectively.
The strongest statistical evidence was observed for coordination_rate_resid and emotion_anchor_resid. Both features remained significant in the KS and permutation tests after FDR correction and showed medium effect sizes according to Cliff’s delta. These findings indicate that they retained the clearest and most consistent class differences after residualization.
A second set of features showed weaker but still statistically supported differences. This group included disfluency_resid, cognitive_effort_resid, sentence_similarity_std_resid, embodied_ratio_resid, direction_resid, subordination_rate_resid, mattr_resid, and hapax_ratio_resid. All remained significant after FDR correction in both the KS and permutation tests, but their effect sizes ranged from negligible to small. This suggests that they still captured informative class differences, although the strength of separation was more modest than in the strongest group.
The remaining features showed limited statistical support after length control. These included topic_drift_score_resid, local_incoherence_rate_resid, sentence_similarity_mean_resid, interpretive_ratio_resid, hesitation_resid, dependency_tree_depth_max_resid, dependency_tree_depth_mean_resid, and experience_resid. Although several of these features remained significant under the KS test, their permutation results did not remain significant after FDR correction, and their effect sizes were negligible. This pattern suggests that any remaining class differences were weak, localized, or more strongly reflected in distributional shape than in robust central separation.
Taken together, these results show that only a limited subset of features retained strong statistical separation after transcript-length control, while a larger group showed weaker or more localized differences.
5.2. Baseline Classification Results and Ablation Study
While the statistical tests in
Section 5.1 evaluated each residualized feature individually, the baseline classifier addresses a different question: whether the feature set, when considered jointly, provides useful discriminative power for classifying authentic and manipulated transcripts. We therefore report baseline classification performance as an assessment of practical predictive utility in a multivariate setting. To further examine the contribution of each feature group, we conduct an ablation study at both the group level and the individual feature level.
Using the full residualized feature set, the linear SVM achieved an Accuracy of 0.7211, a Balanced Accuracy of 0.7211, an F1-score of 0.7338, and an AUC of 0.7808. These results indicate that the proposed transcript-level features retain meaningful discriminative power after length control and allow consistent separation between authentic and manipulated transcripts in a lightweight baseline setting.
At the group level, the linguistic feature set alone achieved performance comparable to the combined model (Accuracy = 0.7177, AUC = 0.7597), whereas the psycholinguistic feature set alone showed notably lower performance (Accuracy = 0.5918, AUC = 0.6337). This suggests that structural and discourse-level cues contribute more strongly to classification in isolation, while psycholinguistic-style features provide complementary information that improves performance when combined with linguistic features.
Table 5 shows the baseline classification results for the combined feature set and the two feature groups.
At the individual feature level, no single feature approached the performance of the full model. The strongest standalone feature was coordination_rate_resid, which achieved an Accuracy of 0.6803 and an AUC of 0.6827, followed by features such as topic_drift_score_resid and cognitive_effort_resid, which showed moderate standalone performance. However, most individual features yielded substantially lower results, indicating that predictive signal is distributed across multiple cues rather than concentrated in a single dominant feature.
Table 6 summarizes the single-feature baseline classification results.
Overall, the ablation results suggest that classification performance benefits from combining heterogeneous feature types. While linguistic features capture strong structural signals, and psycholinguistic features capture more localized behavioral cues, their integration leads to more robust and stable predictive performance.
5.3. Interpretation of Linguistic Patterns
The interpretation integrates statistical evidence and baseline classification results, treating them as complementary indicators of reliable separation and practical utility.
Clause-linking behavior provides the clearest signal. coordination_rate_resid aligns across both statistical and baseline analyses, indicating a stable and useful distinction between classes. This suggests that authentic speech exhibits more incremental idea expansion, whereas manipulated transcripts follow more constrained linking patterns. In contrast, subordination_rate_resid contributes weaklier, indicating that hierarchical structuring plays a more limited and context-dependent role.
Semantic continuity results emphasize variability over averages. sentence_similarity_std_resid retains statistical support but shows limited and imbalanced standalone predictive performance, whereas sentence_similarity_mean_resid does not. This suggests that authentic speech is characterized by uneven semantic progression, while manipulated transcripts are more uniform. Similarly, topic_drift_score_resid shows weak standalone performance despite limited statistical support, indicating that its contribution is primarily contextual within a multivariate setting.
Lexical and structural features show mixed contributions. mattr_resid provides comparatively stronger predictive signal among lexical features, while hapax_ratio_resid is weaker. Dependency depth features show limited statistical and predictive utility, likely due to constraints imposed by short transcript segments.
Overall, the results indicate that linguistic differences are distributed across multiple interacting features, with the most informative signals arising from discourse organization and variability in semantic progression rather than from isolated surface-level properties.
5.4. Interpretation of Psycholinguistic Patterns
The psycholinguistic features reflect production-related cues associated with planning, monitoring, and communicative framing. As in the linguistic analysis, interpretation is based on both statistical evidence and baseline performance, which do not always align.
A key observation is that some features show statistical separation but weak standalone predictive utility. For example, emotion_anchor_resid retains strong statistical evidence but performs poorly in isolation, suggesting that affect-related cues vary across classes but are not sufficient as independent discriminators. Similarly, disfluency_resid and cognitive_effort_resid show statistically supported differences but only limited predictive strength, indicating that production-related signals are present but distributed.
Other features, such as interpretive_ratio_resid, hesitation_resid, and experience_resid, show weak or inconsistent patterns across both analyses. This suggests that these cues are highly context-dependent and influenced by factors such as topic and prompt structure, rather than providing stable class-level distinctions.
Overall, the psycholinguistic results indicate that production-related differences between authentic and manipulated speech are subtle and dispersed across multiple cues. Rather than acting as strong independent signals, these features contribute primarily as complementary information within a multivariate setting. This supports the view that manipulated speech may approximate surface-level fluency, while differing in how production-related variability emerges across the transcript.
5.5. Implications and Limitations
The results highlight that transcript-level cues can provide meaningful signals for distinguishing authentic and manipulated speech, particularly when multiple feature types are combined. Rather than relying on a single dominant indicator, the findings suggest that useful information is distributed across linguistic and psycholinguistic features, with stronger signals emerging from discourse organization and weaker but complementary cues arising from production-related patterns.
At the same time, several limitations should be considered. First, the use of residualization to control transcript length may remove not only trivial length effects but also meaningful variation related to discourse complexity and semantic development. This may partially attenuate differences between classes. Second, the dataset consists of relatively short, segmented interview clips, which constrain the expression of certain features, particularly those related to syntactic depth and extended discourse structure. As a result, some features may appear less informative under these conditions.
Finally, while the proposed features are interpretable and theoretically motivated, they do not capture all possible signals of manipulation. Instead, they provide a structured subset of cues that can be combined with other modeling approaches. Future work could extend this analysis by evaluating longer-form speech, incorporating alternative length-control strategies, and testing generalization across datasets and domains.
6. Conclusions
This study examined whether transcript-level features can reveal reliable differences between authentic and manipulated speech under identity-preserving semantic manipulation. Using FakeSpeech+, we adopted a transcript-first approach and analyzed interpretable linguistic and psycholinguistic features under transcript-length control. The results showed that only a limited subset of cues retained strong statistical separation after residualization, while a broader set of features showed weaker but still informative differences. Baseline experiments further indicated that useful predictive signals are distributed across multiple complementary cues rather than concentrated in a single standalone feature.
Beyond the feature-level findings, FakeSpeech+ constitutes an important contribution by providing a dataset specifically designed for transcript-level analysis under controlled semantic manipulation with preserved speaker identity. This addresses a dataset gap not explicitly covered by existing fake-speech benchmarks, which have largely focused on spoofing, voice cloning, or broad multimodal manipulation. Overall, the findings support the value of transcript-based analysis for identity-preserving fake speech and provide a foundation for future work on robustness, generalization, and transcript-level detection under realistic manipulation settings.
Future work will aim to reduce the reliance on residualization by constructing a more length-balanced dataset so that authentic and manipulated transcripts are more comparable in length from the outset. We also plan to develop and evaluate a dedicated transcript-level detection framework based on the linguistic and psycholinguistic characteristics identified in this study.