Next Article in Journal
“I Was Everything What I Never Wanted to Be”—Exploring Moral Injury Within Forensic Healthcare Settings
Previous Article in Journal
Housing Fragility: Wealth Position, Portfolio Composition, and Education Among Homeowners
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Exploring Ethnicity and Gender Bias in TED Talks: A Study of Audience Online Reactions

by
Meriem El-Yamri
1,
Miguel Ángel Violán
2 and
Borja Manero
1,*
1
Department of Software Engineering and Artificial Intelligence, University Complutense of Madrid, 28040 Madrid, Spain
2
Facultad de Comunicación, Universidad Internacional de Catalunya, 08017 Barcelona, Spain
*
Author to whom correspondence should be addressed.
Soc. Sci. 2026, 15(7), 428; https://doi.org/10.3390/socsci15070428
Submission received: 31 March 2026 / Revised: 24 June 2026 / Accepted: 25 June 2026 / Published: 29 June 2026

Abstract

Audience reactions to oral communication are shaped by both communicative practices and broader social contexts. While elements such as message content, delivery style, and vocal expression can be developed through training, other factors—such as gender and ethnicity—reflect social identities that are often associated with how speakers are perceived and evaluated. This study examines how these contextual attributes are associated with audience engagement in digital public speaking environments. Drawing on an initial dataset of 977 TEDx talks, resulting in two high-confidence subsamples of 610 speakers for gender and 387 for ethnicity, curated through a combination of computational methods with a communication perspective. We analyzed the relationship between the two factors with engagement indicators—including likes, dislikes and interaction rates. The analysis explores whether patterns of audience response differ across demographic groups and at the intersection of gender and ethnicity. The findings reveal that neither gender nor ethnicity, considered on its own, was significantly associated with audience engagement; differences emerged only at the intersection of the two. Specifically, non-Hispanic Black speakers were associated with higher levels of negative feedback in both genders, Hispanic male speakers received more positive engagement than other male speakers, and Asian female speakers showed lower interaction levels—fewer views, likes, and comments—than non-Hispanic White female speakers. These patterns suggest that disparities in how audiences respond to speakers’ social identities in mediated contexts are intersectional, becoming visible only when gender and ethnicity are considered jointly. By providing empirical evidence from a diverse digital corpus, this study contributes to ongoing debates on digital inequalities, representation, and participation in contemporary media environments, highlighting the importance of considering social context in analyses of audience behavior.

1. Introduction

Public communication has undergone a profound transformation with the rise in digital platforms, where oral communication is now shaped by both immediate communicative practices and broader social architectures. While traditional public speaking effectiveness is often measured through verbal content (Englehart 2004), delivery style, and vocal expression (Jasuli et al. 2024; Phutela 2015), digital environments introduce new layers of interaction that reflect and potentially amplify social identities. TEDx talks, functioning as a unique “cultural franchise” system with over 12,000 events in 150 countries, provide an ideal global corpus for studying these dynamics. Unlike the centralized TED stage, TEDx’s decentralized nature allows for a more diverse array of speakers, yet it simultaneously acts as a powerful platform where the underrepresentation or biased reception of certain groups can reinforce or amplify harmful societal stereotypes and negative attitudes among the public (Schwemmer and Jungkunz 2019).
Despite the platform’s goal of democratizing “ideas worth spreading,” audience reactions in digital spaces are often influenced by prediscursive factors such as gender and ethnicity. These attributes shape the perceived credibility and expertise of speakers, triggering interaction patterns—likes, dislikes, and comments—that may mirror existing social hierarchies (Vrij and Winkel 1994). This study provides an original intersectional analysis by examining how these social identities concurrently shape engagement within the TEDx ecosystem. By bridging classical rhetorical concepts of ethos with large-scale computational methods, we fill a critical gap: understanding how the digital architecture of interaction acts as a filter that mirrors or amplifies prediscursive social biases in a globalized platform for public discourse.
Consequently, this research is driven by the following questions:
RQ1: In what ways is the gender of a speaker associated with variations in online audience engagement and specific interaction patterns?
RQ2: How does the perceived ethnicity of a speaker relate to differences in the levels of positive and negative engagement from digital audiences?
To ensure sociological accuracy, we recognize the fundamental distinction between biological sex and gender as a social construct. However, as this study relies on automated tools that infer categories based on naming conventions and audience perception, we use “gender” to refer to the gender identity perceived by the audience, which triggers the prediscursive interaction patterns under analysis (Brooks et al. 2014).

2. State of the Art

2.1. TED Talks

TED Talks represent a unique hybrid genre that bridges the gap between popularization and academic exposition. Originating in the United States in 1984, TED Talks were initially created as a private initiative dedicated to disseminating concise, engaging presentations on a wide array of topics. These talks are designed for a medium- to highly educated audience, delivered in English, and are known for their entertaining style and brief format (Anderson 2016). The primary audience for TED Talks consists largely of non-experts, making them an accessible platform for sharing complex ideas in a way that is understandable and engaging to the general public (Violán Galán and Mendiz 2023; Sugimoto et al. 2013; Larivière et al. 2013). TED Talks are centered around a single, powerful idea and are delivered to both live audiences and remote viewers through online video platforms. Forty years after its creation, TED has grown into a global cultural and educational phenomenon with solid roots and millions of followers around the world (Sugimoto et al. 2013; Cadwalladr 2011; Ferica 2012). By July 2024, the TED platform had surpassed 6600 talks available for free online viewing, further solidifying its role as a significant player in global knowledge dissemination. The TEDx program, created in 2009, is a milestone in the organization’s history, generating a dramatic increase in brand awareness through a unique cultural franchise system open to communities worldwide. To date, more than 12,000 TEDx events have been held in 150 countries, with a total of 50,000 talks. Currently, more than 3000 TEDx events are held each year. Given the widespread popularity, accessibility, and educational value of TED and TEDx Talks, they provide an ideal corpus for studying the role of various factors—such as gender and ethnicity—on audience perception and engagement. The global reach and diverse subject matter of these talks make them particularly suitable for examining how different speaker characteristics influence audience reactions in a broadly applicable and highly relevant context.

2.2. Theoretical Framework: Prediscursive Ethos and Social Signaling

To understand how audience engagement is shaped in digital environments, this study adopts a framework that bridges classical rhetoric with social signaling theory. In The Rhetoric, Aristotle (Cope, 1877) identifies ethos—the speaker’s perceived character—as a fundamental pillar of persuasion. Critically, this concept is bifurcated into discursive ethos (the image constructed through the speech itself) and pre-discursive ethos (the prior reputation or identity attributes that precede the speaker).
In the context of digital public speaking, we argue that demographic markers such as gender and ethnicity function as primary social signals that constitute the speaker’s pre-discursive ethos. These markers activate audience heuristics and stereotypes before the first word of a talk is spoken. The perception of a speaker’s expertise and social status can be significantly influenced by these prediscursive signals (Curtis et al. 2015; Leongómez et al. 2017), acting as a sociotechnical filter in digital spaces (Guyer et al. 2021; Ambady et al. 2002).
This theoretical model allows us to connect computational demographic inference directly to rhetorical analysis. By using automated tools to categorize gender and ethnicity, we are essentially modeling the perceived pre-discursive identity that triggers the audience’s interaction patterns (likes, dislikes, and comments). Thus, the engagement metrics on platforms like YouTube represent a quantified measure of how pre-discursive ethos influences the reception of “ideas worth spreading,” revealing whether certain demographic signals face a systematic “credibility tax” independently of their discursive content.

2.3. Speech Analysis with Automated Tools

In recent years, there has been a surge in interest in using automated tools to analyze the effectiveness of speeches based on the previous categories.
The verbal content of a speech, which includes the presence of specific keywords, the use of rhetorical devices, and the overall structure of the argument, is one area of focus. Natural Language Processing (NLP) (Sharifani and Amini 2023) techniques are commonly used to extract these features. Additionally, machine learning algorithms have been employed in some studies to predict the persuasiveness of a speech based on its verbal content. These methods also encompass the detection of hate speech and negative interactions in social media-based or sentiment analysis based on discourse content (Martins et al. 2018; El-Yamri et al. 2019).
Paraverbal aspects of a speech, such as the speaker’s tone, volume, and valence, are also analyzed using automated tools (El-Yamri et al. 2019). Speech recognition software is often used to extract features like the pitch and intensity of the speaker’s voice. Some studies have also used machine learning algorithms to predict the persuasiveness of a speech based on its paraverbal features.
Nonverbal cues exhibited by a speaker during a speech, such as gestures, facial expressions, and posture, are another area of focus (Azemi 2021). Video analysis software is commonly used to extract features like the frequency and type of gestures and the presence of certain facial expressions. Some studies have also used machine learning algorithms to predict the persuasiveness of a speech based on its nonverbal cues.
Finally, the impact of contextual factors such as the speaker’s age, gender, appearance, social status, knowledge about the subject, and ethnicity on the effectiveness of a speech is also analyzed (Pair et al. 2021; Monteiro et al. 2024). Machine learning algorithms are often used to predict the persuasiveness of speech based on these factors. Some studies have also focused on the use of automated tools to analyze the impact of the audience’s knowledge of the subject matter on the effectiveness of a speech.
In conclusion, the use of automated tools in speech analysis has become increasingly prevalent, offering valuable insights into various aspects of public speaking, further providing an understanding of the complex dynamics of public speaking. In the next section, we will discuss specific examples of studies that focus on gender and ethnicity in public communication.

2.4. Gender and Ethnicity in Communication

2.4.1. Gender

Studies show that the gender of the speaker influences the persuasion process. For example, Strach et al. (2015) analyzed political advertisements and found that female voices are perceived as less credible than male voices, especially on issues considered “masculine” such as national defense. However, women are more credible on “feminine” issues such as education. Subsequent experiments, such as that of Searles et al. (2020), confirm that the male voice is preferred for masculine topics, while the female voice is less effective on these topics. Despite this diversity in the field of science, some studies have revealed that biases against women scientists have extended to women science communicators (Liu 2022). For example, science-related channels hosted by women on YouTube are significantly more likely to receive hostile or negative comments than their male counterparts (Amarasekara and Grant 2019), suggesting that people are often more critical of female science communicators. Another study around economics (Bodea et al. 2021) examines whether gender influences the effectiveness of central bankers as communicators, especially how they shape economic expectations and trust in the European Central Bank. Schwemmer and Jungkunz (2019) further highlight that in the context of TED Talks, women speakers not only remain underrepresented overall, especially when they are also part of an ethnic minority, but they also tend to receive more negative sentiment in online reactions compared to their male peers. This suggests that gender bias persists even in popular science communication platforms and can affect how audiences engage with speakers.

2.4.2. Ethnicity

The role of perceived ethnicity on the credibility and perception of speakers has been the subject of several studies, demonstrating how race and ethnicity influence the way individuals are perceived and evaluated in communication contexts. One study (Haut et al. 2021) explores how credibility is affected in video conferencing if a person changes their ethnicity in real time using deepfakes. The results show that changing race in a static image has minimal impact on credibility, but doing so in a video significantly increases credibility. In addition, a sentiment analysis study revealed that people justify the credibility of an individual of White ethnicity with a more positive sentiment. Another study on perceived ethnicity (Lee and Bailey 2023) investigated how reverse linguistic stereotyping (RLS) affects perceived accent and comprehensibility in speech. There are several studies that analyze a speaker’s accent rating based on the perceived ethnicity (Gnevsheva 2018). Squizzero (2020) investigates how perceived ethnicity influences accent perception in non-native speakers, focusing on a Mandarin-speaking context. Ethnic Chinese participants assessed the personalities and linguistic abilities of ethnic Chinese and non-Chinese Mandarin speakers according to the perceived ethnicity of the speaker.
Although there is some research exploring the influence of gender and ethnicity on public communication, these studies primarily focus on specific contexts such as political communication, science communication, and economic discourse. However, none of these studies specifically address the unique context of public speaking in front of a general audience, nor do they conduct a cross-comparison between both gender and ethnicity factors in this setting. Our study fills this gap by examining how these factors interact to influence audience perception and engagement in the broader and more diverse context of TED Talks, providing new insights into the dynamics of public speaking based on prediscursive conditions.

3. Materials and Methods

3.1. Experimental Design

Using a Python script (3.14.1 version), we analyzed a random sample of 977 YouTube videos of TEDx talks. These videos were collected using the YouTube Data API v3, utilizing the search term “TEDx” and identifying videos with titles formatted as <title of the talk> |<speaker’s name> | <TEDx branch>”, such as “Will AI be able to speak your language?|Linda Hemisdottir|TEDxReykjavik”. Based on the speakers’ names, and the video interactions (views, likes, dislikes, and comments), we used Machine Learning models to extract the ethnicity and gender of each speaker, and subsequently, the dataset was augmented with the video engagement rate, percentages of likes, percentages of dislikes, and percentages of comments based on views. Finally, we conducted Kruskal–Wallis tests on the dataset for gender and ethnicity to find statistically significant data for video interactions based on those variables.
While the TEDx channel hosts hundreds of thousands of videos, sampling was necessary due to usage limitations imposed by the YouTube Data API, as well as the computational cost of extracting and processing metadata at scale. We did not apply a minimum view threshold, so the sample naturally reflects the long tail of talks with low engagement. Although the sample includes TEDx events from around the world, the geographic distribution of talks was not analyzed; this limitation and its implications for interpreting ethnicity classifications are further discussed in Section 6.2. All data will be made available from the authors upon reasonable request to ensure replicability.
To ensure the reliability of the demographic inferences, the initial sample of 977 TEDx talks was filtered based on the confidence scores provided by the automated tools. We only included speakers whose gender could be predicted with a probability higher than 0.8, resulting in a subsample of 610 individuals for the gender analysis (266 female and 344 male). For ethnicity, we applied a 0.7 accuracy threshold, which, combined with the exclusion of the ‘Other’ category due to its low representation, resulted in a final subsample of 387 speakers (231 Non-Hispanic White, 55 Asian, 72 Non-Hispanic Black, and 29 Hispanic). This rigorous filtering process explains the numerical discrepancy between the total videos collected and the specific groups analyzed, prioritizing data precision over sample volume.
The reduction in sample size from 977 to 387 for the ethnicity subsample was a deliberate methodological choice to mitigate the risk of algorithmic ‘hallucination’ or misclassification. By prioritizing precision over volume, we ensured that the categories analyzed were based on high-probability naming conventions. While this introduces a potential selection bias by excluding rare or non-Western naming patterns, it preserves the internal validity of the study by focusing on speakers whose perceived identity is most likely to be clearly categorized by an audience using similar social heuristics.

3.2. Determining Ethnicity

To determine the ethnicity, we used a pytorch (https://github.com/pytorch/pytorch, accessed on 15 November 2025) implementation of a previously trained model called ethnicolr (2024). The package uses the US census data, and the Wikipedia data collected by Ambekar et al. (2009), and the Florida voting registration data to build models to predict the race and ethnicity in the categories Non-Hispanic Whites, Non-Hispanic Blacks, Asians, Hispanics, and Other; based on the first and last name or just the last name.
Namesurnameprediction
JessicaMcCabenon-Hispanic White
IshitaGuptaAsian
BernardoRezendeHispanic
We recognize that ethnicolr (2024) is primarily trained on North American census data. However, its application to the TEDx corpus is justified for two reasons: (1) TEDx talks are predominantly delivered in English and follow a Westernized rhetorical format, aligning the audience’s perceptual framework with the categories provided by the tool; and (2) in digital environments, names act as a primary prediscursive signal. Even if a speaker’s self-identified ethnicity differs from the tool’s classification, the tool effectively models the perceived ethnicity that triggers automated audience reactions (likes/dislikes) in a globalized digital space.

3.3. Determining Gender

To determine the gender of each speaker, we used the API Genderize.io (2024), a tool that maps a given first name to a specific gender. For example, the results for the previous names are:
{ ’ count ’ : 907046 , ’ name ’ : ’Jessica’ , ’ gender ’ : ’ female ’ , ’probability’ : 1 . 0 }
{ ’ count ’ : 7550 , ’ name ’ : ’Ishita’ , ’ gender ’ : ’ female ’ , ’probability’ : 1 . 0 }
{ ’ count ’ : 55445 , ’ name ’ : ’ Bernardo ’ , ’ gender ’ : ’ male ’ , ’ probability’ : 1 . 0 }

3.4. Interaction Data

For this study, we created a dataset using the previous data from the YouTube Data API v3 (https://developers.google.com/youtube/v3, accessed on 12 September 2025), and we included interaction data. As shown in Table 1, we considered 8 measures of interaction with each TEDx Talk video: engagement rate, number of likes, number of dislikes, view count, comment count, percentage of likes, percentage of dislikes, and percentage of comments.
Note: Engagement rate by reach is a common method used in social media that divides the total number of engagements by the total number of people who saw a post. This gives an indication of how engaging the content is for those exposed to it (https://www.linkedin.com/advice/0/whats-your-social-media-engagement-rate, accessed on 30 March 2026).
Next, we conducted Kruskal–Wallis tests to compare the differences in audience reactions based on the speaker’s ethnicity and gender across several metrics (viewCount, likeCount, dislikeCount, commentCount, engagement, pct like, pct dislike, and pct comment).
The engagement formula (likes-dislikes-comments)/views is designed to measure the net valence of interaction. While likes and dislikes represent binary sentiment, we include comments as a measure of high-effort engagement. We acknowledge that these metrics carry different qualitative weights; however, this unified metric provides a standardized baseline to compare the overall ‘friction’ or ‘reception’ across different demographic groups in a way that view counts alone cannot capture.

4. Results

Given the non-normal distribution and unequal sample sizes, we employed the Kruskal–Wallis test, a non-parametric method ideal for assessing statistically significant differences in medians without the stringent assumptions of normality. To determine the practical significance of these findings, we also report effect sizes, allowing us to distinguish between statistical probability and the magnitude of the observed disparities.
To complement the significance tests, we computed epsilon-squared (ε2 = H/(n − 1)) as the effect size for each Kruskal–Wallis test. Following conventional benchmarks, ε2 values of approximately 0.01, 0.06, and 0.14 were interpreted as small, medium, and large effects, respectively.
It is particularly useful for ordinal data or continuous data that do not meet the assumptions required for parametric tests. By using this test, we can robustly determine if the audience’s engagement metrics vary significantly across different speaker contexts, such as gender and ethnicity, without the stringent assumptions of normality and homogeneity of variances.
The following results are structured to move from aggregate demographic factors to their intersectional effects. First, we present the analysis of gender (Table 2) and ethnicity (Table 3) as independent main effects. While these aggregate analyses show no statistically significant differences, they provide the necessary baseline for the subsequent intersectional analysis. Finally, we present the results of ethnicity stratified by gender (Table 4), which reveals the specific disparities that were previously masked at the aggregate level.
Kruskal–Wallis tests revealed no statistically significant differences in metrics based on the speaker’s gender, and the associated effect sizes were negligible across all metrics (ε2 ≤ 0.004). Specifically, the results for viewCount (χ2 = 1.8075, p = 0.1788, ε2 = 0.003; Wilcoxon scores: female = 309.907, male = 329.615), likeCount (χ2 = 1.2843, p = 0.2571, ε2 = 0.002; female = 311.571, male = 328.183), and dislikeCount (χ2 = 0.6881, p = 0.4068, ε2 = 0.001; female = 315.328, male = 324.951) were not significant. Although commentCount (χ2 = 2.3511, p = 0.1252, ε2 = 0.004; female = 308.421, male = 330.894) and engagement (χ2 = 2.6978, p = 0.1005, ε2 = 0.004; female = 333.436, male = 309.369) approached significance, their effect sizes remained negligible, indicating that these trends reflect minimal practical differences.
Considered as a main effect across the full ethnicity sample, perceived ethnicity showed no statistically significant association with any of the engagement metrics tested at the aggregate level (viewCount χ2 = 0.87, p = 0.832; likeCount χ2 = 1.13, p = 0.769; dislikeCount χ2 = 1.63, p = 0.652; commentCount χ2 = 1.96, p = 0.581; engagement χ2 = 6.08, p = 0.108), with negligible effect sizes (all ε2 ≤ 0.016). Likewise, speaker gender alone yielded no significant differences (Table 2).
While neither gender nor ethnicity yielded significant differences when analyzed in isolation (see Table 2 and Table 3), a different pattern emerged during the intersectional analysis. As shown in Table 4, when ethnicity is examined within each gender stratum, significant disparities with medium-to-large effect sizes become visible. This confirms that the audience’s reaction to the speaker’s social identity is not driven by a single factor but by the intersection of both. Among female speakers, Asian women received significantly fewer views, likes, and comments than non-Hispanic White women, while non-Hispanic Black women showed the highest levels of dislikes. Among male speakers, non-Hispanic Black men received markedly more dislikes—both in absolute count and as a percentage (ε2 = 0.095 and 0.107, the largest effects observed)—whereas Hispanic men obtained the highest engagement and the most likes per view.
Although gender showed no significant differences, Figure 1 illustrates the direction and magnitude of the female–male rank differences across metrics. Each point shows the difference in rank scores between female and male speakers, with the vertical dashed line representing no difference; points to the right indicate higher ranks for female speakers, and points to the left, higher ranks for male speakers. None of these differences reached conventional significance (p < 0.05), indicating that the observed variation likely reflects random sample variation rather than systematic gender bias.
Figure 2 summarizes ethnicity differences in engagement separately for female and male speakers. For each metric, the four points show the mean Wilcoxon rank of each ethnic group as a deviation from the average rank within that gender (vertical dashed line at zero): points to the right indicate above-average ranks for that group, and points to the left, below-average ranks. Only metrics with statistically significant ethnicity differences within each gender are shown.

5. Discussion

5.1. RQ1: In What Ways Is the Gender of a Speaker Associated with Variations in Online Audience Engagement and Specific Interaction Patterns?

Our results did not show statistically significant differences in most engagement metrics based solely on the speaker’s gender. While gender alone may not have shown significant differences, we observed exploratory differences when gender was considered alongside ethnicity. These results suggest that gender’s effect on audience perception is intertwined with ethnicity, indicating that the intersection of these factors plays a critical role in shaping audience reactions.
This contrasts with the study by Amarasekara and Grant (2019), which found that women, particularly in science-related YouTube channels, are more likely to receive hostile or negative comments compared to their male counterparts. In our study, female TED Talk speakers did not exhibit comparable levels of negative engagement, suggesting that the TED platform or its audience may help mitigate some of the biases observed in other online environments. Nonetheless, our findings still point to disparities in how different gender and ethnicity combinations are perceived.
In conclusion, while gender alone may not significantly influence audience reactions, its effect becomes clearer when examined in conjunction with ethnicity. These results underscore the importance of considering multiple contextual factors when analyzing audience engagement in online public speaking environments.

5.2. RQ2: How Does the Perceived Ethnicity of a Speaker Relate to Differences in the Levels of Positive and Negative Engagement from Digital Audiences?

At the aggregate level, perceived ethnicity was not significantly associated with any engagement metric (all p > 0.10; Table 3). Meaningful differences emerged only once gender was taken into account (Table 4): among male speakers, non-Hispanic Black individuals received substantially more dislikes than Asian and non-Hispanic White speakers, while Hispanic men obtained the highest engagement and the most likes per view; among female speakers, Asian women received fewer views, likes, and comments than non-Hispanic White women, and non-Hispanic Black women again showed the highest levels of dislikes.
These results indicate that the association between perceived ethnicity and audience engagement is conditional on gender rather than a uniform main effect. While the disparities affecting non-Hispanic Black speakers align with Haut et al. (2021) regarding the influence of race on speaker credibility, we characterize them as observed disparities in digital interaction rather than definitive proof of individual prejudice. This distinction is crucial, as the data reflect aggregate audience behavior within a specific digital architecture.
These disparities underscore that online reactions are not merely a response to the quality of the ‘ideas worth spreading’ but are also shaped by the speaker’s pre-discursive identity. The digital architecture of platforms like YouTube, while seemingly neutral, can function as a sociotechnical filter where existing social hierarchies are re-enacted. By receiving more negative interactions—particularly among male speakers—non-Hispanic Black speakers face a higher ‘entry cost’ for digital visibility, as their pre-discursive ethos, conditioned by racialized perceptions, triggers interaction patterns that differ from those of their White or Hispanic counterparts. This illustrates a ‘digital paradox’: a platform built for the global democratization of knowledge can simultaneously become a space where prediscursive biases surface through automated feedback loops of likes and dislikes. This phenomenon, where external cues influence message reception, aligns with research on how attire and gender interact in the perception of speaker charisma (Brem and Niebuhr 2021).

5.3. Gender and Ethnicity Combined Analysis

The combined analysis of gender and ethnicity showed that the associations between perceived ethnicity and audience engagement were intersectional, emerging within gender strata even though ethnicity displayed no significant main effect across the full sample (Table 4). Among male speakers, Hispanic individuals showed the highest engagement and the most likes per view, significantly exceeding non-Hispanic White speakers, while non-Hispanic Black men received markedly more dislikes—both in absolute number and as a percentage—than Asian and non-Hispanic White men. Among female speakers, a different pattern emerged: Asian women received significantly fewer views, likes, and comments than non-Hispanic White women, whereas non-Hispanic Black women showed the highest levels of dislikes. These results underscore the complex interplay between gender and ethnicity in shaping audience reactions: both the group facing the most negative feedback—non-Hispanic Black speakers, in both genders—and the specific metrics affected differ across gender groups, indicating that these effects cannot be captured by either factor in isolation.
These findings connect with the conclusions of Schwemmer and Jungkunz (2019), who highlighted that while gender and ethnicity shape sentiment on the TED stage, the content and topic of the talk remain the strongest predictors of audience sentiment, especially for negative feedback. Our study did not analyze topic or thematic content directly; instead, it shows that even when controlling only for speaker context, significant disparities in engagement metrics appear, particularly in how non-Hispanic Black speakers receive more negative interactions than other groups, regardless of content. This suggests that biases related to speaker context may persist independently of topic. Future research should therefore combine both perspectives, examining how the interaction between topic, speaker identity, and audience context amplifies or mitigates online biases in public speaking. Doing so could help clarify under what conditions content-related factors outweigh speaker context, and vice versa, providing a more comprehensive understanding of how biases are formed and reinforced in digital public discourse.
These findings provide empirical weight to the Aristotelian concept of pre-discursive ethos (see Section 2.2), suggesting that in digital public speaking, the audience’s judgment begins before the first word is spoken. While the TED platform aims to democratize ‘ideas worth spreading,’ our data suggests that the digital architecture of interactions may function as a filter that reflects existing social disparities. The higher levels of negative engagement associated with non-Hispanic Black speakers suggest that mediated environments may facilitate differential interaction patterns that are less visible in face-to-face academic settings. Rather than claiming a direct amplification of societal hierarchies, our findings highlight a ‘digital paradox’: platforms designed for global inclusion nonetheless exhibit interaction metrics that mirror the complexities of social identity found in broader society.
To conclude the discussion, while our primary focus was on examining the influence of ethnicity on audience engagement, the results also provide insights into how other contextual factors, such as the online environment, might shape audience behavior. The significant disparities in engagement metrics observed here suggest that digital interaction patterns may follow distinct dynamics compared to offline settings. We must be cautious in extrapolating these findings to real-world audiences, as the anonymity and binary feedback mechanisms (likes/dislikes) of YouTube create a unique communicative environment that may not mirror face-to-face professional public speaking. Understanding the potential differences between online and in-person reactions is crucial, as it could inform strategies for more equitable public communication across diverse settings.

6. Conclusions and Future Work

6.1. Conclusions

This study explores how prediscursive factors—specifically perceived gender and ethnicity—are associated with audience engagement in the digital ecosystem of TEDx. Our results indicate that while gender identity alone did not show a statistically significant relationship with engagement metrics in this specific corpus, ethnicity is associated with significant variations in how audiences interact with speakers.
Specifically, when examining the intersection of these factors, we observed that certain groups, such as Hispanic male speakers, achieved higher engagement scores, while others, notably non-Hispanic Black speakers, received higher levels of negative interaction (dislikes). While these findings align with previous literature regarding the impact of identity on speaker perception (Haut et al. 2021; Strach et al. 2015), we characterize these results as observed disparities in digital interaction patterns rather than definitive evidence of universal societal prejudice. These findings also resonate with the patterns of online hostility and engagement disparities observed in other science communication platforms (Amarasekara and Grant 2019; Schwemmer and Jungkunz 2019), reinforcing the idea that digital architectures can mirror societal hierarchies.
Furthermore, our findings suggest that contextual factors related to the speaker’s identity are associated with engagement outcomes independently of the talk’s topic. This indicates that the “ideas worth spreading” on the TEDx stage do not reach the audience in a vacuum; instead, they are filtered through the audience’s prediscursive perceptions. Future research should integrate multivariate models to further isolate these variables and determine the practical significance of these observed disparities across diverse digital and face-to-face environments.

6.2. Limitations

This study has several limitations that must be addressed to ensure a critical interpretation of the findings. First, the automated classification of gender and ethnicity relies on naming conventions, which does not account for the full spectrum of non-binary identities or the complex, self-identified heritage of speakers. This introduces potential classification bias, particularly for multicultural or non-Western names that may not be fully captured by tools trained on North American census data.
Second, our rigorous filtering process—which prioritized high-accuracy probability scores (0.7 and 0.8)—resulted in a significant reduction in the sample size for ethnicity (from 977 to 387 speakers). While this was necessary to ensure data precision, it may introduce a selection bias by systematically excluding individuals with rare or complex names that the algorithms could not categorize with high confidence.
The reliance on automated tools also imposes categorical constraints: gender is inferred as binary and ethnicity is limited to four high-confidence categories (the “Other” category was excluded due to low representation). These choices do not capture non-binary identities or mixed and self-identified heritage, limiting the representativeness of the analyzed sample.
Relatedly, our ethnicity variable raises a construct-validity concern: it reflects perceived ethnicity inferred from names rather than speakers’ self-identified identity. Inferred race is therefore an algorithmic approximation and should not be interpreted as ground truth; we treat it as a prediscursive signal available to audiences, not as a definitive measure of identity.
Furthermore, this research focuses exclusively on YouTube engagement metrics (likes, dislikes, and comments). Although these digital footprints provide valuable insights into online behavior, they represent the reactions of a specific digital audience and may not be fully generalizable to face-to-face academic or professional public speaking settings.
Finally, the study does not account for the specific topic or the physical appearance of the speakers, both of which are known to interact with demographic perceptions to shape audience engagement.

6.3. Future Work

To build upon these findings, future research will move beyond bivariate analyses to implement multivariate regression models. This will allow us to examine the interaction between demographic factors while simultaneously controlling for mediating variables such as talk duration, publication date, and view count. Whereas effect sizes are already reported in the present study, these models will further isolate the practical significance of the observed disparities.
Additionally, we plan to analyze real-time audience reactions using image recognition technology to record facial expressions during speech delivery. This will provide a more granular understanding of how audience engagement unfolds in real-time, bridging the gap between digital metrics and psychological reactions. Ultimately, these insights will be integrated into a virtual reality (VR) trainer for public speaking, designed to help speakers understand and adapt to the impact of their pre-discursive identity on audience perception in diverse settings.

Author Contributions

Conceptualization, M.E.-Y. and B.M.; methodology, M.E.-Y. and B.M.; software, M.E.-Y.; validation, M.E.-Y. and B.M.; investigation, M.E.-Y.; data curation, M.E.-Y.; writing—original draft preparation, M.E.-Y.; writing—review and editing, M.E.-Y., M.Á.V. and B.M.; supervision, B.M.; project administration, B.M.; funding acquisition, B.M. All authors have read and agreed to the published version of the manuscript.

Funding

This research was funded by the Spanish Ministry of Science, Innovation and Universities, grant number PID2024-156187OB-I0.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

The data presented in this study are available on request from the corresponding author due to privacy restrictions.

Acknowledgments

The authors acknowledge the use of large language models (ChatGPT; version 5.5) for linguistic refinement and stylistic editing of the final manuscript.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Amarasekara, Inoka, and Will J. Grant. 2019. Exploring the YouTube science communication gender gap: A sentiment analysis. Public Understanding of Science 28: 68–84. [Google Scholar] [PubMed]
  2. Ambady, Nalini, Debi LaPlante, Thai Nguyen, Robert Rosenthal, Nigel Chaumeton, and Wendy Levinson. 2002. Surgeons’ tone of voice: A clue to malpractice history. Surgery 132: 5–9. [Google Scholar] [CrossRef] [PubMed]
  3. Ambekar, Anurag, Charles Ward, Jahangir Mohammed, Swapna Male, and Steven Skiena. 2009. Name-ethnicity classification from open sources. Paper presented at the 15th ACM SIGKDD international conference on Knowledge Discovery and Data Mining, Paris France, June 28–July 1; pp. 49–58. [Google Scholar]
  4. Anderson, Chris. 2016. TED Talks: The Official TED Guide to Public Speaking: Tips and Tricks for Giving Unforgettable Speeches and Presentations. London: Hachette UK. [Google Scholar]
  5. Azemi, Ilirijana. 2021. Non-verbal communication in public appearance. International Journal of Arts and Social Science 4: 256–67. [Google Scholar]
  6. Bodea, Cristina, Federico Maria Ferrara, Andrew Kerner, and Thomas Sattler. 2021. Gender and Economic Policy: When Do Women Speak with Authority on Economic Issues? Evidence from the Euro Area. Available online: https://www.researchgate.net/profile/Thomas-Sattler-3/publication/353017556_Gender_and_Economic_Policy_When_Do_Women_Speak_with_Authority_on_Economic_Issues_Evidence_from_the_Euro_Area/links/60e463e0458515d6fb02972c/Gender-and-Economic-Policy-When-Do-Women-Speak-with-Authority-on-Economic-Issues-Evidence-from-the-Euro-Area.pdf (accessed on 11 September 2025).
  7. Brem, Alexander, and Oliver Niebuhr. 2021. Dress to impress? On the interaction of attire with prosody and gender in the perception of speaker charisma. In Voice Attractiveness: Studies on Sexy, Likable, and Charismatic Speakers. Singapore: Springer, pp. 183–213. [Google Scholar]
  8. Brooks, Alison Wood, Laura Huang, Sarah Wood Kearney, and Fiona E. Murray. 2014. Investors prefer entrepreneurial ventures pitched by attractive men. Proceedings of the National Academy of Sciences 111: 4427–31. [Google Scholar] [CrossRef]
  9. Cadwalladr, Carole. 2011. TED’s Chris Anderson: The Man Who Made YouTube Clever. London: The Observer. [Google Scholar]
  10. Cope, Edward Meredith. 1877. The Rhetoric of Aristotle. Cambridge: Cambridge University Press, vol. 2. [Google Scholar]
  11. Curtis, Keith, Gareth J.F. Jones, and Nick Campbell. 2015. Effects of good speaking techniques on audience engagement. Paper presented at 2015 ACM International Conference on Multimodal Interaction, Seattle, WA, USA, November 9–13; pp. 35–42. [Google Scholar]
  12. El-Yamri, Meriem, Alejandro Romero-Hernandez, Manuel Gonzalez-Riojo, and Borja Manero. 2019. Emotions-responsive audiences for VR public speaking simulators based on the speakers’ voice. Paper presented at 2019 IEEE 19th International Conference on Advanced Learning Technologies (ICALT), Maceio, Brazil, July 15–18; pp. 349–53. [Google Scholar]
  13. Englehart, Nadine. 2004. Giving effective presentations. Operating Room Nurses Association of Canada Journal 22: 22–24. [Google Scholar]
  14. ethnicolr. 2024. Predict Race and Ethnicity from Name. Available online: https://ethnicolr.readthedocs.io (accessed on 30 March 2026).
  15. Ferica, Imy. 2012. Understanding TED as Alternative Media. Master’s thesis, University of Helsinki, Helsinki, Finland. [Google Scholar]
  16. Genderize.io. 2024. Find Gender from a Name. Available online: https://genderize.io (accessed on 11 September 2025).
  17. Gnevsheva, Ksenia. 2018. The expectation mismatch effect in accentedness perception of Asian and Caucasian non-native speakers of English. Linguistics 56: 581–98. [Google Scholar] [CrossRef]
  18. Guyer, Joshua J., Pablo Briñol, Thomas I. Vaughan-Johnston, Leandre R. Fabrigar, Lorena Moreno, and Richard E. Petty. 2021. Paralinguistic features communicated through voice can affect appraisals of confidence and evaluative judgments. Journal of Nonverbal Behavior 45: 479–504. [Google Scholar] [CrossRef] [PubMed]
  19. Haut, Kurtis, Caleb Wohn, Victor Antony, Aidan Goldfarb, Melissa Welsh, Dillanie Sumanthiran, Ji-ze Jang, Md. Rafayet Ali, and Ehsan Hoque. 2021. Could you become more credible by being white? Assessing impact of race on credibility with deepfakes. arXiv arXiv:2102.08054. [Google Scholar]
  20. Jasuli, Jasuli, Sri Fatmaning Hartatik, and Endang Setiyo Astuti. 2024. The impact of nonverbal communication on effective public speaking in English. Journey: Journal of English Language and Pedagogy 7: 226–32. [Google Scholar] [CrossRef]
  21. Larivière, Vincent, Chaoqun Ni, Yves Gingras, Blaise Cronin, and Cassidy R. Sugimoto. 2013. Bibliometrics: Global gender disparities in science. Nature 504: 211–13. [Google Scholar] [CrossRef] [PubMed]
  22. Lee, Bradford J., and Justin L. Bailey. 2023. Assumptions of speaker ethnicity and the effect on ratings of accentedness, comprehensibility, and intelligibility. Language Awareness 32: 301–22. [Google Scholar]
  23. Leongómez, Juan David, Viktoria R. Mileva, Anthony C. Little, and S. Craig Roberts. 2017. Perceived differences in social status between speaker and listener affect the speaker’s vocal characteristics. PLoS ONE 12: e0179407. [Google Scholar] [CrossRef] [PubMed]
  24. Liu, Sophia Ruijun. 2022. Gendered Science Communication: The Role of Speaker Gender and Pitch in Perceived Credibility and Persuasion of Climate Science. Ph.D. dissertation, University of Pennsylvania, Philadelphia, PA, USA. [Google Scholar]
  25. Martins, Ricardo, Marco Gomes, Jose Joao Almeida, Paulo Novais, and Pedro Henriques. 2018. Hate speech classification in social media using emotional analysis. Paper presented at 2018 7th Brazilian Conference on Intelligent Systems (BRACIS), Sao Paulo, Brazil, October 22–25; pp. 61–66. [Google Scholar] [CrossRef]
  26. Monteiro, Diego, Airong Wang, Luhan Wang, Hongji Li, Alex Barrett, Austin Pack, and Hai-Ning Liang. 2024. Effects of audience familiarity on anxiety in a virtual reality public speaking training tool. Universal Access in the Information Society 23: 23–34. [Google Scholar]
  27. Pair, Emma, Nikitha Vicas, Ann M. Weber, Valerie Meausoone, James Zou, Amos Njuguna, and Gary L. Darmstadt. 2021. Quantification of gender bias and sentiment toward political leaders over 20 years of Kenyan news using natural language processing. Frontiers in Psychology 12: 712646. [Google Scholar] [CrossRef] [PubMed]
  28. Phutela, Deepika. 2015. The importance of non-verbal communication. IUP Journal of Soft Skills 9: 43. [Google Scholar]
  29. Schwemmer, Carsten, and Sebastian Jungkunz. 2019. Whose ideas are worth spreading? The representation of women and ethnic groups in TED Talks. Political Research Exchange 1: 1–23. [Google Scholar] [CrossRef]
  30. Searles, Kathleen, Sophie Spencer, and Adaobi Duru. 2020. Don’t read the comments: The effects of abusive comments on perceptions of women authors’ credibility. Information, Communication & Society 23: 947–62. [Google Scholar]
  31. Sharifani, Koosha, and Mahyar Amini. 2023. Machine learning and deep learning: A review of methods and applications. World Information Technology and Engineering Journal 10: 3897–904. [Google Scholar]
  32. Squizzero, Robert. 2020. Attitudes toward L2 Mandarin speakers of Chinese and non-Chinese ethnicity. Paper presented at 32nd North American Conference on Chinese Linguistics, Hangzhou, China, April 19–21; pp. 521–38. [Google Scholar]
  33. Strach, Patricia, Katherine Zuber, Erika Franklin Fowler, Travis N. Ridout, and Kathleen Searles. 2015. In a different voice? Explaining the use of men and women as voice-over announcers in political advertising. Political Communication 32: 183–205. [Google Scholar] [CrossRef]
  34. Sugimoto, Cassidy R., Mike Thelwall, Vincent Larivière, Andrew Tsou, Philippe Mongeon, and Benoit Macaluso. 2013. Scientists popularizing science: Characteristics and impact of TED Talk presenters. PLoS ONE 8: e62403. [Google Scholar] [CrossRef] [PubMed]
  35. Violán Galán, Miguel Ángel, and Alfonso Mendiz. 2023. Análisis de las Estructuras Retóricas de las TED Talks: Elementos Constitutivos de la Inteligencia Discursiva. Available online: https://www.tdx.cat/handle/10803/688428#page=1 (accessed on 23 July 2025).
  36. Vrij, Aldert, and Frans Willem Winkel. 1994. Perceptual distortions in cross-cultural interrogations: The impact of skin color, accent, speech style, and spoken fluency on impression formation. Journal of Cross-Cultural Psychology 25: 284–95. [Google Scholar]
Figure 1. Forest plot of Wilcoxon rank differences by gender comparison.
Figure 1. Forest plot of Wilcoxon rank differences by gender comparison.
Socsci 15 00428 g001
Figure 2. Ethnicity differences in engagement by gender (deviation from the average Wilcoxon rank).
Figure 2. Ethnicity differences in engagement by gender (deviation from the average Wilcoxon rank).
Socsci 15 00428 g002
Table 1. List of measures considered to determine interaction data.
Table 1. List of measures considered to determine interaction data.
MeasureDescription
viewCountTotal number of views
likeCountTotal number of likes
dislikeCountTotal number of dislikes
commentCountTotal number of comments
engagement(likes-dislikes-comments)/views
pct likePercentage of likes by views
pct dislikePercentage of dislikes by views
pct commentsPercentage of comments by views
Table 2. Kruskal–Wallis relevant results for gender comparison.
Table 2. Kruskal–Wallis relevant results for gender comparison.
MetricComparisonChi-Square (χ2)p-Valueε2Wilcoxon Scores
viewCountfemale vs. male1.80750.17880.003female = 309.907
male = 329.615
likeCountfemale vs. male1.28430.25710.002female = 311.571
male = 328.183
dislikeCountfemale vs. male0.68810.40680.001female = 315.328
male = 324.951
commentCountfemale vs. male2.35110.12520.004female = 308.421
male = 330.894
engagementfemale vs. male2.69780.10050.004female = 333.436
male = 309.369
pct likefemale vs. male2.67090.10220.004female = 333.377
male = 309.420
pct dislikefemale vs. male0.01370.9068<0.001female = 319.770
male = 321.128
pct commentfemale vs. male1.26610.26050.002female = 329.365
male = 312.872
Table 3. Kruskal–Wallis relevant results for ethnicity comparison.
Table 3. Kruskal–Wallis relevant results for ethnicity comparison.
MetricComparisonChi-Square (χ2)p-Valueε2Wilcoxon Scores
viewCountAll ethnic groups0.87240.83210.002Asian = 195.173, Hispanic = 194.759, non-Hispanic Black = 183.021, non-Hispanic White = 197.048
likeCountAll ethnic groups1.13330.76910.003Asian = 198.173, Hispanic = 201.862, non-Hispanic Black = 181.826, non-Hispanic White = 195.814
dislikeCountAll ethnic groups1.63140.65230.004Asian = 196.173, Hispanic = 208.793, non-Hispanic Black = 180.556, non-Hispanic White = 195.816
commentCountAll ethnic groups1.96030.58070.005Asian = 192.927, Hispanic = 211.362, non-Hispanic Black = 179.903, non-Hispanic White = 196.470
engagementAll ethnic groups6.08390.10760.016Asian = 216.127, Hispanic = 226.845, non-Hispanic Black = 191.861, non-Hispanic White = 185.275
Table 4. Ethnicity differences in audience engagement stratified by gender (Kruskal–Wallis tests).
Table 4. Ethnicity differences in audience engagement stratified by gender (Kruskal–Wallis tests).
Metricχ2pε2Significant Pairwise Differences
Female
viewCount10.930.0120.037Asian < non-Hispanic White
likeCount11.940.0080.041Asian < non-Hispanic White
commentCount16.370.0010.056Asian < non-Hispanic White
dislikeCount9.480.0240.033non-Hispanic Black highest
pct dislike9.810.0200.034non-Hispanic Black highest
Male
dislikeCount32.34<0.0010.095non-Hispanic Black > Asian, White
pct dislike36.58<0.0010.107non-Hispanic Black > Asian, White
engagement15.680.0010.046Hispanic > non-Hispanic White
pct like15.960.0010.047Hispanic > non-Hispanic White
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

El-Yamri, M.; Violán, M.Á.; Manero, B. Exploring Ethnicity and Gender Bias in TED Talks: A Study of Audience Online Reactions. Soc. Sci. 2026, 15, 428. https://doi.org/10.3390/socsci15070428

AMA Style

El-Yamri M, Violán MÁ, Manero B. Exploring Ethnicity and Gender Bias in TED Talks: A Study of Audience Online Reactions. Social Sciences. 2026; 15(7):428. https://doi.org/10.3390/socsci15070428

Chicago/Turabian Style

El-Yamri, Meriem, Miguel Ángel Violán, and Borja Manero. 2026. "Exploring Ethnicity and Gender Bias in TED Talks: A Study of Audience Online Reactions" Social Sciences 15, no. 7: 428. https://doi.org/10.3390/socsci15070428

APA Style

El-Yamri, M., Violán, M. Á., & Manero, B. (2026). Exploring Ethnicity and Gender Bias in TED Talks: A Study of Audience Online Reactions. Social Sciences, 15(7), 428. https://doi.org/10.3390/socsci15070428

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop