Bias in Large Language Models: Origin, Evaluation, and Mitigation
Abstract
1. Introduction
2. Intrinsic Bias
2.1. Bias in Training Data
- Bias from over-representativeness and under-representativeness: Certain groups, such as gender, age, race, religion, ethnicity, culture, political, and socioeconomic classes, may be under-represented or over-represented in the corpus [49,50]. For example, men might be over-represented in datasets about leadership or science, while women may be more frequently mentioned in caregiving roles [51], thus leading to biased associations with demographic information. Another example is that the Latinx population, especially Mexican-American students, remains under-represented in U.S. higher education [50].
- Spatial and temporal bias: LLMs trained predominantly on a corpus from certain countries or geographic locations may absorb the cultural norms and values, hence building biases into the underlying LLMs. For instance, a model trained on Western-centric data may have a skewed understanding of non-Western cultures [52,53,54,55,56]. Prior studies have shown such effects in Arabic contexts, where models may prefer Western-associated entities over Arab ones [57], and in broader multilingual settings, where performance and cultural knowledge vary substantially across under-represented languages and regions [58]. Similarly, data collected over different periods may reflect outdated societal norms and values. LLMs trained on these data may be biased by those outdated norms and values. For instance, historical texts may exhibit racist or sexist terminology, which the model may absorb as part of its internal representations [59,60].
2.2. Bias in Data Collection Methods
2.3. Bias in Language Contexts
2.4. Bias in Tokenization and Word Embeddings
3. Extrinsic Bias
3.1. Natural Language Understanding (NLU) Tasks
- Gender bias: Models may incorrectly associate certain professions or roles with a specific gender, leading to errors in tasks like coreference resolution [28]. For example, assuming a doctor is male and a nurse is female, regardless of context.
- Age bias: Models might make assumptions about individuals based on age stereotypes [16]. For instance, associating technological proficiency only with younger people.
- Cultural or regional bias: Models could misinterpret idioms or expressions from different cultures or fail to recognize regional language variants [73]. This can result in misunderstandings in tasks like semantic textual similarity or natural language inference.
3.2. Natural Language Generation (NLG) Tasks
- Gender bias: Models might generate responses that align with gender stereotypes, such as using male pronouns for leaders and female pronouns for nurturing roles [28].
- Age bias: Models may produce content that reflects age-related stereotypes, like suggesting only sedentary activities for older adults [16].
- Cultural or regional bias: Models could favor content from dominant cultures or misrepresent cultural practices, leading to inappropriate or insensitive responses [74].
4. Bias Evaluation
4.1. Data-Level Bias Evaluation Methods
- Data distribution analysis: Data distribution analysis is a critical step in identifying and understanding bias at the data level, as it provides insights into how different demographic groups or categories are represented in the dataset used to train LLMs. The primary focus of this analysis is to ensure that the data is balanced and representative, minimizing the risk of perpetuating or amplifying existing biases in the trained model.
- –
- Representation across demographics: Research has consistently demonstrated that the representation of demographics in training data can significantly influence the outcomes of LLMs. For instance, LLMs trained on datasets with disproportionate representation of certain demographic groups are prone to exhibit biases that favor those groups [121,122,123]. For example, an analysis of political biases in LLMs revealed that most conversational LLMs exhibit left-of-center political preferences when probed with politically charged questions or statements [124]. The study demonstrated that LLMs can be steered towards specific political orientations through supervised fine-tuning (SFT) with modest amounts of politically aligned data.To assess representation, statistical tools can be employed to measure the distribution of demographic attributes within the dataset. Key metrics might include the relative frequency of each demographic group or the variance in representation across different categories. Visual tools such as histograms, bar charts, or demographic distribution tables can be useful for identifying disparities, providing a clear picture of which groups are over-represented or under-represented in the data.
- –
- Imbalance and skewness detection: Imbalance and skewness in data distribution can lead to biased model behavior, where the model is overly influenced by the majority class or demographic group. Detecting these issues is crucial for ensuring that the model performs fairly across all segments of the population.Imbalances occur when certain classes or demographic groups are significantly more prevalent in the dataset than others. This can result in the model being biased towards the majority group, leading to poor performance for minority groups. Techniques such as Gini coefficients [125], entropy measures [126], vocabulary usage [127], and frequency analysis can be used to quantify imbalance. Skewness can be visualized using skew and stereotype quantifying metrics [128,129] box plots, histograms, or cumulative distribution functions (CDFs). These tools help identify the extent of skewness in the data and guide the selection of appropriate mitigation strategies.
- –
- Data source analysis: The quality and bias of a model are heavily influenced by the sources of the data used for training. Biases inherent in the data sources can significantly impact the model’s behavior. For instance, if a model is trained primarily on data from Western countries, it may perform poorly when applied in non-Western contexts.Evaluating the bias of data sources for LLM training involves analyzing the origins, diversity, and quality of the data, as well as its impact on model performance. This includes cataloging the data sources to ensure they cover a wide range of geographical, cultural, and linguistic contexts, and examining the credibility and inherent biases of each source. For instance, research [130] on the SlimPajama dataset, which includes a rigorously deduplicated combination of web text, Wikipedia, GitHub, and books, reveals how different data combinations affect LLM performance. The SlimPajama-DC study highlights two key aspects: the impact of global versus local deduplication on model performance and the importance of data diversity post-deduplication. The findings indicate that models trained with highly deduplicated, diverse datasets outperform those trained on less-refined data, underscoring the significance of comprehensive and balanced data sources. However, duplicate removal is not always neutral with respect to representation. If content from minority, low-resource, or less frequently documented communities is already sparse, aggressive deduplication may disproportionately remove repeated expressions of those voices, thereby reducing their visibility in the corpus and unintentionally increasing representativeness bias. Bias detection tools and adjustments in data weighting can further help in managing these biases, supported by transparent documentation and regular reviews to ensure fairness and accuracy in the model outputs.
- Stereotype and bias detection in text data: Stereotype and bias detection involves analyzing the content of the training data to identify and quantify biases related to stereotypes, offensive language, or prejudiced statements. This can be achieved through the following.
- –
- Lexical analysis: This method focuses on identifying specific words or phrases within the training data that are associated with bias or stereotypes. Lexical analysis relies on predefined lexicons, which are curated lists of terms known to carry biased or stereotypical connotations. Tools like Hatebase, an extensive repository of hate speech terms, provide a valuable resource for identifying harmful language in text data. Hate speech can also be measured by a large-scale study [131]. BiCapsHate [132] is a deep learning model used to detect hate speech in online social media posts. Similarly, the Dictionary of Affect in Language (DAL) [133] categorizes words based on their emotional connotations, offering insights into how language can perpetuate stereotypes. By scanning the training data for occurrences of these terms, researchers can quantify the prevalence of biased language and identify specific areas where the data may reinforce harmful stereotypes. For example, a study by [134] demonstrated how biased language in training data could lead to the propagation of racial and gender stereotypes in NLP models.
- –
- Contextual analysis: While lexical analysis identifies the presence of specific biased terms, contextual analysis delves deeper into how these terms are used within the text. This method employs advanced NLP techniques to examine the context in which potentially biased language appears, allowing for the identification of subtler forms of bias. For instance, the same word might be used neutrally in one context but carry a prejudiced meaning in another. Contextual analysis examines sentence structure, co-occurrence patterns, and the surrounding language to uncover these nuances. A similar study by [27] has shown that even when biased language is not overtly present, underlying patterns can still perpetuate stereotypes, such as gender biases, in word embeddings. This approach is crucial for identifying and mitigating biases that may not be immediately apparent but can significantly impact the behavior of LLMs. Zhao et al. [135] quantified and analyzed the gender bias exhibited in ELMo’s contextualized word vectors. Contextual analysis can be implemented by methods like template-based method [136], applied thematic analysis [137], and contextualized word embedding analysis [138].
- –
- Sentiment analysis: Sentiment analysis is another vital tool in detecting bias in text data. This technique assesses the sentiment—positive, negative, or neutral—associated with different demographic groups or topics within the training data. By analyzing how different groups are described, researchers can identify patterns of negative or biased portrayals that could influence the model’s outputs. For example, if certain demographic groups are consistently associated with negative emotions or connotations, the model might learn to replicate these biases in its predictions or interactions. Research by [31] has highlighted how sentiment analysis can be used to detect and address such biases in text data. This method provides a quantitative measure of bias, enabling targeted interventions to reduce the impact of these biases on the model’s behavior. Sentiment analysis can be implemented using various approaches, including traditional machine learning methods [139,140,141], frameworks based on preprocessed data [141], and transformer-based methods [6,142,143].
- Annotation bias analysis: Annotation bias analysis is a crucial component in evaluating and mitigating bias in LLMs. This process involves examining the biases introduced during the data annotation phase, where human annotators label or categorize training data. Since annotators bring their own biases and perspectives, their subjective decisions can inadvertently introduce skewed or biased annotations, which in turn affect the performance and fairness of the LLMs. To perform annotation bias analysis, researchers must review the annotation guidelines and the training process for annotators, ensuring that they are designed to minimize bias. Additionally, evaluating the consistency and fairness of annotations across different demographic groups and annotators helps identify any disparities. Tools such as inter-annotator agreement metrics [144] and statistical analysis [145] of annotated data can reveal potential biases. For example, research by [146] on gender bias in annotated datasets highlights how annotation practices can reinforce stereotypes and biases, emphasizing the need for rigorous analysis and correction methods. Another way to reduce annotation bias is a human–LLM collaborative approach [147]. By addressing annotation bias, researchers can enhance the quality and fairness of training data, leading to more balanced and unbiased LLMs.
4.2. Model-Level Bias Evaluation Methods
- Fairness metrics: These metrics are pivotal in evaluating the model’s output fairness across different groups. Common metrics include the following.
- –
- –
- –
- Interpretability tools: Tools such as SHAP (Shapley Additive Explanations) [149] and LIME (Local Interpretable Model-agnostic Explanations) [150] provide insights into how specific features influence model predictions, helping identify biases. For instance, SHAP values can reveal that gendered words heavily influence predictions, indicating potential gender bias in the model [120,151]. Similarly, LIME approximates complex models locally with interpretable models to highlight the impact of features on individual predictions, which is crucial for diagnosing bias [150]. However, these tools should be interpreted with caution in LLMs with billions of parameters. In such massive models, feature attributions are often approximate and local, and may be unstable across prompt wording, tokenization, sampling settings, or semantically similar inputs. They may also fail to capture complex high-order interactions and therefore should not be treated as definitive causal explanations of bias. For this reason, SHAP- and LIME-based analyses are best used as diagnostic signals that should be corroborated with counterfactual, robustness, and output-level evaluations.
- Counterfactual fairness: This method generates counterfactual examples by altering sensitive attributes (e.g., changing names from traditionally male to female) to see if the model’s outputs remain invariant. A fair model should ideally produce the same outcomes regardless of these attribute changes [43,71]. Here, counterfactual fairness is used in the model-level sense: it examines whether the model’s decision rule or prediction remains invariant when only a sensitive attribute is changed, rather than comparing differences in the wording or tone of free-form generated responses.
4.3. Output-Level Bias Evaluation Methods
- Counterfactual testing: Counterfactual testing involves modifying input prompts by altering specific attributes, such as gender, race, or ethnicity, while keeping the rest of the context unchanged. This method evaluates whether the LLM’s output varies based on these demographic changes, allowing researchers to isolate the effect of these attributes on the model’s behavior. For instance, swapping “John” with “Maria” in a sentence can reveal if the model responds differently due to the gender of the subject.Recent work [152] emphasizes the importance of robust causal reasoning in LLMs to enhance fairness, arguing that a strong understanding of causal relationships can mitigate biases and reduce hallucinations. Beyond fairness, this connection is also important for understanding hallucinations. When a prompt is ambiguous or underspecified, biased social associations learned from the training corpus may act as anchors, leading the model to generate stereotype-consistent details even when those details are not supported by the input. For example, the model may infer occupations, traits, intentions, or background facts from demographic cues because such associations are statistically frequent in the data, not because they are causally justified in the given context. In this sense, some hallucinations can be viewed not merely as factual errors, but as bias-driven distortions in which stereotype-consistent priors fill evidential gaps. Stronger causal reasoning may reduce this risk by helping the model distinguish socially correlated attributes from causally relevant information. Another study [153] introduces a dynamic framework that compares outputs across different demographic groups in the same context, improving the fairness of generated text without requiring costly model retraining. Additionally, new tools [154] for generating and analyzing counterfactuals allow users to explore LLM behavior interactively, ensuring the counterfactuals are both meaningful and grammatically accurate.These advancements underscore the power of counterfactual testing in identifying biases in LLMs. By systematically altering demographic features in prompts, researchers can evaluate how models treat different groups and ensure equitable outputs across all demographics.
- Stereotype detection in generated text: Stereotype detection in generated text focuses on identifying and mitigating harmful biases that LLMs may perpetuate in their outputs. Since LLMs are trained on vast amounts of publicly available data, which often contain stereotypical narratives related to race, gender, profession, and religion, the risk of these biases appearing in generated content is significant. This method assesses how models reproduce stereotypes, providing insights into implicit prejudices embedded in their responses.One approach [150] introduced a unified dataset combining multiple stereotype detection datasets. Researchers fine-tuned LLMs on this dataset and found that multi-dimensional classifiers were more effective in identifying stereotypes. This study also highlighted the use of explainable AI tools to ensure models align with human understanding. Another study [155] developed a dual framework that combines static evaluations with dynamic, real-world scenario simulations. This dynamic aspect is particularly effective in detecting subtle, context-specific biases that static tests might miss. A qualitative method presented in [156] uses prompting techniques to uncover implicit stereotypes in LLM-generated text. This approach, focusing on biases like gender and ethnicity, employs the Tree of Thoughts technique to systematically reveal hidden prejudices and provide a reproducible method for stereotype detection. These methods collectively help researchers detect and mitigate biases in LLM outputs, offering a deeper understanding of how stereotypes manifest in generated text.
- Sentiment and toxicity analysis: Sentiment and toxicity analysis is crucial for evaluating output-level biases in LLMs, specifically targeting harmful or offensive content and ensuring adherence to ethical standards. This area of evaluation not only addresses toxicity but also helps in identifying subtle biases that may emerge in generated content.Llama Guard [157] introduces a model-based approach to bias evaluation by using a safety risk taxonomy for classifying both prompts and responses in human–AI conversations. This model, fine-tuned on a carefully curated dataset, shows strong performance in detecting various forms of toxicity. It represents a practical application of bias evaluation methods by providing a dynamic tool for assessing and mitigating harmful content in real-world interactions. The definition-based toxicity metric [158] offers a flexible solution for bias evaluation by using LLMs to measure toxicity according to predefined criteria. This method outperforms traditional metrics, enhancing the F1 score significantly and addressing the limitations of existing models that rely on dataset-specific definitions. It exemplifies how bias evaluation methods can be refined to improve the detection of nuanced toxic content. Moderation Using LLM Introspection (MULI) [159] advances bias evaluation by analyzing internal model responses to detect toxic prompts. This approach leverages patterns in response token logits and refusal behaviors, providing a cost-effective way to assess biases without requiring additional training.The OpenAI Moderations Endpoint represents a practical implementation of bias evaluation methods for filtering potentially harmful content. This tool helps developers assess and moderate LLM outputs, ensuring they align with safety and ethical guidelines. It demonstrates how automated tools can be integrated into the bias evaluation process to manage and mitigate risks associated with model-generated text.
- Acceptance and rejection rates: Evaluating acceptance and rejection rates in LLMs is essential for identifying biases that may arise in decision-making tasks. Recent work has explored how these biases manifest, particularly in contexts such as hiring decisions or moral dilemmas, where the outputs of LLMs could reinforce existing societal inequalities.In one study [160], researchers assessed how LLMs respond to job applicants based on perceived race, ethnicity, and gender by manipulating first names. They found that acceptance rates were significantly higher for masculine White names compared to masculine Hispanic names. These acceptance and rejection rates were highly prompt-sensitive, meaning that variations in how the question was framed could lead to different outcomes. This suggests that LLMs can subtly reinforce biases depending on how inputs are structured. Another evaluation method [161] used to study LLM behavior in moral decision making showed that models are consistent in choosing commonsense actions in unambiguous scenarios, while in more ambiguous cases, LLMs exhibited uncertainty and greater variability. Closed-source models often displayed more consistent preferences in these ambiguous scenarios, which suggests that proprietary models might encode specific tendencies in moral decision making. BiasBuster [162] provides a systematic method for uncovering and evaluating cognitive biases in LLMs, particularly in high-stakes decision making. By using a dataset of 16,800 prompts designed to probe various types of biases (e.g., prompt-induced, sequential), this framework allowed researchers to assess how LLMs handled acceptance and rejection decisions under different bias conditions.
4.4. Human-Involved Bias Evaluation Methods
- Human review and assessment: Experts or crowd-sourced reviewers manually assess model outputs for biases. This method is particularly effective in detecting nuanced biases that automated tools might miss, such as subtle stereotypes or culturally specific biases [43].
- Qualitative research methods: Interviews, focus groups, and case studies offer deep insights into how different communities perceive model outputs, providing a qualitative dimension to bias evaluation that complements quantitative approaches [120].
4.5. Conclusions
5. Bias Mitigation
5.1. Pre-Model Debiasing
with“The professor came to the classroom but he forgot to bring his laptop”
“The professor came to the classroom but she forgot to bring her laptop”.
5.2. Intra-Model Debiasing
5.3. Post-Model Debiasing
5.4. Critical Comparison of Mitigation Strategies
6. Ethical Concerns and Legal Challenges
- Stereotyping.
- –
- Definition: Negative abstractions about a labeled social group.
- –
- Result 1: Reinforced social bias against certain social groups.
- –
- Result 2: Toxic language towards certain social groups.
- Misrepresentation.
- –
- Definition: An incomplete sample generalized to a social group.
- –
- Result 1: Reinforced normativity of the dominant social group and implicit exclusion or devaluation of other groups.
- –
- Result 2: Biased outputs that mischaracterize certain social groups.
- Disparate system performance.
- –
- Definition: Degraded understanding or model performance in LLM language processing or generation between social groups.
- –
- Result 1: Allocational harms.
- Allocational harms.
- –
- Definition: Disparate treatment due to membership of a social group or due to proxies associated with a social group.
- –
- LLM-based hiring tools: Resume and cover letter screening.
- –
- LLM-based healthcare tools: AI diagnostics.
7. Conclusions
Funding
Data Availability Statement
Conflicts of Interest
Appendix A. Examples of Extrinsic Biases
Appendix A.1. Natural Language Understanding (NLU) Tasks
Appendix A.1.1. Coreference Resolution
- Gender Bias
- –
- Stereotypical occupation associations: In a sentence like “The doctor finished the surgery, and she went to check on the patient”, a biased model might mistakenly link “she” to a nurse or another female figure, based on the stereotype that doctors are more likely to be male. Similarly, in “The nurse finished the shift, and he went home”, the model might incorrectly associate “he” with a doctor or a male figure rather than the nurse, reflecting a bias that nurses are more likely to be female [28].
- –
- Gendered pronoun resolution: In a sentence like “Sam loves cooking. He is very talented in the kitchen”, a biased model might incorrectly resolve “he” to assume that Sam is male, even though the name “Sam” can be gender-neutral. This reflects a bias that cooking, when associated with talent or professionalism, is more likely to be linked to males, while in other contexts, the same task might be associated with females [76].
- Age Bias
- –
- Assumptions about technological proficiency: In a sentence like “The senior programmer and the young intern were debugging the code. He quickly found the bug”, a biased model might incorrectly resolve “He” to refer to the senior programmer, based on the stereotype that senior individuals are more skilled with technology, even though both the senior programmer and the young intern are equally likely candidates [77].
- –
- Bias in linking pronouns to age-related roles: In a sentence like “The retiree and the young employee discussed the new policies. He suggested some changes”, a biased model might incorrectly resolve “he” to refer to the young employee, based on the stereotype that younger individuals are more likely to suggest changes or new ideas, whereas retirees might be stereotypically viewed as less active in professional settings [69].
- Cultural or Regional Bias
- –
- Regional variants and pronoun use: In different cultural or regional contexts, the use of pronouns can vary significantly. For instance, in some languages, pronouns might be omitted altogether (pro-drop languages), while in others, they are used extensively. A coreference resolution system trained predominantly on non-pro-drop language data might struggle with correctly resolving entities in pro-drop languages, where subjects are often implied rather than explicitly mentioned [77].
- –
- Cultural context in family roles: In some cultures, family roles are strictly defined, with certain responsibilities typically assigned to specific genders or ages. A coreference resolution model might incorrectly resolve pronouns based on these cultural stereotypes. For example, in a sentence like “The eldest son took care of his siblings. She prepared dinner”, a biased model might incorrectly resolve “She” to the mother rather than the eldest daughter, based on cultural assumptions about family roles [76].
Appendix A.1.2. Semantic Textual Similarity (STS)
- Gender Bias
- –
- Gendered language and pronoun resolution: When comparing the similarity between sentences with gendered pronouns, such as: sentence A: “She is a leader.”; sentence B: “He is a leader”. A biased model might incorrectly rate the similarity between these sentences lower than it should, due to underlying gender associations, despite the fact that the sentences are nearly identical except for the pronoun [79].
- –
- Gender-stereotyped professions: Consider two sentences: sentence A: “The nurse administered the medication.”; sentence B: “She took care of the patient.” If an LLM trained with gender biases in its data associates nursing primarily with women, it might assess a higher similarity between these sentences compared to another pair where the nurse is referred to as “he”, even though the gender of the nurse should not affect the similarity score [27].
- Age Bias
- –
- Age-based expectations in language use: The sentences “She is full of energy and enthusiasm” and “She is calm and experienced” might be rated as less similar if the first sentence is associated with a young person and the second with an older person, despite both sentences describing positive qualities. A biased model might reinforce the stereotype that energy is associated with youth and calmness with older age, skewing the similarity score [16].
- –
- Age stereotypes in sentiment and perception: Sentences like “She is wise” and “She is elderly” might be rated as similar by a biased model due to the stereotype that elderly people are wise. This could lead to an overestimation of the similarity between texts that describe older individuals, even if the context suggests different meanings [80].
- Cultural or Regional Bias
- –
- Cultural idioms and expressions: Consider the sentences “It’s raining cats and dogs” (an English idiom meaning heavy rain) and “Il pleut des cordes” (a French idiom meaning the same). A culturally biased model trained predominantly on English data might rate these sentences as less similar because the literal words are different, even though they convey the same meaning in their respective cultures [73].
- –
- Regional dialects and variants: In English, the word “apartment” is commonly used in American English, while “flat” is used in British English. A model trained predominantly on one variant might incorrectly assess the similarity between “I live in an apartment” and “I live in a flat” as low, due to not recognizing the regional synonymy [81].
Appendix A.1.3. Natural Language Inference
- Gender Bias
- –
- Stereotypical gender roles in professions: Given a premise like “The person is a doctor”, and a hypothesis “She is caring”, a biased model might be more likely to infer that the hypothesis is true because of the stereotype associating women with caring professions. Conversely, if the premise is “The person is an engineer”, and the hypothesis is “She is analytical”, the model might incorrectly infer this as less likely due to the stereotype that engineering is male-dominated [76].
- –
- Gendered language and assumptions: In a scenario where the premise is “Alex received a promotion”, and the hypothesis is “She worked very hard”, a biased model might rate this inference as less likely due to the name “Alex”, which can be gender-neutral but may be more commonly associated with males. The model may incorrectly prefer “He worked very hard” based on gender assumptions [79].
- Age Bias
- –
- Bias toward younger individuals in dynamic roles: Given the premise “The person is a dynamic entrepreneur”, and the hypothesis “The person is young”, a biased model might overly favor the hypothesis due to the stereotype that entrepreneurship and dynamism are traits associated with younger people, disregarding the possibility that older individuals can also embody these characteristics [16].
- –
- Age-related stereotypes in professional contexts: If the premise is “The CEO announced the company’s new strategic direction”, a biased model might infer that “The CEO is in their 50s” is more likely to be true (entailment), reflecting the stereotype that leadership roles are typically held by older individuals. On the other hand, it might incorrectly classify “The CEO is in their 20s” as a contradiction [28].
- Cultural or Regional Bias
- –
- Cultural norms and social roles: Consider the premise “She is a nurse” and the hypothesis “She is caring”. A model influenced by cultural stereotypes might incorrectly infer that the hypothesis is true based on the cultural assumption that nurses are inherently caring, which reflects a bias rather than a logical inference based on the text [76].
- –
- Regional political and historical contexts: A premise might state, “The Berlin Wall fell in 1989”, with a hypothesis, “Germany was reunified shortly after”. A model trained predominantly on Western narratives might correctly infer entailment, but could struggle with similar inferences in different cultural or geopolitical contexts where regional historical knowledge is less emphasized [83].
Appendix A.1.4. Classification
- Gender Bias
- –
- Gendered language in job title classification: When classifying job-related text, an LLM might associate certain professions with specific genders based on stereotypical norms. For instance, “nurse” might be more frequently classified as female, while “engineer” might be classified as male, even when the text does not explicitly mention gender. This bias can lead to skewed recommendations in hiring algorithms or job description parsing [27].
- –
- Gendered pronouns and name classification: When classifying names or pronouns, LLMs may exhibit bias by associating certain names or pronouns with specific gendered stereotypes. For instance, names like “Alex” or “Jordan” might be classified with a gendered label (male or female) based on historical or cultural associations rather than the context provided [79].
- Age Bias
- –
- Sentiment classification bias based on age: A model might classify text written by older individuals as more negative or less enthusiastic compared to text written by younger individuals. For instance, if an older person discusses new technology, the model might incorrectly classify the sentiment as negative or apprehensive, based on the stereotype that older people are less tech-savvy or more resistant to change [86].
- –
- Bias in job application classification: In job application screening, an LLM might classify older applicants as less suitable for certain roles based on age-related stereotypes. For example, the model might down-rank resumes of older applicants for tech-related jobs due to biases that associate youth with innovation and adaptability, while assuming older candidates are less capable of learning new skills [84].
- Cultural or Regional Bias
- –
- Misclassification due to cultural language variants: A text classification model might incorrectly classify content written in African American Vernacular English (AAVE) as informal, unprofessional, or even toxic, due to biases against non-standard English dialects. This can result in discriminatory outcomes, such as the unfair flagging of social media posts by Black users [30].
- –
- Regional bias in political text classification: A political text classifier might be biased toward the dominant political ideology of the region where the training data was sourced. For instance, texts supporting socialist policies might be classified as “radical” or “extreme” if the training data predominantly reflects a region where socialism is less accepted [11].
Appendix A.1.5. Reading Comprehension
- Gender Bias
- –
- Stereotypes in gendered activities: In a scenario where the passage mentions, “Alex is an excellent cook”, a biased model might assume that Alex is female and reflect this in its answers to questions about the passage. For example, when asked, “What are her skills?” the model might incorrectly use a gendered pronoun, despite “Alex” being a gender-neutral name, thereby reflecting gender stereotypes associated with cooking [88].
- –
- Assumptions about gender roles in family settings: A passage might describe both a mother and a father taking care of children, but when asked, “Who is responsible for preparing dinner?” a biased model might infer that the mother is responsible, reflecting traditional gender roles [16].
- Age Bias
- –
- Bias in health-related content: In passages discussing health, a model might incorrectly infer that certain conditions or concerns (e.g., memory loss, frailty) are more likely associated with older characters, even when the text is neutral or unrelated to age. For instance, a passage about a character feeling tired might lead the model to infer an age-related cause if the character is older [27].
- –
- Misinterpretation of age-related roles: In a passage describing a family scenario, a reading comprehension model might incorrectly assume that the younger character is less responsible or that the older character is in a caregiving role, even if the text does not support these assumptions. This could lead to biased answers to questions about the characters’ roles or behaviors [89].
- Cultural or Regional Bias
- –
- Cultural context misinterpretation: A reading comprehension model might misinterpret a text that involves culturally specific practices or norms. For example, if a passage describes a traditional Japanese tea ceremony, a model trained predominantly on Western texts might misunderstand the significance of the ceremony, interpreting it as a simple social gathering rather than a ritual with deep cultural meaning [89].
- –
- Regional language varieties and dialects: A reading comprehension model might struggle with understanding regional dialects or non-standard language varieties. For instance, a passage written in African American Vernacular English (AAVE) might be misinterpreted or considered less coherent by the model, leading to incorrect or biased responses to comprehension questions [12].
Appendix A.1.6. Sentiment Analysis
- Gender Bias
- –
- Bias in sentiment toward gendered products or topics: Sentiment analysis models might exhibit bias when analyzing reviews or discussions about gendered products. For example, reviews of products marketed toward women, such as beauty products, might be categorized with more extreme sentiment (either overly positive or negative) compared to neutral categorizations for products marketed toward men [79].
- –
- Sentiment analysis in gendered contexts: Sentiment analysis models may misinterpret the sentiment of sentences involving gendered contexts. For example, statements about women’s rights or feminism might be more likely to be categorized as negative, reflecting a bias against topics that challenge traditional gender norms [91].
- Age Bias
- –
- Sentiment analysis on age-related topics: When analyzing sentiment on topics related to aging, retirement, or health, sentiment analysis models might display bias by assuming a more negative sentiment in texts written by older individuals. For example, discussions about retirement might be categorized as negative due to societal biases about aging, even if the text is neutral or positive in tone [92].
- –
- Stereotyping language use by older adults: Sentiment analysis models might misinterpret text written by older adults as more negative or less enthusiastic due to stereotypical views that older individuals are more conservative or less expressive. For instance, a review from an older person saying “The movie was good” might be rated as less positive compared to a more exuberant review from a younger person, even though both reviews express positive sentiment [93].
- Cultural or Regional Bias
- –
- Cultural bias in sentiment toward social norms: Sentiment analysis might misinterpret text related to social norms differently across cultures. For instance, in some cultures, discussing money openly is seen as positive and associated with success, while in others, it might be considered impolite or negative. A model trained on Western data might categorize open discussions about money in a non-Western text as negative [94].
- –
- Sentiment in multilingual contexts: In multilingual regions, a sentiment analysis model might incorrectly categorize sentiment when it encounters code-switching (the practice of alternating between two or more languages or dialects). For instance, in Latin America, code-switching between Spanish and indigenous languages might lead to incorrect sentiment assessments if the model is biased towards Spanish and fails to interpret the sentiment conveyed in the indigenous language correctly [93].
Appendix A.2. Natural Language Generation (NLG) Tasks
Appendix A.2.1. Question Answering
- Gender Bias
- –
- Gender bias in answer generation: When asked, “What should a good leader do?”, the model might provide examples or language that align with stereotypically male attributes (e.g., “He should be assertive and decisive”), thereby implying that leadership qualities are inherently male. If asked about nurturing roles, the model might default to female pronouns and qualities, reinforcing gendered stereotypes [28].
- –
- Bias in answering ambiguous gender questions: For a question like “What does a typical manager do?” where gender is not specified, a biased model might generate an answer using male pronouns (“He manages the team…”), reflecting an assumption that managers are typically male. This bias can also appear in reverse, where certain roles like teaching or nursing might default to female pronouns [88].
- Age Bias
- –
- Stereotypical answers about aging: If asked, “What are the best activities for elderly people?” a model might focus on sedentary activities like knitting or watching TV, neglecting more active pursuits like hiking, traveling, or volunteering. This reflects a bias that older adults are less capable of engaging in physically demanding or adventurous activities [16].
- –
- Negative bias toward youth: If a question is posed like, “Are young people capable of managing a company?” a biased model might respond with, “They may lack the experience needed for such a role”, reflecting stereotypes that associate youth with inexperience, despite many young people successfully managing companies [27].
- Cultural or Regional Bias
- –
- Bias in answering culturally specific questions: When asked, “What is the most popular sport in the world?” a biased QA system might answer “American football” if trained predominantly on data from the United States, ignoring that globally, soccer (football) is more widely popular. This reflects a bias toward regional popularity rather than global knowledge [96].
- –
- Language and regional bias in answer accuracy: A QA system might perform better when answering questions in or about regions and languages that are well-represented in its training data. For example, questions about European history or in English might be answered more accurately than questions about African history or in less commonly spoken languages [97].
Appendix A.2.2. Sentence Completions
- Gender Bias
- –
- Gender bias in descriptions of physical appearance: If the sentence starts with “She looked at herself in the mirror and…” a biased model might complete it with “…adjusted her makeup”, whereas “He looked at himself in the mirror and…” might be completed with “…straightened his tie”. This reflects the stereotype that women are more concerned with appearance and men with professionalism [16].
- –
- Gender bias in personal attributes: If prompted with “She is very…” a model might complete the sentence with adjectives like “emotional”, “beautiful”, or “caring”, while “He is very…” might be completed with “strong”, “intelligent”, or “ambitious”. These completions reflect stereotypical views of women as being more emotional and men as being more rational or powerful [79].
- Age Bias
- –
- Activity and lifestyle assumptions: For a prompt like “At 25, Jenny enjoys”, a biased model might complete the sentence with “partying and going out every night”, based on the stereotype that young adults are primarily interested in nightlife and socializing, ignoring the diversity of interests in this age group [69].
- –
- Learning and education stereotypes: Given the prompt “At 50, Mark decided to”, a biased model might complete it with “go back to school to finally get his degree”, assuming that older adults are “catching up” on education rather than pursuing lifelong learning or new academic challenges [26].
- Cultural or Regional Bias
- –
- Cultural stereotyping in sentence completion: If the model is given a sentence like “In Japan, people often eat”, it might complete with “sushi”, reflecting a cultural stereotype that overemphasizes a specific aspect of Japanese cuisine, while ignoring the diversity of food in Japanese culture. Similarly, completing “In Mexico, people celebrate” with “Cinco de Mayo” might reinforce a limited and stereotypical understanding of Mexican culture [74].
- –
- Regional bias in place-based completions: For the sentence “In Africa, many people live in”, the model might complete with “villages”, reflecting a regional bias that assumes rural living conditions are more common across an entire continent, despite the presence of large urban areas. This completion overlooks the diversity of living environments in different African countries and regions [99].
Appendix A.2.3. Conversational
- Gender Bias
- –
- Stereotypical responses based on gendered prompts: When given a prompt like “Describe a nurse” versus “Describe a doctor”, an LLM might generate responses that reinforce gender stereotypes. For example, it might describe a nurse as “caring, nurturing, and female”, while describing a doctor as “authoritative, knowledgeable, and male”, despite no gender being specified in the prompts [27].
- –
- Bias in gendered interactions: In a customer service chatbot, the model might respond more politely or deferentially to queries assumed to be from female users based on gender cues in the text (like names or pronouns), while responding more assertively or formally to male users. This reflects gendered expectations of communication styles [101].
- Age Bias
- –
- Assumptions about being tech-savvy: In a conversation where the user is discussing technology, a biased model might assume the user is younger if they express familiarity with tech jargon or concepts. Conversely, if the user asks for help with basic technology, the model might assume the user is older and potentially provide overly simplistic explanations, which could be patronizing [93].
- –
- Bias in addressing age-related concerns: In a scenario where a user asks for advice on starting a new career later in life, a biased model might discourage them by emphasizing the challenges of changing careers at an older age. For instance, if the user says, “I’m thinking of switching careers at 50”, the model might respond with comments like, “It might be difficult at your age”, reflecting a bias that older individuals face more barriers in the job market [102].
- Cultural or Regional Bias
- –
- Culturally inappropriate responses: A conversational LLM might provide responses that are culturally inappropriate or insensitive due to a lack of understanding of cultural norms. For example, if a user from Japan mentions “visiting a shrine”, a culturally biased model might respond with a suggestion that is more aligned with Western religious practices, failing to acknowledge the cultural significance of shrines in Japanese culture [103].
- –
- Bias in handling regional topics: A conversational AI might be biased in how it handles topics related to certain regions. For instance, if asked about news in Africa, the model might focus disproportionately on negative topics like conflict or poverty, reflecting a bias in the training data, while similar questions about Europe might elicit more varied and positive topics [104].
Appendix A.2.4. Recommender Systems
- Gender Bias
- –
- Product recommendations based on gender stereotypes: A recommender system might suggest beauty products, fashion items, or household goods primarily to women, while recommending electronics, tools, or sports equipment predominantly to men. This bias reinforces traditional gender roles and can lead to irrelevant or unappealing recommendations for users whose interests do not align with these stereotypes [106].
- –
- Career and education recommendations: A job or education recommender system might suggest STEM (Science, Technology, Engineering, and Mathematics) careers more often to male users while recommending roles in healthcare, education, or the arts more frequently to female users. This bias can perpetuate gender disparities in career choices and educational opportunities [107].
- Age Bias
- –
- Age-related product recommendations: A recommender system might automatically suggest health-related products, such as supplements or exercise equipment, to older users while recommending tech gadgets or trendy fashion items to younger users. These recommendations may not accurately reflect the individual’s interests but are based on age-related stereotypes [108].
- –
- Media and entertainment recommendations: A recommender system might assume that older users prefer classic movies or oldies music, while younger users prefer contemporary pop culture content. This can result in older users being recommended content that does not reflect their actual tastes if they have a preference for contemporary media, and younger users missing out on discovering older, classic content they might enjoy [109].
- Cultural or Regional Bias
- –
- Regional bias in news and information recommendations: A news recommender system might prioritize local or regional news from dominant cultures, under-representing or completely ignoring news from minority regions or less dominant cultures. For instance, users in a country might predominantly receive news recommendations about urban centers and mainstream political issues, while rural or indigenous issues are marginalized [110].
- –
- Bias in language and cultural content recommendations: A recommender system might prioritize content in the dominant language of a region, leading to a lack of recommendations for content in minority languages. For example, in a multilingual country, the system might recommend mostly English-language content, marginalizing content in local languages like Tamil, Welsh, or Basque [111].
Appendix A.2.5. Machine Translation
- Gender Bias
- –
- Gendered language mismatch: When translating from a gender-neutral language like Turkish or Finnish into a gendered language like English or Spanish, the model might introduce gender bias by assigning gendered pronouns or roles based on stereotypes. For example, the Turkish sentence “O bir doktor” (which means “He/She is a doctor”) might be translated into English as “He is a doctor”, reflecting the stereotype that doctors are male [113].
- –
- Gender stereotyping in occupational translations: When translating sentences that involve professions, the model might incorrectly assign gendered pronouns based on stereotypes. For instance, translating the phrase “The nurse” from a gender-neutral language might result in “La enfermera” (female nurse) in Spanish, while “The engineer” might be translated as “El ingeniero” (male engineer), even if the original language did not specify gender [114].
- Age Bias
- –
- Bias in addressing older adults: A machine translation system might translate sentences in a way that condescends to older adults, reflecting societal biases. For example, translating “The elderly person learned to use a smartphone” might introduce a tone or wording in the target language that implies surprise or patronization, even if the original sentence was neutral [16].
- –
- Translation of age-related idioms: When translating age-related idioms or expressions, a biased translation system might reinforce negative stereotypes. For instance, translating a phrase like “old people are slow” into another language might retain the negative connotation or even intensify it if the target language has a stronger cultural bias against the elderly [93].
- Cultural or Regional Bias
- –
- Cultural nuance loss: Cultural idioms, proverbs, or colloquial phrases often lose their meaning or are mistranslated when a model does not consider the cultural context. For example, the English idiom “It’s raining cats and dogs” could be literally translated into another language, resulting in a confusing or meaningless phrase in that target culture. A culturally aware model would translate it into a local equivalent, such as “Il pleut des cordes” (It’s raining ropes) in French [115].
- –
- Bias toward dominant cultures: A machine translation system might favor translations that align with Western cultural norms over those of less dominant cultures. For instance, translating phrases related to food, clothing, or customs might reflect Western standards, even when the source text belongs to a non-Western culture. An example could be translating a traditional Chinese clothing item, “Qipao”, simply as “dress”, which dilutes the cultural significance [116].
Appendix A.2.6. Summarization
- Gender Bias
- –
- Differential emphasis on roles: In summarizing biographies or obituaries, a model might emphasize traditional gender roles. For example, summaries of women’s lives might focus more on their roles as wives and mothers, while summaries of men’s lives might emphasize their careers or public achievements, even if the original text gave equal importance to both aspects [89].
- –
- Selective emphasis on gendered information: In summarizing a news article, a biased model might disproportionately emphasize gender-specific details that are not central to the story. For instance, when summarizing an article about a successful entrepreneur, the model might overemphasize details about the entrepreneur’s gender or appearance if the subject is female while focusing on achievements and business strategies if the subject is male. This kind of bias reinforces gender stereotypes and can skew the perceived importance of certain information [117].
- Age Bias
- –
- Bias in summarizing content for different age groups: When summarizing content aimed at different age groups, a model might simplify or alter the tone of content intended for younger audiences in a way that reflects condescension or a lack of complexity. For instance, a model might overly simplify a summary of educational content intended for teenagers, assuming they cannot handle complex information [32].
- –
- Omission of contributions based on age: When summarizing a report or article that includes contributions from both younger and older individuals, a biased summarization model might disproportionately highlight the contributions of younger people while downplaying or omitting the contributions of older individuals. For example, in summarizing a collaborative project, the model might emphasize the innovative ideas of younger team members while neglecting the experience and insights provided by older members [89].
- Cultural or Regional Bias
- –
- Omission of culturally significant details: A summarization model might omit culturally significant details that are not widely understood in the model’s primary training data. For instance, if summarizing a text about a traditional festival like Diwali, the model might focus on generic elements like “a festival with lights” while omitting important cultural and religious aspects such as the significance of the festival in Hinduism [118].
- –
- Bias toward western narratives: When summarizing global news articles, a model might prioritize Western perspectives or narratives, underplaying or omitting the viewpoints from non-Western cultures. For example, in summarizing an article about a political conflict in the Middle East, the model might focus on the perspectives of Western governments and omit the perspectives of local populations [119].
References
- Keskar, N.S.; McCann, B.; Varshney, L.R.; Xiong, C.; Socher, R. CTRL: A Conditional Transformer Language Model for Controllable Generation. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing, Hong Kong, China, 3–7 November 2019; pp. 111–129. [Google Scholar]
- Yang, Y.; Yuan, Y.; Liu, L. FinBERT: A Pretrained Language Model for Financial Communications. arXiv 2020, arXiv:2006.08097. [Google Scholar] [CrossRef]
- Okonkwo, C.W.; Ade-Ibijola, A. The role of artificial intelligence in education: Prospects and challenges. J. Educ. Technol. Soc. 2023, 26, 1–12. [Google Scholar]
- Jiang, L.; Yang, M.; Li, X. Smart Music Player: User-Adaptive Music Recommendation System. IEEE Trans. Multimed. 2020, 22, 666–675. [Google Scholar]
- Brown, T.B.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. Language models are few-shot learners. Adv. Neural Inf. Process. Syst. 2020, 33, 1877–1901. [Google Scholar]
- Devlin, J.; Chang, M.W.; Lee, K.; Toutanova, K. BERT: Pre-training of deep bidirectional transformers for language understanding. arXiv 2018, arXiv:1810.04805. [Google Scholar]
- Lewis, M.; Liu, Y.; Goyal, N.; Ghazvininejad, M.; Mohamed, A.; Levy, O.; Stoyanov, V.; Zettlemoyer, L. BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Online, 5–10 July 2020; pp. 7871–7880. [Google Scholar]
- Nallapati, R.; Zhou, B.; Gulcehre, C.; Xiang, B. Abstractive text summarization using sequence-to-sequence rnns and beyond. arXiv 2016, arXiv:1602.06023. [Google Scholar]
- Zhang, L.; Wang, S.; Liu, B. Deep Learning for Sentiment Analysis: A Survey. IEEE Trans. Affect. Comput. 2018, 10, 235–255. [Google Scholar] [CrossRef]
- Raffel, C.; Shazeer, N.; Roberts, A.; Lee, K.; Narang, S.; Matena, M.; Zhou, Y.; Li, W.; Liu, P.J. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer. J. Mach. Learn. Res. 2020, 21, 1–67. [Google Scholar]
- Bender, E.M.; Gebru, T.; McMillan-Major, A.; Shmitchell, S. On the dangers of stochastic parrots: Can language models be too big. In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, Virtual, 3–10 March 2021; pp. 610–623. [Google Scholar]
- Blodgett, S.L.; Barocas, S.; Daumé, H., III; Wallach, H. Language (Technology) is Power: A Critical Survey of “Bias” in NLP. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics; Jurafsky, D., Chai, J., Schluter, N., Tetreault, J., Eds.; Association for Computational Linguistics: Stroudsburg, PA, USA, 2020; pp. 5454–5476. [Google Scholar]
- Rajkomar, A.; Hardt, M.; Howell, M.D.; Corrado, G.; Chin, M.H. Ensuring Fairness in Machine Learning to Advance Health Equity. Ann. Intern. Med. 2018, 169, 866–872. [Google Scholar] [CrossRef]
- Angwin, J.; Larson, J.; Mattu, S.; Kirchner, L. Machine Bias. ProPublica 2016, 23, 139–159. [Google Scholar]
- Chen, M.; Ma, Z.; Hannak, A.; Wilson, C. My Fair LADY: Detecting and Mitigating Bias in Job Advertisements. In Proceedings of the 2018 World Wide Web Conference, Lyon, France, 23–27 April 2018; pp. 991–1000. [Google Scholar]
- Caliskan, A.; Bryson, J.J.; Narayanan, A. Semantics derived automatically from language corpora contain human-like biases. Science 2017, 356, 183–186. [Google Scholar] [CrossRef] [PubMed]
- Hastie, T.; Tibshirani, R.; Friedman, J. The Elements of Statistical Learning: Data Mining, Inference, and Prediction, 2nd ed.; Springer Science & Business Media: Berlin/Heidelberg, Germany, 2009. [Google Scholar]
- Akaike, H. A New Look at the Statistical Model Identification. IEEE Trans. Autom. Control 1974, 19, 716–723. [Google Scholar] [CrossRef]
- Heckman, J.J. Sample Selection Bias as a Specification Error. Econometrica 1979, 47, 153–161. [Google Scholar] [CrossRef] [PubMed]
- Vardi, Y. Nonparametric Estimation in the Presence of Length Bias. Ann. Stat. 1982, 10, 616–620. [Google Scholar] [CrossRef]
- Tibshirani, R. Regression Shrinkage and Selection via the Lasso. J. R. Stat. Soc. Ser. B (Methodol.) 1996, 58, 267–288. [Google Scholar] [CrossRef]
- Nemes, S.; Jonasson, J.M.; Genell, A.; Steineck, G. Bias in odds ratios by logistic regression modelling and sample size. BMC Med. Res. Methodol. 2009, 9, 56. [Google Scholar] [CrossRef] [PubMed]
- Li, H.; Qian, H.; Li, C.T.; Hou, K. Issues in cox proportional hazards model with unequal randomization. J. Biopharm. Stat. 2026, 36, 330–335. [Google Scholar] [CrossRef]
- Rubin, D.B. Estimating Causal Effects of Treatments in Randomized and Nonrandomized Studies. J. Educ. Psychol. 1974, 66, 688–701. [Google Scholar] [CrossRef]
- Schick, T.; Udupa, S.; Schütze, H. Self-Diagnosis and Self-Debiasing: A Proposal for Reducing Corpus-Based Bias in NLP. Trans. Assoc. Comput. Linguist. 2021, 9, 1408–1424. [Google Scholar] [CrossRef]
- Sun, T.; Gaut, A.; Tang, S.; Huang, Y.; ElSherief, M.; Zhao, J.; Mirza, D.; Belding, E.; Chang, K.W.; Wang, W.Y. Mitigating gender bias in natural language processing: Literature review. arXiv 2019, arXiv:1906.08976. [Google Scholar] [CrossRef]
- Bolukbasi, T.; Chang, K.W.; Zou, J.Y.; Saligrama, V.; Kalai, A.T. Man is to computer programmer as woman is to homemaker? Debiasing word embeddings. Adv. Neural Inf. Process. Syst. 2016, 29, 4356–4364. [Google Scholar]
- Zhao, J.; Wang, T.; Yatskar, M.; Ordonez, V.; Chang, K.W. Gender Bias in Coreference Resolution: Evaluation and Debiasing Methods. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics, New Orleans, LA, USA, 1–6 June 2018; pp. 15–20. [Google Scholar]
- Hovy, D.; Prabhumoye, S. Five Sources of Bias in Natural Language Processing. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics, Online, 1–6 August 2021; pp. 878–888. [Google Scholar]
- Sap, M.; Card, D.; Gabriel, S.; Choi, Y.; Smith, N.A. The Risk of Racial Bias in Hate Speech Detection. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Florence, Italy, 28 July–2 August 2019; pp. 1668–1678. [Google Scholar]
- Kiritchenko, S.; Mohammad, S.M. Examining Gender and Race Bias in Two Hundred Sentiment Analysis Systems. arXiv 2018, arXiv:1805.04508. [Google Scholar] [CrossRef]
- Obermeyer, Z.; Powers, B.; Vogeli, C.; Mullainathan, S. Dissecting Racial Bias in an Algorithm Used to Manage the Health of Populations. Science 2019, 366, 447–453. [Google Scholar] [CrossRef] [PubMed]
- Pariser, E. The Filter Bubble: How the New Personalized Web is Changing What We Read and How We Think; Penguin: London, UK, 2011. [Google Scholar]
- Sap, M.; Gabriel, S.; Qin, L.; Jurafsky, D.; Smith, N.A.; Choi, Y. Social Bias Frames: Reasoning about Social and Power Implications of Language. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Florence, Italy, 28 July–2 August 2019; pp. 5477–5483. [Google Scholar]
- Sheng, E.; Chang, K.W.; Natarajan, P.; Peng, N. Societal Biases in Language Generation: Progress and Challenges. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics, Online, 1–6 August 2021; pp. 4275–4293. [Google Scholar]
- Nadeem, M.; Bethke, A.; Reddy, S. StereoSet: Measuring stereotypical bias in pretrained language models. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), Online, 1–6 August 2021; pp. 5356–5371. [Google Scholar]
- Hewitt, J.; Manning, C.D. A Structural Probe for Finding Syntax in Word Representations. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics, Minneapolis, MN, USA, 2–7 June 2019; pp. 4129–4138. [Google Scholar]
- Zhao, J.; Wang, T.; Yatskar, M.; Ordonez, V.; Chang, K.W. Learning Gender-Neutral Word Embeddings. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Brussels, Belgium, 31 October–4 November 2018; pp. 4847–4853. [Google Scholar]
- Liang, P.P.; Zheng, L.; Salakhutdinov, R.; Morency, L.P. Towards Debiasing Sentence Representations. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Online, 5–10 July 2020; pp. 5502–5515. [Google Scholar]
- Mehrabi, N.; Morstatter, F.; Saxena, N.; Lerman, K.; Galstyan, A. A Survey on Bias and Fairness in Machine Learning. ACM Comput. Surv. 2021, 54, 1–35. [Google Scholar] [CrossRef]
- Navigli, R.; Conia, S.; Ross, B. Biases in Large Language Models: Origins, Inventory, and Discussion. ACM J. Data Inf. Qual. 2023, 15, 1–21. [Google Scholar] [CrossRef]
- Gallegos, I.O.; Rossi, R.A.; Barrow, J.; Tanjim, M.M.; Kim, S.; Dernoncourt, F.; Yu, T.; Zhang, R.; Ahmed, N.K. Bias and fairness in large language models: A survey. Comput. Linguist. 2024, 50, 1097–1179. [Google Scholar] [CrossRef]
- Doan, T.V.; Wang, Z.; Nguyen, M.N.; Zhang, W. Fairness in Large Language Models in three hours. In Proceedings of the 33rd ACM International Conference on Information & Knowledge Management, Boise, ID, USA, 21–25 October 2024. [Google Scholar]
- May, C.; Wang, A.; Bordia, S.; Bowman, S.R.; Rudinger, R. On measuring social biases in sentence encoders. arXiv 2019, arXiv:1903.10561. [Google Scholar] [CrossRef]
- Pagano, T.P.; Loureiro, R.B.; Lisboa, F.V.; Peixoto, R.M.; Guimarães, G.A.; Cruz, G.O.; Araujo, M.M.; Santos, L.L.; Cruz, M.A.; Oliveira, E.L.; et al. Bias and unfairness in machine learning models: A systematic review on datasets, tools, fairness metrics, and identification and mitigation methods. Big Data Cogn. Comput. 2023, 7, 15. [Google Scholar] [CrossRef]
- Ray, P.P. ChatGPT: A comprehensive review on background, applications, key challenges, bias, ethics, limitations and future scope. Internet Things -Cyber-Phys. Syst. 2023, 3, 121–154. [Google Scholar] [CrossRef]
- Goldfarb-Tarrant, S. Fairness in Transfer Learning for Natural Language Processing. Ph.D. Thesis, Institute for Language, Cognition and Computation, School of Informatics, University of Edinburgh, Edinburgh, UK, 2024. Available online: https://hdl.handle.net/1842/41857 (accessed on 15 November 2024).
- Goldman, J.; Tsotsos, J.K. Statistical Challenges with Dataset Construction: Why You Will Never Have Enough Images. arXiv 2024, arXiv:2408.11160. [Google Scholar] [CrossRef]
- Das, D.; De Langis, K.; Martin, A.; Kim, J.; Lee, M.; Kim, Z.M.; Hayati, S.; Owan, R.; Hu, B.; Parkar, R.; et al. Under the surface: Tracking the artifactuality of llm-generated data. arXiv 2024, arXiv:2401.14698. [Google Scholar] [CrossRef]
- Alvero, A.; Lee, J.; Regla-Vargas, A.; Kizilec, R.; Joachims, T.; Antonio, A.L. Large Language Models, Social Demography, and Hegemony: Comparing Authorship in Human and Synthetic Text. J. Big Data 2024, 11, 138. [Google Scholar] [CrossRef]
- UNESCO; IRCAI. Challenging Systematic Prejudices: An Investigation into Bias Against Women and Girls in Large Language Models. 2024. 20p. Available online: https://unesdoc.unesco.org/ark:/48223/pf0000388971 (accessed on 15 November 2024).
- Ahmad, A.; Bhattacharyya, P. Bias in Language Models: A Survey. CFILT, Indian Institute of Technology Bombay, Mumbai, India, 2024. Available online: https://www.cfilt.iitb.ac.in/resources/surveys/2024/Bias_Survey.pdf (accessed on 15 September 2024).
- Talat, Z.; Névéol, A.; Biderman, S.; Clinciu, M.; Dey, M.; Longpre, S.; Luccioni, S.; Masoud, M.; Mitchell, M.; Radev, D.; et al. You reap what you sow: On the challenges of bias evaluation under multilingual settings. In Proceedings of the BigScience Episode# 5–Workshop on Challenges & Perspectives in Creating Large Language Models, Virtual, 27 May 2022; pp. 26–41. [Google Scholar]
- Mukherjee, A.; Caliskan, A.; Zhu, Z.; Anastasopoulos, A. Global Gallery: The Fine Art of Painting Culture Portraits through Multilingual Instruction Tuning. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), Mexico City, Mexico, 16–21 June 2024; pp. 6398–6415. [Google Scholar]
- Lee, B.; Aiyappa, R.; Ahn, Y.Y.; Kwak, H.; An, J. Neural embedding of beliefs reveals the role of relative dissonance in human decision-making. arXiv 2024, arXiv:2408.07237. [Google Scholar] [CrossRef]
- Shi, W.; Li, R.; Zhang, Y.; Ziems, C.; Horesh, R.; de Paula, R.A.; Yang, D. Culturebank: An online community-driven knowledge base towards culturally aware language technologies. arXiv 2024, arXiv:2404.15238. [Google Scholar]
- Naous, T.; Ryan, M.J.; Ritter, A.; Xu, W. Having beer after prayer? Measuring cultural bias in large language models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Bangkok, Thailand, 11–16 August 2024; pp. 16366–16393. [Google Scholar]
- Wang, W.; Jiao, W.; Huang, J.; Dai, R.; Huang, J.T.; Tu, Z.; Lyu, M. Not all countries celebrate thanksgiving: On the cultural dominance in large language models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Bangkok, Thailand, 11–16 August 2024; pp. 6349–6384. [Google Scholar]
- Salinas, A.; Penafiel, L.; McCormack, R.; Morstatter, F. “Im not Racist but…”: Discovering Bias in the Internal Knowledge of Large Language Models. arXiv 2023, arXiv:2310.08780. [Google Scholar]
- Bussaja, J. Evaluating Racial Bias in Large Language Models: The Necessity for “SMOKY”. Int. J. Soft Comput. (IJSC) 2024, 15, 1–15. [Google Scholar] [CrossRef]
- Kuntz, J.B.; Silva, E.C. Who Authors the Internet? Analyzing Gender Diversity in ChatGPT-3 Training Data; University of Pittsburgh: Pittsburgh, PA, USA, 2023. [Google Scholar]
- Qu, Y.; Wang, J. Performance and biases of Large Language Models in public opinion simulation. Humanit. Soc. Sci. Commun. 2024, 11, 1–13. [Google Scholar] [CrossRef]
- Kotek, H.; Dockum, R.; Sun, D. Gender bias and stereotypes in large language models. In Proceedings of the ACM Collective Intelligence Conference, Delft, The Netherlands, 6–9 November 2023; pp. 12–24. [Google Scholar]
- Dwivedi, S.; Ghosh, S.; Dwivedi, S. Breaking the bias: Gender fairness in LLMs using prompt engineering and in-context learning. Rupkatha J. Interdiscip. Stud. Humanit. 2023, 15, 1–18. [Google Scholar] [CrossRef]
- de Sá, J.M.C.; Da Silveira, M.; Pruski, C. Semantic Change Characterization with LLMs using Rhetorics. arXiv 2024, arXiv:2407.16624. [Google Scholar] [CrossRef]
- Petrov, A.; La Malfa, E.; Torr, P.; Bibi, A. Language model tokenizers introduce unfairness between languages. Adv. Neural Inf. Process. Syst. 2023, 36, 36963–36990. [Google Scholar]
- Ahia, O.; Kumar, S.; Gonen, H.; Kasai, J.; Mortensen, D.R.; Smith, N.A.; Tsvetkov, Y. Do all languages cost the same? tokenization in the era of commercial language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, Singapore, 6–10 December 2023; pp. 9904–9923. [Google Scholar]
- Ovalle, A.; Mehrabi, N.; Goyal, P.; Dhamala, J.; Chang, K.W.; Zemel, R.; Galstyan, A.; Pinter, Y.; Gupta, R. Tokenization matters: Navigating data-scarce tokenization for gender inclusive language technologies. In Proceedings of the Findings of the Association for Computational Linguistics: NAACL 2024, Mexico City, Mexico, 16–21 June 2024; pp. 1739–1756. [Google Scholar]
- Garg, N.; Schiebinger, L.; Jurafsky, D.; Zou, J. Word embeddings quantify 100 years of gender and ethnic stereotypes. Proc. Natl. Acad. Sci. USA 2018, 115, E3635–E3644. [Google Scholar] [CrossRef] [PubMed]
- Rakshit, A.; Singh, S.; Keshari, S.; Chowdhury, A.G.; Jain, V.; Chadha, A. From prejudice to parity: A new approach to debiasing large language model word embeddings. In Proceedings of the 31st International Conference on Computational Linguistics, Abu Dhabi, United Arab Emirates, 19–24 January 2025; pp. 6718–6747. [Google Scholar]
- Li, Y.; Du, M.; Song, R.; Wang, X.; Wang, Y. A survey on fairness in large language models. arXiv 2023, arXiv:2308.10149. [Google Scholar]
- Chang, Y.; Wang, X.; Wang, J.; Wu, Y.; Yang, L.; Zhu, K.; Chen, H.; Yi, X.; Wang, C.; Wang, Y.; et al. A survey on evaluation of large language models. ACM Trans. Intell. Syst. Technol. 2024, 15, 1–45. [Google Scholar] [CrossRef]
- Vulic, I.; Moens, M.F. Cross-lingual semantic similarity of words as the similarity of their semantic word responses. In Proceedings of the 2013 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT 2013), Atlanta, GA, USA, 9–14 June 2013; ACL: East Stroudsburg, PA, USA, 2013; pp. 106–116. [Google Scholar]
- Hendricks, L.A.; Burns, K.; Saenko, K.; Darrell, T.; Rohrbach, A. Women also snowboard: Overcoming bias in captioning models. In Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany, 8–14 September 2018; pp. 771–787. [Google Scholar]
- Lee, K.; He, L.; Lewis, M.; Zettlemoyer, L. End-to-end neural coreference resolution. arXiv 2017, arXiv:1707.07045. [Google Scholar]
- Rudinger, R.; Naradowsky, J.; Leonard, B.; Van Durme, B. Gender bias in coreference resolution. arXiv 2018, arXiv:1804.09301. [Google Scholar]
- Hovy, D.; Yang, D. The importance of modeling social factors of language: Theory and practice. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Online, 6–11 June 2021; pp. 588–602. [Google Scholar]
- Cer, D.; Diab, M.; Agirre, E.; Lopez-Gazpio, I.; Specia, L. Semeval-2017 task 1: Semantic textual similarity-multilingual and cross-lingual focused evaluation. arXiv 2017, arXiv:1708.00055. [Google Scholar]
- Zhao, J.; Wang, T.; Yatskar, M.; Ordonez, V.; Chang, K.W. Men also like shopping: Reducing gender bias amplification using corpus-level constraints. arXiv 2017, arXiv:1707.09457. [Google Scholar] [CrossRef]
- Diaz, F.; Mitra, B.; Craswell, N. Query expansion with locally-trained word embeddings. arXiv 2016, arXiv:1605.07891. [Google Scholar] [CrossRef]
- Tan, L.; Bond, F. Building and annotating the linguistically diverse NTU-MC (NTU-multilingual corpus). In Proceedings of the 25th Pacific Asia Conference on Language, Information and Computation, Singapore, 16–18 December 2011; pp. 362–371. [Google Scholar]
- Bowman, S.R.; Angeli, G.; Potts, C.; Manning, C.D. A large annotated corpus for learning natural language inference. arXiv 2015, arXiv:1508.05326. [Google Scholar] [CrossRef]
- Chen, Q.; Zhu, X.; Ling, Z.; Wei, S.; Jiang, H.; Inkpen, D. Enhanced LSTM for natural language inference. arXiv 2016, arXiv:1609.06038. [Google Scholar]
- De-Arteaga, M.; Romanov, A.; Wallach, H.; Chayes, J.; Borgs, C.; Chouldechova, A.; Geyik, S.; Kenthapadi, K.; Kalai, A.T. Bias in bios: A case study of semantic representation bias in a high-stakes setting. In Proceedings of the Conference on Fairness, Accountability, and Transparency, Atlanta, GA, USA, 29–31 January 2019; pp. 120–128. [Google Scholar]
- Meng, Y.; Zhang, Y.; Huang, J.; Xiong, C.; Ji, H.; Zhang, C.; Han, J. Text classification using label names only: A language model self-training approach. arXiv 2020, arXiv:2010.07245. [Google Scholar] [CrossRef]
- Díaz, M.; Johnson, I.; Lazar, A.; Piper, A.M.; Gergle, D. Addressing age-related bias in sentiment analysis. In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems, Montréal, QC, Canada, 21–26 April 2018; pp. 1–14. [Google Scholar]
- Rajpurkar, P.; Jia, R.; Liang, P. Know what you don’t know: Unanswerable questions for SQuAD. arXiv 2018, arXiv:1806.03822. [Google Scholar]
- Webster, K.; Recasens, M.; Axelrod, V.; Baldridge, J. Mind the GAP: A balanced corpus of gendered ambiguous pronouns. Trans. Assoc. Comput. Linguist. 2018, 6, 605–617. [Google Scholar] [CrossRef]
- Bender, E.M.; Friedman, B. Data statements for natural language processing: Toward mitigating system bias and enabling better science. Trans. Assoc. Comput. Linguist. 2018, 6, 587–604. [Google Scholar] [CrossRef]
- Pang, B.; Lee, L. Opinion mining and sentiment analysis. Found. Trends® Inf. Retr. 2008, 2, 1–135. [Google Scholar] [CrossRef]
- Misiunas, K.; Keyser, U.F. Density-dependent speed-up of particle transport in channels. Phys. Rev. Lett. 2019, 122, 214501. [Google Scholar] [CrossRef] [PubMed]
- Burnap, P.; Williams, M.L. Cyber hate speech on twitter: An application of machine classification and statistical modeling for policy and decision making. Policy Internet 2015, 7, 223–242. [Google Scholar] [CrossRef]
- Hovy, D.; Søgaard, A. Tagging performance correlates with author age. In Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 2: Short Papers), Beijing, China, 26–31 July 2015; pp. 483–488. [Google Scholar]
- Blodgett, S.L.; O’Connor, B. Racial disparity in natural language processing: A case study of social media african-american english. arXiv 2017, arXiv:1707.00061. [Google Scholar] [CrossRef]
- Rajpurkar, P.; Zhang, J.; Lopyrev, K.; Liang, P. Squad: 100,000+ questions for machine comprehension of text. arXiv 2016, arXiv:1606.05250. [Google Scholar]
- Gardner, M.; Grus, J.; Neumann, M.; Tafjord, O.; Dasigi, P.; Liu, N.; Peters, M.; Schmitz, M.; Zettlemoyer, L. Allennlp: A deep semantic natural language processing platform. arXiv 2018, arXiv:1803.07640. [Google Scholar] [CrossRef]
- Kwiatkowski, T.; Palomaki, J.; Redfield, O.; Collins, M.; Parikh, A.; Alberti, C.; Epstein, D.; Polosukhin, I.; Devlin, J.; Lee, K.; et al. Natural questions: A benchmark for question answering research. Trans. Assoc. Comput. Linguist. 2019, 7, 453–466. [Google Scholar] [CrossRef]
- Radford, A.; Wu, J.; Child, R.; Luan, D.; Amodei, D.; Sutskever, I. Language models are unsupervised multitask learners. OpenAI Blog 2019, 1, 9. [Google Scholar]
- Birhane, A.; Prabhu, V.U. Large image datasets: A pyrrhic win for computer vision? In Proceedings of the 2021 IEEE Winter Conference on Applications of Computer Vision (WACV), Waikoloa, HI, USA, 3–8 January 2021; pp. 1536–1546. [Google Scholar]
- Zhang, Y.; Sun, S.; Galley, M.; Chen, Y.C.; Brockett, C.; Gao, X.; Gao, J.; Liu, J.; Dolan, B. Dialogpt: Large-scale generative pre-training for conversational response generation. arXiv 2019, arXiv:1911.00536. [Google Scholar]
- Hitsch, G.J.; Hortaçsu, A.; Ariely, D. Matching and sorting in online dating. Am. Econ. Rev. 2010, 100, 130–163. [Google Scholar] [CrossRef]
- Wagner, C.; Graells-Garrido, E.; Garcia, D.; Menczer, F. Women through the glass ceiling: Gender asymmetries in Wikipedia. EPJ Data Sci. 2016, 5, 5. [Google Scholar] [CrossRef]
- Hovy, D.; Spruit, S.L. The social impact of natural language processing. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), Berlin, Germany, 7–12 August 2016; pp. 591–598. [Google Scholar]
- Gebru, T.; Morgenstern, J.; Vecchione, B.; Vaughan, J.W.; Wallach, H.; Iii, H.D.; Crawford, K. Datasheets for datasets. Commun. ACM 2021, 64, 86–92. [Google Scholar] [CrossRef]
- Zhang, S.; Yao, L.; Sun, A.; Tay, Y. Deep learning based recommender system: A survey and new perspectives. ACM Comput. Surv. (CSUR) 2019, 52, 1–38. [Google Scholar] [CrossRef]
- Ekstrand, M.D.; Tian, M.; Azpiazu, I.M.; Ekstrand, J.D.; Anuyah, O.; McNeill, D.; Pera, M.S. All the cool kids, how do they fit in?: Popularity and demographic biases in recommender evaluation and effectiveness. In Proceedings of the Conference on Fairness, Accountability and Transparency, PMLR, New York, NY, USA, 23–24 February 2018; pp. 172–186. [Google Scholar]
- Chen, R.C.; Ai, Q.; Jayasinghe, G.; Croft, W.B. Correcting for recency bias in job recommendation. In Proceedings of the 28th ACM International Conference on Information and Knowledge Management, Beijing, China, 3–7 November 2019; pp. 2185–2188. [Google Scholar]
- Biega, A.J.; Gummadi, K.P.; Weikum, G. Equity of attention: Amortizing individual fairness in rankings. In Proceedings of the 41st International ACM SIGIR Conference on Research & Development in Information Retrieval, Ann Arbor, MI, USA, 8–12 July 2018; pp. 405–414. [Google Scholar]
- Burke, R.; Sonboli, N.; Ordonez-Gauger, A. Balanced neighborhoods for multi-sided fairness in recommendation. In Proceedings of the Conference on Fairness, Accountability and Transparency, PMLR, New York, NY, USA, 23–24 February 2018; pp. 202–214. [Google Scholar]
- Karimi, M.; Jannach, D.; Jugovac, M. News recommender systems–Survey and roads ahead. Inf. Process. Manag. 2018, 54, 1203–1227. [Google Scholar] [CrossRef]
- Lakew, S.M.; Federico, M.; Negri, M.; Turchi, M. Multilingual neural machine translation for low-resource languages. IJCoL Ital. J. Comput. Linguist. 2018, 4, 11–25. [Google Scholar] [CrossRef]
- Vaswani, A. Attention is all you need. Adv. Neural Inf. Process. Syst. 2017, 30, 5998–6008. [Google Scholar]
- Prates, M.O.; Avelar, P.H.; Lamb, L.C. Assessing gender bias in machine translation: A case study with google translate. Neural Comput. Appl. 2020, 32, 6363–6381. [Google Scholar] [CrossRef]
- Neri, D.; Soldani, J.; Zimmermann, O.; Brogi, A. Design principles, architectural smells and refactorings for microservices: A multivocal review. SICS Softw.-Intensive-Cyber-Phys. Syst. 2020, 35, 3–15. [Google Scholar] [CrossRef]
- Vanmassenhove, E.; Shterionov, D.; Way, A. Lost in translation: Loss and decay of linguistic richness in machine translation. arXiv 2019, arXiv:1906.12068. [Google Scholar] [CrossRef]
- Sennrich, R. Neural Machine Translation; Institute for Language, Cognition and Computation University of Edinburgh: Edinburgh, UK, 2016; Volume 18. [Google Scholar]
- Otterbacher, J.; Bates, J.; Clough, P. Competent men and warm women: Gender stereotypes and backlash in image search results. In Proceedings of the 2017 CHI Conference on Human Factors in Computing Systems, Denver, CO, USA, 6–11 May 2017; pp. 6620–6631. [Google Scholar]
- Shen, T.; Lei, T.; Barzilay, R.; Jaakkola, T. Style transfer from non-parallel text by cross-alignment. Adv. Neural Inf. Process. Syst. 2017, 30, 6833–6844. [Google Scholar]
- Perez-Beltrachini, L.; Lapata, M. Models and datasets for cross-lingual summarisation. arXiv 2022, arXiv:2202.09583. [Google Scholar] [CrossRef]
- Liang, P.P.; Wu, C.; Morency, L.P.; Salakhutdinov, R. Towards understanding and mitigating social biases in language models. In Proceedings of the International Conference on Machine Learning, PMLR, Online, 18–24 July 2021; pp. 6565–6576. [Google Scholar]
- Simmons, G.; Hare, C. Large language models as subpopulation representative models: A review. arXiv 2023, arXiv:2310.17888. [Google Scholar] [CrossRef]
- Wang, A.; Morgenstern, J.; Dickerson, J.P. Large language models cannot replace human participants because they cannot portray identity groups. arXiv 2024, arXiv:2402.01908. [Google Scholar] [CrossRef]
- Gorti, A.; Gaur, M.; Chadha, A. Unboxing Occupational Bias: Grounded Debiasing LLMs with US Labor Data. arXiv 2024, arXiv:2408.11247. [Google Scholar] [CrossRef]
- Rozado, D. The political preferences of LLMs. PLoS ONE 2024, 19, e0306621. [Google Scholar] [CrossRef]
- Khanuja, S.; Ruder, S.; Talukdar, P. Evaluating the Diversity, Equity and Inclusion of NLP Technology: A Case Study for Indian Languages. arXiv 2022, arXiv:2205.12676. [Google Scholar]
- Csáky, R.; Purgai, P.; Recski, G. Improving neural conversational models with entropy-based data filtering. arXiv 2019, arXiv:1905.05471. [Google Scholar] [CrossRef]
- Durward, M.; Thomson, C. Evaluating Vocabulary Usage in LLMs. In Proceedings of the 19th Workshop on Innovative Use of NLP for Building Educational Applications (BEA 2024), Mexico City, Mexico, 20 June 2024; pp. 266–282. [Google Scholar]
- de Vassimon Manela, D.; Errington, D.; Fisher, T.; van Breugel, B.; Minervini, P. Stereotype and skew: Quantifying gender bias in pre-trained and fine-tuned language models. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, Online, 19–23 April 2021; pp. 2232–2242. [Google Scholar]
- Dang, V.M.H.; Verma, R.M. Data quality in NLP: Metrics and a comprehensive taxonomy. In Proceedings of the International Symposium on Intelligent Data Analysis, Stockholm, Sweden, 24–26 April 2024; Springer: Berlin/Heidelberg, Germany, 2024; pp. 217–229. [Google Scholar]
- Shen, Z.; Tao, T.; Ma, L.; Neiswanger, W.; Hestness, J.; Vassilieva, N.; Soboleva, D.; Xing, E. Slimpajama-dc: Understanding data combinations for llm training. arXiv 2023, arXiv:2309.10818. [Google Scholar]
- Silva, L.; Mondal, M.; Correa, D.; Benevenuto, F.; Weber, I. Analyzing the targets of hate in online social media. In Proceedings of the International AAAI Conference on Web and Social Media, Cologne, Germany, 17–20 May 2016; Volume 10, pp. 687–690. [Google Scholar]
- Kamal, A.; Anwar, T.; Sejwal, V.K.; Fazil, M. BiCapsHate: Attention to the linguistic context of hate via bidirectional capsules and hatebase. IEEE Trans. Comput. Soc. Syst. 2023, 11, 1781–1792. [Google Scholar] [CrossRef]
- Whissell, C.M. The dictionary of affect in language. In The Measurement of Emotions; Elsevier: Berlin/Heidelberg, Germany, 1989; pp. 113–131. [Google Scholar]
- Dev, S.; Sheng, E.; Zhao, J.; Amstutz, A.; Sun, J.; Hou, Y.; Sanseverino, M.; Kim, J.; Nishi, A.; Peng, N.; et al. On measures of biases and harms in NLP. arXiv 2021, arXiv:2108.03362. [Google Scholar]
- Zhao, J.; Wang, T.; Yatskar, M.; Cotterell, R.; Ordonez, V.; Chang, K.W. Gender bias in contextualized word embeddings. arXiv 2019, arXiv:1904.03310. [Google Scholar]
- Kurita, K.; Vyas, N.; Pareek, A.; Black, A.W.; Tsvetkov, Y. Measuring bias in contextualized word representations. arXiv 2019, arXiv:1906.07337. [Google Scholar] [CrossRef]
- Mackieson, P.; Shlonsky, A.; Connolly, M. Increasing rigor and reducing bias in qualitative research: A document analysis of parliamentary debates using applied thematic analysis. Qual. Soc. Work. 2019, 18, 965–980. [Google Scholar] [CrossRef]
- Basta, C.; Costa-Jussà, M.R.; Casas, N. Evaluating the underlying gender bias in contextualized word embeddings. arXiv 2019, arXiv:1904.08783. [Google Scholar] [CrossRef]
- Wankhade, M.; Rao, A.C.S.; Kulkarni, C. A survey on sentiment analysis methods, applications, and challenges. Artif. Intell. Rev. 2022, 55, 5731–5780. [Google Scholar] [CrossRef]
- Saxena, A.; Reddy, H.; Saxena, P. Introduction to sentiment analysis covering basics, tools, evaluation metrics, challenges, and applications. In Principles of Social Networking: The New Horizon and Emerging Challenges; Springer: Singapore, 2022; pp. 249–277. [Google Scholar]
- Hasan, M.R.; Maliha, M.; Arifuzzaman, M. Sentiment analysis with NLP on Twitter data. In Proceedings of the 2019 International Conference on Computer, Communication, Chemical, Materials and Electronic Engineering (IC4ME2), Rajshahi, Bangladesh, 11–12 July 2019; pp. 1–4. [Google Scholar]
- Naseem, U.; Razzak, I.; Musial, K.; Imran, M. Transformer based deep intelligent contextual embedding for twitter sentiment analysis. Future Gener. Comput. Syst. 2020, 113, 58–69. [Google Scholar] [CrossRef]
- Abdullah, T.; Ahmet, A. Deep learning in sentiment analysis: Recent architectures. ACM Comput. Surv. 2022, 55, 1–37. [Google Scholar] [CrossRef]
- Artstein, R. Inter-annotator agreement. In Handbook of Linguistic Annotation; Springer: Dordrecht, The Netherlands, 2017; pp. 297–313. [Google Scholar]
- Paun, S.; Artstein, R.; Poesio, M. Statistical Methods for Annotation Analysis; Springer Nature: Berlin/Heidelberg, Germany, 2022. [Google Scholar]
- Havens, L.; Terras, M.; Bach, B.; Alex, B. Uncertainty and inclusivity in gender bias annotation: An annotation taxonomy and annotated datasets of British English text. In Proceedings of the 4th Workshop on Gender Bias in Natural Language Processing, Seattle, WA, USA, 15 July 2022; pp. 30–57. [Google Scholar]
- Wang, X.; Kim, H.; Rahman, S.; Mitra, K.; Miao, Z. Human-LLM collaborative annotation through effective verification of LLM labels. In Proceedings of the CHI Conference on Human Factors in Computing Systems, Honolulu, HI, USA, 11–16 May 2024; pp. 1–21. [Google Scholar]
- Esiobu, D.; Tan, X.; Hosseini, S.; Ung, M.; Zhang, Y.; Fernandes, J.; Dwivedi-Yu, J.; Presani, E.; Williams, A.; Smith, E. ROBBIE: Robust bias evaluation of large generative language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, Singapore, 6–10 December 2023; pp. 3764–3814. [Google Scholar]
- Mosca, E.; Szigeti, F.; Tragianni, S.; Gallagher, D.; Groh, G. SHAP-based explanation methods: A review for NLP interpretability. In Proceedings of the 29th International Conference on Computational Linguistics, Gyeongju, Republic of Korea, 12–17 October 2022; pp. 4593–4603. [Google Scholar]
- Wu, Z.; Bulathwela, S.; Perez-Ortiz, M.; Koshiyama, A.S. Auditing Large Language Models for Enhanced Text-Based Stereotype Detection and Probing-Based Bias Evaluation. arXiv 2024, arXiv:2404.01768. [Google Scholar]
- Lin, Z.; Guan, S.; Zhang, W.; Zhang, H.; Li, Y.; Zhang, H. Towards trustworthy LLMs: A review on debiasing and dehallucinating in large language models. Artif. Intell. Rev. 2024, 57, 1–50. [Google Scholar] [CrossRef]
- Wang, Z. CausalBench: A Comprehensive Benchmark for Evaluating Causal Reasoning Capabilities of Large Language Models. In Proceedings of the 10th SIGHAN Workshop on Chinese Language Processing (SIGHAN-10), Bangkok, Thailand, 16 August 2024; pp. 143–151. [Google Scholar]
- Banerjee, P.; Java, A.; Jandial, S.; Shahid, S.; Furniturewala, S.; Krishnamurthy, B.; Bhatia, S. All Should Be Equal in the Eyes of LMs: Counterfactually Aware Fair Text Generation. In Proceedings of the AAAI Conference on Artificial Intelligence, Vancouver, BC, Canada, 20–27 February 2024; Volume 38, pp. 17673–17681. [Google Scholar]
- Cheng, F.; Zouhar, V.; Chan, R.S.M.; Fürst, D.; Strobelt, H.; El-Assady, M. Interactive Analysis of LLMs using Meaningful Counterfactuals. arXiv 2024, arXiv:2405.00708. [Google Scholar] [CrossRef]
- Bai, Y.; Zhao, J.; Shi, J.; Xie, Z.; Wu, X.; He, L. FairMonitor: A Dual-framework for Detecting Stereotypes and Biases in Large Language Models. arXiv 2024, arXiv:2405.03098. [Google Scholar]
- Babonnaud, W.; Delouche, E.; Lahlouh, M. The Bias that Lies Beneath: Qualitative Uncovering of Stereotypes in Large Language Models. Swed. Artif. Intell. Soc. 2024, 195–203. [Google Scholar] [CrossRef]
- Inan, H.; Upasani, K.; Chi, J.; Rungta, R.; Iyer, K.; Mao, Y.; Tontchev, M.; Hu, Q.; Fuller, B.; Testuggine, D.; et al. Llama guard: Llm-based input-output safeguard for human-ai conversations. arXiv 2023, arXiv:2312.06674. [Google Scholar]
- Koh, H.; Kim, D.; Lee, M.; Jung, K. Can LLMs Recognize Toxicity? Structured Toxicity Investigation Framework and Semantic-Based Metric. arXiv 2024, arXiv:2402.06900. [Google Scholar] [CrossRef]
- Hu, Z.; Piet, J.; Zhao, G.; Jiao, J.; Wagner, D. Toxicity Detection for Free. arXiv 2024, arXiv:2405.18822. [Google Scholar] [CrossRef]
- An, H.; Acquaye, C.; Wang, C.; Li, Z.; Rudinger, R. Do Large Language Models Discriminate in Hiring Decisions on the Basis of Race, Ethnicity, and Gender? arXiv 2024, arXiv:2406.10486. [Google Scholar] [CrossRef]
- Scherrer, N.; Shi, C.; Feder, A.; Blei, D. Evaluating the moral beliefs encoded in llms. Adv. Neural Inf. Process. Syst. 2024, 36, 51778–51809. [Google Scholar]
- Echterhoff, J.; Liu, Y.; Alessa, A.; McAuley, J.; He, Z. Cognitive bias in high-stakes decision-making with llms. arXiv 2024, arXiv:2403.00811. [Google Scholar]
- Sheng, E.; Chang, K.W.; Natarajan, P.; Peng, N. The Woman Worked as a Babysitter: On Biases in Language Generation. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), Hong Kong, China, 3–7 November 2019; pp. 3407–3412. [Google Scholar]
- Garimella, A.; Amarnath, A.; Kumar, K.; Yalla, A.P.; Anandhavelu, N.; Chhaya, N.; Srinivasan, B.V. He is very intelligent, she is very beautiful? on mitigating social biases in language modelling and generation. In Proceedings of the Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, Online, 1–6 August 2021; pp. 4534–4545. [Google Scholar]
- Garimella, A.; Mihalcea, R.; Amarnath, A. Demographic-Aware Language Model Fine-tuning as a Bias Mitigation Technique. In Proceedings of the 2nd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 12th International Joint Conference on Natural Language Processing (Volume 2: Short Papers), Online, 20–23 November 2022; pp. 311–319. [Google Scholar] [CrossRef]
- Liu, R.; Jia, C.; Wei, J.; Xu, G.; Vosoughi, S. Quantifying and alleviating political bias in language models. Artif. Intell. 2022, 304, 103654. [Google Scholar] [CrossRef]
- Nie, S.; Fromm, M.; Welch, C.; Görge, R.; Karimi, A.; Plepi, J.; Mowmita, N.; Flores-Herr, N.; Ali, M.; Flek, L. Do Multilingual Large Language Models Mitigate Stereotype Bias? In Proceedings of the 2nd Workshop on Cross-Cultural Considerations in NLP, Bangkok, Thailand, 16 August 2024; pp. 65–83. [Google Scholar] [CrossRef]
- Ferrara, E. Should chatgpt be biased? Challenges and risks of bias in large language models. First Monday 2023, 28. [Google Scholar] [CrossRef]
- Zmigrod, R.; Mielke, S.J.; Wallach, H.; Cotterell, R. Counterfactual Data Augmentation for Mitigating Gender Stereotypes in Languages with Rich Morphology. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Florence, Italy, 28 July–2 August 2019; pp. 1651–1661. [Google Scholar] [CrossRef]
- Webster, K.; Wang, X.; Tenney, I.; Beutel, A.; Pitler, E.; Pavlick, E.; Chen, J.; Chi, E.; Petrov, S. Measuring and reducing gendered correlations in pre-trained models. arXiv 2020, arXiv:2010.06032. [Google Scholar]
- Barikeri, S.; Lauscher, A.; Vulić, I.; Glavaš, G. RedditBias: A Real-World Resource for Bias Evaluation and Debiasing of Conversational Language Models. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), Online, 1–6 August 2021; pp. 1941–1955. [Google Scholar] [CrossRef]
- Mondal, D.; Lipizzi, C. Mitigating Large Language Model Bias: Automated Dataset Augmentation and Prejudice Quantification. Computers 2024, 13, 141. [Google Scholar] [CrossRef]
- Maudslay, R.H.; Gonen, H.; Cotterell, R.; Teufel, S. It’s All in the Name: Mitigating Gender Bias with Name-Based Counterfactual Data Substitution. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), Hong Kong, China, 3–7 November 2019; pp. 5267–5275. [Google Scholar] [CrossRef]
- Meade, N.; Poole-Dayan, E.; Reddy, S. An Empirical Survey of the Effectiveness of Debiasing Techniques for Pre-trained Language Models. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Dublin, Ireland, 22–27 May 2022; pp. 1878–1898. [Google Scholar] [CrossRef]
- Serouis, I.M.; Sèdes, F. Exploring large language models for bias mitigation and fairness. In Proceedings of the 1st International Workshop on AI Governance (AIGOV) in Conjunction with the Thirty-Third International Joint Conference on Artificial Intelligence, Jeju, Republic of Korea, 3 August 2024. [Google Scholar]
- Qian, Y.; Muaz, U.; Zhang, B.; Hyun, J.W. Reducing Gender Bias in Word-Level Language Models with a Gender-Equalizing Loss Function. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics: Student Research Workshop, Florence, Italy, 28 July–2 August 2019; pp. 223–228. [Google Scholar] [CrossRef]
- Joniak, P.; Aizawa, A. Gender Biases and Where to Find Them: Exploring Gender Bias in Pre-Trained Transformer-based Language Models Using Movement Pruning. In Proceedings of the 4th Workshop on Gender Bias in Natural Language Processing (GeBNLP), Seattle, WA, USA, 15 July 2022; pp. 67–73. [Google Scholar] [CrossRef]
- Zhuang, F.; Qi, Z.; Duan, K.; Xi, D.; Zhu, Y.; Zhu, H.; Xiong, H.; He, Q. A comprehensive survey on transfer learning. Proc. IEEE 2020, 109, 43–76. [Google Scholar] [CrossRef]
- Azunre, P. Transfer Learning for Natural Language Processing; Simon and Schuster: New York, NY, USA, 2021. [Google Scholar]
- Ge, Y.; Hua, W.; Mei, K.; Tan, J.; Xu, S.; Li, Z.; Zhang, Y. OpenAGI: When llm meets domain experts. In Proceedings of the 37th International Conference on Neural Information Processing Systems, New Orleans, LA, USA, 10–16 December 2023. [Google Scholar]
- Liu, S.S. Unified Transfer Learning in High-Dimensional Linear Regression. In Proceedings of the 27th International Conference on Artificial Intelligence and Statistics, PMLR, Valencia, Spain, 2–4 May 2024; pp. 1036–1044. [Google Scholar]
- Delobelle, P.; Berendt, B. Fairdistillation: Mitigating stereotyping in language models. In Proceedings of the Joint European Conference on Machine Learning and Knowledge Discovery in Databases, Grenoble, France, 19–23 September 2022; Springer: Berlin/Heidelberg, Germany, 2022; pp. 638–654. [Google Scholar]
- Ahn, J.; Oh, A. Mitigating Language-Dependent Ethnic Bias in BERT. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, Online and Punta Cana, Dominican Republic, 7–11 November 2021; pp. 533–549. [Google Scholar] [CrossRef]
- Levy, S.; John, N.; Liu, L.; Vyas, Y.; Ma, J.; Fujinuma, Y.; Ballesteros, M.; Castelli, V.; Roth, D. Comparing Biases and the Impact of Multilingual Training across Multiple Languages. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, Singapore, 6–10 December 2023; pp. 10260–10280. [Google Scholar] [CrossRef]
- Srivastava, N.; Hinton, G.; Krizhevsky, A.; Sutskever, I.; Salakhutdinov, R. Dropout: A Simple Way to Prevent Neural Networks from Overfitting. J. Mach. Learn. Res. 2014, 15, 1929–1958. [Google Scholar]
- Da, Y.; Bossa, M.N.; Berenguer, A.D.; Sahli, H. Reducing Bias in Sentiment Analysis Models Through Causal Mediation Analysis and Targeted Counterfactual Training. IEEE Access 2024, 12, 10120–10134. [Google Scholar] [CrossRef]
- Cai, Y.; Cao, D.; Guo, R.; Wen, Y.; Liu, G.; Chen, E. Locating and mitigating gender bias in large language models. In Proceedings of the International Conference on Intelligent Computing, Tianjin, China, 5–8 August 2024; Springer: Berlin/Heidelberg, Germany, 2024; pp. 471–482. [Google Scholar]
- Vig, J.; Gehrmann, S.; Belinkov, Y.; Qian, S.; Nevo, D.; Singer, Y.; Shieber, S. Investigating gender bias in language models using causal mediation analysis. Adv. Neural Inf. Process. Syst. 2020, 33, 12388–12401. [Google Scholar]
- Liu, R.; Jia, C.; Wei, J.; Xu, G.; Wang, L.; Vosoughi, S. Mitigating political bias in language models through reinforced calibration. In Proceedings of the AAAI Conference on Artificial Intelligence, Virtual, 2–9 February 2021; Volume 35, pp. 14857–14866. [Google Scholar] [CrossRef]
- Park, J.H.; Shin, J.; Fung, P. Reducing Gender Bias in Abusive Language Detection. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Brussels, Belgium, 31 October–4 November 2018; pp. 2799–2804. [Google Scholar] [CrossRef]
- Bordia, S.; Bowman, S.R. Identifying and Reducing Gender Bias in Word-Level Language Models. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Student Research Workshop, Minneapolis, MN, USA, 2–7 June 2019; Kar, S., Nadeem, F., Burdick, L., Durrett, G., Han, N.R., Eds.; Association for Computational Linguistics: Stroudsburg, PA, USA, 2019; pp. 7–15. [Google Scholar] [CrossRef]
- Ravfogel, S.; Elazar, Y.; Gonen, H.; Twiton, M.; Goldberg, Y. Null It Out: Guarding Protected Attributes by Iterative Nullspace Projection. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Online, 5–10 July 2020; pp. 7237–7256. [Google Scholar] [CrossRef]
- Li, J.; Tang, Z.; Liu, X.; Spirtes, P.; Zhang, K.; Leqi, L.; Liu, Y. Steering LLMs Towards Unbiased Responses: A Causality-Guided Debiasing Framework. arXiv 2024, arXiv:2403.08743. [Google Scholar] [CrossRef]
- Zhang, C.; Zhang, L.; Zhou, D.; Xu, G. Causal Prompting: Debiasing Large Language Model Prompting based on Front-Door Adjustment. arXiv 2024, arXiv:2403.02738. [Google Scholar] [CrossRef]
- Pearl, J.; Glymour, M.; Jewell, N.P. Causal Inference in Statistics: A Primer; John Wiley & Sons: Hoboken, NJ, USA, 2016. [Google Scholar]
- Abid, A.; Farooqi, M.; Zou, J. Persistent Anti-Muslim Bias in Large Language Models. In Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society, New York, NY, USA, 19–21 May 2021; pp. 298–306. [Google Scholar]
- Gehman, S.; Gururangan, S.; Sap, M.; Choi, Y.; Smith, N.A. RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models. In Proceedings of the Findings of the Association for Computational Linguistics: EMNLP 2020; Cohn, T., He, Y., Liu, Y., Eds.; Association for Computational Linguistics: Stroudsburg, PA, USA, 2020; pp. 3356–3369. [Google Scholar]
- Koenecke, A.; Nam, A.; Lake, E.; Nudell, J.; Quartey, M.; Mengesha, Z.; Toups, C.; Rickford, J.R.; Jurafsky, D.; Goel, S. Racial disparities in automated speech recognition. Proc. Natl. Acad. Sci. USA 2020, 117, 7684–7689. [Google Scholar] [CrossRef]
- Saunders, D.; Byrne, B. Reducing Gender Bias in Neural Machine Translation as a Domain Adaptation Problem. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics; Jurafsky, D., Chai, J., Schluter, N., Tetreault, J., Eds.; Association for Computational Linguistics: Stroudsburg, PA, USA, 2020; pp. 7724–7736. [Google Scholar]
- EEOC. EEOC Guidance on AI in Employment. 2023. Available online: https://www.eeoc.gov (accessed on 15 September 2024).
- New York City Council. Local Law 144: Automated Employment Decision Tools. 2021. Available online: https://legistar.council.nyc.gov (accessed on 15 September 2024).
- U.S. Department of Health and Human Services. Section 1557 Final Rule. 2024. Available online: https://www.hhs.gov (accessed on 15 September 2024).
- European Commission. Artificial Intelligence Act. 2023. Available online: https://digital-strategy.ec.europa.eu/en/policies/european-approach-artificial-intelligence (accessed on 20 April 2026).
- Zack, T.; Lehman, E.; Suzgun, M.; Rodriguez, J.A. Assessing the potential of GPT-4 to perpetuate racial and gender biases in health care: A model evaluation study. Lancet Digit. Health 2024, 6, E12–E22. [Google Scholar] [CrossRef]
- MacCartney, B. Natural Language Inference; Stanford University: Stanford, CA, USA, 2009. [Google Scholar]
- Dev, S.; Li, T.; Phillips, J.M.; Srikumar, V. On measuring and mitigating biased inferences of word embeddings. In Proceedings of the AAAI Conference on Artificial Intelligence, New York, NY, USA, 7–12 February 2020; Volume 34, pp. 7659–7666. [Google Scholar]
- Zadeh, L.A. From search engines to question answering systems—The problems of world knowledge, relevance, deduction and precisiation. In Capturing Intelligence; Elsevier: Amsterdam, The Netherlands, 2006; Volume 1, pp. 163–210. [Google Scholar]
- Li, W.; Li, J.; Ma, W.; Liu, Y. Citation-Enhanced Generation for LLM-based Chatbot. arXiv 2024, arXiv:2402.16063. [Google Scholar]
- Hua, W.; Ge, Y.; Xu, S.; Ji, J.; Zhang, Y. Up5: Unbiased foundation model for fairness-aware recommendation. arXiv 2023, arXiv:2305.12090. [Google Scholar] [CrossRef]





| NLU Task | Gender Bias | Age Bias | Cultural or Regional Bias |
|---|---|---|---|
| Coreference resolution Identifying instances where different expressions refer to the same entity [75] | (1) Stereotypical occupation associations: Mislinking pronouns based on gender stereotypes in professions, e.g., doctor → male [28] (2) Gendered pronoun resolution: Incorrectly resolving pronouns for gender-neutral names, reflecting underlying gender biases [76] | (1) Assumptions about technological proficiency: Attributing technological skill to certain ages during pronoun resolution [77] (2) Bias in linking pronouns to age-related roles: Misassigning actions to younger or older individuals based on stereotypes [69] | (1) Regional variants and pronoun use: Difficulty resolving pronouns in pro-drop languages due to training on non-pro-drop data [77] (2) Cultural context in family roles: Misresolving pronouns based on cultural stereotypes about family responsibilities [76] |
| Semantic textual similarity Evaluating how similar the meanings of two texts are [78] | (1) Gendered language and pronoun resolution: Lower similarity scores for sentences differing only by gendered pronouns [79] (2) Gender-stereotyped professions: Overestimating similarity when sentences align with gender stereotypes [27] | (1) Age-based expectations in language use: Bias in similarity assessments due to age-related stereotypes [16] (2) Age stereotypes in sentiment and perception: Overestimating similarity based on stereotypes about age and wisdom [80] | (1) Cultural idioms and expressions: Underestimating similarity between culturally equivalent idioms [73] (2) Regional dialects and variants: Misjudging similarity due to differences in regional vocabulary [81] |
| Natural language inference Determining the relationship between a premise and a hypothesis [82] | (1) Stereotypical gender roles in professions: Incorrect inferences based on gender stereotypes in job roles [76] (2) Gendered language and assumptions: Bias in inference due to gender assumptions with neutral names [79] | (1) Bias toward younger individuals in dynamic roles: Overemphasizing youth in roles like entrepreneurship [16] (2) Age-related stereotypes in professional contexts: Assuming leadership roles are held by older individuals [28] | (1) Cultural norms and social roles: Inferences influenced by cultural stereotypes about professions [76] (2) Regional political and historical contexts: Difficulty with inferences requiring diverse historical knowledge [83] |
| Classification Assigning categories or labels to input text [84,85] | (1) Gendered language in job title classification: Associating certain professions with specific genders [27] (2) Gendered pronouns and name classification: Misclassifying names or pronouns based on gender stereotypes [79] | (1) Sentiment classification bias based on age: Misclassifying sentiment due to age-related stereotypes [86] (2) Bias in job application classification: Discriminating against older applicants in job screening [84] | (1) Misclassification due to cultural language variants: Incorrectly classifying non-standard dialects as negative [30] (2) Regional bias in political text classification: Bias toward dominant regional political ideologies [11] |
| Reading comprehension Answering questions based on a given text [87] | (1) Stereotypes in gendered activities: Assuming gender based on activities, e.g., cooking → female [88] (2) Assumptions about gender roles in family settings: Inferring roles based on traditional gender norms [16] | (1) Bias in health-related content: Associating certain health issues with age stereotypes [27] (2) Misinterpretation of age-related roles: Incorrectly assuming roles based on age [89] | (1) Cultural context misinterpretation: Misunderstanding culturally specific practices in texts [89] (2) Regional language varieties and dialects: Struggling with comprehension of non-standard dialects [12] |
| Sentiment analysis Identifying the emotional tone in text [90] | (1) Bias in sentiment toward gendered products or topics: Skewed sentiment analysis for gender-specific items [79] (2) Sentiment analysis in gendered contexts: Misclassifying sentiment in discussions challenging gender norms [91] | (1) Sentiment analysis on age-related topics: Assuming negative sentiment in texts by older individuals [92] (2) Stereotyping language use by older adults: Misinterpreting sentiment due to age-related language stereotypes [93] | (1) Cultural bias in sentiment toward social norms: Misclassifying sentiment across different cultural norms [94] (2) Sentiment in multilingual contexts: Incorrect sentiment assessment in code-switching texts [93] |
| NLG Task | Gender Bias | Age Bias | Cultural or Regional Bias |
|---|---|---|---|
| Question answering Providing accurate answers based on a given text or knowledge base [95] | (1) Gender bias in answer generation: When asked, “What should a good leader do?”, the model might use stereotypically male attributes, implying leadership qualities are inherently male [28] (2) Bias in answering ambiguous gender questions: For questions without specified gender, the model might default to male pronouns, reinforcing gender assumptions [88] | (1) Stereotypical answers about aging: When asked about activities for elderly people, the model might focus on sedentary activities, neglecting active pursuits [16] (2) Negative bias toward youth: Suggesting that young people lack experience for roles like managing a company [27] | (1) Bias in answering culturally specific questions: Providing answers based on regional popularity rather than global knowledge, e.g., stating “American football” as the most popular sport [96] (2) Language and regional bias in answer accuracy: Better performance on questions related to well-represented regions and languages [97] |
| Sentence completion Predicting and generating the next word or sequence to complete a sentence [98] | (1) Gender bias in descriptions of physical appearance: Completing sentences in ways that reinforce stereotypes about women’s and men’s concerns [16] (2) Gender bias in personal attributes: Associating women with “emotional” and men with “strong” in sentence completions [79] | (1) Activity and lifestyle assumptions: Completing sentences with stereotypical activities based on age [69] (2) Learning and education stereotypes: Assuming older adults are “catching up” on education [26] | (1) Cultural stereotyping in sentence completion: Overemphasizing specific aspects of a culture, e.g., “In Japan, people often eat sushi” [74] (2) Regional bias in place-based completions: Reinforcing stereotypes about regions, e.g., “In Africa, many people live in villages” [99] |
| Conversational Generating coherent dialogue and maintaining conversation context [100] | (1) Stereotypical responses based on gendered prompts: Reinforcing gender stereotypes in descriptions, e.g., nurses as female and caring [27] (2) Bias in gendered interactions: Responding differently based on assumed user gender [101] | (1) Assumptions about being tech-savvy: Assuming tech-savvy users are younger and providing simplistic explanations to older users [93] (2) Bias in addressing age-related concerns: Discouraging older individuals from pursuing new careers [102] | (1) Culturally inappropriate responses: Failing to understand cultural norms in responses [103] (2) Bias in handling regional topics: Focusing on negative topics for certain regions [104] |
| Recommender systems Suggesting personalized content or items based on user data [105] | (1) Product recommendations based on gender stereotypes: Suggesting products reinforcing traditional gender roles [106] (2) Career and education recommendations: Suggesting STEM careers more to men and arts to women [107] | (1) Age-related product recommendations: Recommending products based on age stereotypes [108] (2) Media and entertainment recommendations: Assuming preferences based on age, limiting content diversity [109] | (1) Regional bias in news and information recommendations: Under-representing news from minority regions [110] (2) Bias in language and cultural content recommendations: Prioritizing content in dominant languages [111] |
| Machine translation Translating text from one language to another [112] | (1) Gendered language mismatch: Introducing gender bias when translating from gender-neutral to gendered languages [113] (2) Gender stereotyping in occupational translations: Assigning gendered pronouns based on stereotypes [114] | (1) Bias in addressing older adults: Translations that condescend to older adults [16] (2) Translation of age-related idioms: Reinforcing negative stereotypes in translations [93] | (1) Cultural nuance loss: Mistranslating idioms and expressions without cultural context [115] (2) Bias toward dominant cultures: Favoring translations aligning with dominant (e.g., Western) norms [116] |
| Summarization Generating concise summaries of longer texts [8] | (1) Differential emphasis on roles: Emphasizing traditional gender roles in summaries [89] (2) Selective emphasis on gendered information: Overemphasizing gender-specific details not central to the story [117] | (1) Bias in summarizing content for different age groups: Simplifying content in a condescending way [32] (2) Omission of contributions based on age: Highlighting contributions of younger people over older individuals [89] | (1) Omission of culturally significant details: Omitting culturally important information in summaries [118] (2) Bias toward western narratives: Prioritizing Western perspectives in global news summaries [119] |
| Method | Description |
|---|---|
| Fairness metrics | Equal opportunity: Ensures similar true positive rates (TPRs) across sensitive groups. |
| Predictive parity: Checks for consistent prediction accuracy across groups. | |
| Calibration: Aligns predicted probabilities with actual outcomes for fairness. | |
| Interpretability tools | SHAP: Explains the impact of individual features on model predictions. |
| LIME: Provides local approximations of complex models for feature impact analysis. | |
| Counterfactual fairness | Scenario testing: Alters sensitive attributes (e.g., gender) to test output consistency. |
| Equity check: Verifies that changes do not affect model outcomes unfairly. |
| Debiasing | Methods | Summary |
|---|---|---|
| Pre-model | Resampling; data augmentation; expert intervention | Time saving without model training; time consuming to annotate bias cases, ineffectiveness with a self-biased model, privacy issue of resampling |
| Intra-model | Equalization and declustering; movement pruning; transfer learning; dropout regularization; causal inference | Flexibility in mitigating various types of biases, strong performance in empirical studies; time consuming for modifying and training models |
| Post-model | Reinforced calibration; Self-Debias; projection-based methods; causal prompting | Time saving without model training; demand of a large amount of data |
| Category | Method | Key Mechanism | Relative Cost | GPU Requirement |
|---|---|---|---|---|
| Pre-model | CDA/CDS | Augmenting training data | Low | Minimal |
| Pre-model | Expert curation | Human review of data | Medium | None |
| Intra-model | Equalization loss | Modified training objective | High | Full training |
| Intra-model | Movement pruning | Pruning biased subnetworks | High | Full training |
| Intra-model | FairDistillation | Knowledge distillation | Medium–high | Distillation run |
| Intra-model | Dropout regularization | Additional pretraining | Medium | Pretraining pass |
| Post-model | Self-Debias | Modified decoding | Low | Inference only |
| Post-model | SENT-DEBIAS/INLP | Subspace projection | Low–medium | Inference + PCA |
| Post-model | Reinforced calibration | RL-based generation | Medium | RL training |
| Post-model | Causal prompting | Prompt modification | Low–medium | Multiple queries |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Guo, Y.; Guo, M.; Su, J.; Yang, Z.; Zhu, M.; Li, H.; Qiu, M.; Liu, S.S. Bias in Large Language Models: Origin, Evaluation, and Mitigation. Electronics 2026, 15, 1824. https://doi.org/10.3390/electronics15091824
Guo Y, Guo M, Su J, Yang Z, Zhu M, Li H, Qiu M, Liu SS. Bias in Large Language Models: Origin, Evaluation, and Mitigation. Electronics. 2026; 15(9):1824. https://doi.org/10.3390/electronics15091824
Chicago/Turabian StyleGuo, Yufei, Muzhe Guo, Juntao Su, Zhou Yang, Mengqiu Zhu, Hongfei Li, Mengyang Qiu, and Shuo Shuo Liu. 2026. "Bias in Large Language Models: Origin, Evaluation, and Mitigation" Electronics 15, no. 9: 1824. https://doi.org/10.3390/electronics15091824
APA StyleGuo, Y., Guo, M., Su, J., Yang, Z., Zhu, M., Li, H., Qiu, M., & Liu, S. S. (2026). Bias in Large Language Models: Origin, Evaluation, and Mitigation. Electronics, 15(9), 1824. https://doi.org/10.3390/electronics15091824

