Next Article in Journal
Pedagogical Transformation Using Large Language Models in a Cybersecurity Course
Next Article in Special Issue
Could You Be Wrong: Metacognitive Prompts for Improving Human Decision Making Help LLMs Identify Their Own Biases
Previous Article in Journal
Machine Learning and Deep Learning in Lung Cancer Diagnostics: A Systematic Review of Technical Breakthroughs, Clinical Barriers, and Ethical Imperatives
Previous Article in Special Issue
Benchmarking Psychological Lexicons and Large Language Models for Emotion Detection in Brazilian Portuguese
 
 
Article
Peer-Review Record

No Free Lunch in Language Model Bias Mitigation? Targeted Bias Reduction Can Exacerbate Unmitigated LLM Biases

by Shireen Chand †, Faith Baca *,† and Emilio Ferrara *
Reviewer 1: Anonymous
Reviewer 2: Anonymous
Submission received: 23 November 2025 / Revised: 24 December 2025 / Accepted: 8 January 2026 / Published: 13 January 2026

Round 1

Reviewer 1 Report

Comments and Suggestions for Authors

The paper addresses an important and timely problem, and the “No Free Lunch” is engaging 😊. That said, several aspects of the manuscript would benefit from tightening and clarification.

The abstract and introduction make quite strong claims (“consistent and statistically significant evidence,” etc.) before the reader has seen any detail on how results were obtained. Phrases like “inevitably” and “necessary” overstate what can realistically be supported by a finite empirical study on specific models and benchmarks.

The connection to the “No Free Lunch” theorem is currently more rhetorical than rigorous. NFL is originally about average performance over problem distributions. As here you apply it to bias dimensions in LLMs, but that mapping is not clearly spelled out. You might frame NFL as an inspiration rather than a direct theoretical foundation, explain explicitly what counts as a problem class in your analogy, and situate your work alongside existing fairness trade-off and impossibility results, which are in some ways closer to what you are studying.

The central research question and hypothesis also need slight alignment. The question (“How does mitigation along one axis affect several axes?”) is descriptive, while the hypothesis asserts that targeted interventions will “inevitably” cause side effects. Reformulating the hypothesis in less absolute terms would make it both more realistic and more clearly testable.

Methodologically, the introduction is too vague given how central the empirical setup is to your claims. You mention four debiasing techniques by name, ten models from seven families, and StereoSet as your benchmark, but the reader does not yet understand what distinguishes these techniques, how the models were chosen or grouped, or what exactly StereoSet measures. Briefly characterizing each method (prompting vs activation vs logit vs parameter editing), clarifying what types and sizes of models you include, and stating explicitly what “stereotypical preference” and “coherence” mean in terms of StereoSet metrics would significantly improve readability. It is also important to acknowledge limitations and to be explicit that your conclusions are about this particular operationalization of bias, not bias in all forms.

The interpretation of results would benefit from more precise language. Terms like “frequently,” “often negative,” and “more harm than the original intervention sought to fix” should be grounded in numbers  and compared a baseline. Even in the abstract, a single sentence that sketches the overall pattern would make the “No Free Lunch” story more concrete and less anecdotal.

In the related work section, the current emphasis on general trade-offs and NFL could be complemented with more direct engagement with fairness and bias literature. Work on fairness accuracy trade-offs, incompatibility of fairness criteria, subgroup fairness and fairness gerrymandering, and previous studies on debiasing LMs and their unintended side effects. That would better situate your contribution as extending known issues into the specific context of LLM bias mitigation across multiple axes.

Finally, some terminology and stylistic details could be sharpened. Concepts like “entangled biases,” “trading one bias for another,” “coherence,” and “harm” should be explicitly defined in terms of your metrics rather than left at an intuitive level.

Author Response

We kindly thank the referee for the thoughtful feedback: please see attached response letter and happy holidays!

Author Response File: Author Response.pdf

Reviewer 2 Report

Comments and Suggestions for Authors

Reviewer Report

The manuscript is interesting, and the analysis is comprehensive, but all sections lack depth in terms of what this work contributes to previous studies. The methodological contribution is poorly defined, and the proposed framework is not clearly connected to literature. Furthermore, the results need to be better specified to understand their real scope.

Below, section by section, I indicate possible improvements:

Abstract

The abstract identifies a relevant problem, but the scientific contribution should be expressed more clearly. The study measures the multidimensional effects of bias mitigation, simply indicating in one line what the work contributes in relation to previous work. 

Introduction

The introduction presents the general problem of bias in LLMs, opening the door to the need to examine the multidimensional effects of bias removal. The section does not indicate what methodological innovation is, what new contributions are made, nor does it describe the specific phases of the research process that the article will follow.

My recommendations to the authors are as follows:

  • Clarify scientific contribution by explicitly stating what methodological or analytical innovation the work introduces beyond existing bias-mitigation evaluations.
  • Outline the structure or phases of the proposed auditing framework within the introduction, so readers understand how the study is organized and what steps the research process will follow.
  • Situate the contribution more clearly within the literature, identifying the specific gap this work addresses and how it advances current understanding of cross-dimensional bias effects.

Related Work

The section presents a well-referenced and comprehensive overview, perhaps at an overly descriptive level. The article should be positioned in relation to previous work and clearly explain the methodological gap that the research aims to fill. I think the literature review is very good, but the manuscript's contribution needs to emerge clearly from it.

My recommendations to the authors are as follows:

  • Clarify the methodological gap by explicitly pointing out that previous studies have not evaluated the cross-effects of mitigation techniques on non-target dimensions.
  • Better position the contribution by showing how the proposed framework advances on previous benchmarks, techniques, or evaluations.
  • Link each block of the state of the art (trade-offs, benchmarks, techniques) to the problem addressed in the article and justify why existing approaches are insufficient.
  • Reduce descriptive content and reinforce the critical analysis that contextualizes the novelty of the work.

Methodology

The methodology section is clear and describes the experimental procedure in detail. The text is sometimes overly descriptive and does not adequately justify design decisions. It is necessary to highlight the methodological contribution, the choice of metrics, the layers involved, the parameters used, and the selection of StereoSet as the main benchmark.

My recommendations to the authors are as follows:

  • Justify experimental decisions, explaining why specific layers are chosen, the value α=1.0, or the intervention in five layers in the case of Activation Patching.
  • Clarify the methodological contribution, indicating how the proposed audit framework represents an advance over previous work.
  • Justify the exclusive use of StereoSet, discussing its known limitations and explaining why it remains appropriate for this study rather than recent benchmarks.
  • Consider the imbalance between dimensions, especially the disparity in sizes (e.g., 78 examples in religion vs. 976 in race), and comment on how this may influence the stability of the metrics.
  • Integrate a brief conceptual reflection, explaining how the proposed phases are articulated as a coherent methodological process and not just as an operational sequence.

Results

I believe that the results section provides a solid statistical analysis; its approach is descriptive, and a critical interpretation of the patterns observed would be necessary. The findings are clear, but these results need to be linked to the initial hypotheses, benchmark limitations, and previous literature.

My recommendations to the authors are as follows:

  • Delve deeper into the interpretation of the results, explaining why certain patterns emerge and how they relate to the ‘No Free Lunch’ hypothesis.
  • Explicitly discuss the impact of imbalance in StereoSet, especially in dimensions with very few examples, to better contextualize the observed variability.
  • Expand on the statistical justification, detailing the approach used (tests applied, corrections, significance criteria).
  • Incorporate a clear summary of the findings, for example through tables or summaries that allow key patterns between techniques, dimensions, and models to be quickly identified.
  • Link the results to previous studies, showing to what extent they confirm, contradict, or expand existing literature on bias mitigation and spillover between dimensions.

 

Discussion

The discussion presents important ideas, but they should be further refined. It identifies the interdependence of representations as a probable cause of spillovers, but does not delve into a more rigorous analysis linking the results to previous research, nor to the limitations of the proposed experimental design. The critical reflection on StereoSet and its impact on the findings is somewhat superficial given its relevance to the interpretation of the results.

My recommendations to the authors are as follows:

  • Connect the results more closely to existing literature, explaining how the findings confirm, contradict, or qualify previous research on intertwined representations, catastrophic forgetting, and bias mitigation.
  • Delve deeper into the role of StereoSet as a limitation, detailing how its cultural biases, errors, and imbalances may influence the observed spillover patterns.
  • Offer more technical explanations of the phenomenon, beyond a conceptual description (‘entanglement’), incorporating mechanistic hypotheses or internal evidence where possible.
  • Qualify conclusions about generalization, making it clear that the results are obtained from a single benchmark and with specific techniques, which limit extrapolation.
  • Expand the reflection on practical implications, indicating what researchers or developers should do in the face of spillovers and how to better evaluate future mitigation techniques.

Conclusions

The conclusions adequately summarize the main results but remain general and do not emphasize the specific contribution of the study or its methodological implications. Nor is it clearly highlighted how the work advances the state of the art, and future recommendations remain too broad.

  • Clarify the final contribution of the study, explicitly explaining what the work adds to previous research on debiasing.
  • Delimit the scope of the results, emphasizing that the conclusions depend heavily on a single benchmark (StereoSet) and specific techniques.
  • Propose more concrete future lines of research, aimed at methodological or experimental improvements directly derived from the findings.
  • Reinforce the methodological message, highlighting how the audit framework can be adopted, extended, or integrated into current bias assessment practices.
  • Better differentiate between empirical observations and general statements, avoiding extrapolation beyond what the data allows.

Author Response

We kindly thank the referee for the thoughtful feedback: please see attached response letter and happy holidays!

Author Response File: Author Response.pdf

Round 2

Reviewer 1 Report

Comments and Suggestions for Authors

good work on corrections

Reviewer 2 Report

Comments and Suggestions for Authors

Thank you for the revised version of the manuscript and for your detailed responses to the reviewer comments. All points have been satisfactorily addressed, and the revisions clearly strengthen the clarity, methodological positioning, and interpretation of the results.

From my side, everything is now in order.

Back to TopTop