When Agents Disagree: The Selection Bottleneck in Multi-Agent LLM Pipelines
Featured Application
Abstract
1. Introduction
2. Related Work
2.1. Multi-Agent Debate and Frameworks
2.2. Mixture-of-Agents and Self-MoA
2.3. Selection and Routing
2.4. Diversity Theory
2.5. LLM-as-Judge
2.6. The Gap This Paper Fills
3. Materials and Methods
3.1. Setup and Notation
- Team mean: , the expected quality of a randomly chosen candidate.
- Team oracle: , the expected quality of the best candidate.
3.2. The Selection Quality Model
3.3. The Selection Bottleneck
- (i)
- (the diverse team’s mean is lower),
- (ii)
- (the diverse team’s oracle is higher).
3.4. Optimal Team Size
3.5. Connection to Classical Results
3.6. Design Overview
3.7. Team Compositions
- homo_opus: Claude Opus × 3. All-strong homogeneous baseline.
- diverse_strong: Claude Opus + GPT-5.4 + Gemini 2.5 Pro. Three frontier models from different families.
- diverse_mixed: Claude Opus + Gemini 2.5 Pro + Claude Haiku. Two strong models plus one weaker model from a different capability tier.
3.8. Selector Mechanisms
- Judge-based selection: An external judge panel—Claude Sonnet, GPT-5-mini, and DeepSeek-V3p2—reads all candidate outputs and selects the best via pairwise comparison with Bradley–Terry scoring (Section 3.11). Candidate order was randomized for each pairwise judge comparison to control for position bias [20].
- Majority vote: Each agent independently selects the best candidate from the pool. The candidate receiving the most votes is chosen; ties are broken randomly.
- MoA synthesis: Claude Sonnet reads all candidate outputs and produces a single synthesized response that blends elements from each candidate, following the Mixture-of-Agents protocol [2].
3.9. Agent–Judge Separation
3.10. Task Battery
- Coding (6): streaming pipeline design, race condition debugging, multi-tenant architecture, security/performance code review, API migration, flaky test stabilization.
- Creative extended (6): polyphonic narrative, epistolary fiction, memory-themed poetry cycle, worldbuilding charter, courtroom dialogue, myth retelling.
- Ethics and policy (6): facial recognition policy, AI tutor data ethics, autonomous weapons export, organ allocation, ventilator triage, carbon border adjustment.
- Math and logic (6): probability paradox, integer optimization, logic grid puzzles, Bayesian diagnostics, scheduling with dependencies, game-theoretic resource division.
- Reasoning (6): causal policy analysis, counterfactual outbreak response, argument evaluation, root cause analysis, strategic negotiation, uncertainty assessment.
- Science (6): heat dome mechanisms, memory consolidation, adaptive clinical trials, battery degradation, ecosystem restoration, and epidemiological modeling.
- Summarization (6): board packet crisis brief, incident timeline, expert panel comparison, customer feedback synthesis, multi-opinion legal summary, policy roundtable digest.
3.11. Evaluation Metric
3.12. Pre-Registration and Analysis Plan
3.13. Threats to Validity
4. Results
4.1. Confirmatory Results
4.2. Exploratory Findings
4.3. Regression Analysis
4.4. Calibrating the Crossover Threshold
4.5. Replication Stability
5. Discussion
5.1. Reconciling Conflicting Prior Work
5.2. Why Synthesis Fails
5.3. Why Weak Models May Help
5.4. Practical Decision Framework
5.5. Limitations
- 1.
- Targeted design. Our five-cell design maximizes power for specific contrasts but does not estimate all possible interactions (e.g., homogeneous + vote, diverse_mixed + synthesis). A full factorial would enable richer interaction analyses.
- 2.
- Model specificity and baseline scope. Results are demonstrated for a specific set of frontier models using an Opus-only homogeneous run as the single-model baseline. Comparing against the strongest individual model across the full diverse candidate pool might yield different quantitative estimates; whether the selector advantage survives that stricter comparison is left to future work. Whether the same patterns hold for other model families or future generations is also an open question.
- 3.
- Fixed generation temperature. All generation runs use . Varying temperature could alter within-model output variance and potentially affect the relative performance of homogeneous versus diverse teams. The conclusion that homogeneous sampling variation is negligible is therefore specific to moderate fixed-temperature settings.
- 4.
- LLM-as-judge and subjective tasks. LLM-based evaluation remains a proxy for human judgment. For open-ended subjective categories—creative writing and ethics/policy in our task battery—LLM judges are known to exhibit stylistic preferences and may agree less with human raters than on analytical tasks [19]. We do not have human evaluation data for these categories, and the extent to which our results generalize to human preferences in subjective domains is unknown. Human evaluation on a representative subset of tasks, particularly creative and policy tasks, is a priority for future work. Our decoupled evaluation pass (Table 2) partially addresses selection–evaluation circularity, confirming all directional contrasts under independent judges, but the independent panel itself exhibited limitations: one of three judges (GPT-4o-mini) proved degenerate (99.6% tie rate), reducing the effective independent panel to two judges. Per-judge tie rates varied substantially (GPT-4o-mini: 99.6%, Gemini Flash: 76.7%, GLM-5: 50.3%), suggesting that weaker models may lack the discriminative capacity for reliable pairwise evaluation. The synthesis–judge overlap (Claude Sonnet serving as both synthesizer and one of three judges) remains a specific concern, though the very low synthesis win rate argues against self-enhancement bias as a primary driver.
- 5.
- Bounded dependent variable. Win rates are bounded in , yet we model them linearly. Our observed values (0.13–0.96) avoid extreme floor/ceiling effects, and we verified that logit-transformed results are qualitatively identical.
- 6.
- Static topology. All experiments use a single-round generate-then-select pipeline. Iterative topologies (multi-round debate, recursive refinement) may exhibit different dynamics.
- 7.
- Distinguishability not measured. Our framework invokes output distinguishability (s) as the key mediator, but we do not measure it directly. Future work should operationalize distinguishability via embedding-space distances.
- 8.
- Diverse_mixed confound. The diverse_mixed cell simultaneously changes capability (replacing GPT-5.4 with Claude Haiku) and family diversity (two Anthropic models instead of one). We cannot isolate these effects and flag this as a design limitation of the exploratory comparison.
5.6. Future Work
6. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- Hong, L.; Page, S.E. Groups of Diverse Problem Solvers Can Outperform Groups of High-Ability Problem Solvers. Proc. Natl. Acad. Sci. USA 2004, 101, 16385–16389. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Wang, J.; Wang, J.; Athiwaratkun, B.; Zhang, C.; Zou, J. Mixture-of-Agents Yields State-of-the-Art on AlpacaEval 2.0, MT-Bench, and FLASK. arXiv 2024, arXiv:2406.04692. [Google Scholar]
- Li, X.; Zhang, L.; Zhang, Z.; Yang, Y.; Wang, Z. More Agents Is All You Need: Self-MoA Outperforms Mixed-MoA. arXiv 2025, arXiv:2502.00674. [Google Scholar]
- Wu, Q.; Bansal, G.; Zhang, J.; Wu, Y.; Li, B.; Zhu, E.; Jiang, L.; Zhang, X.; Zhang, S.; Liu, J.; et al. AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation. arXiv 2023, arXiv:2308.08155. [Google Scholar]
- Li, G.; Hammoud, H.A.A.K.; Itani, H.; Khizbullin, D.; Ghanem, B. CAMEL: Communicative Agents for “Mind” Exploration of Large Language Model Society. arXiv 2023, arXiv:2303.17760. [Google Scholar] [CrossRef] [Scilit]
- Hong, S.; Zheng, X.; Chen, J.; Cheng, Y.; Wang, J.; Zhang, C.; Wang, Z.; Yau, S.K.S.; Lin, Z.; Zhou, L.; et al. MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework. arXiv 2023, arXiv:2308.00352. [Google Scholar]
- Du, Y.; Li, S.; Torralba, A.; Tenenbaum, J.B.; Mordatch, I. Improving Factuality and Reasoning in Language Models through Multiagent Debate. In Proceedings of the 40th International Conference on Machine Learning (ICML), Honolulu, HI, USA, 23–29 July 2023. [Google Scholar]
- Smit, K.; Keane, I.; Mao, W. When Models Think Alike: The Limits of Multi-Agent Debate. arXiv 2023, arXiv:2311.17371. [Google Scholar]
- Choi, J.; Lee, S.; Ok, J. Identity Bias in Large Language Model Debate. arXiv 2025, arXiv:2510.07517. [Google Scholar]
- Jiang, D.; Ren, X.; Lin, B.Y. LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative Fusion. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (ACL), Toronto, ON, Canada, 9–14 July 2023. [Google Scholar]
- Chen, J.; Wang, X.; Xu, R.; Yuan, S.; Chen, L.; Xiao, Y. LLMSelector: Selecting the Right LLM for Any Task. arXiv 2025, arXiv:2502.14815. [Google Scholar]
- Stiennon, N.; Ouyang, L.; Wu, J.; Ziegler, D.M.; Lowe, R.; Voss, C.; Radford, A.; Amodei, D.; Christiano, P. Learning to Summarize from Human Feedback. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), Virtual, 6–12 December 2020; Volume 33. [Google Scholar]
- Snell, C.; Lee, J.; Xu, K.; Kumar, A. Scaling LLM Test-Time Compute Optimally Can Be More Effective than Scaling Model Parameters. arXiv 2024, arXiv:2408.03314. [Google Scholar] [CrossRef] [Scilit]
- Wang, X.; Wei, J.; Schuurmans, D.; Le, Q.; Chi, E.; Narang, S.; Chowdhery, A.; Zhou, D. Self-Consistency Improves Chain of Thought Reasoning in Language Models. In Proceedings of the International Conference on Learning Representations (ICLR), Kigali, Rwanda, 1–5 May 2023. [Google Scholar]
- Tang, Z.; Kharlapenko, D.; Li, B.; Li, F.; Chen, H.; Cai, D. The Value of Diversity in Multi-Agent Systems. arXiv 2026, arXiv:2602.07186. [Google Scholar]
- de Condorcet, M. Essai sur l’Application de l’Analyse à la Probabilité des Décisions Rendues à la Pluralité des Voix; Royale: Paris, France, 1785. [Google Scholar]
- Ladha, K.K. The Condorcet Jury Theorem, Free Speech, and Correlated Votes. Am. J. Political Sci. 1992, 36, 617–634. [Google Scholar] [CrossRef] [Scilit]
- Li, J.; Zhang, Q.; Yu, Y.; Fu, Q.; Ye, D. More Agents Is All You Need. arXiv 2024, arXiv:2402.05120. [Google Scholar]
- Zheng, L.; Chiang, W.L.; Sheng, Y.; Zhuang, S.; Wu, Z.; Zhuang, Y.; Lin, Z.; Li, Z.; Li, D.; Xing, E.P.; et al. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), New Orleans, LA, USA, 10–16 December 2023; Volume 36. [Google Scholar]
- Stureborg, R.; Alikaniotis, D.; Suhara, Y. Large Language Models are Inconsistent and Biased Evaluators. arXiv 2024, arXiv:2405.01724. [Google Scholar] [CrossRef] [Scilit]
- Panickssery, N.; Bowman, S.R.; Feng, S. LLM Evaluators Recognize and Favor Their Own Generations. arXiv 2024, arXiv:2404.13076. [Google Scholar] [CrossRef] [Scilit]
- Verga, P.; Hofstätter, S.; Althammer, S.; Su, Y.; Piktus, A.; Arkhangorodsky, A.; Xu, M.; White, N.; Lewis, P. Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models. arXiv 2024, arXiv:2404.18796. [Google Scholar] [CrossRef] [Scilit]
- Chiang, W.L.; Zheng, L.; Sheng, Y.; Angelopoulos, A.N.; Li, T.; Li, D.; Zhang, H.; Zhu, B.; Jordan, M.; Gonzalez, J.E.; et al. Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference. arXiv 2024, arXiv:2403.04132. [Google Scholar] [CrossRef] [Scilit]
- Bradley, R.A.; Terry, M.E. Rank Analysis of Incomplete Block Designs: I. The Method of Paired Comparisons. Biometrika 1952, 39, 324–345. [Google Scholar] [CrossRef] [Scilit]
- Barr, D.J.; Levy, R.; Scheepers, C.; Tily, H.J. Random Effects Structure for Confirmatory Hypothesis Testing: Keep It Maximal. J. Mem. Lang. 2013, 68, 255–278. [Google Scholar] [CrossRef] [Scilit] [PubMed]



| Cell | Agents | Selector | Tests |
|---|---|---|---|
| div_strong + judge | Opus + GPT-5.4 + Gem.-Pro | Judge panel | Reference |
| homo_opus + judge | Opus | Judge panel | Diversity effect |
| div_mixed + judge | Opus + Gem.-Pro + Haiku | Judge panel | Weak-model effect |
| div_strong + vote | Opus + GPT-5.4 + Gem.-Pro | Majority vote | Judge vs. vote |
| div_strong + synth | Opus + GPT-5.4 + Gem.-Pro | MoA synthesis | Selection vs. synthesis |
| Cell | Original WR | Decoupled WR (2J) | Decoupled WR (3J) |
|---|---|---|---|
| div_mixed + judge | 0.929 | 0.726 | 0.722 |
| div_strong + judge | 0.810 | 0.611 | 0.500 |
| div_strong + vote | 0.496 | 0.506 | 0.389 |
| homo_opus + judge | 0.512 | 0.500 † | 0.000 † |
| div_strong + synth | 0.179 | 0.312 | 0.119 |
| Cell | Sonnet | GPT-5m | DeepSeek | Mean | |
|---|---|---|---|---|---|
| div_strong + judge | 0.958 | 0.756 | 0.595 | 0.770 | 0.095 |
| homo_opus + judge | 0.500 | 0.518 | 0.619 | 0.546 | 0.667 |
| div_mixed + judge | 0.994 | 0.893 | 0.631 | 0.839 | 0.175 |
| div_strong + vote | 0.492 | 0.496 | 0.512 | 0.500 | 0.236 |
| div_strong + synth | 0.234 | 0.131 | 0.345 | 0.237 | 0.263 |
| Configuration | Rel. Cost | BT-WR | 95% CI |
|---|---|---|---|
| div_mixed + judge † | 0.929 | [0.887, 0.964] | |
| div_strong + judge | 0.810 | [0.768, 0.851] | |
| homo_opus + judge | 0.512 | [0.500, 0.530] | |
| div_strong + vote | 0.496 | [0.425, 0.563] | |
| div_strong + synth | 0.179 | [0.127, 0.234] |
| OLS (HC3) | MixedLM | |
|---|---|---|
| Intercept (homo_opus + judge) | *** (0.008) | *** (0.024) |
| diverse_strong + judge | *** (0.024) | *** (0.035) |
| diverse_mixed + judge | *** (0.022) | *** (0.035) |
| diverse_strong + synth | *** (0.029) | *** (0.035) |
| diverse_strong + vote | (0.037) | (0.035) |
| Task var. () | — | ≈0 |
| 0.740 | — | |
| F/ | — |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Maryanskyy, A.; Budnikov, D.; Kaliyev, A.T. When Agents Disagree: The Selection Bottleneck in Multi-Agent LLM Pipelines. Appl. Sci. 2026, 16, 4914. https://doi.org/10.3390/app16104914
Maryanskyy A, Budnikov D, Kaliyev AT. When Agents Disagree: The Selection Bottleneck in Multi-Agent LLM Pipelines. Applied Sciences. 2026; 16(10):4914. https://doi.org/10.3390/app16104914
Chicago/Turabian StyleMaryanskyy, Artem, Dmitry Budnikov, and Alibek T. Kaliyev. 2026. "When Agents Disagree: The Selection Bottleneck in Multi-Agent LLM Pipelines" Applied Sciences 16, no. 10: 4914. https://doi.org/10.3390/app16104914
APA StyleMaryanskyy, A., Budnikov, D., & Kaliyev, A. T. (2026). When Agents Disagree: The Selection Bottleneck in Multi-Agent LLM Pipelines. Applied Sciences, 16(10), 4914. https://doi.org/10.3390/app16104914

