Valid but Not Always Runnable: An Open, Reproducible Benchmark of Large Language Models Drafting Gherkin Scenarios
Abstract
1. Introduction
2. Related Work
2.1. LLM Generation of Gherkin and Acceptance Tests
2.2. Requirements-to-Test Generation and Classical Precursors
2.3. Data Scarcity and Evaluation Methods
2.4. Positioning
3. Materials and Methods
3.1. Research Questions
3.2. Experimental Design
3.3. Inputs
3.4. Metrics
3.5. Prototype and Reproducibility
3.6. Analysis
4. Results
4.1. RQ1: Differences Across Models
4.2. RQ2: Differences Across the Three Requirement Corpora
4.3. RQ3: Cost, Tokens, and Cost-Efficiency
4.4. RQ4: Output Stability
4.5. Anti-Patterns and Gherkin Constructs
4.6. Budget Sensitivity: Output Token Limit as a Confound
4.7. Does the Quality Lead Survive a Length Control?
4.8. Judge Robustness: Independent Off-Panel Judges
4.9. Human Calibration of the Judge
4.10. A Judge-Free Coverage Check on the Requirement List Arm
4.11. Similarity to Hand-Authored Gold Standard, Anchored to a Human–Human Ceiling
4.12. From Parsing to Running: A Cucumber Runner Probe
4.13. Which Rubric Dimensions Separate the Models?
4.14. Prompt Sensitivity
4.15. Panel Refresh: Two Current Efficient Models
5. Discussion
5.1. Threats to Validity
5.2. Limitations and Future Work
6. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| API | Application Programming Interface |
| BDD | Behaviour-Driven Development |
| CI | Confidence Interval |
| LLM | Large Language Model |
| RFP | Request for Proposal |
| SD | Standard Deviation |
| Draftability | Whether syntactically valid Gherkin can be produced, as distinct from executability (binding to step definitions and passing as a test). |
| Well-formed | An output in which every scenario carries at least one step. |
| Grounding | Lexical overlap (0–1) of the generated steps with the source text. |
| Singularity | Rubric dimension: one behaviour per scenario. |
| Output-volume stability | Run-to-run SD of scenario count, as distinct from content-level reproducibility (step-set overlap across repetitions). |
| Off-panel judge | A judge model not among the eight evaluated and sharing no vendor lineage with them. |
| Neutral consensus | The mean judged score across the three off-panel judges. |
Appendix A. Prompts and Judge Rubric
System: You are an expert QA engineer. Convert the given software requirements into Gherkin (Cucumber) acceptance scenarios using Feature/Scenario/Given/When/Then. Write declarative, unambiguous scenarios. Output ONLY Gherkin.
User (base): You are given a software {input_type}. Convert it into Gherkin (Cucumber) acceptance scenarios that a BDD test suite could execute. Rules: use Feature, optional Background, and Scenario blocks with Given/When/Then; write declarative, business-focused steps, not UI mechanics; cover every requirement; invent nothing extra; output ONLY Gherkin, no code fences.
Appendix B. Model Snapshots, Pricing, and Configuration
| Model | Provider Identifier | In ($/M) | Out ($/M) |
|---|---|---|---|
| GPT-5.5 | openai/gpt-5.5 | 5.00 | 30.00 |
| Claude Opus 4.8 | anthropic/claude-opus-4-8 | 5.00 | 25.00 |
| GPT-4o | openai/gpt-4o | 2.50 | 10.00 |
| Claude Sonnet 5 | anthropic/claude-sonnet-5 | 3.00 | 15.00 |
| Gemini 2.5 Pro | google/gemini-2.5-pro | 1.25 | 10.00 |
| GPT-4o-mini | openai/gpt-4o-mini | 0.15 | 0.60 |
| Llama-3.1-70B | meta-llama/llama-3.1-70b-instruct | 0.40 | 0.40 |
| DeepSeek-V3 | deepseek/deepseek-chat | 0.20 | 0.80 |
| Panel refresh, generated 22 August 2026; pricing snapshot 22 August 2026 | |||
| * Gemini 3.5 Flash | google/gemini-3.5-flash | 1.50 | 9.00 |
| * DeepSeek-V4-Flash | deepseek/deepseek-v4-flash | 0.0713 | 0.1425 |
| Model | Prompt Tok. | Completion Tok. | Latency (s) | Cost ($/gen) |
|---|---|---|---|---|
| Claude Sonnet 5 | 550 | 796 | 7.4 | 0.0136 |
| Claude Opus 4.8 | 550 | 726 | 8.9 | 0.0209 |
| GPT-5.5 | 343 | 1084 | 13.5 | 0.0342 |
| Gemini 2.5 Pro | 345 | 1907 | 17.5 | 0.0195 |
| DeepSeek-V3 | 339 | 300 | 9.4 | 0.0003 |
| GPT-4o | 344 | 275 | 3.0 | 0.0036 |
| GPT-4o-mini | 344 | 276 | 5.6 | 0.0002 |
| Llama-3.1-70B | 347 | 300 | 8.9 | 0.0003 |
Appendix C. Output Normalisation and Software
Appendix D. Anti-Pattern, Structural, and Grounding Detection Rules
References
- North, D. Introducing BDD. Better Software. 2006. Available online: https://dannorth.net/introducing-bdd/ (accessed on 4 July 2026).
- Adzic, G. Specification by Example: How Successful Teams Deliver the Right Software; Manning Publications: Shelter Island, NY, USA, 2011. [Google Scholar]
- Solís, C.; Wang, X. A Study of the Characteristics of Behaviour Driven Development. In Proceedings of the 2011 37th EUROMICRO Conference on Software Engineering and Advanced Applications (SEAA); IEEE: New York, NY, USA, 2011; pp. 383–387. [Google Scholar] [CrossRef] [Scilit]
- Rathnayake, A.; Shahin, M.; Abaei, G. Behaviour Driven Development Scenario Generation with Large Language Models. arXiv 2026, arXiv:2603.04729. [Google Scholar] [CrossRef] [Scilit]
- Ferreira, M.; Viegas, L.; Faria, J.; Lima, B. Acceptance Test Generation with Large Language Models: An Industrial Case Study. arXiv 2025, arXiv:2504.07244. [Google Scholar] [CrossRef] [Scilit]
- Fonseca, P.; Lima, B.; Faria, J. Streamlining Acceptance Test Generation for Mobile Applications Through Large Language Models: An Industrial Case Study. arXiv 2025, arXiv:2510.18861. [Google Scholar] [CrossRef] [Scilit]
- Fernandes, H.; Perkusich, M.; Albuquerque, D.; Silva, I.; Santos, D.; Gorgônio, K.; Perkusich, A. A Comparative Study of LLMs for Gherkin Generation. In Proceedings of the Anais do XXXIX Simpósio Brasileiro de Engenharia de Software (SBES), Recife, Brazil, 22–26 September 2025. [Google Scholar] [CrossRef] [Scilit]
- dos Santos, S.R.R.; dos Santos, L.F.C.; Silva, M.; dos Santos, M.C.B.; Mendonça, M.F.; Santos, M.V.; da Silva Santos, M.F.; de Souza Bastos, A.L.; Marczak, S.; Soares, M.; et al. Automated Test Generation Using LLM Based on BDD: A Comparative Study. In Proceedings of the 21st International Conference on Web Information Systems and Technologies (WEBIST); SciTePress: Setúbal, Portugal, 2025; pp. 47–58. [Google Scholar] [CrossRef] [Scilit]
- Folorunsho, O.; Reza, H. AI-Driven Test Case Generation from Natural Language Requirements: A Survey of Techniques and Research Gaps. arXiv 2026, arXiv:2606.06563. [Google Scholar] [CrossRef] [Scilit]
- Tasarsu, M.; Tokmak, A.; Catal, C. Test Case Generation Using Large Language Models: A Systematic Literature Review. Clust. Comput. 2026, 29, 227. [Google Scholar] [CrossRef] [Scilit]
- Hassani, S.; Sabetzadeh, M.; Amyot, D. From Law to Gherkin: A Human-Centred Quasi-Experiment on the Quality of LLM-Generated Behavioural Specifications from Food-Safety Regulations. arXiv 2025, arXiv:2508.20744. [Google Scholar] [CrossRef] [Scilit]
- Deininger, P.; Slany, W. Replication Package—Valid but Not Always Runnable: An Open, Reproducible Benchmark of Large Language Models Drafting Gherkin Scenarios [Data Set and Software]. Zenodo. 2026. Available online: https://zenodo.org/records/22653864 (accessed on 7 September 2026).
- Karpurapu, S.; Myneni, S.; Nettur, U.; Gajja, L.S.; Burke, D.; Stiehm, T.; Payne, J. Comprehensive Evaluation and Insights Into the Use of Large Language Models in the Automation of Behavior-Driven Development Acceptance Test Formulation. IEEE Access 2024, 12, 58715–58721. [Google Scholar] [CrossRef] [Scilit]
- Siddeeq, S.; Abbasi, M.; Rasku, J.; Zhang, Z.; Christophe, F.; Mikkonen, T.; Abrahamsson, P. Epic-Organized vs. Requirement-Aligned Gherkin: An Empirical Evaluation of LLM-Based Acceptance Criteria Generation. arXiv 2026, arXiv:2607.01980. [Google Scholar] [CrossRef] [Scilit]
- Korraprolu, B.; Pinninti, P.; Reddy, Y. Test Case Generation for Requirements in Natural Language, An LLM Comparison Study. In Proceedings of the 18th Innovations in Software Engineering Conference (ISEC); ACM: New York, NY, USA, 2025. [Google Scholar] [CrossRef] [Scilit]
- Alagarsamy, S.; Tantithamthavorn, C.; Takerngsaksiri, W.; Arora, C.; Aleti, A. Enhancing Large Language Models for Text-to-Testcase Generation. J. Syst. Softw. 2025, 230, 112531. [Google Scholar] [CrossRef] [Scilit]
- Bob, R.; Storer, T. Behave Nicely! Automatic Generation of Code for Behaviour Driven Development Test Suites. In Proceedings of the IEEE 19th International Working Conference on Source Code Analysis and Manipulation (SCAM), Cleveland, OH, USA, 30 September–1 October 2019. [Google Scholar] [CrossRef] [Scilit]
- Gröpler, R.; Sudhi, V.; Calleja García, E.; Bergmann, A. NLP-Based Requirements Formalization for Automatic Test Case Generation. In Proceedings of the 29th International Workshop on Concurrency, Specification and Programming (CS&P), Berlin, Germany, 27–28 September 2021; CEUR-WS Vol-2951. pp. 18–30. [Google Scholar]
- Galloy, M.; Balfroid, M.; Vanderose, B.; Devroey, X. SelfBehave: Generating a Synthetic Behaviour-Driven Development Dataset Using SELF-INSTRUCT. In Proceedings of the A-MOST Workshop at IEEE ICST, Naples, Italy, 31 March–4 April 2025. [Google Scholar] [CrossRef] [Scilit]
- Zhang, T.; Kishore, V.; Wu, F.; Weinberger, K.; Artzi, Y. BERTScore: Evaluating Text Generation with BERT. arXiv 2020, arXiv:1904.09675. [Google Scholar] [CrossRef] [Scilit]
- Blackwell, R.; Barry, J.; Cohn, A. Towards Reproducible LLM Evaluation: Quantifying Uncertainty in LLM Benchmark Scores. arXiv 2024, arXiv:2410.03492. [Google Scholar] [CrossRef] [Scilit]
- Zheng, L.; Chiang, W.L.; Sheng, Y.; Zhuang, S.; Wu, Z.; Zhuang, Y.; Lin, Z.; Li, Z.; Li, D.; Xing, E.; et al. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. arXiv 2023, arXiv:2306.05685. [Google Scholar] [CrossRef] [Scilit]
- Panickssery, A.; Bowman, S.; Feng, S. LLM Evaluators Recognize and Favor Their Own Generations. arXiv 2024, arXiv:2404.13076. [Google Scholar] [CrossRef] [Scilit]
- Dubois, Y.; Galámbosi, B.; Liang, P.; Hashimoto, T. Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators. arXiv 2024, arXiv:2404.04475. [Google Scholar] [CrossRef] [Scilit]
- Huang, D.; Chew, S.; Dutkiewicz, A.; Wang, Z. LLM-as-a-Judge for Scalable Test Coverage Evaluation: Accuracy, Operational Reliability, and Cost. arXiv 2025, arXiv:2512.01232. [Google Scholar] [CrossRef] [Scilit]
- Binamungu, L.; Embury, S.; Konstantinou, N. Characterising the Quality of Behaviour Driven Development Specifications. In Proceedings of the Agile Processes in Software Engineering and Extreme Programming (XP 2020); Lecture Notes in Business Information Processing; Springer: Berlin/Heidelberg, Germany, 2020; Volume 383, pp. 87–102. [Google Scholar] [CrossRef] [Scilit]
- Binamungu, L.; Embury, S.; Konstantinou, N. Detecting Duplicate Examples in Behaviour Driven Development Specifications. In Proceedings of the 2018 IEEE Workshop on Validation, Analysis and Evolution of Software Tests (VST), Campobasso, Italy, 20 March 2018; pp. 6–10. [Google Scholar] [CrossRef] [Scilit]
- Wang, J.; Huang, Y.; Chen, C.; Liu, Z.; Wang, S.; Wang, Q. Software Testing with Large Language Models: Survey, Landscape, and Vision. IEEE Trans. Softw. Eng. 2024, 50, 911–936. [Google Scholar] [CrossRef] [Scilit]
- DeepSeek-AI. DeepSeek-V3 Technical Report. arXiv 2024, arXiv:2412.19437. [Google Scholar] [CrossRef] [Scilit]
- Ferrari, A.; Spagnolo, G.; Gnesi, S. PURE: A Dataset of Public Requirements Documents. In Proceedings of the 2017 IEEE 25th International Requirements Engineering Conference (RE), Lisbon, Portugal, 4–8 September 2017; pp. 502–505. [Google Scholar] [CrossRef] [Scilit]
- Dalpiaz, F. Requirements Data Sets (User Stories), Version 2 [Data Set]. Mendeley Data. 2024. Available online: https://data.mendeley.com/datasets/7zbk8zsd8y/2 (accessed on 7 September 2026).
- Lucassen, G.; Dalpiaz, F.; van der Werf, J.M.E.M.; Brinkkemper, S. Improving Agile Requirements: The Quality User Story Framework and Tool. Requir. Eng. 2016, 21, 383–403. [Google Scholar] [CrossRef] [Scilit]
- Cucumber Ltd. Gherkin: A Parser and Compiler for the Gherkin Language [Computer Software]. Python Package Gherkin-Official. Available online: https://github.com/cucumber/gherkin (accessed on 3 July 2026).
- Wynne, M.; Hellesøy, A. The Cucumber Book: Behaviour-Driven Development for Testers and Developers; Pragmatic Bookshelf: Dallas, TX, USA, 2012. [Google Scholar]
- Femmer, H.; Méndez Fernández, D.; Wagner, S.; Eder, S. Rapid Quality Assurance with Requirements Smells. J. Syst. Softw. 2017, 123, 190–213. [Google Scholar] [CrossRef] [Scilit]
- ISO/IEC/IEEE 29148:2018; Systems and Software Engineering—Life Cycle Processes—Requirements Engineering. International Organization for Standardization: Geneva, Switzerland, 2018.





| Study | Models/Corpora | Automated Evaluation | Human Evaluation | Public Release |
|---|---|---|---|---|
| Rathnayake et al. [4] | 3 models; 500 user-story/BDD pairs from four proprietary products of one company | Text similarity, semantic similarity, LLM-as-judge | Yes: human-expert preference | Yes: dataset and code (MIT) |
| Fernandes et al. [7] | 7 models; free-form test descriptions | METEOR; repeated-measures variability | Not reported | No dataset or prototype |
| Dos Santos et al. [8] | 4 commercial assistants; user stories with Gherkin acceptance criteria | Similarity, coverage, accuracy, efficiency | Not reported | No dataset or prototype |
| This study | 10 models (8 primary, 2 refresh); 74 artifacts from PURE, Dalpiaz, and constructed RFP excerpts; 3700 generations | Parser validity raw and normalised, Cucumber runner acceptance, primary plus three off-panel LLM judges, chrF against human gold standard, anti-patterns, stability, cost and Pareto | Yes: two blinded raters on 72 items, reported against a measured human–human ceiling (Section 4.9); 30-item hand-authored gold standard with an external coverage audit | Yes: corpora, gold standard, prototype, and analysis scripts |
| Model | Parse | Parse | Well- | Cover- | Quality | Scen. | Latency | Cost | Stab. |
|---|---|---|---|---|---|---|---|---|---|
| Raw (%) | Norm. (%) | Formed (%) | Age (%) | (1–5) | /Item | (s) | ($/gen) | (SD) | |
| Claude Sonnet 5 | 72.7 | 99.7 | 94.9 | 99.9 [99.7, 100] | 4.68 [4.62, 4.73] | 7.2 | 7.4 | 0.0136 | 0.78 |
| Claude Opus 4.8 | 40.8 | 99.5 | 99.5 | 99.8 [99.4, 100] | 4.66 [4.60, 4.71] | 7.5 | 8.9 | 0.0209 | 0.74 |
| GPT-5.5 | 99.7 | 100.0 | 100.0 | 99.1 [97.6, 100] | 4.65 [4.59, 4.71] | 6.3 | 13.5 | 0.0342 | 0.58 |
| Gemini 2.5 Pro | 50.8 | 98.4 | 98.4 | 91.3 [87.8, 94.5] | 4.52 [4.44, 4.60] | 5.9 | 17.5 | 0.0195 | 0.86 |
| DeepSeek-V3 | 74.6 | 96.5 | 95.7 | 98.7 [97.7, 99.4] | 4.38 [4.32, 4.44] | 5.3 | 9.4 | 0.0003 | 0.59 |
| GPT-4o | 98.6 | 98.9 | 98.9 | 98.8 [97.6, 99.6] | 4.37 [4.31, 4.42] | 5.3 | 3.0 | 0.0036 | 0.16 |
| GPT-4o-mini | 94.9 | 98.9 | 98.9 | 98.2 [97.2, 99.1] | 4.35 [4.29, 4.41] | 5.4 | 5.6 | 0.0002 | 0.22 |
| Llama-3.1-70B | 91.1 | 97.3 | 97.3 | 96.5 [94.9, 97.8] | 4.32 [4.25, 4.39] | 5.2 | 8.9 | 0.0003 | 0.50 |
| Model | Quality [95% CI] | Cost $/gen [95% CI] | On Front | P(non-dom.) |
|---|---|---|---|---|
| GPT-4o-mini | 4.35 [4.29, 4.41] | 0.00022 [0.00021, 0.00022] | yes | 1.00 |
| Llama-3.1-70B | 4.32 [4.24, 4.39] | 0.00026 [0.00025, 0.00027] | – | 0.17 |
| DeepSeek-V3 | 4.38 [4.32, 4.43] | 0.00031 [0.00030, 0.00032] | yes | 0.79 |
| GPT-4o | 4.37 [4.31, 4.42] | 0.00362 [0.00350, 0.00373] | – | 0.29 |
| Claude Sonnet 5 | 4.68 [4.62, 4.73] | 0.01359 [0.01277, 0.01446] | yes | 1.00 |
| Gemini 2.5 Pro | 4.52 [4.44, 4.60] | 0.01950 [0.01870, 0.02028] | – | 0.00 |
| Claude Opus 4.8 | 4.66 [4.61, 4.71] | 0.02089 [0.01960, 0.02214] | – | 0.10 |
| GPT-5.5 | 4.65 [4.59, 4.71] | 0.03423 [0.03228, 0.03631] | – | 0.12 |
| Model | Halluc. | Ground. | GWT (%) | UI-voc (%) | Repeat. (%) | Outline (%) | Backgr. (%) |
|---|---|---|---|---|---|---|---|
| Claude Sonnet 5 | 0.28 | 0.44 | 98.0 | 0.2 | 11.3 | 8.9 | 67.8 |
| Claude Opus 4.8 | 0.18 | 0.50 | 97.3 | 0.2 | 14.0 | 13.2 | 26.2 |
| GPT-5.5 | 0.11 | 0.48 | 97.9 | 0.1 | 17.7 | 38.9 | 25.9 |
| Gemini 2.5 Pro | 0.26 | 0.34 | 97.6 | 0.4 | 9.5 | 18.1 | 60.8 |
| DeepSeek-V3 | 0.16 | 0.49 | 91.5 | 0.3 | 8.9 | 0.0 | 0.0 |
| GPT-4o | 0.02 | 0.64 | 95.3 | 0.3 | 8.7 | 0.0 | 0.0 |
| GPT-4o-mini | 0.06 | 0.59 | 96.0 | 0.2 | 7.7 | 0.0 | 72.2 |
| Llama-3.1-70B | 0.11 | 0.54 | 91.5 | 0.3 | 10.2 | 0.0 | 84.6 |
| Model | Sonnet 5 (Primary) | Mistral Large | Qwen3-235B | Grok 4.3 | Neutral Mean |
|---|---|---|---|---|---|
| Claude Sonnet 5 | 4.68 | 4.85 | 4.97 | 4.70 | 4.84 |
| Claude Opus 4.8 | 4.66 | 4.82 | 4.96 | 4.67 | 4.82 |
| GPT-5.5 | 4.65 | 4.85 | 4.98 | 4.72 | 4.85 |
| Gemini 2.5 Pro | 4.52 | 4.71 | 4.87 | 4.53 | 4.70 |
| DeepSeek-V3 | 4.38 | 4.76 | 4.95 | 4.50 | 4.74 |
| GPT-4o | 4.37 | 4.75 | 4.93 | 4.44 | 4.71 |
| GPT-4o-mini | 4.35 | 4.70 | 4.89 | 4.37 | 4.65 |
| Llama-3.1-70B | 4.32 | 4.69 | 4.91 | 4.37 | 4.66 |
| Comparison | Dimension | ICC(2,1) | MAE | qw- | |
|---|---|---|---|---|---|
| Human A vs. human B | Quality | +0.73 (+0.69) | +0.64 (+0.60) | 0.354 (0.395) | +0.56 (+0.54) |
| (ceiling) | Coverage | +0.72 (+0.70) | +0.76 (+0.75) | 0.033 (0.037) | |
| Hallucinations | +0.79 (+0.75) | +0.80 (+0.78) | 0.250 (0.281) | ||
| Judge vs. human | Quality | +0.60 (+0.57) | +0.47 (+0.44) | 0.330 (0.334) | +0.46 (+0.45) |
| consensus | Coverage | +0.53 (+0.56) | +0.21 (+0.21) | 0.047 (0.049) | |
| Hallucinations | +0.43 (+0.42) | +0.45 (+0.46) | 0.403 (0.375) |
| Model | chrF (Multi-Ref) | [95% CI] | % of Range | Judged-Quality Rank |
|---|---|---|---|---|
| Claude Opus 4.8 | 72.16 | [69.79, 74.56] | 107 | 3 |
| Claude Sonnet 5 | 69.01 | [66.40, 71.65] | 100 | 1 |
| GPT-4o | 68.35 | [65.80, 70.86] | 91 | 7 |
| GPT-5.5 | 67.97 | [65.57, 70.37] | 97 | 2 |
| Llama-3.1-70B | 66.54 | [64.04, 69.06] | 87 | 8 |
| GPT-4o-mini | 66.12 | [63.61, 68.63] | 85 | 6 |
| DeepSeek-V3 | 64.35 | [61.80, 66.82] | 80 | 5 |
| Gemini 2.5 Pro | 58.13 | [55.46, 60.80] | 66 | 4 |
| Human–human ceiling | 67.40 | 100 | ||
| Skeleton floor (unrelated items) | 35.07 | 0 |
| Model | Parse Valid (%) | Runner Loads (%) | Multi-Feature (%) | Defs/Scenario |
|---|---|---|---|---|
| GPT-5.5 | 100.0 | 99.7 | 0.3 | 3.74 |
| GPT-4o | 98.9 | 98.6 | 0.3 | 3.19 |
| GPT-4o-mini | 98.9 | 94.9 | 4.1 | 3.11 |
| Llama-3.1-70B | 97.3 | 91.1 | 6.2 | 3.26 |
| DeepSeek-V3 | 96.5 | 74.6 | 23.2 | 3.40 |
| Claude Sonnet 5 | 99.7 | 73.5 | 26.5 | 3.51 |
| Gemini 2.5 Pro | 98.4 | 50.8 | 47.6 | 3.89 |
| Claude Opus 4.8 | 99.5 | 40.8 | 58.9 | 3.26 |
| All | 98.6 | 78.0 | 20.9 | 3.40 |
| Human gold standard (A+B) | 100.0 | 100.0 | 0.0 | 3.17 |
| Model | Relevance | Clarity | Completeness | Singularity |
|---|---|---|---|---|
| Claude Sonnet 5 | 4.89 | 4.64 | 4.43 | 4.72 |
| Claude Opus 4.8 | 4.97 | 4.61 | 4.28 | 4.73 |
| GPT-5.5 | 4.92 | 4.73 | 4.36 | 4.61 |
| Gemini 2.5 Pro | 4.74 | 4.78 | 4.04 | 4.65 |
| DeepSeek-V3 | 4.81 | 4.26 | 4.08 | 4.45 |
| GPT-4o | 4.93 | 4.04 | 3.91 | 4.54 |
| GPT-4o-mini | 4.82 | 4.08 | 3.92 | 4.51 |
| Llama-3.1-70B | 4.74 | 4.09 | 3.81 | 4.49 |
| Model | Parse (%) | Runner (%) | Multi-F. (%) | Quality | Coverage | Cost ($/gen) | P(non-dom.) |
|---|---|---|---|---|---|---|---|
| * DeepSeek-V4-Flash | 97.3 | 85.1 | 25.7 | 4.42 | 0.966 | 0.00014 | 1.00 |
| GPT-4o-mini | 98.9 | 94.9 | 5.4 | 4.35 | 0.982 | 0.00022 | 0.02 |
| Llama-3.1-70B | 97.3 | 91.1 | 10.8 | 4.32 | 0.965 | 0.00026 | 0.00 |
| DeepSeek-V3 | 96.5 | 74.6 | 40.5 | 4.38 | 0.987 | 0.00031 | 0.07 |
| GPT-4o | 98.9 | 98.6 | 1.4 | 4.37 | 0.988 | 0.00362 | 0.02 |
| Claude Sonnet 5 | 99.7 | 73.5 | 41.9 | 4.68 | 0.999 | 0.01359 | 1.00 |
| * Gemini 3.5 Flash | 97.3 | 97.3 | 0.0 | 4.53 | 0.983 | 0.01424 | 0.06 |
| Gemini 2.5 Pro | 98.4 | 50.8 | 59.5 | 4.52 | 0.913 | 0.01950 | 0.00 |
| Claude Opus 4.8 | 99.5 | 40.8 | 64.9 | 4.66 | 0.998 | 0.02089 | 0.10 |
| GPT-5.5 | 100.0 | 99.7 | 1.4 | 4.65 | 0.991 | 0.03423 | 0.12 |
| Team Priority | Suggested Model(s) | Principal Caveat |
|---|---|---|
| General first-draft generation (default) | Any low-cost model: GPT-4o-mini, DeepSeek-V3, Llama-3.1-70B | All draft valid Gherkin at near-ceiling judged quality. The per-draft dollar difference among them is immaterial once a human reviews them, so the choice should be based on convenience or ecosystem fit. GPT-4o-mini and Llama also need file-splitting least often (Section 4.12). |
| Drop-in runner readiness (objective) | GPT-5.5 (99.7%), GPT-4o (98.6%), GPT-4o-mini (94.9%) | Fraction of raw outputs the Cucumber runner loads without repair. Claude Opus 4.8 (40.8%) and Gemini 2.5 Pro (50.8%) emit several Feature blocks in one file and need a splitting step before any runner accepts them. This is objective, mechanically repairable, and invisible to parse-validity alone. |
| Best fidelity signal (provisional) | GPT-4o, GPT-4o-mini (cluster) | Lowest-hallucination cluster (GPT-4o 0.02 invented per output), but provisional: hallucination is judge-scored, and it is the dimension on which the judge and human ranked outputs most differently ( against a human–human , Section 4.9). Grounding (GPT-4o 0.64) is a copy-sensitive lexical floor, not a quality merit. |
| Data-driven scenarios (Scenario Outline) | Any model except GPT-4o-mini, prompted permissively | Under the base prompt, only frontier models emit Scenario Outline, but a permissive prompt elicits it from most low-cost models too in a 24-item probe (Llama 100%, DeepSeek 29%, GPT-4o 13%; GPT-4o-mini still 0%): prompt for it explicitly. |
| On-prem/data governance | Open-weight: Llama-3.1-70B, DeepSeek-V3 | Self-hostable, keeping requirements and RFP text off third-party APIs. Same Scenario Outline caveat as the data-driven row. |
| Minimal cost at scale | DeepSeek-V4-Flash if available (Section 4.15), else GPT-4o-mini, then DeepSeek-V3, Llama-3.1-70B | The dollar edge dominates only in high-volume, light- or no-review pipelines (unlike the default row, where review cost washes it out). With light review, the provisional frontier clarity/completeness edge (Section 4.13) re-enters. Same Scenario Outline caveat as the data-driven row. |
| Output-volume determinism | GPT-4o | Lowest scenario-count SD (0.16), though content-level reproducibility is low across the whole panel. |
| Low-latency/interactive use | GPT-4o (3.0 s), GPT-4o-mini (5.6 s) | Fastest in the panel, but latency is a single-session wall-clock through non-uniform API routes (Section 5.1), so it should be read as indicative. The reasoning-heavy GPT-5.5 (13.5 s) and Gemini 2.5 Pro (17.5 s) are the slowest. |
| Maximum quality, cost no object | Claude Sonnet 5 | Of the frontier tier, it is the only Pareto-non-dominated member: it dominates Claude Opus 4.8 in 90% of bootstrap replicates and GPT-5.5 in 87%, so the two costlier models buy nothing measurable (Table 3). The quality lead itself is small, direction-only, within judge re-scoring noise, and not separable from length (Section 4.7), so the premium is not firmly established (judge-scored; the judge resolves gaps this small poorly, Section 4.9). It also needs file-splitting for 27% of outputs. |
| Not recommended on current evidence | Gemini 2.5 Pro | The only model dominated in every bootstrap replicate (Claude Sonnet 5 gives higher judged quality at 70% of the cost). It is also lowest on requirement coverage, lowest on similarity to hand-authored gold standard (66% of the human ceiling), the least stable, the slowest, and the second-worst on runner acceptance. Four instruments with different failure modes agree. |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Deininger, P.; Slany, W. Valid but Not Always Runnable: An Open, Reproducible Benchmark of Large Language Models Drafting Gherkin Scenarios. AI 2026, 7, 359. https://doi.org/10.3390/ai7090359
Deininger P, Slany W. Valid but Not Always Runnable: An Open, Reproducible Benchmark of Large Language Models Drafting Gherkin Scenarios. AI. 2026; 7(9):359. https://doi.org/10.3390/ai7090359
Chicago/Turabian StyleDeininger, Patrick, and Wolfgang Slany. 2026. "Valid but Not Always Runnable: An Open, Reproducible Benchmark of Large Language Models Drafting Gherkin Scenarios" AI 7, no. 9: 359. https://doi.org/10.3390/ai7090359
APA StyleDeininger, P., & Slany, W. (2026). Valid but Not Always Runnable: An Open, Reproducible Benchmark of Large Language Models Drafting Gherkin Scenarios. AI, 7(9), 359. https://doi.org/10.3390/ai7090359

