Bench of Euler: A Benchmark for Evaluating the Problem-Solving Abilities of Large Language Models
Abstract
1. Introduction
2. Problem Definition
3. Dataset
3.1. Dataset Construction
- The dataset includes fields for (a) the problem ID (aligned as Project Euler enumeration for brevity), (b) the standardized problem statement (in HTML format), (c) the correct answer, and the other difficulty metrics. The HTML encodings were further pre-processed to remove erroneous
markup, normalize the mathematical expressions, and simplify phrasing for clarity and token efficiency when ingested by LLMs. The raw dataset (in JSON format) can be obtained from the corresponding author upon reasonable request. To augment this dataset with the LLM generated final answers, prompts were given to the LLM APIs with sleep(15) (being gentle to both the LLM and the wallet). For Pass@1, each LLM was provided with the problem statement and was expected to generate code in its first attempt without any retries or external hints. The prompt for Pass@1 was as follows,
- You are a mathematical problem solver. Read the problem carefully and write an optimized Python script to compute the exact numeric answer.
- Problem: [HTML Problem Statement]
- Answer:
- The generated Python code was further used for getting the numeric value, which was later compared to the ground truth for evaluation. For Pass@2, the same prompt was given to the model with same configuration, except that the decoding was allowed to be non-deterministic by enabling temperature-based sampling (temp = 0.7).
3.2. Topic Descriptions
3.3. Code Evaluation Judgment
4. Experimentation
5. Limitations
5.1. Data Contamination
5.2. Sampling and Statistical Limitations
5.3. Failure-Mode Granularity
6. Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Mahdavi, S.; Li, M.; Liu, K.; Thrampoulidis, C.; Sigal, L.; Liao, R. Leveraging online olympiad-level math problems for LLMs training and contamination-resistant evaluation. arXiv 2025, arXiv:2501.14275. [Google Scholar]
- Liu, J.; Jiang, B. Scaffolding Students’ Ill-Structured Problem Solving via LLM—Multi-Armed Bandit Problem as a Case. In Proceedings of the 2024: ICCE 2024: The 32nd International Conference on Computers in Education, Quezon City, Philippines, 25–29 November 2024. [Google Scholar]
- Cheng, F.; Li, H.; Liu, F.; van Rooij, R.; Zhang, K.; Lin, Z. Empowering LLMs with logical reasoning: A comprehensive survey. arXiv 2025, arXiv:2502.15652. [Google Scholar]
- Xia, S.; Li, X.; Liu, Y.; Wu, T.; Liu, P. Evaluating mathematical reasoning beyond accuracy. Proc. Aaai Conf. Artif. Intell. 2025, 39, 27723–27730. [Google Scholar] [CrossRef] [Scilit]
- Patil, R.; Gudivada, V. A review of current trends, techniques, and challenges in large language models (LLMs). Appl. Sci. 2024, 14, 2074. [Google Scholar] [CrossRef] [Scilit]
- Pouransari, H.; Li, C.-L.; Chang, J.-H.; Anasosalu Vasu, P.K.; Koc, C.; Shankar, V.; Tuzel, O. Dataset decomposition: Faster LLM training with variable sequence length curriculum. Adv. Neural Inf. Process. Syst. 2024, 37, 36121–36147. [Google Scholar] [CrossRef] [Scilit]
- Kampelopoulos, D.; Tsanousa, A.; Vrochidis, S.; Kompatsiaris, I. A review of LLMs and their applications in the architecture, engineering and construction industry. Artif. Intell. Rev. 2025, 58, 250. [Google Scholar] [CrossRef] [Scilit]
- Naveed, H.; Khan, A.U.; Qiu, S.; Saqib, M.; Anwar, S.; Usman, M.; Akhtar, N.; Barnes, N.; Mian, A. A comprehensive overview of large language models. ACM Trans. Intell. Syst. Technol. 2025, 16, 106. [Google Scholar] [CrossRef] [Scilit]
- Maini, P.; Jia, H.; Papernot, N.; Dziedzic, A. LLM Dataset Inference: Did you train on my dataset? Adv. Neural Inf. Process. Syst. 2024, 37, 124069–124092. [Google Scholar] [CrossRef] [Scilit]
- Zhang, B.; Liu, Z.; Cherry, C.; Firat, O. When scaling meets LLM finetuning: The effect of data, model and finetuning method. arXiv 2024, arXiv:2402.17193. [Google Scholar]
- Mahapatra, J.; Garain, U. Impact of model size on fine-tuned LLM performance in data-to-text generation: A state-of-the-art investigation. arXiv 2024, arXiv:2407.14088. [Google Scholar]
- Chen, A.; Phang, J.; Parrish, A.; Padmakumar, V.; Zhao, C.; Bowman, S.R.; Cho, K. Two failures of self-consistency in the multi-step reasoning of LLMs. arXiv 2023, arXiv:2305.14279. [Google Scholar]
- Schnitzler, J.; Ho, X.; Huang, J.; Boudin, F.; Sugawara, S.; Aizawa, A. MorehopQA: More than multi-hop reasoning. arXiv 2024, arXiv:2406.13397. [Google Scholar]
- Hendrycks, D.; Burns, C.; Kadavath, S.; Arora, A.; Basart, S.; Tang, E.; Song, D.; Steinhardt, J. Measuring mathematical problem solving with the MATH dataset. arXiv 2021, arXiv:2103.03874. [Google Scholar]
- Cobbe, K.; Kosaraju, V.; Bavarian, M.; Chen, M.; Jun, H.; Kaiser, L.; Plappert, M.; Tworek, J.; Hilton, J.; Nakano, R.; et al. Training verifiers to solve math word problems. arXiv 2021, arXiv:2110.14168. [Google Scholar]
- Fan, J.; Martinson, S.; Wang, E.Y.; Hausknecht, K.; Brenner, J.; Liu, D.; Peng, N.; Wang, C.; Brenner, M.P. HardMath: A benchmark dataset for challenging problems in applied mathematics. arXiv 2024, arXiv:2410.09988. [Google Scholar]
- Liu, H.; Zheng, Z.; Qiao, Y.; Duan, H.; Fei, Z.; Zhou, F.; Zhang, W.; Zhang, S.; Lin, D.; Chen, K. MathBench: Evaluating the theory and application proficiency of LLMs with a hierarchical mathematics benchmark. arXiv 2024, arXiv:2405.12209. [Google Scholar]
- Arora, D.; Singh, H.G. Have LLMs advanced enough? A challenging problem solving benchmark for large language models. arXiv 2023, arXiv:2305.15074. [Google Scholar]
- Frieder, S.; Pinchetti, L.; Griffiths, R.-R.; Salvatori, T.; Lukasiewicz, T.; Petersen, P.; Berner, J. Mathematical capabilities of ChatGPT. Adv. Neural Inf. Process. Syst. 2023, 36, 27699–27744. [Google Scholar] [CrossRef] [Scilit]
- Sawada, T.; Paleka, D.; Havrilla, A.; Tadepalli, P.; Vidas, P.; Kranias, A.; Nay, J.J.; Gupta, K.; Komatsuzaki, A. ARB: Advanced reasoning benchmark for large language models. arXiv 2023, arXiv:2307.13692. [Google Scholar]
- Project Euler. Available online: https://projecteuler.net/ (accessed on 27 July 2025).
- Yuan, Y.; Zhao, L.; Zhang, K.; Zheng, G.; Liu, Q. Do LLMs overcome shortcut learning? An evaluation of shortcut challenges in large language models. arXiv 2024, arXiv:2410.13343. [Google Scholar]
- Finzi, M.; Kapoor, S.; Granziol, D.; Gu, A.; De Sa, C.; Kolter, J.Z.; Wilson, A.G. Compute-Optimal LLMs Provably Generalize Better With Scale. arXiv 2025, arXiv:2504.15208. [Google Scholar]
- Haller, P.; Golde, J.; Akbik, A. PECC: Problem Extraction and Coding Challenges. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), Torino, Italy, 20–25 May 2024; pp. 12690–12699. [Google Scholar]
- Balunović, M.; Dekoninck, J.; Petrov, I.; Jovanović, N.; Vechev, M. MathArena: Evaluating LLMs on Uncontaminated Math Competitions. arXiv 2025, arXiv:2505.23281. [Google Scholar]
- Chen, M.; Tworek, J.; Jun, H.; Yuan, Q.; Pinto, H.P.D.O.; Kaplan, J.; Edwards, H.; Burda, Y.; Joseph, N.; Brockman, G.; et al. Evaluating large language models trained on code. arXiv 2021, arXiv:2107.03374. [Google Scholar]
- Austin, J.; Odena, A.; Nye, M.; Bosma, M.; Michalewski, H.; Dohan, D.; Jiang, E.; Cai, C.; Terry, M.; Le, Q.; et al. Program Synthesis with Large Language Models. arXiv 2021, arXiv:2108.07732. [Google Scholar]
- Hendrycks, D.; Basart, S.; Kadavath, S.; Mazeika, M.; Arora, A.; Guo, E.; Burns, C.; Puranik, S.; He, H.; Song, D.; et al. Measuring Coding Challenge Competence with APPS. Adv. Neural Inf. Process. Syst. 2021, 34, 21143–21161. [Google Scholar]
- Li, Y.; Choi, D.; Chung, J.; Kushman, N.; Schrittwieser, J.; Leblond, R.; Eccles, T.; Keeling, J.; Gimeno, F.; Dal Lago, A.; et al. Competition-Level Code Generation with AlphaCode. Science 2022, 378, 1092–1097. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Jain, N.; Han, K.; Gu, A.; Li, W.-D.; Yan, F.; Zhang, T.; Wang, S.; Solar-Lezama, A.; Sen, K.; Stoica, I. LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code. arXiv 2024, arXiv:2403.07974. [Google Scholar]
- Jimenez, C.E.; Yang, J.; Wettig, A.; Yao, S.; Pei, K.; Press, O.; Narasimhan, K. SWE-bench: Can Language Models Resolve Real-World GitHub Issues? arXiv 2023, arXiv:2310.06770. [Google Scholar]
- Wei, L.; Fu, N.; Song, Y.; Wang, Q.; Hu, J. Probabilistic generative transformer language models for generative design of molecules. J. Cheminform. 2023, 15, 88. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Meister, C.; Vieira, T.; Cotterell, R. If beam search is the answer, what was the question? arXiv 2020, arXiv:2010.02650. [Google Scholar]
- Meister, C.I. Algorithms for Decoding Probabilistic Language Generators: Beam Search and Locally Typical Sampling. Ph.D. Thesis, ETH Zurich, Zurich, Switzerland, 2024. [Google Scholar]
- Borec, L.; Sadler, P.; Schlangen, D. The Unreasonable Ineffectiveness of Nucleus Sampling on Mitigating Text Memorization. arXiv 2024, arXiv:2408.16345. [Google Scholar]
- Banerjee, S.; Agarwal, A.; Singh, E. The Vulnerability of Language Model Benchmarks: Do They Accurately Reflect True LLM Performance? arXiv 2024, arXiv:2412.03597. [Google Scholar]
- Ivanov, T.; Penchev, V. AI Benchmarks and Datasets for LLM Evaluation. arXiv 2024, arXiv:2412.01020. [Google Scholar]
- Farchi, E.; Froimovich, S.; Katan, R.; Raz, O. Automatic generation of benchmarks and reliable LLM judgment for code tasks. arXiv 2024, arXiv:2410.21071. [Google Scholar]
- McNemar, Q. Note on the sampling error of the difference between correlated proportions or percentages. Psychometrika 1947, 12, 153–157. [Google Scholar] [CrossRef] [Scilit] [PubMed]






| Authors, Year | Dataset | # Problems | Problem Sourcing | Difficulty |
|---|---|---|---|---|
| Hendrycks et al., 2021 | Math [14] | 12,500 | Manual only | High School |
| Cobbe et al., 2021 | Gsm8k [15] | 8500 | Manual only | Grade School |
| Fan et al., 2024 | HARDMath [16] | 1400 | Algorithmic only | Graduate |
| Liu et al., 2024 | MathBench-T [17] | 632 | Manual + Algorithmic | Undergraduate |
| Arora et al., 2023 | JeeBench [18] | 236 | Manual only | High School |
| Frieder et al., 2024 | Ghosts [19] | 190 | Manual only | Graduate |
| Sawada et al., 2023 | ARB [20] | 34 | Manual only | Graduate |
| (Ours, 2025 +) | Bench of Euler | 954 | Algorithmic only | Graduate + |
| Topic | % of Problems | Description |
|---|---|---|
| Number Theory | 37.20% | Focused on primes, modular arithmetic, factorization, etc. |
| Combinatorics | 19.92% | Includes permutations, combinations, partitions, and counting. |
| Geometry | 13.00% | Covers triangles, circles, lattice points, and spatial relationships. |
| Others | 10.59% | Mixed or uncategorized problems that do not fit standard topics. |
| Logic | 6.18% | Involves deductive reasoning, puzzles, and constraint satisfaction. |
| Arithmetic | 5.87% | Covers operations on integers, digit manipulation, etc. |
| Algebra | 2.94% | Involves equations, series, polynomials, and algebraic identities. |
| Optimization | 2.83% | Focused on finding maximum or minimum under constraints. |
| Probability | 1.47% | Based on expected value, randomness, and outcome distributions. |
| Provider | Model Name | Reasons? | Pass@1 Accuracy | Pass@2 Accuracy |
|---|---|---|---|---|
| Anthropic | claude-opus-4-thinking | ✓ | 0.4185 | 0.4429 |
| claude-3-7-sonnet-20250219 | ✗ | 0.3431 | 0.3691 | |
| claude-sonnet-4-thinking | ✓ | 0.3368 | 0.3573 | |
| claude-sonnet-4 | ✗ | 0.3348 | 0.3426 | |
| claude-3-5-sonnet-20241022 | ✗ | 0.3281 | 0.3487 | |
| claude-3-5-haiku-20241022 | ✗ | 0.2756 | 0.2963 | |
| DeepSeek | deepseek-r1 | ✗ | 0.3165 | 0.3384 |
| deepseek-chat | ✗ | 0.2832 | 0.3043 | |
| gemini-ultra-preview | ✗ | 0.2351 | 0.2789 | |
| gemini-2.5-pro-preview | ✓ | 0.1602 | 0.1777 | |
| gemini-2.5-flash-preview | ✗ | 0.1187 | 0.1341 | |
| gemini-2.5-flash-lite-preview | ✗ | 0.0908 | 0.1032 | |
| OpenAI | gpt-4o | ✗ | 0.4321 | 0.4557 |
| o3-pro | ✓ | 0.4260 | 0.4481 | |
| o3-2025-04-16 | ✗ | 0.4043 | 0.4294 | |
| o3 | ✗ | 0.4020 | 0.4226 | |
| gpt-4.1-mini-2025-04-14 | ✗ | 0.3475 | 0.3691 | |
| o3-mini | ✗ | 0.3251 | 0.3453 | |
| gpt-4.1-nano-2025-04-14 | ✗ | 0.2879 | 0.3087 | |
| xAI | grok-3-reasoner | ✓ | 0.4005 | 0.4312 |
| grok-4 | ✗ | 0.2382 | 0.2447 | |
| grok-3 | ✗ | 0.1754 | 0.1928 |
| Provider | Model Name | TLE | Logical | Syntax | Wall Clock (s) |
|---|---|---|---|---|---|
| Anthropic | claude-opus-4-thinking | 29.45% | 27.51% | 1.19% | 25.39 |
| claude-3-7-sonnet-20250219 | 22.06% | 37.61% | 6.02% | 21.40 | |
| claude-sonnet-4-thinking | 26.01% | 36.75% | 3.56% | 25.87 | |
| claude-sonnet-4 | 22.89% | 37.61% | 6.02% | 21.81 | |
| claude-3-5-sonnet-20241022 | 22.18% | 38.72% | 6.28% | 21.45 | |
| claude-3-5-haiku-20241022 | 15.49% | 45.28% | 11.67% | 20.48 | |
| DeepSeek | deepseek-r1 | 23.19% | 38.95% | 6.21% | 21.96 |
| deepseek-chat | 15.59% | 44.47% | 11.62% | 20.52 | |
| gemini-ultra-preview | 16.73% | 48.02% | 11.74% | 21.10 | |
| gemini-2.5-pro-preview | 14.88% | 49.74% | 19.36% | 24.37 | |
| gemini-2.5-flash-preview | 11.81% | 52.67% | 23.65% | 20.21 | |
| gemini-2.5-flash-lite-preview | 12.16% | 54.35% | 24.41% | 20.27 | |
| OpenAI | gpt-4o | 26.06% | 27.98% | 2.75% | 21.30 |
| o3-pro | 29.04% | 26.80% | 1.56% | 25.14 | |
| o3-2025-04-16 | 27.65% | 29.20% | 2.73% | 22.15 | |
| o3 | 27.34% | 29.16% | 3.30% | 21.95 | |
| gpt-4.1-mini-2025-04-14 | 22.57% | 37.25% | 5.43% | 21.71 | |
| o3-mini | 22.64% | 38.86% | 5.99% | 21.71 | |
| gpt-4.1-nano-2025-04-14 | 15.86% | 44.02% | 11.33% | 20.67 | |
| xAI | grok-3-reasoner | 30.37% | 28.22% | 1.36% | 25.84 |
| grok-4 | 16.72% | 47.04% | 12.42% | 20.98 | |
| grok-3 | 10.44% | 49.53% | 22.48% | 19.75 |
| Provider | Model | B1 Accuracy | B2 Accuracy | B3 Accuracy | B4 Accuracy |
|---|---|---|---|---|---|
| OpenAI | gpt-4o | 95.67% | 72.10% | 5.07% | 0.00% |
| Anthropic | claude-opus-4-thinking | 99.00% | 61.74% | 6.66% | 0.00% |
| DeepSeek | deepseek-r1 | 77.10% | 38.80% | 10.70% | 0.00% |
| xAI | grok-3-reasoner | 94.10% | 50.35% | 15.75% | 0.00% |
| gemini-ultra-preview | 84.08% | 4.98% | 4.98% | 0.00% |
| Provider | Model | B1 Accuracy | B2 Accuracy | B3 Accuracy | B4 Accuracy |
|---|---|---|---|---|---|
| OpenAI | gpt-4o | 96.46% | 79.13% | 6.69% | 0.00% |
| Anthropic | claude-opus-4-thinking | 99.15% | 67.99% | 10.02% | 0.00% |
| DeepSeek | deepseek-r1 | 79.07% | 42.72% | 13.57% | 0.00% |
| xAI | grok-3-reasoner | 97.35% | 55.27% | 19.86% | 0.00% |
| gemini-ultra-preview | 91.73% | 11.80% | 8.03% | 0.00% |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Dutta, A.; Priya, S.S.; Ramamoorthy, A.; Kumar, P.K. Bench of Euler: A Benchmark for Evaluating the Problem-Solving Abilities of Large Language Models. AppliedMath 2026, 6, 143. https://doi.org/10.3390/appliedmath6090143
Dutta A, Priya SS, Ramamoorthy A, Kumar PK. Bench of Euler: A Benchmark for Evaluating the Problem-Solving Abilities of Large Language Models. AppliedMath. 2026; 6(9):143. https://doi.org/10.3390/appliedmath6090143
Chicago/Turabian StyleDutta, Anurag, S. Shanmuga Priya, A. Ramamoorthy, and Pijush Kanti Kumar. 2026. "Bench of Euler: A Benchmark for Evaluating the Problem-Solving Abilities of Large Language Models" AppliedMath 6, no. 9: 143. https://doi.org/10.3390/appliedmath6090143
APA StyleDutta, A., Priya, S. S., Ramamoorthy, A., & Kumar, P. K. (2026). Bench of Euler: A Benchmark for Evaluating the Problem-Solving Abilities of Large Language Models. AppliedMath, 6(9), 143. https://doi.org/10.3390/appliedmath6090143


