A Spanish Language Proficiency Dataset for AI Evaluation
Abstract
1. Introduction
2. Related Work
2.1. Existing Datasets for Reading Comprehension Evaluation
2.2. Datasets in Spanish
3. Methodology for the Elaboration of Exams
3.1. Specifications for Exam Design
3.2. Difficulty and Quality Control in DELE
- 1.
- Task authoring under the test specifications. Item writers create tasks that match the blueprint (task models, number of items, and time constraints) so that each form targets the same construct.
- 2.
- Use of reference documents. Writers anchor content selection to the Instituto Cervantes Curriculum Plan, which lists grammatical, lexical, pragmatic, and functional inventories expected at each CEFR level.
- 3.
- Successive expert revisions. The coordinating author performs an initial review and then coordinates a validation cycle with expert reviewers. Depending on the test, the workflow includes multiple review rounds (typically 4–7) until the tasks meet quality criteria.
- 4.
- Pre-testing with learners. Instituto Cervantes pilots each form with approximately 100–200 learners who resemble the target candidate population. The pilot includes an anchor task (kept constant across pilots but excluded from operational administrations) to link results across forms.
- 5.
- Psychometric analysis. Analysts evaluate item functioning using standard indices. The difficulty coefficient estimates how easy/hard an item is relative to others, while the discrimination index measures whether the item separates higher- from lower-proficiency examinees.
- 6.
- Item revision. When an item is too easy, too hard, or insufficiently discriminative, writers revise it and re-check it for consistency with the blueprint.
- 7.
- Post-administration monitoring. After an operational call, Instituto Cervantes analyzes candidate results at scale to monitor item functioning and stores validated tasks in a task bank for potential reuse.
3.3. Scoring and Pass/Fail Criteria
3.4. Exercise Types and Mapping to IC-UNED-RC-ES
- 1.
- A single text followed by several questions.
- 2.
- Several short texts with one question each.
- 3.
- Several texts that must be related to several statements.
- 4.
- Reconstruction of a text.
- 5.
- Multiple texts paired with a larger set of statements/questions; each statement must be assigned to the correct text (with distractors).
- 6.
- Gap-filling: a text with removed fragments and response options.
- 1.
- Multiple-choice. The system must answer questions about a text by selecting the correct option among a small set of candidates (typically three).
- 2.
- Matching. The system must match each query text (or statement) to the best candidate text in a second set. The candidate pool may include distractors, and each query has a single correct match.
- 3.
- Fill-in-the-gap. The system must reconstruct a text by placing removed fragments into the correct gaps. Each gap has exactly one correct fragment, and the candidate pool may contain more fragments than gaps.
3.5. Selection of Exams for IC-UNED-RC-ES and Conversion to a Machine-Readable Format
4. Description of IC-UNED-RC-ES Dataset
- Lexical complexity [33], determined by the number of different content words per sentence (Lexical Complexity Index) and the number of low-frequency words per 100 content words (Index of Low Frequency Words). The higher the former, the greater the difficulty in reading comprehension.
- Complexity of sentences [33], focuses on measuring the number of words per sentence, thus obtaining the sentence length index (average sentence length), and the number of complex sentences per sentence, from a complex sentence index (complex sentences).
- Readability of Fernández-Huerta [34] is an adaptation of the Flesch Reading Ease score [35] into Spanish, based on the number of syllables and the number of sentences. We use the corrected version of the formula as suggested by [32]. A higher score indicates that the text is easier to read: 0–30 is very hard, 30–50 is hard, 50–60 is slightly hard, 60–70 is normal for an adult, 70–80 is slightly easy, 80–90 is easy, and 90–100 is very easy.
- Mean dependency tree depth [37] aims at capturing syntactic complexity in terms of recursive or nested structures. Higher depth means the text is more complex to read.
- First, it is important to highlight that fill-in-the-gap and matching exercises are uncommon in other collections, regardless of the language. Results indicate that the probability of randomly choosing the correct answer in these types of exercises (around 0.1) is lower than in multiple-choice exercises (around 0.3). Thus, in addition to the difficulty of their own content, matching and fill-in-the-gap exercises introduce the added complexity of more combinations, making it more challenging to succeed through random guessing. This makes our dataset especially relevant for research purposes in the field.
- Comparing the complexity measures across the different exercise types reveals consistent trends and values at each level, except for readability measures in the fill-in-the-gap tasks, which show the lowest scores. This indicates that the texts used in this exercise type are more difficult to read.
- In general, our results align with the exam levels, validating them and opening the door for assessing the expected exercise difficulty for each level.
- Although these measures have proven effective as indicators of proficiency, certain aspects relevant to the most advanced levels may be only partially captured, as examiners may look for the use of specific syntactic constructions that do not necessarily imply more complex texts but rather deeper language understanding. Additional mechanisms may be required to account for specific lexical or grammatical structures in order to more accurately determine the complexity of texts at those levels. This is reflected in the last rows of the complexity metrics tables, where several cases show similar values for B2 to C2 levels, compared with the previous ones.
4.1. Multiple-Choice
4.2. Matching
4.3. Fill-in-the-Gap
5. Evaluation Metrics
5.1. Human Evaluation
5.2. Machine Evaluation
- Multiple-choice exercises: We measure accuracy as the proportion of questions correctly answered.
- Matching exercises: We measure accuracy as the proportion of correctly matched texts.
- Fill-in-the-gap exercises: We measure accuracy as the proportion of correctly filled gaps.
- At the question level, where correct answers are counted across all exercises associated with the same task, providing an assessment of task-specific performance.
- At the exam level, where scores for each exam are considered. Each exam contains several exercises of different types. An exam is deemed to be passed if an accuracy score (calculated as the proportion of correct answers) above 0.6 is reached. Then, the proportion of passed exams is given as a global score. The idea is to summarize functional performance across task types with varying cognitive demands, providing an indicator of overall assessment difficulty.
6. Experimental Analysis
6.1. Question Level Results
6.2. Exam Level Results
7. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- European-Council. Relating Language Examinations to the CEFR. Highlights from the Manual. 2011. Available online: https://www.ecml.at/Portals/1/documents/ECML-resources/2011_10_10_relex._E_web.pdf?ver=2018-03-21-100940-823 (accessed on 9 December 2025).
- De Rodrigo, I.; Sanchez-Cuadrado, A.; Boal, J.; Lopez-Lopez, A.J. The MERIT Dataset: Modelling and Efficiently Rendering Interpretable Transcripts. Pattern Recognit. 2025, 172, 112502. [Google Scholar] [CrossRef]
- Alawwad, H.A.; Alhothali, A.; Naseem, U.; Alkhathlan, A.; Jamal, A. Enhancing textual textbook question answering with large language models and retrieval augmented generation. Pattern Recognit. 2025, 162, 111332. [Google Scholar] [CrossRef]
- Wang, L.; Zheng, K.; Qian, L.; Li, S. A Survey of Extractive Question Answering. In Proceedings of the 2022 International Conference on High Performance Big Data and Intelligent Systems (HDIS), Tianjin, China, 10–11 December 2022; pp. 147–153. [Google Scholar] [CrossRef]
- Aydin, B.I.; Yilmaz, Y.S.; Li, Y.; Li, Q.; Gao, J.; Demirbas, M. Crowdsourcing for multiple-choice question answering. In Proceedings of the Twenty-Eighth AAAI Conference on Artificial Intelligence, Québec City, QC, Canada, 27–31 July 2014; AAAI Press: Palo Alto, CA, USA, 2014; pp. 2946–2953. [Google Scholar]
- Chen, A.; Stanovsky, G.; Singh, S.; Gardner, M. Evaluating Question Answering Evaluation. In Proceedings of the 2nd Workshop on Machine Reading for Question Answering, Hong Kong, China, 4 November 2019; pp. 119–124. [Google Scholar] [CrossRef]
- Voorhees, E.M.; Tice, D.M. Building a Question Answering Test Collection. In Proceedings of the 23rd Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, New York, NY, USA, 24–28 July 2000; pp. 200–207. [Google Scholar] [CrossRef]
- Nguyen, T.; Rosenberg, M.; Song, X.; Gao, J.; Tiwary, S.; Majumder, R.; Deng, L. MS MARCO: A Human Generated MAchine Reading COmprehension Dataset. In Proceedings of the Workshop on Cognitive Computation: Integrating Neural and Symbolic Approaches 2016 Co-Located with the 30th Annual Conference on Neural Information Processing Systems (NIPS 2016), Barcelona, Spain, 9 December 2016; Volume 1773, p. 112278. [Google Scholar]
- Joshi, M.; Choi, E.; Weld, D.S.; Zettlemoyer, L. TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Vancouver, BC, Canada, 30 July–4 August 2017; pp. 1601–1611. [Google Scholar]
- Trischler, A.; Wang, T.; Yuan, X.; Harris, J.; Sordoni, A.; Bachman, P.; Suleman, K. NewsQA: A Machine Comprehension Dataset. In Proceedings of the 2nd Workshop on Representation Learning for NLP, Vancouver, BC, Canada, 3 August 2017; pp. 191–200. [Google Scholar] [CrossRef]
- Rajpurkar, P.; Jia, R.; Liang, P. Know What You Don’t Know: Unanswerable Questions for SQuAD. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (ACL), Melbourne, Australia, 15–20 July 2018; Volume 2, pp. 784–789. [Google Scholar]
- Gururangan, S.; Swayamdipta, S.; Levy, O.; Schwartz, R.; Bowman, S.; Smith, N.A. Annotation Artifacts in Natural Language Inference Data. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers), New Orleans, LA, USA, 1–6 June 2018; pp. 107–112. [Google Scholar] [CrossRef]
- Kočiský, T.; Schwarz, J.; Blunsom, P.; Dyer, C.; Hermann, K.M.; Melis, G.; Grefenstette, E. The NarrativeQA Reading Comprehension Challenge. Trans. Assoc. Comput. Linguist. 2018, 6, 317–328. [Google Scholar] [CrossRef]
- Rogers, A.; Kovaleva, O.; Downey, M.; Rumshisky, A. Getting Closer to AI Complete Question Answering: A Set of Prerequisite Real Tasks. In Proceedings of the AAAI Conference on Artificial Intelligence, Atlanta, GA, USA, 11–15 July 2020; Volume 34, pp. 8722–8731. [Google Scholar] [CrossRef]
- Dua, D.; Wang, Y.; Dasigi, P.; Stanovsky, G.; Singh, S.; Gardner, M. DROP: A Reading Comprehension Benchmark Requiring Discrete Reasoning Over Paragraphs. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), Minneapolis, MN, USA, 4 June 2019; pp. 2368–2378. [Google Scholar] [CrossRef]
- Lai, G.; Xie, Q.; Liu, H.; Yang, Y.; Hovy, E. RACE: Large-scale ReAding Comprehension Dataset From Examinations. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, Copenhagen, Denmark, 9–11 September 2017; pp. 785–794. [Google Scholar] [CrossRef]
- Peñas, A.; Rodrigo, Á.; Magnini, B.; Forner, P.; Hovy, E.H.; Sutcliffe, R.F.E.; Giampiccolo, D. Results and Lessons of the Question Answering Track at CLEF. In Information Retrieval Evaluation in a Changing World-Lessons Learned from 20 Years of CLEF; The Information Retrieval Series; Springer: Berlin/Heidelberg, Germany, 2019; Volume 41, pp. 441–460. [Google Scholar] [CrossRef]
- Rodrigo, Á.; Peñas, A.; Miyao, Y.; Kando, N. Do systems pass university entrance exams? Inf. Process. Manag. 2018, 54, 564–575. [Google Scholar] [CrossRef]
- Chandrasekaran, D.; Mago, V. Evolution of Semantic Similarity—A Survey. ACM Comput. Surv. 2021, 54, 1–37. [Google Scholar] [CrossRef]
- Jensen, K.; Elbro, C. Clozing in on reading comprehension: A deep cloze test of global inference making. Read. Writ. Interdiscip. J. 2022, 35, 1221–1237. [Google Scholar] [CrossRef]
- Pradeesh, N.; Remya, T.; MG, T.; K Arun, K.; Pranav, V. Retrieval-Augmented Generation for Multiple-Choice Questions and Answers Generation. Procedia Comput. Sci. 2025, 259, 504–511. [Google Scholar] [CrossRef]
- Saadaoui, S.; Alonso, E. Coordinated LLM multi-agent systems for collaborative question-answer generation. Knowl.-Based Syst. 2025, 330, 114627. [Google Scholar] [CrossRef]
- Lewis, P.; Oguz, B.; Rinott, R.; Riedel, S.; Schwenk, H. MLQA: Evaluating Cross-lingual Extractive Question Answering. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Online, 5–10 July 2020; pp. 7315–7330. [Google Scholar] [CrossRef]
- Artetxe, M.; Ruder, S.; Yogatama, D. On the Cross-lingual Transferability of Monolingual Representations. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Online, 5–10 July 2020; pp. 4623–4637. [Google Scholar] [CrossRef]
- Carrino, C.P.; Costa-jussà, M.R.; Fonollosa, J.A.R. Automatic Spanish Translation of SQuAD Dataset for Multi-lingual Question Answering. In Proceedings of the Twelfth Language Resources and Evaluation Conference, Marseille, France, 11–16 May 2020; pp. 5515–5523. [Google Scholar]
- Rosá, A.; Chiruzzo, L.; Bouza, L.; Dragonetti, A.; Castro, S.; Etcheverry, M.; Góngora, S.; Goycoechea, S.; Machado, J.; Moncecchi, G.; et al. Overview of QuALES at IberLEF 2022: Question Answering Learning from Examples in Spanish. Proces. Leng. Nat. 2022, 69, 273–280. [Google Scholar]
- Gutiérrez-Fandiño, A.; Armengol-Estapé, J.; Pàmies, M.; Llop-Palao, J.; Silveira-Ocampo, J.; Carrino, C.P.; Armentano-Oller, C.; Penagos, C.R.; Gonzalez-Agirre, A.; Villegas, M. MarIA: Spanish Language Models. Proces. Leng. Nat. 2022, 68, 39–60. [Google Scholar]
- Taulé, M.; Martí, M.A.; Recasens, M. AnCora: Multilevel Annotated Corpora for Catalan and Spanish. In Proceedings of the Sixth International Conference on Language Resources and Evaluation (LREC’08), Marrakech, Morocco, 26 May–1 June 2008; pp. 96–101. [Google Scholar]
- Cervantes, I. Plan Curricular del Instituto Cervantes. Niveles de Referencia Para el Español. 2006. Available online: https://cvc.cervantes.es/ensenanza/biblioteca_ele/plan_curricular/ (accessed on 9 December 2025).
- European-Council. Marco Común Europeo de Referencia Para las Lenguas: Enseñanza, Aprendizaje y Evaluación. 2001. Available online: https://cvc.cervantes.es/ensenanza/biblioteca_ele/marco/cvc_mer.pdf (accessed on 9 December 2025).
- López-Anguita, R.; Collado-Montañez, J.; Montejo-Ráez, A. The Text Complexity Library. Proces. Leng. Nat. 2020, 65, 127–130. [Google Scholar]
- Law, G. Error in the Fernandez Huerta Readability Formula. 2011. Available online: https://linguistlist.org/issues/22/2332/ (accessed on 9 December 2025).
- Anula, A. Lecturas adaptadas a la enseñanza del español como L2: Variables lingüísticas para la determinación del nivel de legibilidad. Dialnet 2008, 2, 162–170. [Google Scholar]
- Fernández Huerta, J. Medidas sencillas de lecturabilidad. Consigna 1959, 214, 29–32. [Google Scholar]
- Flesch, R. A new readability yardstick. J. Appl. Psychol. 1948, 32, 221. [Google Scholar] [CrossRef] [PubMed]
- Pazos, F.S. Sistemas Predictivos de Legilibilidad del Mensaje Escrito: Fórmula de Perspicuidad. Ph.D. Thesis, Universidad Complutense de Madrid, Madrid, Spain, 1993. [Google Scholar]
- Saggion, H.; Štajner, S.; Bott, S.; Mille, S.; Rello, L.; Drndarevic, B. Making it simplext: Implementation and evaluation of a text simplification system for spanish. ACM Trans. Access. Comput. TACCESS 2015, 6, 1–36. [Google Scholar] [CrossRef]
- García Lopez, J.A. Legibilidad de los folletos informativos. Pharm. Care Esp. 2001, 1, 49–56. [Google Scholar]
- Akyol, P.; Key, J.; Krishna, K. Hit or Miss? Test Taking Behavior in Multiple Choice Exams; NBER Working Papers 22401; National Bureau of Economic Research, Inc.: Cambridge, MA, USA, 2016. [Google Scholar]
- Valadares, J. Correcting scores of tests taking into account the guessing factor: Yes or no. Int. J. Cross-Discip. Subj. Educ. 2011, 2, 381–387. [Google Scholar] [CrossRef]
- Liang, Y.; Li, J.; Yin, J. A New Multi-choice Reading Comprehension Dataset for Curriculum Learning. In Proceedings of the Eleventh Asian Conference on Machine Learning, PMLR, Nagoya, Japan, 17–19 November 2019; Volume 101, pp. 742–757. [Google Scholar]
- Pal, A.; Umapathi, L.K.; Sankarasubbu, M. MedMCQA: A Large-scale Multi-Subject Multi-Choice Dataset for Medical domain Question Answering. In Proceedings of the Conference on Health, Inference, and Learning, PMLR, Virtually, 7–8 April 2022; Volume 174, pp. 248–260. [Google Scholar]


| Exercise Type | Number of Responses |
|---|---|
| Multiple-choice | 3544 |
| Matching | 2309 |
| Fill-in-the-gap | 293 |
| Total | 6146 |
| Level | Exercises | Questions | Options | Probabilities | Tokens | ||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Mean | Median | Min | Max | Instructions | Text | Question | Options | ||||
| A1 | 39 | 164 | 632 | 0.262 | 0.250 | 0.250 | 0.333 | 42.59 | 199.641 | 6.295 | 3.936 |
| A1E | 18 | 108 | 324 | 0.333 | 0.333 | 0.333 | 0.333 | 46.056 | 204.333 | 5.701 | 3.405 |
| A2 | 490 | 976 | 2928 | 0.333 | 0.333 | 0.333 | 0.333 | 23.324 | 132.406 | 6.791 | 5.944 |
| A2B1E | 8 | 48 | 144 | 0.333 | 0.333 | 0.333 | 0.333 | 37.125 | 479.125 | 6.854 | 7.528 |
| B1 | 422 | 1002 | 2867 | 0.356 | 0.333 | 0.333 | 0.500 | 19.654 | 134.962 | 9.251 | 3.099 |
| B1E | 81 | 180 | 512 | 0.359 | 0.333 | 0.333 | 0.500 | 18.951 | 121.062 | 9.642 | 3.116 |
| B2 | 194 | 612 | 1704 | 0.370 | 0.333 | 0.250 | 0.500 | 18.954 | 407.072 | 11.421 | 4.473 |
| C1 | 14 | 84 | 252 | 0.333 | 0.333 | 0.333 | 0.333 | 31.071 | 694.929 | 9.690 | 7.992 |
| C2 | 108 | 324 | 972 | 0.333 | 0.333 | 0.333 | 0.333 | 18.000 | 664.935 | 10.151 | 6.907 |
| Level | Lexical Complexity | Complexity of Sentences | Fernández-Huerta Readability | IFSZ Readability | Mean Dep. Tree Depth | Min. Age to Understand |
|---|---|---|---|---|---|---|
| A1 | 2.402 | 9.571 | 79.977 | 75.782 | 5.996 | 8.129 |
| A1E | 3.740 | 19.844 | 73.050 | 69.227 | 5.667 | 10.381 |
| A2 | 4.056 | 14.655 | 59.888 | 55.264 | 6.716 | 11.103 |
| A2B1E | 4.105 | 18.661 | 61.559 | 57.206 | 7.292 | 11.409 |
| B1 | 4.459 | 15.946 | 65.076 | 60.710 | 6.496 | 10.680 |
| B1E | 4.885 | 18.759 | 65.580 | 61.383 | 6.651 | 10.980 |
| B2 | 6.267 | 28.480 | 46.080 | 41.716 | 8.015 | 14.454 |
| C1 | 5.314 | 26.632 | 39.483 | 34.756 | 8.226 | 14.907 |
| C2 | 5.893 | 28.257 | 46.880 | 42.518 | 7.980 | 14.299 |
| Level | Lexical Complexity | Complexity of Sentences | Fernández-Huerta Readability | IFSZ Readability | Mean Dep. Tree Depth | Min. Age to Understand |
|---|---|---|---|---|---|---|
| A1 | 1.896 | 3.917 | 90.632 | 86.546 | 5.183 | 6.271 |
| A1E | 1.695 | 3.642 | 87.212 | 82.943 | 4.986 | 6.518 |
| A2 | 1.772 | 4.052 | 80.938 | 76.491 | 5.057 | 7.344 |
| A2B1E | 1.703 | 3.406 | 74.193 | 69.484 | 4.938 | 8.063 |
| B1 | 2.505 | 7.047 | 68.862 | 64.097 | 5.813 | 8.989 |
| B1E | 2.609 | 7.389 | 71.964 | 67.349 | 5.894 | 8.728 |
| B2 | 2.835 | 8.263 | 76.908 | 72.603 | 6.165 | 8.477 |
| C1 | 2.505 | 7.125 | 63.286 | 58.343 | 5.940 | 9.674 |
| C2 | 2.482 | 7.210 | 80.737 | 76.507 | 5.954 | 7.896 |
| Level | Lexical Complexity | Complexity of Sentences | Fernández-Huerta Readability | IFSZ Readability | Mean Dep. Tree Depth | Min. Age to Understand |
|---|---|---|---|---|---|---|
| A1 | 1.278 | 2.918 | 80.947 | 76.354 | 4.479 | 6.997 |
| A1E | 1.119 | 2.428 | 77.004 | 72.231 | 4.058 | 7.353 |
| A2 | 1.834 | 4.799 | 70.951 | 66.087 | 5.380 | 8.343 |
| A2B1E | 2.298 | 6.074 | 68.918 | 64.073 | 6.004 | 8.789 |
| B1 | 1.323 | 2.156 | 46.050 | 40.066 | 4.049 | 10.639 |
| B1E | 1.300 | 2.081 | 44.576 | 38.533 | 3.997 | 10.793 |
| B2 | 1.661 | 3.564 | 50.189 | 44.450 | 4.744 | 10.395 |
| C1 | 2.362 | 6.050 | 54.262 | 48.886 | 5.924 | 10.445 |
| C2 | 2.182 | 5.788 | 60.945 | 55.766 | 5.816 | 9.584 |
| Level | Exercises | Questions | Answers | Probabilities | Tokens | |||||
|---|---|---|---|---|---|---|---|---|---|---|
| Mean | Median | Min | Max | Instructions | Questions | Answers | ||||
| A1 | 37 | 222 | 333 | 0.111 | 0.111 | 0.111 | 0.111 | 37.730 | 5.982 | 24.444 |
| A1E | 11 | 66 | 99 | 0.111 | 0.111 | 0.111 | 0.111 | 50.455 | 6.788 | 28.818 |
| A2 | 104 | 676 | 987 | 0.105 | 0.100 | 0.100 | 0.111 | 44.971 | 8.058 | 43.033 |
| A2B1E | 8 | 48 | 72 | 0.111 | 0.111 | 0.111 | 0.111 | 50.250 | 29.646 | 44.569 |
| B1 | 17 | 102 | 153 | 0.111 | 0.111 | 0.111 | 0.111 | 53.529 | 22.971 | 53.843 |
| B2 | 10 | 100 | 40 | 0.250 | 0.250 | 0.250 | 0.250 | 41.700 | 11.210 | 148.975 |
| C1 | 7 | 56 | 42 | 0.167 | 0.167 | 0.167 | 0.167 | 41.571 | 15.107 | 135.381 |
| C2 | 40 | 392 | 384 | 0.105 | 0.100 | 0.100 | 0.167 | 26.600 | 16.339 | 52.139 |
| Level | Lexical Complexity | Complexity of Sentences | Readability of Fernández-Huerta | IFSZ Readability | Mean Dep. Tree Depth | Min. Age to Understand |
|---|---|---|---|---|---|---|
| A1 | 1.923 | 4.62 | 82.788 | 78.385 | 5.592 | 7.082 |
| A1E | 2.132 | 5.398 | 79.308 | 74.795 | 5.515 | 7.511 |
| A2 | 2.416 | 6.557 | 75.496 | 70.941 | 6.285 | 8.168 |
| A2B1E | 3.602 | 11.679 | 67.804 | 63.248 | 6.946 | 9.691 |
| B1 | 4.027 | 12.829 | 64.874 | 60.276 | 7.125 | 10.173 |
| B2 | 2.854 | 9.48 | 63.489 | 58.661 | 7.63 | 9.905 |
| C1 | 4.197 | 13.723 | 48.514 | 43.344 | 7.536 | 12.071 |
| C2 | 3.385 | 11.057 | 66.365 | 61.746 | 6.871 | 9.827 |
| Level | Lexical Complexity | Complexity of Sentences | Fernández-Huerta Readability | IFSZ Readability | Mean Dep. Tree Depth | Min. Age to Understand |
|---|---|---|---|---|---|---|
| A1 | 2.642 | 7.456 | 76.427 | 72.012 | 5.385 | 8.314 |
| A1E | 2.425 | 6.781 | 79.022 | 74.649 | 5.432 | 7.898 |
| A2 | 3.555 | 11.645 | 62.239 | 57.494 | 6.015 | 10.351 |
| A2B1E | 3.121 | 10.064 | 66.68 | 61.981 | 5.89 | 9.578 |
| B1 | 4.57 | 14.937 | 58.131 | 53.456 | 7.04 | 11.328 |
| B2 | 5.246 | 21.353 | 56.031 | 51.61 | 7.968 | 12.347 |
| C1 | 6.077 | 23.105 | 42.022 | 37.194 | 7.51 | 14.165 |
| C2 | 4.59 | 18.175 | 61.255 | 56.855 | 7.223 | 11.36 |
| Level | Exercises | Gaps | Options | Probabilities | Tokens | |||||
|---|---|---|---|---|---|---|---|---|---|---|
| Mean | Median | Min | Max | Instructions | Text | Options | ||||
| B1 | 17 | 102 | 136 | 0.125 | 0.125 | 0.125 | 0.125 | 53.000 | 318.941 | 18.919 |
| B2 | 10 | 60 | 80 | 0.125 | 0.125 | 0.125 | 0.125 | 51.200 | 319.200 | 19.425 |
| C1 | 7 | 42 | 49 | 0.143 | 0.143 | 0.143 | 0.143 | 51.571 | 420.000 | 40.224 |
| C2 | 4 | 24 | 28 | 0.143 | 0.143 | 0.143 | 0.143 | 51.000 | 439.750 | 41.500 |
| Level | Lexical Complexity | Complexity of Sentences | Fernández-Huerta Readability | IFSZ Readability | Mean Dep. Tree Depth | Min. Age to Understand |
|---|---|---|---|---|---|---|
| B1 | 4.53 | 19.126 | 56.481 | 51.971 | 7.875 | 12.046 |
| B2 | 4.762 | 20.004 | 54.775 | 50.258 | 7.873 | 12.369 |
| C1 | 5.551 | 24.099 | 45.889 | 41.294 | 7.819 | 13.945 |
| C2 | 6.028 | 26.58 | 42.074 | 37.443 | 8.403 | 14.62 |
| Level | Lexical Complexity | Complexity of Sentences | Fernández-Huerta Readability | IFSZ Readability | Mean Dep. Tree Depth | Min. Age to Understand |
|---|---|---|---|---|---|---|
| B1 | 4.714 | 16.474 | 58.093 | 53.471 | 7.699 | 11.463 |
| B2 | 4.846 | 17.334 | 56.875 | 52.264 | 7.938 | 11.73 |
| C1 | 6.877 | 26.367 | 47.238 | 42.818 | 7.966 | 14.093 |
| C2 | 7.968 | 29.411 | 37.874 | 33.313 | 8.661 | 15.619 |
| Task | Metric | Level | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| A1 | A1E | A2 | A2B1E | B1 | B1E | B2 | C1 | C2 | ||
| Multiple-choice | Exam min questions | 4.00 | 4.00 | 5.00 | 6.00 | 3.00 | 3.00 | 6.00 | 6.00 | 9.00 |
| Exam max questions | 8.00 | 8.00 | 9.00 | 6.00 | 10.00 | 10.00 | 12.00 | 6.00 | 9.00 | |
| Exam avg questions | 4.20 | 6.00 | 6.02 | 6.00 | 6.59 | 6.67 | 10.92 | 6.00 | 9.00 | |
| Accuracy (%) | 71.95 | 63.89 | 64.04 | 68.75 | 63.07 | 73.89 | 68.63 | 58.33 | 56.17 | |
| Matching | Exam min questions | 6.00 | 6.00 | 6.00 | 6.00 | 6.00 | - | 10.00 | 8.00 | 8.00 |
| Exam max questions | 6.00 | 6.00 | 7.00 | 6.00 | 6.00 | - | 10.00 | 8.00 | 10.00 | |
| Exam avg questions | 6.00 | 6.00 | 6.50 | 6.00 | 6.00 | - | 10.00 | 8.00 | 9.80 | |
| Accuracy (%) | 57.21 | 50.00 | 39.64 | 56.25 | 33.33 | - | 35.00 | 32.14 | 44.39 | |
| Fill-in-the-gap | Exam min questions | - | - | - | - | 5 | - | 5 | 5 | 6 |
| Exam max questions | - | - | - | - | 6 | - | 6 | 6 | 6 | |
| Exam avg questions | - | - | - | - | 5.94 | - | 5.9 | 6.0 | 5.25 | |
| Accuracy (%) | - | - | - | - | 32.67 | - | 35.59 | 40.48 | 28.57 | |
| Metric | Level | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| A1 | A1E | A2 | A2B1E | B1 | B1E | B2 | C1 | C2 | |
| Exams | 38 | 14 | 158 | 8 | 131 | 19 | 64 | 17 | 56 |
| Passed exams | 21 | 8 | 75 | 5 | 73 | 17 | 35 | 4 | 27 |
| Accuracy (%) | 55.26 | 57.14 | 47.47 | 62.50 | 55.73 | 89.47 | 54.59 | 23.53 | 48.21 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Peñas, A.; Rodrigo, Á.; Fruns-Jiménez, J.; Soria-Pastor, I.; Moreno-Álvarez, S.; Pérez, A.; Reyes-Montesinos, J. A Spanish Language Proficiency Dataset for AI Evaluation. Information 2026, 17, 159. https://doi.org/10.3390/info17020159
Peñas A, Rodrigo Á, Fruns-Jiménez J, Soria-Pastor I, Moreno-Álvarez S, Pérez A, Reyes-Montesinos J. A Spanish Language Proficiency Dataset for AI Evaluation. Information. 2026; 17(2):159. https://doi.org/10.3390/info17020159
Chicago/Turabian StylePeñas, Anselmo, Álvaro Rodrigo, Javier Fruns-Jiménez, Inés Soria-Pastor, Sergio Moreno-Álvarez, Alberto Pérez, and Julio Reyes-Montesinos. 2026. "A Spanish Language Proficiency Dataset for AI Evaluation" Information 17, no. 2: 159. https://doi.org/10.3390/info17020159
APA StylePeñas, A., Rodrigo, Á., Fruns-Jiménez, J., Soria-Pastor, I., Moreno-Álvarez, S., Pérez, A., & Reyes-Montesinos, J. (2026). A Spanish Language Proficiency Dataset for AI Evaluation. Information, 17(2), 159. https://doi.org/10.3390/info17020159

