Extending a Vision–Language Model with Audio Understanding: Introducing Qolda-AVL for the Kazakh Language
Abstract
1. Introduction
- Qolda-AVL, the first open-source tri-modal (audio, vision, language) model designed specifically for Kazakh, released together with model weights, training code (https://github.com/IS2AI/ms-swift-Qolda-AVL (accessed on 1 June 2026)), and inference code (https://huggingface.co/issai/Qolda-AVL-5B (accessed on 30 April 2026)).
- Integration of Audio DeepStack into a vision–language backbone. We adapt the hierarchical multi-level injection mechanism, originally proposed for vision in DeepStack [15] and Qwen3-VL [14], to the audio modality, routing features from three intermediate Whisper encoder layers into the early LLM decoder layers.
- A staged training pipeline for integrating an audio branch into a pretrained vision–language model, together with a synthetic chain-of-thought (CoT) generation methodology that constructs audio reasoning data for Kazakh through two complementary approaches: direct generation with a reasoning audio model, and refinement of intermediate-checkpoint outputs by a stronger external LLM.
- A publicly released Kazakh audio understanding benchmark suite, covering spoken attribute reasoning, spoken mathematical question answering, and audio captioning with QA. The suite is the first multi-task audio benchmark collection targeted at Kazakh and is released as a Hugging Face collection (https://huggingface.co/collections/issai/qolda-avl-audio-benchmarks (accessed on 29 April 2026)).
2. Related Work
2.1. Vision–Language Models
2.2. Audio–Language Models
2.3. Omni-Modal Systems
2.4. Language and Multi-Modal Resources for Kazakh
3. Methodology
3.1. Model Architecture
3.1.1. Vision Branch
3.1.2. Audio Branch
3.1.3. Multi-Modal Integration
3.2. Training Pipeline
3.2.1. Stage 1.1: Whisper Fine-Tuning
3.2.2. Stage 1.2: LLM Fine-Tuning
3.2.3. Model Assembly
3.2.4. Stage 2: Audio Alignment
- Automatic Speech Recognition (ASR): The model receives a speech input and is asked to produce the corresponding text transcript. The ASR mixture is predominantly Kazakh (∼3000 h), complemented by English (∼1000 h) and Russian (∼600 h), with a smaller residual portion (under 200 h) covering other languages.
- Speech-to-Text Translation (S2TT): The model is given speech in a source language and is asked to generate a written translation in a target language. Training covers directional pairs among English, Russian, Kazakh, and Turkish, thereby exposing the model to cross-lingual understanding of spoken input.
- Language Identification (LangID): Given a short speech clip, the model is asked to predict the spoken language. The training set spans 17 languages drawn from a curated subset of VoxLingua107 [46], with a deliberate emphasis on the Turkic family (Kazakh, Azerbaijani, Turkish, Turkmen, Tatar, Uzbek), alongside a selection of major world languages (Arabic, German, English, Spanish, Persian, French, Italian, Japanese, Korean, Russian, Chinese).
- Audio Question Answering (audio QA): Given an audio clip and a natural-language question, the model is asked to produce a short factual answer. The data is drawn from the MECAT dataset [47], whose questions probe different aspects of the provided audio, including direct perception of events, sound characteristics, acoustic environment, and contextual reasoning. The text annotations are in English.
- Audio Captioning: The model is asked to produce a natural-language description of an audio clip. The data is again based on MECAT, and the audio clips are shared with the audio QA subset. Six prompt variants target complementary aspects of the recording: a long-form description of the overall content, a short one-sentence summary, the acoustic environment, the speech content, the musical content, and the discrete sound events. This ensures that a single clip can be described from multiple perspectives.
- Gender and Emotion Recognition: The model is asked to classify the speaker’s gender and/or expressed emotion from a single utterance. The data for this task is entirely in English.
- Audio Classification: The model is asked to assign a categorical label to a non-speech audio clip, for example, in the classification of animal sounds.
3.2.5. Stage 3: Multi-Modal Fine-Tuning
3.2.6. Training Configuration
3.3. Evaluation
3.3.1. Language Evaluation
3.3.2. Vision Evaluation
3.3.3. Audio Evaluation
3.3.4. Model Selection and Comparison
4. Results
4.1. Language Results
4.2. Vision Results
4.3. Audio Results
4.4. Ablation Study
5. Discussion
6. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Liu, H.; Li, C.; Wu, Q.; Lee, Y.J. Visual Instruction Tuning. In Proceedings of the Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2023; Volume 36, pp. 34892–34916. [Google Scholar]
- Chen, Z.; Wu, J.; Wang, W.; Su, W.; Chen, G.; Xing, S.; Zhong, M.; Zhang, Q.; Zhu, X.; Lu, L.; et al. InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 17–21 June 2024; pp. 24185–24198. [Google Scholar]
- Deitke, M.; Clark, C.; Lee, S.; Tripathi, R.; Yang, Y.; Park, J.S.; Salehi, M.; Muennighoff, N.; Lo, K.; Soldaini, L.; et al. Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA, 11–15 June 2025; pp. 91–104. [Google Scholar]
- Bai, J.; Bai, S.; Yang, S.; Wang, S.; Tan, S.; Wang, P.; Lin, J.; Zhou, C.; Zhou, J. Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond. arXiv 2023, arXiv:2308.12966. [Google Scholar]
- OpenAI. GPT-4o System Card. arXiv 2024, arXiv:2410.21276. [Google Scholar] [CrossRef]
- Xu, J.; Guo, Z.; He, J.; Hu, H.; He, T.; Bai, S.; Chen, K.; Wang, J.; Fan, Y.; Dang, K.; et al. Qwen2.5-Omni Technical Report. arXiv 2025, arXiv:2503.20215. [Google Scholar] [CrossRef]
- Xu, J.; Guo, Z.; Hu, H.; Chu, Y.; Wang, X.; He, J.; Wang, Y.; Shi, X.; He, T.; Zhu, X.; et al. Qwen3-Omni Technical Report. arXiv 2025, arXiv:2509.17765. [Google Scholar] [CrossRef]
- Radford, A.; Kim, J.W.; Xu, T.; Brockman, G.; Mcleavey, C.; Sutskever, I. Robust Speech Recognition via Large-Scale Weak Supervision. In Proceedings of the 40th International Conference on Machine Learning, 23–29 July 2023; Proceedings of Machine Learning Research; PMLR: New York, NY, USA, 2023; Volume 202, pp. 28492–28518. Available online: https://proceedings.mlr.press/v202/radford23a.html (accessed on 30 April 2026).
- Yang, Z.; Shimizu, S.; Yu, Y.; Chu, C. When Large Language Models Meet Speech: A Survey on Integration Approaches. In Proceedings of the Findings of the Association for Computational Linguistics: ACL 2025, Vienna, Austria, 27 July–1 August 2025; pp. 20298–20315. [Google Scholar] [CrossRef]
- Chu, Y.; Xu, J.; Zhou, X.; Yang, Q.; Zhang, S.; Yan, Z.; Zhou, C.; Zhou, J. Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models. arXiv 2023, arXiv:2311.07919. [Google Scholar]
- Tang, C.; Yu, W.; Sun, G.; Chen, X.; Tan, T.; Li, W.; Lu, L.; MA, Z.; Zhang, C. SALMONN: Towards Generic Hearing Abilities for Large Language Models. In Proceedings of the Twelfth International Conference on Learning Representations, Vienna, Austria, 7–11 May 2024. [Google Scholar]
- Veitsman, Y.; Hartmann, M. Recent Advancements and Challenges of Turkic Central Asian Language Processing. In Proceedings of the First Workshop on Language Models for Low-Resource Languages, Abu Dhabi, United Arab Emirates, 20 January 2025; pp. 309–324. [Google Scholar]
- Arystanbekov, B.; Nurimanov, A.; Maxutov, A.; Albrekht, V.; Kuzdeuov, A.; Varol, H.A. Qolda: A Small Vision-Language Model for the Kazakh Language. IEEE Access 2026, 14, 46392–46414. [Google Scholar] [CrossRef]
- Bai, S.; Cai, Y.; Chen, R.; Chen, K.; Chen, X.; Cheng, Z.; Deng, L.; Ding, W.; Gao, C.; Ge, C.; et al. Qwen3-VL Technical Report. arXiv 2025, arXiv:2511.21631. [Google Scholar] [CrossRef]
- Meng, L.; Yang, J.; Tian, R.; Dai, X.; Wu, Z.; Gao, J.; Jiang, Y.G. DeepStack: Deeply stacking visual tokens is surprisingly simple and effective for LMMs. In Proceedings of the 38th International Conference on Neural Information Processing Systems, Vancouver, BC, Canada, 10–15 December 2024. NIPS ’24. [Google Scholar]
- Radford, A.; Kim, J.W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. Learning Transferable Visual Models From Natural Language Supervision. In Proceedings of the 38th International Conference on Machine Learning, 18–24 July 2021; Proceedings of Machine Learning Research; PMLR: New York, NY, USA, 2021; Volume 139, pp. 8748–8763. [Google Scholar]
- Wang, W.; Lv, Q.; Yu, W.; Hong, W.; Qi, J.; Wang, Y.; Ji, J.; Yang, Z.; Zhao, L.; Song, X.; et al. CogVLM: Visual expert for pretrained language models. In Proceedings of the 38th International Conference on Neural Information Processing Systems NIPS ’24, Vancouver, BC, Canada, 10–15 December 2024. [Google Scholar]
- Lu, H.; Liu, W.; Zhang, B.; Wang, B.; Dong, K.; Liu, B.; Sun, J.; Ren, T.; Li, Z.; Yang, H.; et al. DeepSeek-VL: Towards Real-World Vision-Language Understanding. arXiv 2024, arXiv:2403.05525. [Google Scholar]
- Wang, P.; Bai, S.; Tan, S.; Wang, S.; Fan, Z.; Bai, J.; Chen, K.; Liu, X.; Wang, J.; Ge, W.; et al. Qwen2-VL: Enhancing Vision-Language Model’s Perception of the World at Any Resolution. arXiv 2024, arXiv:2409.12191. [Google Scholar]
- Bai, S.; Chen, K.; Liu, X.; Wang, J.; Ge, W.; Song, S.; Dang, K.; Wang, P.; Wang, S.; Tang, J.; et al. Qwen2.5-VL Technical Report. arXiv 2025, arXiv:2502.13923. [Google Scholar] [CrossRef]
- Baevski, A.; Zhou, H.; Mohamed, A.; Auli, M. wav2vec 2.0: A framework for self-supervised learning of speech representations. In Proceedings of the 34th International Conference on Neural Information Processing Systems NIPS ’20, Red Hook, NY, USA, 6–12 December 2020. [Google Scholar]
- Hsu, W.N.; Bolte, B.; Tsai, Y.H.H.; Lakhotia, K.; Salakhutdinov, R.; Mohamed, A. HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units. IEEE/ACM Trans. Audio Speech Lang. Proc. 2021, 29, 3451–3460. [Google Scholar] [CrossRef]
- Yang, S.W.; Chi, P.H.; Chuang, Y.S.; Lai, C.I.J.; Lakhotia, K.; Lin, Y.Y.; Liu, A.T.; Shi, J.; Chang, X.; Lin, G.T.; et al. SUPERB: Speech Processing Universal PERformance Benchmark. In Proceedings of the Interspeech 2021, Brno, Czech Republic, 30 August–3 September 2021; pp. 1194–1198. [Google Scholar] [CrossRef]
- Chen, S.; Wu, Y.; Wang, C.; Liu, S.; Tompkins, D.; Chen, Z.; Che, W.; Yu, X.; Wei, F. BEATs: Audio pre-training with acoustic tokenizers. In Proceedings of the 40th International Conference on Machine Learning ICML’23, Honolulu, HI, USA, 23–29 July 2023. [Google Scholar]
- Elizalde, B.; Deshmukh, S.; Ismail, M.A.; Wang, H. CLAP Learning Audio Concepts from Natural Language Supervision. In Proceedings of the ICASSP 2023—2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Rhodes Island, Greece, 4–10 June 2023; pp. 1–5. [Google Scholar] [CrossRef]
- Wu, Y.; Chen, K.; Zhang, T.; Hui, Y.; Berg-Kirkpatrick, T.; Dubnov, S. Large-scale Contrastive Language-Audio Pretraining with Feature Fusion and Keyword-to-Caption Augmentation. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP, Rhodes Island, Greece, 4–10 June 2023. [Google Scholar]
- Chu, Y.; Xu, J.; Yang, Q.; Wei, H.; Wei, X.; Guo, Z.; Leng, Y.; Lv, Y.; He, J.; Lin, J.; et al. Qwen2-Audio Technical Report. arXiv 2024, arXiv:2407.10759. [Google Scholar] [CrossRef]
- Gong, Y.; Khurana, S.; Karlinsky, L.; Glass, J. Whisper-AT: Noise-Robust Automatic Speech Recognizers are Also Strong General Audio Event Taggers. In Proceedings of the Interspeech 2023, Dublin, Ireland, 20–24 August 2023; pp. 2798–2802. [Google Scholar] [CrossRef]
- Huang, R.; Li, M.; Yang, D.; Shi, J.; Chang, X.; Ye, Z.; Wu, Y.; Hong, Z.; Huang, J.; Liu, J.; et al. AudioGPT: Understanding and generating speech, music, sound, and talking head. In Proceedings of the Thirty-Eighth AAAI Conference on Artificial Intelligence and Thirty-Sixth Conference on Innovative Applications of Artificial Intelligence and Fourteenth Symposium on Educational Advances in Artificial Intelligence; AAAI′24/IAAI′24/EAAI′24; AAAI Press: Washington, DC, USA, 2024. [Google Scholar] [CrossRef]
- Ghosh, S.; Kumar, S.; Seth, A.; Evuru, C.K.R.; Tyagi, U.; Sakshi, S.; Nieto, O.; Duraiswami, R.; Manocha, D. GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Miami, FL, USA, 12–16 November 2024; pp. 6288–6313. Available online: https://aclanthology.org/2024.emnlp-main.361 (accessed on 29 April 2026).
- OpenMOSS Team. MOSS-Audio Technical Report. GitHub Repository. 2026. Available online: https://github.com/OpenMOSS/MOSS-Audio (accessed on 20 April 2026).
- Google DeepMind. Gemini 3 Pro Model Card. 2025. Available online: https://storage.googleapis.com/deepmind-media/Model-Cards/Gemini-3-Pro-Model-Card.pdf (accessed on 6 April 2026).
- Fu, C.; Lin, H.; Long, Z.; Shen, Y.; Dai, Y.; Zhao, M.; Zhang, Y.F.; Dong, S.; Li, Y.; Wang, X.; et al. VITA: Towards Open-Source Interactive Omni Multimodal LLM. arXiv 2025, arXiv:2408.05211. [Google Scholar] [CrossRef]
- Li, Y.; Sun, H.; Lin, M.; Li, T.; Dong, G.; Zhang, T.; Ding, B.; Song, W.; Cheng, Z.; Huo, Y.; et al. Baichuan-Omni Technical Report. arXiv 2024, arXiv:2410.08565. [Google Scholar] [CrossRef]
- Luo, R.; Lin, T.E.; Zhang, H.; Wu, Y.; Liu, X.; Li, Y.; Chen, L.; Li, J.; Zhang, L.; Xia, X.; et al. OpenOmni: Advancing Open-Source Omnimodal Large Language Models with Progressive Multimodal Alignment and Real-time Emotional Speech Synthesis. In Proceedings of the Advances in Neural Information Processing Systems; Belgrave, D., Zhang, C., Lin, H., Pascanu, R., Koniusz, P., Ghassemi, M., Chen, N., Eds.; Curran Associates, Inc.: Red Hook, NY, USA, 2025; Volume 38, pp. 158925–158953. [Google Scholar]
- Liu, Z.; Dong, Y.; Wang, J.; Liu, Z.; Hu, W.; Lu, J.; Rao, Y. Ola: Pushing the Frontiers of Omni-Modal Language Model. arXiv 2025, arXiv:2502.04328. [Google Scholar] [CrossRef]
- Institute of Smart Systems and Artificial Intelligence. Kazakh Large Language Model (ISSAI KAZ-LLM). 2024. Available online: https://huggingface.co/collections/issai/issai-kazllm-10-6732d58c81bcaf177442c362 (accessed on 29 April 2026).
- Koto, F.; Joshi, R.; Mukhituly, N.; Wang, Y.; Xie, Z.; Pal, R.; Orel, D.; Mullah, P.; Turmakhan, D.; Goloburda, M.; et al. Sherkala-Chat: Building a State-of-the-Art LLM for Kazakh in a Moderately Resourced Setting. In Proceedings of the Second Conference on Language Modeling, Montreal, QC, Canada, 7–10 October 2025. [Google Scholar]
- Maxutov, A.; Myrzakhmet, A.; Braslavski, P. Do LLMs Speak Kazakh? A Pilot Evaluation of Seven Models. In Proceedings of the First Workshop on Natural Language Processing for Turkic Languages (SIGTURK 2024), Bangkok, Thailand, 15 August 2024; pp. 81–91. Available online: https://aclanthology.org/2024.sigturk-1.8/ (accessed on 30 April 2026).
- Mussakhojayeva, S.; Khassanov, Y.; Atakan Varol, H. KSC2: An Industrial-Scale Open-Source Kazakh Speech Corpus. In Proceedings of the Interspeech 2022, Incheon, Republic of Korea, 18–22 September 2022; pp. 1367–1371. [Google Scholar] [CrossRef]
- Li, J.; Pu, Y.; Sun, Q.; Zhang, W.Q. Improving Whisper’s Recognition Performance for Under-Represented Language Kazakh Leveraging Unpaired Speech and Text. In Proceedings of the Interspeech 2024, Kos, Greece, 1–5 September 2024; pp. 2514–2518. [Google Scholar] [CrossRef]
- Mussakhojayeva, S.; Dauletbek, K.; Yeshpanov, R.; Varol, H.A. Multilingual Speech Recognition for Turkic Languages. Information 2023, 14, 74. [Google Scholar] [CrossRef]
- Mussakhojayeva, S.; Gilmullin, R.; Khakimov, B.; Galimov, M.; Orel, D.; Abilbekov, A.; Varol, H.A. Noise-Robust Multilingual Speech Recognition and the Tatar Speech Corpus. In Proceedings of the 2024 International Conference on Artificial Intelligence in Information and Communication (ICAIIC), Osaka, Japan, 19–22 February 2024; pp. 732–737. [Google Scholar] [CrossRef]
- Zubitskii, P.; Morozov, V.; Murzakhmetov, S.; Sagyndyk, B.; Umbet, S. HordeVision: An Open-Source Kazakh Vision–Language Model. TechRxiv 2026. [Google Scholar] [CrossRef]
- Ardila, R.; Branson, M.; Davis, K.; Kohler, M.; Meyer, J.; Henretty, M.; Morais, R.; Saunders, L.; Tyers, F.; Weber, G. Common Voice: A Massively-Multilingual Speech Corpus. In Proceedings of the Twelfth Language Resources and Evaluation Conference, Marseille, France, 11–16 May 2020; pp. 4218–4222. Available online: https://aclanthology.org/2020.lrec-1.520/ (accessed on 29 April 2026).
- Valk, J.; Alumäe, T. VOXLINGUA107: A Dataset for Spoken Language Recognition. In 2021 IEEE Spoken Language Technology Workshop (SLT); IEEE: New York, NY, USA, 2020; pp. 652–658. Available online: https://api.semanticscholar.org/CorpusID:227209238 (accessed on 25 April 2026).
- Niu, Y.; Wang, T.; Dinkel, H.; Sun, X.; Zhou, J.; Li, G.; Liu, J.; Liu, X.; Zhang, J.; Luan, J. MECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks. arXiv 2025, arXiv:2507.23511. [Google Scholar]
- Abilbekov, A.; Mussakhojayeva, S.; Yeshpanov, R.; Varol, H.A. KazEmoTTS: A Dataset for Kazakh Emotional Text-to-Speech Synthesis. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), Torino, Italy, 20–25 May 2024; pp. 9626–9632. Available online: https://aclanthology.org/2024.lrec-main.841/ (accessed on 26 April 2026).
- Karpov, N.; Denisenko, A.; Minkin, F. Golos: Russian Dataset for Speech Research. In Proceedings of the Interspeech 2021, Brno, Czech Republic, 30 August–3 September 2021; pp. 1419–1423. [Google Scholar] [CrossRef]
- Yamagishi, J.; Veaux, C.; MacDonald, K. CSTR VCTK Corpus: English Multi-Speaker Corpus for CSTR Voice Cloning Toolkit (Version 0.92); University of Edinburgh, The Centre for Speech Technology Research (CSTR): Edinburgh, UK, 2019. [Google Scholar] [CrossRef]
- Bakhturina, E.; Lavrukhin, V.; Ginsburg, B.; Zhang, Y. Hi-Fi Multi-Speaker English TTS Dataset. In Proceedings of the Interspeech 2021, Brno, Czech Republic, 30 August–3 September 2021; pp. 2776–2780. [Google Scholar] [CrossRef]
- Busso, C.; Bulut, M.; Lee, C.C.; Kazemzadeh, A.; Mower, E.; Kim, S.; Chang, J.; Lee, S.; Narayanan, S.S. IEMOCAP: Interactive emotional dyadic motion capture database. J. Lang. Resour. Eval. 2008, 42, 335–359. [Google Scholar] [CrossRef]
- Piczak, K.J. ESC: Dataset for Environmental Sound Classification. In Proceedings of the 23rd ACM International Conference on Multimedia MM ’15, Brisbane, Australia, 26–30 October 2015; pp. 1015–1018. [Google Scholar] [CrossRef]
- Tian, F.; Zhang, X.T.; Zhang, Y.; Zhang, H.; Li, Y.; Liu, D.; Deng, Y.; Wu, D.; Chen, J.; Zhao, L.; et al. Step-Audio-R1 Technical Report. arXiv 2025, arXiv:2511.15848. [Google Scholar]
- Qwen Team. Qwen3.5: Towards Native Multimodal Agents. 2026. Available online: https://qwen.ai/blog?id=qwen3.5 (accessed on 30 April 2026).
- Liu, H.; Li, C.; Li, Y.; Lee, Y.J. Improved Baselines with Visual Instruction Tuning. In Proceedings of the 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 17–21 June 2024; pp. 26286–26296. [Google Scholar] [CrossRef]
- Zhao, Y.; Huang, J.; Hu, J.; Wang, X.; Mao, Y.; Zhang, D.; Jiang, Z.; Wu, Z.; Ai, B.; Wang, A.; et al. SWIFT: A Scalable lightWeight Infrastructure for Fine-Tuning. arXiv 2024, arXiv:2408.05517. [Google Scholar] [CrossRef]
- Hendrycks, D.; Burns, C.; Basart, S.; Zou, A.; Mazeika, M.; Song, D.; Steinhardt, J. Measuring Massive Multitask Language Understanding. In Proceedings of the 9th International Conference on Learning Representations (ICLR), Vienna, Austria, 4 May 2021. [Google Scholar]
- Wang, Y.; Ma, X.; Zhang, G.; Ni, Y.; Chandra, A.; Guo, S.; Ren, W.; Arulraj, A.; He, X.; Jiang, Z.; et al. MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark. In Proceedings of the Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2024; Volume 37, pp. 95266–95290. [Google Scholar] [CrossRef]
- Rein, D.; Hou, B.L.; Stickland, A.C.; Petty, J.; Pang, R.Y.; Dirani, J.; Michael, J.; Bowman, S.R. GPQA: A Graduate-Level Google-Proof Q&A Benchmark. In Proceedings of the First Conference on Language Modeling, Philadelphia, PA, USA, 7–9 October 2024. [Google Scholar]
- Clark, P.; Cowhey, I.; Etzioni, O.; Khot, T.; Sabharwal, A.; Schoenick, C.; Tafjord, O. Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge. arXiv 2018, arXiv:1803.05457. [Google Scholar] [CrossRef]
- Cobbe, K.; Kosaraju, V.; Bavarian, M.; Chen, M.; Jun, H.; Kaiser, L.; Plappert, M.; Tworek, J.; Hilton, J.; Nakano, R.; et al. Training Verifiers to Solve Math Word Problems. arXiv 2021, arXiv:2110.14168. [Google Scholar] [CrossRef]
- Bandarkar, L.; Liang, D.; Muller, B.; Artetxe, M.; Shukla, S.N.; Husa, D.; Goyal, N.; Krishnan, A.; Zettlemoyer, L.; Khabsa, M. The Belebele Benchmark: A Parallel Reading Comprehension Dataset in 122 Language Variants. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Bangkok, Thailand, 11–16 August 2024; pp. 749–775. [Google Scholar] [CrossRef]
- Team, N.; Costa-jussà, M.R.; Cross, J.; Çelebi, O.; Elbayad, M.; Heafield, K.; Heffernan, K.; Kalbassi, E.; Lam, J.; Licht, D.; et al. No Language Left Behind: Scaling Human-Centered Machine Translation. arXiv 2022, arXiv:2207.04672. [Google Scholar] [CrossRef]
- Togmanov, M.; Mukhituly, N.; Turmakhan, D.; Mansurov, J.; Goloburda, M.; Sakip, A.; Xie, Z.; Wang, Y.; Syzdykov, B.; Laiyk, N.; et al. KazMMLU: Evaluating Language Models on Kazakh, Russian, and Regional Knowledge of Kazakhstan. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Vienna, Austria, 27 July–1 August 2025; pp. 14403–14416. [Google Scholar] [CrossRef]
- Maxutov, A.; Arystanbekov, B.; Makhataeva, Z.; Yergen, A.; Taizhanov, N.; Nauryzbaikyzy, G.; Varol, H.A. Introducing Cultural Knowledge in Language Models: KazCulture Dataset for Kazakh Culture. IEEE Access 2026, 14, 44027–44042. [Google Scholar] [CrossRef]
- Yeshpanov, R.; Efimov, P.; Boytsov, L.; Shalkarbayuli, A.; Braslavski, P. KazQAD: Kazakh Open-Domain Question Answering Dataset. In Proceedings of the Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), Torino, Italy, 20–25 May 2024; pp. 9645–9656. Available online: https://aclanthology.org/2024.lrec-main.843 (accessed on 28 April 2026).
- Guerreiro, N.M.; Rei, R.; Stigt, D.v.; Coheur, L.; Colombo, P.; Martins, A.F.T. xCOMET: Transparent Machine Translation Evaluation through Fine-grained Error Detection. Trans. Assoc. Comput. Linguist. 2024, 12, 979–995. [Google Scholar] [CrossRef]
- Kembhavi, A.; Salvato, M.; Kolve, E.; Seo, M.; Hajishirzi, H.; Farhadi, A. A Diagram is Worth a Dozen Images. In Proceedings of the Computer Vision—ECCV 2016, Amsterdam, The Netherlands, 11–14 October 2016; pp. 235–251. [Google Scholar]
- Chen, L.; Li, J.; Dong, X.; Zhang, P.; Zang, Y.; Chen, Z.; Duan, H.; Wang, J.; Qiao, Y.; Lin, D.; et al. Are We on the Right Way for Evaluating Large Vision-Language Models? In Proceedings of the Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2024; Volume 37, pp. 27056–27087. [Google Scholar] [CrossRef]
- xAI. Grok-1.5 Vision Preview. 2024. Available online: https://x.ai/news/grok-1.5v (accessed on 20 May 2024).
- Lu, P.; Bansal, H.; Xia, T.; Liu, J.; Li, C.; Hajishirzi, H.; Cheng, H.; Chang, K.W.; Galley, M.; Gao, J. MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts. In Proceedings of the Twelfth International Conference on Learning Representations, Vienna, Austria, 7–11 May 2024. [Google Scholar]
- Liu, Y.; Li, Z.; Huang, M.; Yang, B.; Yu, W.; Li, C.; Yin, X.C.; Liu, C.L.; Jin, L.; Bai, X. OCRBench: On the hidden mystery of OCR in large multi-modal models. Sci. China Inf. Sci. 2024, 67, 220102. [Google Scholar] [CrossRef]
- MTS AI Research. MWS-Vision-Bench: Russian Multimodal OCR Benchmark. 2025. Available online: https://huggingface.co/datasets/MTSAIR/MWS-Vision-Bench (accessed on 29 April 2026).
- Conneau, A.; Ma, M.; Khanuja, S.; Zhang, Y.; Axelrod, V.; Dalmia, S.; Riesa, J.; Rivera, C.; Bapna, A. FLEURS: Few-shot Learning Evaluation of Universal Representations of Speech. arXiv 2022, arXiv:2205.12446. [Google Scholar] [CrossRef]
- Yang, C.K.; Ho, N.; Piao, Y.T.; Lee, H.-Y. SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information. In Proceedings of the Interspeech 2025, Rotterdam, The Netherlands, 17–21 August 2025; pp. 1788–1792. [Google Scholar] [CrossRef]
- Wei, C.; Wang, B.; Kim, J.J.; Chen, N.F. Towards Spoken Mathematical Reasoning: Benchmarking Speech-based Models over Multi-faceted Math Problems. arXiv 2025, arXiv:2505.15000. [Google Scholar]
- Institute of Smart Systems and Artificial Intelligence. Mangisoz. 2026. Available online: https://mangisoz.nu.edu.kz/soyle (accessed on 22 March 2026).
- Mei, X.; Meng, C.; Liu, H.; Kong, Q.; Ko, T.; Zhao, C.; Plumbley, M.D.; Zou, Y.; Wang, W. WavCaps: A ChatGPT-Assisted Weakly-Labelled Audio Captioning Dataset for Audio-Language Multimodal Research. IEEE/ACM Trans. Audio Speech Lang. Process. 2024, 32, 3339–3354. [Google Scholar] [CrossRef]
- Wang, B.; Zou, X.; Lin, G.; Sun, S.; Liu, Z.; Zhang, W.; Liu, Z.; Aw, A.; Chen, N.F. AudioBench: A Universal Benchmark for Audio Large Language Models. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), Albuquerque, NM, USA, 29 April–4 May 2025; pp. 4297–4316. [Google Scholar] [CrossRef]
- InflexionLab. Sybyrla. 2026. Available online: https://huggingface.co/InflexionLab/sybyrla (accessed on 28 May 2026).
- Rust, P.; Pfeiffer, J.; Vulić, I.; Ruder, S.; Gurevych, I. How Good is Your Tokenizer? On the Monolingual Performance of Multilingual Language Models. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), Online, 1–6 August 2021; pp. 3118–3135. [Google Scholar] [CrossRef]
- Petrov, A.; Malfa, E.L.; Torr, P.; Bibi, A. Language Model Tokenizers Introduce Unfairness Between Languages. In Proceedings of the Thirty-Seventh Conference on Neural Information Processing Systems, New Orleans, LA, USA, 10–16 December 2023. [Google Scholar]
- Gemmeke, J.F.; Ellis, D.P.W.; Freedman, D.; Jansen, A.; Lawrence, W.; Moore, R.C.; Plakal, M.; Ritter, M. Audio Set: An ontology and human-labeled dataset for audio events. In Proceedings of the 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), New Orleans, LA, USA, 5–9 March 2017; pp. 776–780. [Google Scholar] [CrossRef]
- Fonseca, E.; Favory, X.; Pons, J.; Font, F.; Serra, X. FSD50K: An Open Dataset of Human-Labeled Sound Events. IEEE/ACM Trans. Audio Speech Lang. Proc. 2021, 30, 829–852. [Google Scholar] [CrossRef]
- Kim, C.D.; Kim, B.; Lee, H.; Kim, G. AudioCaps: Generating Captions for Audios in The Wild. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), Minneapolis, MN, USA, 2–7 June 2019; pp. 119–132. [Google Scholar] [CrossRef]



| Task | Samples | Row-Wise Hours | Unique Hours | Primary Audio Sources |
|---|---|---|---|---|
| ASR | 2780K | 4750 | 3620 | KSC2 [40], CommonVoice [45], KazEmoTTS [48], Golos [49], YouTube transcripts |
| S2TT | 833K | 1078 | 1078 | CommonVoice [45], KazEmoTTS [48], VCTK [50], Hi-Fi TTS [51] |
| LangID | 362K | 978 | 978 | VoxLingua107 [46] (17-language subset of major world and Turkic languages) |
| Audio QA | 100K | 279 | 56 | MECAT-QA [47] (train subset) |
| Audio Captioning | 89K | 249 | 56 | MECAT-Caption [47] (shares audio clips with MECAT-QA) |
| Gender/Emotion | 37K | 28 | 10 | IEMOCAP [52], in-house gender + emotion set |
| Audio Classification | 2K | 3 | 3 | ESC-50 [53] |
| Total | 4.2M | 7365 | 5800 |
| Modality | Prompt Text | Prompt Media | Prompt Total | Completion | Full Seq. |
|---|---|---|---|---|---|
| Text-only | 139 | 0 | ≈139 | 1674 | 1813 |
| Audio–text | 25 | 310 (audio) | ≈336 | 881 | 1217 |
| Image–text | 90 | 367 (image) | ≈459 | 1410 | 1869 |
| Task | Benchmark | Audio Source | Text Source |
|---|---|---|---|
| ASR | FLEURS [75] | Natural | Native |
| S2TT | FLEURS [75] | Natural | Native |
| Spoken attribute reasoning | SAKURA [76] | Natural | MT into KK/RU |
| Spoken mathematical QA | Spoken MQA [77] | Natural | Native for EN |
| Synth. (KK via TTS) | MT into KK | ||
| Audio captioning | WavCaps [79] | Natural | MT into KK/RU |
| Audio captioning QA | WavCaps-QA [80] | Natural | MT into KK/RU |
| Lang. | Model | Average | MMLU-Pro | MMLU | GPQA | ARC-easy | ARC-chg. | GSM8K | Belebele |
|---|---|---|---|---|---|---|---|---|---|
| Kazakh | KazLLM-8B | 54.39 | 25.98 | 47.65 | 28.61 | 79.92 | 65.19 | 59.14 | 74.22 |
| Qolda-nothink (4.3B) | 62.95 | 31.60 | 55.81 | 30.87 | 88.09 | 77.30 | 77.66 | 79.35 | |
| Qolda-think (4.3B) | 73.45 | 59.49 | 68.63 | 39.41 | 94.82 | 87.29 | 84.60 | 79.93 | |
| Qolda-AVL-5B (Ours) | 77.44 | 63.39 | 74.45 | 43.95 | 95.75 | 90.87 | 87.79 | 85.88 | |
| Qwen3-VL-4B-Instruct | 59.18 | 39.11 | 52.70 | 36.10 | 75.93 | 66.16 | 75.59 | 68.67 | |
| Qwen3-VL-4B-Thinking | 74.60 | 62.64 | 67.81 | 56.62 | 89.66 | 81.82 | 85.43 | 78.21 | |
| Qwen2.5-Omni-7B | 38.06 | 23.57 | 39.54 | 31.16 | 46.78 | 41.75 | 30.96 | 52.67 | |
| Qwen2.5-Omni-3B | 26.66 | 16.75 | 30.94 | 23.80 | 32.54 | 30.98 | 11.38 | 40.22 | |
| Qwen3-Omni-30B-Instruct | 73.20 | 54.55 | 67.26 | 47.36 | 92.71 | 84.85 | 85.43 | 80.22 | |
| Qwen3-Omni-30B-Thinking | 83.24 | 72.66 | 79.41 | 66.37 | 95.51 | 93.27 | 90.59 | 84.86 | |
| English | KazLLM-8B | 67.64 | 37.24 | 63.75 | 29.36 | 93.31 | 82.51 | 77.10 | 90.22 |
| Qolda-nothink (4.3B) | 72.59 | 43.92 | 67.77 | 31.41 | 96.14 | 87.32 | 88.21 | 93.33 | |
| Qolda-think (4.3B) | 80.83 | 66.88 | 76.75 | 45.43 | 96.27 | 90.24 | 93.71 | 96.50 | |
| Qolda-AVL-5B (Ours) | 83.14 | 70.60 | 79.41 | 49.54 | 98.15 | 94.11 | 94.84 | 95.31 | |
| Qwen3-VL-4B-Instruct | 80.61 | 65.74 | 74.59 | 46.94 | 96.53 | 92.59 | 94.77 | 93.11 | |
| Qwen3-VL-4B-Thinking | 87.95 | 75.87 | 82.75 | 71.59 | 98.22 | 95.26 | 95.30 | 96.65 | |
| Qwen2.5-Omni-7B | 69.48 | 43.56 | 67.22 | 33.03 | 96.27 | 89.56 | 64.04 | 92.65 | |
| Qwen2.5-Omni-3B | 62.42 | 29.87 | 59.94 | 27.62 | 91.53 | 80.47 | 60.68 | 86.86 | |
| Qwen3-Omni-30B-Instruct | 84.53 | 69.27 | 81.77 | 54.53 | 98.39 | 96.46 | 95.30 | 96.00 | |
| Qwen3-Omni-30B-Thinking | 91.31 | 81.90 | 87.50 | 79.61 | 98.73 | 96.63 | 97.27 | 97.54 | |
| Russian | KazLLM-8B | 62.21 | 27.96 | 50.47 | 27.18 | 88.95 | 76.95 | 75.99 | 88.00 |
| Qolda-nothink (4.3B) | 69.58 | 38.13 | 59.90 | 29.96 | 93.09 | 86.36 | 90.27 | 89.33 | |
| Qolda-think (4.3B) | 77.78 | 61.94 | 70.27 | 43.43 | 95.77 | 90.00 | 93.62 | 89.40 | |
| Qolda-AVL-5B (Ours) | 80.70 | 67.17 | 76.95 | 46.33 | 96.84 | 94.89 | 92.65 | 90.07 | |
| Qwen3-VL-4B-Instruct | 76.21 | 57.83 | 67.05 | 46.65 | 94.49 | 88.38 | 89.53 | 89.56 | |
| Qwen3-VL-4B-Thinking | 82.83 | 68.48 | 74.98 | 61.46 | 97.88 | 92.59 | 92.56 | 91.78 | |
| Qwen2.5-Omni-7B | 63.18 | 36.45 | 59.28 | 29.96 | 90.08 | 78.96 | 60.39 | 87.11 | |
| Qwen2.5-Omni-3B | 55.34 | 24.33 | 49.68 | 28.89 | 82.63 | 71.04 | 51.59 | 79.19 | |
| Qwen3-Omni-30B-Instruct | 81.23 | 64.35 | 75.73 | 49.73 | 97.97 | 94.44 | 94.39 | 92.00 | |
| Qwen3-Omni-30B-Thinking | 88.33 | 78.02 | 83.81 | 71.75 | 98.81 | 96.30 | 95.90 | 93.75 |
| Model | Average | en–kk | kk–en | ru–kk | kk–ru | en–ru | ru–en |
|---|---|---|---|---|---|---|---|
| KazLLM-8B | 0.7156 | 0.5542 | 0.7426 | 0.4945 | 0.7139 | 0.8612 | 0.9272 |
| Qolda-nothink (4.3B) | 0.7858 | 0.7230 | 0.7590 | 0.7081 | 0.7529 | 0.8509 | 0.9210 |
| Qolda-think (4.3B) | 0.8040 | 0.7406 | 0.7851 | 0.7111 | 0.7843 | 0.8714 | 0.9313 |
| Qolda-AVL-5B (Ours) | 0.7783 | 0.6534 | 0.7845 | 0.6146 | 0.7926 | 0.8875 | 0.9370 |
| Qwen3-VL-4B-Instruct | 0.5933 | 0.2342 | 0.6474 | 0.2400 | 0.6330 | 0.8637 | 0.9412 |
| Qwen3-VL-4B-Thinking | 0.7105 | 0.4895 | 0.6784 | 0.5042 | 0.7250 | 0.9199 | 0.9461 |
| Qwen2.5-Omni-7B | 0.4667 | 0.1387 | 0.4086 | 0.1408 | 0.3817 | 0.7965 | 0.9341 |
| Qwen2.5-Omni-3B | 0.3963 | 0.1270 | 0.2526 | 0.1425 | 0.2545 | 0.6907 | 0.9103 |
| Qwen3-Omni-30B-Instruct | 0.7625 | 0.6007 | 0.7647 | 0.5601 | 0.7712 | 0.9284 | 0.9499 |
| Qwen3-Omni-30B-Thinking | 0.8371 | 0.7505 | 0.7933 | 0.7354 | 0.8440 | 0.9424 | 0.9572 |
| Model | Average | KazMMLU | KazQAD | KazCulture | |
|---|---|---|---|---|---|
| PQ | Q | ||||
| KazLLM-8B | 64.72 | 51.32 | 73.30 | 96.85 | 37.41 |
| Qolda-nothink (4.3B) | 68.64 | 52.95 | 74.19 | 98.27 | 49.13 |
| Qolda-think (4.3B) | 72.62 | 66.12 | 77.58 | 97.90 | 48.87 |
| Qolda-AVL-5B (Ours) | 72.16 | 69.47 | 76.19 | 97.68 | 45.28 |
| Qwen3-VL-4B-Instruct | 68.10 | 61.73 | 76.47 | 97.15 | 37.03 |
| Qwen3-VL-4B-Thinking | 68.42 | 68.44 | 76.47 | 97.15 | 31.63 |
| Qwen2.5-Omni-7B | 55.14 | 49.12 | 52.65 | 93.40 | 25.38 |
| Qwen2.5-Omni-3B | 49.29 | 41.11 | 41.96 | 87.71 | 26.39 |
| Qwen3-Omni-30B-Instruct | 72.44 | 70.38 | 78.44 | 98.35 | 42.58 |
| Qwen3-Omni-30B-Thinking | 73.56 | 77.80 | 80.46 | 98.50 | 37.48 |
| Lang. | Model | Average | AI2D | MMStar | RealWorldQA | MathVista | OCRBench |
|---|---|---|---|---|---|---|---|
| Kazakh | Qolda-nothink (4.3B) | 44.85 | 46.73 | 37.30 | 46.01 | 37.54 | 56.67 |
| Qolda-think (4.3B) | 53.20 | 55.18 | 59.37 | 49.02 | 45.09 | 57.36 | |
| Qolda-AVL-5B (Ours) | 62.73 | 72.49 | 67.59 | 55.95 | 67.24 | 50.39 | |
| Qwen3-VL-4B-Instruct | 54.02 | 61.32 | 50.60 | 51.24 | 50.25 | 56.69 | |
| Qwen3-VL-4B-Thinking | 60.62 | 67.16 | 65.70 | 55.03 | 66.70 | 48.53 | |
| Qwen2.5-Omni-7B | 44.31 | 52.03 | 45.46 | 35.42 | 44.90 | 43.76 | |
| Qwen2.5-Omni-3B | 35.29 | 39.57 | 33.80 | 21.31 | 37.80 | 43.99 | |
| Qwen3-Omni-30B-Instruct | 60.50 | 73.15 | 61.06 | 54.90 | 54.42 | 58.96 | |
| Qwen3-Omni-30B-Thinking | 64.48 | 75.37 | 68.85 | 54.77 | 71.01 | 52.38 | |
| English | Qolda-nothink (4.3B) | 52.16 | 69.14 | 36.96 | 53.33 | 40.20 | 61.16 |
| Qolda-think (4.3B) | 64.23 | 73.07 | 60.88 | 60.00 | 68.10 | 59.08 | |
| Qolda-AVL-5B (Ours) | 75.45 | 81.41 | 70.62 | 68.63 | 74.57 | 82.00 | |
| Qwen3-VL-4B-Instruct | 73.75 | 81.41 | 65.12 | 71.11 | 65.90 | 85.20 | |
| Qwen3-VL-4B-Thinking | 76.39 | 83.47 | 73.88 | 67.97 | 75.72 | 80.90 | |
| Qwen2.5-Omni-7B | 69.78 | 80.94 | 61.17 | 65.10 | 58.60 | 83.10 | |
| Qwen2.5-Omni-3B | 64.46 | 76.23 | 53.13 | 56.21 | 51.95 | 84.80 | |
| Qwen3-Omni-30B-Instruct | 76.98 | 86.59 | 70.57 | 73.33 | 69.10 | 85.30 | |
| Qwen3-Omni-30B-Thinking | 80.91 | 88.64 | 74.48 | 74.77 | 79.76 | 86.90 | |
| Russian | Qolda-nothink (4.3B) | 37.98 | – | 36.87 | 48.10 | – | 28.96 |
| Qolda-think (4.3B) | 47.91 | – | 54.95 | 56.07 | – | 32.72 | |
| Qolda-AVL-5B (Ours) | 56.52 | – | 67.23 | 64.18 | – | 38.14 | |
| Qwen3-VL-4B-Instruct | 52.54 | – | 57.68 | 61.70 | – | 38.25 | |
| Qwen3-VL-4B-Thinking | 54.68 | – | 71.33 | 60.92 | – | 31.80 | |
| Qwen2.5-Omni-7B | 49.09 | – | 58.87 | 46.01 | – | 42.40 | |
| Qwen2.5-Omni-3B | 42.52 | – | 50.13 | 42.48 | – | 34.95 | |
| Qwen3-Omni-30B-Instruct | 60.30 | – | 68.11 | 65.49 | – | 47.31 | |
| Qwen3-Omni-30B-Thinking | 61.88 | – | 71.83 | 65.36 | – | 48.46 |
| Model | Kazakh | English | Russian | |||
|---|---|---|---|---|---|---|
| Mean | Cleaned | Mean | Cleaned | Mean | Cleaned | |
| Sybyrla | 0.2244|0.1951 | 0.1968|0.1681 | 0.1555|0.1423 | 0.1140|0.0991 | 0.1253|0.1142 | 0.1131|0.0999 |
| Qolda-AVL-5B (Ours) | 0.2027|0.1874 | 0.1801|0.1648 | 0.1130|0.0982 | 0.0995|0.0876 | 0.1388|0.1293 | 0.1222|0.1129 |
| Qwen2.5-Omni-7B | 3.6722|3.6671 | 1.0413|1.0362 | 0.1407|0.1160 | 0.1317|0.1033 | 0.3014|0.2395 | 0.2755|0.2103 |
| Qwen2.5-Omni-3B | 7.5563|7.5534 | 1.0828|1.0796 | 0.7357|0.7096 | 0.1626|0.1418 | 0.2892|0.2537 | 0.2653|0.2277 |
| Qwen3-Omni-30B-Instruct | 0.8350|0.8173 | 0.6217|0.6037 | 0.1118|0.0963 | 0.0996|0.0804 | 0.1074|0.0944 | 0.0928|0.0811 |
| Qwen3-Omni-30B-Thinking | 0.7300|0.7175 | 0.7025|0.6890 | 0.1105|0.0540 | 0.0964|0.0532 | 0.1048|0.0953 | 0.0939|0.0821 |
| Model | Average | kk–en | kk–ru | en–kk | en–ru | ru–kk | ru–en |
|---|---|---|---|---|---|---|---|
| Qolda-AVL-5B (Ours) | 0.7349 | 0.6924 | 0.6870 | 0.6364 | 0.8420 | 0.6461 | 0.9053 |
| Qwen2.5-Omni-7B | 0.4457 | 0.1430 | 0.1440 | 0.2508 | 0.8018 | 0.4447 | 0.8898 |
| Qwen2.5-Omni-3B | 0.4090 | 0.1411 | 0.1449 | 0.2046 | 0.7103 | 0.3938 | 0.8595 |
| Qwen3-Omni-30B-Instruct | 0.5718 | 0.1924 | 0.2061 | 0.6024 | 0.9095 | 0.5888 | 0.9315 |
| Qwen3-Omni-30B-Thinking | 0.6067 | 0.2728 | 0.2707 | 0.6377 | 0.9066 | 0.6199 | 0.9324 |
| Lang. | Model | Average | Gender | Animal | Language | Emotion | ||||
|---|---|---|---|---|---|---|---|---|---|---|
| Single | Multi | Single | Multi | Single | Multi | Single | Multi | |||
| Kazakh | Qolda-AVL-5B (Ours) | 64.02 | 81.20 | 78.80 | 60.00 | 52.40 | 86.14 | 79.36 | 36.47 | 37.80 |
| Qwen2.5-Omni-7B | 46.38 | 52.20 | 52.60 | 46.00 | 33.80 | 92.60 | 43.40 | 24.80 | 25.60 | |
| Qwen2.5-Omni-3B | 32.80 | 50.40 | 48.80 | 25.80 | 28.20 | 33.20 | 25.00 | 27.20 | 23.80 | |
| Qwen3-Omni-30B-Instruct | 62.58 | 79.00 | 52.80 | 95.00 | 50.60 | 97.00 | 56.80 | 36.80 | 32.60 | |
| Qwen3-Omni-30B-Thinking | 72.44 | 87.00 | 66.80 | 93.40 | 71.54 | 96.40 | 90.60 | 40.80 | 33.00 | |
| English | Qolda-AVL-5B (Ours) | 68.04 | 88.80 | 84.60 | 57.20 | 69.40 | 86.75 | 88.20 | 34.80 | 34.60 |
| Qwen2.5-Omni-7B | 61.18 | 64.60 | 51.00 | 94.60 | 63.00 | 94.20 | 58.80 | 35.00 | 28.20 | |
| Qwen2.5-Omni-3B | 49.15 | 51.80 | 49.80 | 94.40 | 50.40 | 62.00 | 32.60 | 27.00 | 25.20 | |
| Qwen3-Omni-30B-Instruct | 69.33 | 88.20 | 54.60 | 94.40 | 65.40 | 97.80 | 74.80 | 44.60 | 34.80 | |
| Qwen3-Omni-30B-Thinking | 78.70 | 91.80 | 76.20 | 93.00 | 85.60 | 98.40 | 96.40 | 50.00 | 38.20 | |
| Russian | Qolda-AVL-5B (Ours) | 66.34 | 82.80 | 80.80 | 60.32 | 59.40 | 90.98 | 84.80 | 35.40 | 36.20 |
| Qwen2.5-Omni-7B | 58.45 | 57.60 | 46.80 | 96.20 | 59.60 | 92.60 | 51.80 | 34.60 | 28.40 | |
| Qwen2.5-Omni-3B | 45.10 | 47.20 | 50.80 | 87.00 | 44.80 | 47.80 | 30.80 | 28.20 | 24.20 | |
| Qwen3-Omni-30B-Instruct | 65.70 | 76.80 | 46.20 | 94.80 | 67.20 | 99.20 | 63.60 | 43.40 | 34.40 | |
| Qwen3-Omni-30B-Thinking | 76.85 | 88.00 | 65.80 | 94.00 | 85.80 | 99.00 | 93.60 | 48.00 | 40.60 | |
| Lang. | Model | Average | Audio Cap. | Audio Cap. QA |
|---|---|---|---|---|
| Kazakh | Qolda-AVL-5B (Ours) | 25.89 | 12.95 | 38.82 |
| Qwen2.5-Omni-7B | 4.30 | 1.68 | 6.91 | |
| Qwen2.5-Omni-3B | 3.55 | 4.80 | 2.30 | |
| Qwen3-Omni-30B-Instruct | 18.83 | 11.33 | 26.32 | |
| Qwen3-Omni-30B-Thinking | 25.20 | 16.85 | 33.55 | |
| English | Qolda-AVL-5B (Ours) | 26.47 | 16.76 | 36.18 |
| Qwen2.5-Omni-7B | 37.15 | 22.97 | 51.32 | |
| Qwen2.5-Omni-3B | 30.24 | 19.02 | 41.45 | |
| Qwen3-Omni-30B-Instruct | 36.18 | 23.99 | 48.36 | |
| Qwen3-Omni-30B-Thinking | 37.69 | 24.06 | 51.32 | |
| Russian | Qolda-AVL-5B (Ours) | 27.22 | 13.64 | 40.79 |
| Qwen2.5-Omni-7B | 26.29 | 21.33 | 31.25 | |
| Qwen2.5-Omni-3B | 17.78 | 12.20 | 23.36 | |
| Qwen3-Omni-30B-Instruct | 29.89 | 19.31 | 40.46 | |
| Qwen3-Omni-30B-Thinking | 33.71 | 23.01 | 44.41 |
| Lang. | Model | Average | Short Digit | Long Digit | Single-Step Reason. | Multi-Step Reason. |
|---|---|---|---|---|---|---|
| Kazakh | Qolda-AVL-5B (Ours) | 87.24 | 90.00 | 88.37 | 93.58 | 77.01 |
| Qwen2.5-Omni-7B | 1.31 | 2.00 | 0.00 | 2.36 | 0.86 | |
| Qwen2.5-Omni-3B | 1.24 | 1.00 | 0.00 | 3.21 | 0.74 | |
| Qwen3-Omni-30B-Instruct | 14.72 | 14.00 | 4.65 | 26.35 | 13.86 | |
| Qwen3-Omni-30B-Thinking | 19.10 | 12.12 | 2.40 | 36.99 | 24.89 | |
| English | Qolda-AVL-5B (Ours) | 88.65 | 88.00 | 84.88 | 94.93 | 86.78 |
| Qwen2.5-Omni-7B | 71.63 | 86.00 | 48.26 | 89.19 | 63.07 | |
| Qwen2.5-Omni-3B | 65.41 | 75.00 | 43.02 | 83.28 | 60.34 | |
| Qwen3-Omni-30B-Instruct | 83.44 | 93.00 | 84.30 | 96.11 | 60.34 | |
| Qwen3-Omni-30B-Thinking | 94.34 | 93.00 | 94.77 | 95.44 | 94.53 |
| Task | Audio DeepStack Ablation | Whisper Fine-Tuning Ablation | ||||
|---|---|---|---|---|---|---|
| Baseline | Audio DS | Whisper-Orig. | Whisper-FT | |||
| ASR (WER ↓) | −0.12% | −1.88% | ||||
| S2TT (xCOMET ↑) | +0.0098 | |||||
| QA Animal (Acc ↑) | +12.00% | +2.08% | ||||
| QA Language (Acc ↑) | +6.53% | |||||
| QA Emotion (Acc ↑) | +3.30% | |||||
| QA Gender (Acc ↑) | +0.07% | |||||
| Reasoning (Acc ↑) | +3.81% | |||||
| Language | Words | Tokens | Tokens/Word | Ratio vs. English |
|---|---|---|---|---|
| English | 22,503 | 28,303 | 1.26 | 1.00× |
| Russian | 21,503 | 52,873 | 2.46 | 1.95× |
| Kazakh | 20,693 | 98,961 | 4.78 | 3.79× |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Arystanbekov, B.; Maxutov, A.; Nurimanov, A.; Varol, H.A. Extending a Vision–Language Model with Audio Understanding: Introducing Qolda-AVL for the Kazakh Language. Big Data Cogn. Comput. 2026, 10, 192. https://doi.org/10.3390/bdcc10060192
Arystanbekov B, Maxutov A, Nurimanov A, Varol HA. Extending a Vision–Language Model with Audio Understanding: Introducing Qolda-AVL for the Kazakh Language. Big Data and Cognitive Computing. 2026; 10(6):192. https://doi.org/10.3390/bdcc10060192
Chicago/Turabian StyleArystanbekov, Batyr, Akylbek Maxutov, Aspandiyar Nurimanov, and Huseyin Atakan Varol. 2026. "Extending a Vision–Language Model with Audio Understanding: Introducing Qolda-AVL for the Kazakh Language" Big Data and Cognitive Computing 10, no. 6: 192. https://doi.org/10.3390/bdcc10060192
APA StyleArystanbekov, B., Maxutov, A., Nurimanov, A., & Varol, H. A. (2026). Extending a Vision–Language Model with Audio Understanding: Introducing Qolda-AVL for the Kazakh Language. Big Data and Cognitive Computing, 10(6), 192. https://doi.org/10.3390/bdcc10060192

