Vision-Language Models in Teaching and Learning: A Systematic Literature Review
Abstract
1. Introduction
- Educational motivation: Learners typically gain understanding from more than one channel, such as linguistic input (e.g., speech and text) and imagery input (e.g., pictures and diagrams) (Paivio, 2013). The first language pathway excels at sequential and symbolic processing, whereas the second visual pathway supports holistic, spatial reasoning. From a constructivist perspective, VLMs act as scaffolds to help learners actively build knowledge and provide feedback. From a connectivist lens, VLMs retrieve, filter, and translate multimodal resources (e.g., slides, and lecture video recordings), enabling just-in-time connections among learners, instructors, and learning materials.
- Technological motivation: Beyond text-only large language models, VLMs combine inputs across formats such as text, images, and video (Danish et al., 2026; Shinde et al., 2025; Zhang et al., 2024). This integration allows systems to interpret and generate pedagogy-relevant artifacts (e.g., slides and figures) within one integrated framework, aligning naturally with how students learn. Traditionally, these abilities were pursued in separate communities, including computer vision for understanding images and natural language processing for reasoning over language. However, the multimodal architectures used in VLMs can integrate both modalities, enabling new practical opportunities for teaching, assessment, and feedback.
- Research gap: A few systematic review articles have studied the application of large language models (LLMs) in education (Agbo et al., 2025; Ali et al., 2024; Kostopoulos et al., 2025; H. Y. Lee et al., 2025; Raihan et al., 2025; P. Wang et al., 2025). However, they have only addressed the single modality model (i.e., language only), which is different from the focus of this paper on two modalities (i.e., both vision and language), as shown in Table 2. On the other hand, multimodal large language models (MLLMs) improve learner engagement, support personalized pathways, and deepen understanding by processing and producing context-aware content (G. Lee et al., 2025). The integration of diverse cognitive channels motivates the integration of MLLMs to enhance learners’ educational experience (Xing et al., 2024). It outlines both the opportunities and the challenges created by bringing MLLMs into educational practice (Küchemann et al., 2025). Different from these studies, this review has a focus on a set of new research questions curated for VLMs, including where VLMs are applied, which VLM solutions are applied, what the input–output modalities are, what the pedagogical roles of VLMs and their associated outcomes for learners and instructors are, and how they are evaluated.
2. Methodology
2.1. Literature Search Process
2.2. Quality Assessment
2.3. Research Questions
- RQ1. In which educational levels and academic disciplines have VLMs been applied to teaching? It is important to map the deployment landscape to clarify where VLMs are already used. Therefore, we need to extract discipline and education level from each article.
- RQ2. What type of VLM solutions are used, what modalities do they handle, and what do they generate? It is critical to apply a user-facing taxonomy (off-the-shelf pre-trained vs tuned) that supports practical adoption decisions, and understands input/output modalities, which indicate how to fit VLMs with real teaching artifacts (e.g., diagrams, slides). Therefore, we need to record the solution type, input modalities, and outputs from each article.
- RQ3. What is the role of the VLMs in the teaching workflow? It is essential to clarify the role of VLMs in the workflow for effective deployment. Therefore, we need to record the VLM role from each article.
- RQ4. What benefits are reported for learners and instructors? VLMs might introduce different benefits for different stakeholders (e.g., learners and instructors). Therefore, we need to extract learner and instructor benefits from each article.
- RQ5. What challenges and risks are reported for learners and instructors, and what mitigation strategies are described or evaluated? The adoption challenges and tested mitigation strategies would benefit other educators in integrating VLMs into their practice. Therefore, we need to record technical, pedagogical, and ethical issues, and identify which mitigation is empirically evaluated in each article.
- RQ6. How are studies designed and evaluated? To verify the educational impacts of VLMs on teaching and learning, it is required to assess methodological quality. Therefore, we need to record the dataset and reproducibility from each article.
3. Results
3.1. RQ1: Education Levels and Disciplines
3.2. RQ2: VLMs Solution and Modality
3.3. RQ3: Pedagogical Functions
- Analyst (13 articles). In this role, VLMs analyze educational data and provide insights to augment instructors’ decisions, rather than making decisions automatically. For example, as classroom analytics tools, VLMs analyze visual classroom data (images/video) to infer engagement and participation, and can also support academic integrity monitoring in online settings (e.g., activity detection). These functions provide instructors with timely and actionable insights.
- Assessor (10 articles). In this role, VLMs support assessment by interpreting learners’ submissions, such as scanned worksheets, and automatically producing scores.
- Content curator (9 articles). VLMs support content authoring and augmentation by transforming raw instructional assets (e.g., slides, textbooks, and lecture videos) into richer learning materials. They can also generate assessment content from multimodal inputs, turning lecture videos into structured questions (MCQs, short-answer open-ended questions).
- Simulator (1 article). VLMs help craft interactive, persona-driven experiences that simulate learning in concrete visual contexts. They combine visual perception with dialogue to create situated practice opportunities for learners.
- Tutor (9 articles). For this role, VLMs deliver just-in-time guidance grounded in visual context, such as interpreting diagrams or student sketches, and responding with hints or clarifications. Such approaches can be deployed synchronously (during class) or asynchronously (homework support). This tutor role differs from the analyst role because it decides to provide direct feedback to learners, while the analyst only supports users’ decision-making through analyzed insights.
3.4. RQ4: Learners and Instructors’ Benefits from VLMs
3.5. RQ5: Challenge and Mitigation
- Technical challenges. VLMs face several technical difficulties in educational settings. Low-quality inputs can degrade performance. For example, poor images lead to inaccurate color identification and object counting (Tapia-Mandiola & Araya, 2025), and heterogeneous transcript layouts cause recognition errors (Bhaskaran & Pardos, 2025). Robust pre-processing (e.g., segmentation and layout normalization) is recommended to improve accuracy. The inherent hallucination risk of VLMs might produce irrelevant content; targeted fine-tuning was used to improve question relevance (Nguyen & Park, 2025; Stamatakis et al., 2025). The performance of VLMs is affected by how they are called via various prompts. For that, specific prompt designs are reported in (Xie et al., 2025) for enhancing grading accuracy and consistency. Finally, VLMs can be computationally expensive to use; (J. Chen et al., 2024) highlights the need for computationally efficient fine-tuning methods.
- Pedagogical challenges. In content generation, (Kunuku & Dehbozorgi, 2025) notes that outputs must align with higher-order thinking in Bloom’s taxonomy (e.g., application, analysis, and evaluation) and proposes an LLM-as-Judge framework to automate evaluation of cognitive alignment. The type of feedback significantly influences students’ motivation to revise their work. The direct and informative feedback conditions were found to be more effective in encouraging students’ revisions compared to the general feedback conditions (Zhuang et al., 2025). Therefore, VLMs should provide more detailed and guided feedback.
- Many ethical concerns have been reported in the reviewed articles; ethical safeguards are essential for the VLM-enabled applications in teaching (Marquez-Carpintero et al., 2025). They emphasize the need for informed consent, recommend establishing regulatory frameworks, and preventing surveillance or misuse of sensitive information. For example, one reported concern is the risk of unauthorized disclosure of student information and the privacy risks when learner data are processed with proprietary VLM services (U. Lee et al., 2024). Furthermore, cultural misalignment is also reported; they often struggle to interpret visual content in non-Western contexts, potentially misrepresenting under-resourced cultures (Tan et al., 2025). To mitigate these risks, many strategies are reported. One typical strategy is to implement a human-in-the-loop approach, where instructors provide essential oversight to ensure pedagogical validity (Shu et al., 2025). Retrieval-Augmented Generation (RAG) is also implemented to reduce hallucinations and ensure factual, context-aware outputs (Tan et al., 2025). Lastly, compared with proprietary VLMs, open-weight models might enhance privacy by allowing for on-device deployment (U. Lee et al., 2024). These measures should be integrated into the implementation plans to ensure responsible use of VLMs in educational settings.
3.6. RQ6: Validation and Evaluation
4. Discussions
4.1. Future Research Directions
4.2. Limitations
5. Conclusions
- Opportunities. VLMs are shown in the reviewed articles to provide opportunities to transform teaching and learning practices by enabling more interactive and tailored educational environments, even for instructors and learners without computer science backgrounds. For learners, VLMs offer personalized support and interactive learning activities. Instructors benefit from VLMs through automatic content generation and analytics insights from educational data.
- Challenges. It is important to understand the new challenges that are inherent in this emerging technology. Learners face challenges of verifying the accuracy and managing potential biases of VLM outputs, which could affect critical thinking and independent learning. On the other hand, instructors also face challenges of developing effective prompting strategies to use VLMs and ensuring the alignment between the generated content with learning outcomes.
Supplementary Materials
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
Appendix A
| Reference | Description |
|---|---|
| Abdelhadi et al. (2025) | A real-time classroom and examination system verifies identity, monitors behaviors, and recommends teacher actions. |
| Ambali Parambil et al. (2025) | It classifies student emotions from online-classroom images. |
| Anderer et al. (2024) | It analyzes lecture slides for visually impaired access by converting visual elements to multilayer textual descriptions. |
| Anderer et al. (2025) | It provides a lecture assistant to support accessible lecture video navigation and visual question answering. |
| Asseri et al. (2025) | It applies VLMs on emotion recognition for Arabic children’s storybook images under multiple prompting schemes. |
| Bhaskaran and Pardos (2025) | It compares OCR- and VLM-based pipelines to extract structured course and grade information from heterogeneous student transcript documents. |
| Bossema et al. (2025) | A course with older adults to explore human–robot co-creativity with a multimodal LLM. |
| Busic et al. (2024) | It applies multimodal LLMs to automate grading in GUI design courses. |
| Cao et al. (2025) | A classroom-video analytics system that integrates LLM prompting to identify teaching behaviors and produce interpretable feedback. |
| J. Chen et al. (2024) | A VLM-based approach generates primary-school multiple-choice questions from cartoon images by combining image captioning with text generation. |
| Dang et al. (2025) | It studies learner–agent interaction patterns with an embodied AI assistant in mixed reality. |
| Edwards et al. (2025) | A statistical procedure to evaluate whether VLMs grade engineering sketches equivalently to human experts. |
| Fahmi and Bousmah (2025) | An AI virtual teacher assistant for student diagnosis and remediation. |
| Han et al. (2025) | A tangible storytelling pipeline where multimodal LLMs generate paper-cut style assets and support in crafting stage-based narratives. |
| Hang and Man Ho (2025) | It applies multimodal LLM workflows to generate personalized vocabulary flashcards for early childhood education. |
| Ibanez et al. (2025) | It applies multimodal LLMs as graders of UML class diagrams. |
| Kunuku and Dehbozorgi (2025) | A multimodal framework creates quiz questions from video lectures by fusing text, visuals, and audio. |
| U. Lee et al. (2024) | It develops a personalized art-appreciation tutor, including the creation of a GPT-generated dialogue dataset and benchmarking. |
| G. G. Lee and Zhai (2025) | It provides educational visual question answering for researchers to query and analyze image data through textural prompts. |
| J. Lee et al. (2025) | A multimodal tutoring system offers step-by-step textual and visual guidance by generating diagrams through code-assisted reasoning. |
| Marquez-Carpintero et al. (2025) | A two-phase method leverages VLMs to estimate student attention and emotions from classroom imagery for STEM education. |
| Mittal et al. (2025) | A VLM-enabled system provides real-time, context-aware answers to student questions during live lectures using live content and retrieved course materials. |
| Nguyen and Hayward (2025) | It uses a multimodal LLM to annotate K-12 science assessments and suggest revisions. |
| Nguyen and Park (2025) | It applies multimodal LLMs for scoring and feedback on multimodal science assessments. |
| Pang et al. (2026) | A tuned VLM to recognize student engagement cues in still images. |
| Picard et al. (2025) | It evaluates VLMs across engineering design tasks and provides benchmark datasets for continued assessment. |
| Rahmanian et al. (2025) | It applies VLMs to assess student ER diagrams under different input contexts and prompting strategies. |
| Sheng et al. (2025) | It combines digital-pen traces with a multimodal LLM to reconstruct students’ step-by-step reasoning chains. |
| Shu et al. (2025) | It extracts learning outcomes from lecture notes and uses a multimodal LLM to generate multiple-choice questions with solutions and explanations. |
| Singh et al. (2023) | A VLM-driven pipeline retrieves and assigns web images to e-textbooks via a text–image matching optimization. |
| Stamatakis et al. (2025) | It applies VLMs to generate learning-oriented questions from educational videos. |
| Su et al. (2025) | An intelligent tutoring system for learning Chinese characters that leverages a multimodal LLM to deliver corrective feedback. |
| Tan et al. (2025) | A picture-guided conversational chatbot for early childhood language learning. |
| Tapia-Mandiola and Araya (2025) | A two-step approach segments key regions in students’ coloring-task images and then applies a VLM to analyze the cropped sections for automated grading support. |
| Teotia et al. (2024) | It evaluates VLMs on classroom learning-engagement detection using behavior and emotion datasets. |
| Tschope et al. (2025) | It applies VLMs for recognizing activities in nursing-procedure training videos. |
| Y. Wang et al. (2025) | A modular system generates lecture scripts from multimodal slide inputs using instruction-guided VLM workflows. |
| X. Wang et al. (2025) | It uses a multimodal LLM to perform automated essay scoring. |
| X. Wang et al. (2025) | A wearable system uses computer vision, VLMs, and speech technologies to equip everyday objects with conversational personas for interactive guidance. |
| Xie et al. (2025) | It studies prompt-engineering strategies for automated K-12 exam grading with a VLM across six question types and proposes an evaluation framework for grading behavior. |
| Zheng et al. (2025) | It conducts art-evaluation dialogues with multimodal LLMs for teacher support. |
| Zhuang et al. (2025) | It scores picture-cued student writing against images and provides feedback for middle-school language learning. |
References
- Abdelhadi, Z., Naseif, M., Alhejali, W., & Elhayek, A. (2025, January 15–16). TeacherEye: An AI-powered system for monitoring student engagement in online education. International Learning and Technology Conference (pp. 25–30), Jeddah, Saudi Arabia. [Google Scholar] [CrossRef]
- Agbo, F. J., Olivia, C., Oguibe, G., Sanusi, I. T., & Sani, G. (2025). Computing education using generative artificial intelligence tools: A systematic literature review. Computers and Education Open, 9, 100266. [Google Scholar] [CrossRef]
- Ali, D., Fatemi, Y., Boskabadi, E., Nikfar, M., Ugwuoke, J., & Ali, H. B. (2024). ChatGPT in teaching and learning: A systematic review. Education Sciences, 14(6), 643. [Google Scholar] [CrossRef]
- Ambali Parambil, M. M., Bouktif, S., Gochoo, M., & Alnajjar, F. S. K. (2025, April 22–25). Comparing emotion detection methods in online classrooms: YOLO models, multimodal LLM, and human baseline. IEEE Global Engineering Education Conference (pp. 1–7), London, UK. [Google Scholar] [CrossRef]
- Anderer, K., Muller, K. E., Strobel, L., Wolfel, M., Niehues, J. M., & Gerling, K. M. (2025, October 26–29). Making lecture videos accessible for students who are blind or have low vision through AI-assisted navigation and visual question answering. International ACM SIGACCESS Conference on Computers and Accessibility, Denver, CO, USA. [Google Scholar] [CrossRef]
- Anderer, K., Wölfel, M., & Niehues, J. M. (2024, September 2–6). Identifying the information gap for visually impaired students during lecture talks. IEEE Symposium on Visual Languages and Human-Centric Computing (pp. 168–173), Liverpool, UK. [Google Scholar] [CrossRef]
- Anthropic. (2024). The Claude 3 model family: Opus, sonnet, haiku. Available online: https://www.anthropic.com/claude-3-model-card (accessed on 24 December 2025).
- Asseri, B., Abaker, E., Al Mogren, M., Alhefdhi, T., & Al-Wabil, A. (2025). Deciphering emotions in children’s storybooks: A comparative analysis of multimodal LLMs in educational applications. AI, 6(9), 211. [Google Scholar] [CrossRef]
- Bhaskaran, M., & Pardos, Z. A. (2025, July 21–23). Automating academic transcript evaluation: A comparative study of OCR techniques for course and grade evaluation. ACM Conference on Learning and Scale (pp. 366–370), Palermo, Italy. [Google Scholar] [CrossRef]
- Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., Brynjolfsson, E., Buch, S., Card, D., Castellon, R., Chatterji, N., Chen, A., Creel, K., Davis, J. Q., Demszky, D., … Liang, P. (2022). On the opportunities and risks of foundation models. arXiv. [Google Scholar] [CrossRef]
- Bossema, M., Ben Allouch, S., Plaat, A., & Saunders, R. (2025, August 25–29). LLM-enhanced interactions in human-robot collaborative drawing with older adults. IEEE International Conference on Robot and Human Interactive Communication (pp. 700–707), Eindhoven, The Netherlands. [Google Scholar] [CrossRef]
- Busic, B., Leventic, H., Romic, K., & Habijan, M. (2024, September 16–18). Towards using multimodal LLMs as graders in a GUI design course. International Symposium ELMAR (pp. 97–100), Zadar, Croatia. [Google Scholar] [CrossRef]
- Cao, Y., Xiong, X., Shao, X., Chen, R., Hou, Y., Li, B., Zhao, P., & Guo, K. (2025, February 21–23). Research on teaching video monitoring platform based on large language model prompt engineering. International Conference on Computer Science, Engineering, and Education (pp. 118–125), Nanjing, China. [Google Scholar] [CrossRef]
- Chen, J., Atmosukarto, I., & Bin Abbas, M. F. (2024, December 1–4). Image question-distractors generation as a conversational model. IEEE Region 10 Annual International Conference (pp. 39–42), Singapore. [Google Scholar] [CrossRef]
- Chen, L., Chen, P., & Lin, Z. (2020). Artificial Intelligence in education: A review. IEEE Access, 8, 75264–75278. [Google Scholar] [CrossRef]
- Comanici, G., Bieber, E., Schaekermann, M., Pasupat, I., Sachdeva, N., Dhillon, I., Blistein, M., Ram, O., Zhang, D., Rosen, E., Marris, L., Petulla, S., Gaffney, C., Aharoni, A., Lintz, N., Pais, T. C., Jacobsson, H., Szpektor, I., Jiang, N.-J., … Helmholz, W. (2025). Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities. Available online: https://arxiv.org/abs/2507.06261 (accessed on 24 December 2025).
- Dang, B., Huynh, L., Gul, F., Rose, C. P., Jarvela, S. M., & Nguyen, A. (2025). Human–AI collaborative learning in mixed reality: Examining the cognitive and socio-emotional interactions. British Journal of Educational Technology, 56(5), 2078–2101. [Google Scholar] [CrossRef]
- Danish, S., Sadeghi-Niaraki, A., Khan, S. U., Dang, L. M., Tightiz, L., & Moon, H. (2026). A comprehensive survey of Vision–Language Models: Pretrained models, fine-tuning, prompt engineering, adapters, and benchmark datasets. Information Fusion, 126, 103623. [Google Scholar] [CrossRef]
- Edwards, K. M., Tehranchi, F., Miller, S. R., & Ahmed, F. (2025, August 17–20). AI judges in design: Statistical perspectives on achieving human expert equivalence with vision-language models. International Design Engineering Technical Conferences and Computers and Information in Engineering Conference, Anaheim, CA, USA. [Google Scholar] [CrossRef]
- Fahmi, Y., & Bousmah, M. (2025, October 4–10). Designing an educational virtual assistant based on agentic AI for student diagnosis and individualized remediation. IEEE Congress on Information Science and Technology (pp. 406–415), Marrakech, Morocco. [Google Scholar] [CrossRef]
- Han, K., Tang, K., & Wang, M. (2025, March 4–7). Stage wizard: Enhancing tangible storytelling with multimodal LLMs. International Conference on Tangible, Embedded, and Embodied Interaction (pp. 1–13), Bordeaux/Talence Colorado, France. [Google Scholar] [CrossRef]
- Hang, C. N., & Man Ho, S. (2025, March 15). Personalized vocabulary learning through images: Harnessing multimodal large language models for early childhood education. IEEE Integrated STEM Education Conference (pp. 1–7), Princeton, NJ, USA. [Google Scholar] [CrossRef]
- Hong, Q. N., Pluye, P., Fabregues, S., Bartlett, G., Boardman, F., Cargo, M., Dagenais, P., Gagnon, M., Griffiths, F., Nicolau, B., O’Cathain, A., Rousseau, M., & Vedel, I. (2018). Mixed methods appraisal tool (mmat), version 2018. Available online: https://medschool.cuanschutz.edu/docs/librariesprovider94/di-docs/methods-%28design%29-docx/mmat_2018_criteria-manual_2018-08-01_eng.pdf (accessed on 24 December 2025).
- Ibanez, M. B., Barron-Estrada, M. L., & Zatarain-Cabada, R. (2025). Can multimodal large language models grade like an expert? A study on UML class diagram assessment accuracy. Computer Applications in Engineering Education, 33(5), e70080. [Google Scholar] [CrossRef]
- Kostopoulos, G., Vasileios, G., Rigou, M., & Kotsiantis, S. B. (2025). Agentic AI in education: State of the art and future directions. IEEE Access, 13, 177467–177491. [Google Scholar] [CrossRef]
- Kunuku, M. T., & Dehbozorgi, N. (2025, July 22–26). Exploring multimodal quiz generation and evaluation aligned with higher-order learning objectives in bloom’s taxonomy. International Conference on Artificial Intelligence in Education (pp. 433–438), Palermo, Italy. [Google Scholar] [CrossRef]
- Küchemann, S., Avila, K. E., Dinc, Y., Hortmann, C., Revenga, N., Ruf, V., Stausberg, N., Steinert, S., Fischer, F., Fischer, M. R., Kasneci, E., Kasneci, G., Kuhr, T., Kutyniok, G., Malone, S., Sailer, M., Schmidt, A., Stadler, M. J., Weller, J., & Kuhn, J. (2025). On opportunities and challenges of large multimodal foundation models in education. npj Science of Learning, 10(1), 11. [Google Scholar] [CrossRef]
- Lee, G., Shi, L., Latif, E., Gao, Y., Bewersdorff, A., Nyaaba, M., Guo, S., Liu, Z., Mai, G., Liu, T., & Zhai, X. (2025). Multimodality of AI for education: Toward artificial general intelligence. IEEE Transactions on Learning Technologies, 18, 666–683. [Google Scholar] [CrossRef]
- Lee, G. G., & Zhai, X. (2025). Realizing visual question answering for education: GPT-4V as a multimodal AI. TechTrends, 69(2), 271–287. [Google Scholar] [CrossRef]
- Lee, H. Y., Huang, Y. M., & Wu, T. T. (2025). ChatGPT in education: A systematic review of current landscape, limitations and future directions through general system theory lens. European Journal of Education, 60(4), e70262. [Google Scholar] [CrossRef]
- Lee, J., Chen, S. S., & Liang, P. P. (2025, April 26–May 1). Interactive sketchpad: A multimodal tutoring system for collaborative, visual problem-solving. CHI Conference on Human Factors in Computing Systems (pp. 1–14), Yokohama, Japan. [Google Scholar] [CrossRef]
- Lee, U., Jeon, M., Lee, Y., Byun, G., Son, Y., Shin, J., Ko, H., & Kim, H. (2024). LLaVA-docent: Instruction tuning with multimodal large language model to support art appreciation education. Computers and Education: Artificial Intelligence, 7, 100297. [Google Scholar] [CrossRef]
- Liu, H., Li, C., Li, Y., & Lee, Y.-J. (2024, June 16–22). Improved baselines with visual instruction tuning. IEEE Conference on Computer Vision and Pattern Recognition (pp. 26286–26296), Seattle, WA, USA. [Google Scholar] [CrossRef]
- Marquez-Carpintero, L., Viejo, D., & Cazorla, M. (2025). Enhancing engineering and STEM education with vision and multimodal large language models to predict student attention. IEEE Access, 13, 114681–114695. [Google Scholar] [CrossRef]
- Mittal, M., Tyagi, G., Bailey, A., Ranade, G. V., & Norouzi, N. (2025, July 22–26). Askademia: A real-time AI system for automatic responses to student questions. International Conference on Artificial Intelligence in Education (pp. 105–118), Palermo, Italy. [Google Scholar] [CrossRef]
- Ng, D. T. K., Chan, E. K. C., & Lo, C. K. (2025). Opportunities, challenges and school strategies for integrating generative AI in education. Computers and Education: Artificial Intelligence, 8, 100373. [Google Scholar] [CrossRef]
- Nguyen, H., & Hayward, J. (2025). Applying generative artificial intelligence to critiquing science assessments. Journal of Science Education and Technology, 34(1), 199–214. [Google Scholar] [CrossRef]
- Nguyen, H., & Park, S. (2025, March 3–7). Providing automated feedback on formative science assessments: Uses of multimodal large language models. International Conference on Learning Analytics and Knowledge (pp. 803–809), Dublin, Ireland. [Google Scholar] [CrossRef]
- Page, M. J., McKenzie, J. E., Bossuyt, P. M., Boutron, I., Hoffmann, T. C., Mulrow, C. D., Shamseer, L., Tetzlaff, J. M., Akl, E. A., Brennan, S. E., Chou, R., Glanville, J., Grimshaw, J. M., Hróbjartsson, A., Lalu, M. M., Li, T., Loder, E. W., Mayo-Wilson, E., McDonald, S., … Moher, D. (2021). The PRISMA 2020 statement: An updated guideline for reporting systematic reviews. Systematic Reviews, 10, 89. [Google Scholar] [CrossRef]
- Paivio, A. (2013). Imagery and verbal processes. Taylor & Francis. [Google Scholar] [CrossRef]
- Pang, L., Siu, T., Alazzawe, A., Kant, K., & Latecki, L. J. (2026, September 22–25). Generalizable detection of student engagement in online learning environments. International Conference on Computer Analysis of Images and Patterns (pp. 207–219), Las Palmas de Gran Canaria, Spain. [Google Scholar] [CrossRef]
- Picard, C., Edwards, K. M., Doris, A. C., Man, B., Giannone, G., Alam, M. F., & Ahmed, F. (2025). From concept to manufacturing: Evaluating vision-language models for engineering design. Artificial Intelligence Review, 58(9), 288. [Google Scholar] [CrossRef]
- Radford, A., Kim, J.-W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., & Sutskever, I. (2021, July 18–24). Learning transferable visual models from natural language supervision. International Conference on Machine Learning (pp. 8748–8763), Virtual. Available online: https://proceedings.mlr.press/v139/radford21a.html (accessed on 25 November 2025).
- Rahmanian, M., Sami, A., & Yu, Y. (2025). Challenges and feasibility of multimodal LLMs in ER diagram evaluation. Cogent Education, 12(1), 2590901. [Google Scholar] [CrossRef]
- Raihan, M. N., Siddiq, M. L., Santos, J. C., & Zampieri, M. (2025, February 26–March 1). Large language models in computer science education: A systematic literature review. ACM Technical Symposium on Computer Science Education (Vol. 1, pp. 938–944), Pittsburgh, PA, USA. [Google Scholar] [CrossRef]
- Sheng, Z., Shen, S., Shen, L., Duan, Q., Tang, N., Hui, P., Qu, H., & Luo, Y. (2025, July 22–26). Automatic modeling and analysis of students’ problem-solving handwriting trajectories. International Conference on Artificial Intelligence in Education (pp. 221–235), Palermo, Italy. [Google Scholar] [CrossRef]
- Shinde, G., Ravi, A., Dey, E., Sakib, S., Rampure, M., & Roy, N. (2025). A survey on efficient vision-language models. Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery, 15(3), e70036. [Google Scholar] [CrossRef]
- Shu, C., Yao, N., Chen, Y., Wijeratne, V., Ma, L., Loo, J. K. K., Chai, K. K., Alam, A. S., & Abuelmaatti, A. (2025, April 22–25). AI-assisted multiple-choice questions generation with multimodal large language models in engineering higher education. IEEE Global Engineering Education Conference (pp. 1–9), London, UK. [Google Scholar] [CrossRef]
- Singh, J., Zouhar, V., & Sachan, M. (2023, December 6–10). Enhancing textbooks with visuals from the web for improved learning. International Conference on Empirical Methods in Natural Language Processing (pp. 11931–11944), Singapore. [Google Scholar] [CrossRef]
- Stamatakis, M., Berger, J., Wartena, C., Ewerth, R., & Hoppe, A. (2025, July 22–26). Enhancing the learning experience: Using vision-language models to generate questions for educational videos. International Conference on Artificial Intelligence in Education (pp. 305–319), Palermo, Italy. [Google Scholar] [CrossRef]
- Su, B., Chen, Q., Peng, J., Tan, W., & Wang, L. (2025, May 14–16). Enhancing chinese character writing learning: The role of MLLM-based intelligent tutoring systems. International Conference on Artificial Intelligence and Education (pp. 194–200), Suzhou, China. [Google Scholar] [CrossRef]
- Tan, H., Gu, Y., Li, L., Leong, M. C., & Chen, N. F. (2025, October 13–17). Contextualized visual storytelling for conversational chatbot in education. International Conference on Multimodal Interaction (pp. 185–189), Canberra, Australia. [Google Scholar] [CrossRef]
- Tapia-Mandiola, S., & Araya, R. (2025, June 25–27). From palette to reasoning: Improving LLM’s visual recognition capabilities in children’s coloring tasks. International Conference in Methodologies and intelligent Systems for Techhnology Enhanced Learning (pp. 172–183), Lille, France. [Google Scholar] [CrossRef]
- Teotia, J., Zhang, X., Mao, R., & Cambria, E. (2024, December 9). Evaluating vision language models in detecting learning engagement. IEEE International Conference on Data Mining (pp. 496–502), Abu Dhabi, United Arab Emirates. [Google Scholar] [CrossRef]
- Tian, J. (2025). Integrating artificial intelligence into the cybersecurity curriculum in higher education: A systematic literature review. Education Sciences, 15(11), 1540. [Google Scholar] [CrossRef]
- Tschope, M., Fritsch, S. G., Fortes Rey, V., Nandurkar, N. N., Trevenna, S., Monger, E. J., & Lukowicz, P. (2025, April 21–25). NEEDLE: Nurse education enhanced by vision-based deep learning evaluation. International Conference on Activity and Behavior Computing (pp. 1–10), Al Ain, United Arab Emirates. [Google Scholar] [CrossRef]
- Wang, P., Bai, S., Tan, S., Wang, S., Fan, Z., Bai, J., Chen, K., Liu, X., Wang, J., Ge, W., Fan, Y., Dang, K., Du, M., Ren, X., Men, R., Liu, D., Zhou, C., Zhou, J., & Lin, J. (2024). Qwen2-VL: Enhancing vision-language model’s perception of the world at any resolution. Available online: https://arxiv.org/abs/2409.12191 (accessed on 25 November 2025).
- Wang, P., Jing, Y., & Shen, S. (2025). A systematic literature review on the application of generative artificial intelligence (GAI) in teaching within higher education: Instructional contexts, process, and strategies. The Internet and Higher Education, 65, 100996. [Google Scholar] [CrossRef]
- Wang, X., Pang, C. C., & Hui, P. (2025, September 28–October 1). Talking spell: A wearable system enabling real-time anthropomorphic voice interaction with everyday objects. ACM Symposium on User Interface Software and Technology, Busan, Republic of Korea. [Google Scholar] [CrossRef]
- Wang, X., Yu, R., Zhang, Y., & Xu, Y. (2025, November 1–2). English composition image automatic scoring based on multi-modal large language models. International Conference on Artificial Intelligence and Future Education (pp. 247–254), Shanghai, China. [Google Scholar] [CrossRef]
- Wang, Y., Yu, J., Zhang-Li, D., Lim, J. J. Y., Tu, S., Li, H., Liu, Z., Liu, H., Hou, L., Li, J., & Xu, B. (2025, November 10–14). EduCraft: A system for generating pedagogical lecture scripts from long-context multimodal presentations. ACM International Conference on Information and Knowledge Management (pp. 6153–6160), Seoul, Republic of Korea. [Google Scholar] [CrossRef]
- Xie, T., Wang, X., & Li, J. (2025, July 11–13). A study on prompt engineering for K12 exam paper correction using Qwen2.5-VL-72B-Instruct. International Conference on Educational Knowledge and Informatization (pp. 138–142), Chongqing, China. [Google Scholar] [CrossRef]
- Xing, W., Zhu, T., Wang, J., & Liu, B. (2024). A survey on MLLMs in education: Application and future directions. Future Internet, 16(12), 467. [Google Scholar] [CrossRef]
- Yang, Z., Li, L., Lin, K., Wang, J., Lin, C.-C., Liu, Z., & Wang, L. (2023). The dawn of LMMs: Preliminary explorations with GPT-4V(ision). Available online: https://arxiv.org/abs/2309.17421 (accessed on 25 November 2025).
- Zawacki-Richter, O., Marín, V. I., Bond, M., & Gouverneur, F. (2019). Systematic review of research on artificial intelligence applications in higher education—Where are the educators? International Journal of Educational Technology in Higher Education, 16(1), 39. [Google Scholar] [CrossRef]
- Zhan, Z., Tong, Y., Lan, X., & Zhong, B. (2024). A systematic literature review of game-based learning in Artificial Intelligence education. Interactive Learning Environments, 32(3), 1137–1158. [Google Scholar] [CrossRef]
- Zhang, J., Huang, J., Jin, S., & Lu, S. (2024). Vision-Language models for vision tasks: A survey. IEEE Transactions on Pattern Analysis and Machine Intelligence, 46(8), 5625–5644. [Google Scholar] [CrossRef]
- Zheng, C., Yu, Z., Jiang, Y., Zhang, M., Lu, X., Jin, J., & Gao, L. (2025, April 26–May 1). ArtMentor: AI-assisted evaluation of artworks to explore multimodal large language models capabilities. CHI Conference on Human Factors in Computing Systems, Yokohama, Japan. [Google Scholar] [CrossRef]
- Zhou, K., Liu, Z., & Gao, P. (2025). Large vision-language models: Pre-training, prompting, and applications. Springer Nature. [Google Scholar]
- Zhuang, Y., Zhao, R., Xie, Z., & Yu, P. L. H. (2025). Enhancing language learning through generative AI feedback on picture-cued writing tasks. Computers and Education: Artificial Intelligence, 9, 100450. [Google Scholar] [CrossRef]



| VLM Solution | Year | Model Weights | Open-Weight or Proprietary | Reference |
|---|---|---|---|---|
| CLIP | 2021 | 63 M–355 M | Open-weight | Radford et al. (2021) |
| LLaVa-v1.5 | 2023 | 13 B | Open-weight | Liu et al. (2024) |
| Qwen2-VL | 2024 | 7 B–14 B | Open-weight | P. Wang et al. (2024) |
| GPT-4V | 2023 | − | Proprietary | Yang et al. (2023) |
| Claude-3 | 2024 | − | Proprietary | Anthropic (2024) |
| Gemini 2.5 | 2025 | − | Proprietary | Comanici et al. (2025) |
| Reference | Year | LLM or VLM | Sources | Period Covered |
|---|---|---|---|---|
| Ali et al. (2024) | 2024 | LLM | Academic Search Premier Web of Science, IEEE Xplore | 30 November 2022– 12 October 2023 |
| Agbo et al. (2025) | 2025 | LLM | ACM Digital Library Scopus, Web of Science IEEE Xplore | up to 13 May 2024 |
| Kostopoulos et al. (2025) | 2025 | LLM | Scopus, Web of Science IEEE Xplore, ACM Digital Library | January 2015– August 2025 |
| H. Y. Lee et al. (2025) | 2025 | LLM | Web of Science, ScienceDirect IEEE Xplore, ACM Digital Library Taylor and Francis, Scopus, Wiley | November 2022– March 2024 |
| Raihan et al. (2025) | 2025 | LLM | ACM Digital Library, Scopus IEEE Xplore, ACL Anthology Web of Science, Springer Link ScienceDirect, ArXiv | January 2019– June 2024 |
| P. Wang et al. (2025) | 2025 | LLM | Web of Science EBSCO, Scopus | November 2022– August 2024 |
| Ours | VLM | ACM Digital Library, Scopus Web of Science, IEEE Xplore Engineering Village | January 2020– December 2025 |
| Database | Search Syntax |
|---|---|
| Scopus | (TITLE-ABS-KEY (“VLM *” OR “vision language model *” OR “MLLM *” OR “multimodal large language model *”) AND TITLE-ABS-KEY (“teaching” OR “education” OR “classroom” OR “course”)) AND PUBYEAR > 2019 AND PUBYEAR < 2027 AND (LIMIT-TO (DOCTYPE, “ar”) OR LIMIT-TO (DOCTYPE, “cp”) ) AND (LIMIT-TO (LANGUAGE, “English”)) |
| Web of Science | ((AB = ((“VLM *” OR “vision language model *” OR “MLLM *” OR “multimodal large language model *”) AND (“teaching” OR “education” OR “classroom” OR “course”))) AND LA = (English)) AND DT = (Article OR Proceedings Paper) AND PY = 2020–2026 |
| IEEE Xplore | (“Abstract”: “VLM *” OR “Abstract”: “vision language model *” OR “Abstract”: “MLLM *” OR “Abstract”: “multimodal large language model *”) AND (“Abstract”: “teaching” OR “Abstract”: “education” OR “Abstract”: “course” OR “Abstract”: “classroom”) Filters Applied: Conferences Journals 2020–2026 |
| ACM Digital Library | [[Abstract: “VLM *”] OR [Abstract: “vision language model *”] OR [Abstract: “MLLM*”] OR [Abstract: “multimodal large language model *”]] AND [[Abstract: “teaching”] OR [Abstract: “education”] OR [Abstract: “classroom”] OR [Abstract: “course”]] AND [E-Publication Date: (1 January 2020 TO 31 December 2026)] |
| Engineering Village | (((“VLM *” OR “vision language model *” OR “MLLM *” OR “multimodal large language model *”) WN AB) AND ((“teaching” OR “education” OR “classroom” OR “course”) WN AB)) AND (English WN LA) + (ca OR ja) WN DT |
| Inclusion | Peer-reviewed articles. |
| criteria | Research articles, such as journal and conference articles. |
| Full texts are available. | |
| Published in the English language. | |
| Published between 2020 and 2025. | |
| Exclusion | Non-peer-reviewed publications, |
| criteria | Not research articles, such as tutorials, abstracts, book chapters, etc. |
| Articles not directly related to teaching and learning. | |
| Articles solely on educational data evaluation rather than enhancing teaching and learning. |
| Year | 2023 | 2024 | 2025 | Total |
|---|---|---|---|---|
| Journal articles | 0 | 1 | 9 | 10 |
| Conference articles | 1 | 4 | 27 | 32 |
| Reference | Education Level | Discipline |
|---|---|---|
| Bossema et al. (2025) | Adult learner | Human–robot collaborative creativity |
| X. Wang et al. (2025) | Adult learner | − |
| Su et al. (2025) | College | Chinese language |
| Hang and Man Ho (2025) | Early childhood | Vocabulary learning |
| Nguyen and Park (2025) | Grade 6 | Science |
| Sheng et al. (2025) | High school | Mathematics, Physics, Chemistry |
| Fahmi and Bousmah (2025) | K-12 | − |
| Han et al. (2025) | K-12 | − |
| Nguyen and Hayward (2025) | K-12 | Science |
| Tapia-Mandiola and Araya (2025) | K-12 | Coloring activity |
| Xie et al. (2025) | K-12 | Chinese language, Mathematics, English, Physics, Chemistry, Biology, Geography, History |
| J. Chen et al. (2024) | Primary school | English language |
| Tan et al. (2025) | Primary school | Chinese language |
| Zhuang et al. (2025) | Secondary school | Picture-cued writing |
| Ibanez et al. (2025) | Undergraduate | Engineering |
| Kunuku and Dehbozorgi (2025) | Undergraduate | Computer science course |
| Picard et al. (2025) | Undergraduate | Engineering design |
| Rahmanian et al. (2025) | Undergraduate | Software engineering |
| Shu et al. (2025) | Undergraduate | Cryptography and Cyber Security, Software Engineering, Communications and Networks, Broadband and Fibre Optics |
| Anderer et al. (2024) | University | Psychology, Computer science, Climate science, Game design |
| Anderer et al. (2025) | University | − |
| Bhaskaran and Pardos (2025) | University | − |
| Busic et al. (2024) | University | UX design |
| Cao et al. (2025) | University | − |
| Dang et al. (2025) | University | − |
| J. Lee et al. (2025) | University | Math |
| Mittal et al. (2025) | University | Data science course |
| Y. Wang et al. (2025) | University | − |
| Edwards et al. (2025) | − | Engineering design |
| U. Lee et al. (2024) | − | Art education |
| Marquez-Carpintero et al. (2025) | − | STEM, Engineering courses |
| Singh et al. (2023) | − | Math, science, Social science, Business |
| Tschope et al. (2025) | − | Nursing education |
| X. Wang et al. (2025) | − | English language |
| Zheng et al. (2025) | − | Art |
| Reference | Input Modality | Output Modality |
|---|---|---|
| Abdelhadi et al. (2025) | Video | Student engagement |
| Ambali Parambil et al. (2025) | Online class image | Emotion |
| Anderer et al. (2024) | Visual slides | Text descriptions |
| Anderer et al. (2025) | Lecture video, questions | Text answers |
| Asseri et al. (2025) | Story book | Emotion |
| Bhaskaran and Pardos (2025) | Scanned transcript | OCR and grade |
| Bossema et al. (2025) | Drawing picture | Description and co-drawing |
| Busic et al. (2024) | GUI design | Grade |
| Cao et al. (2025) | Online teaching video | Teaching content and behavior |
| J. Chen et al. (2024) | Cartoon image | Quiz questions |
| Dang et al. (2025) | Human–AI interaction | Interaction type |
| Edwards et al. (2025) | Design image, description | Grade |
| Fahmi and Bousmah (2025) | Students’ hand writing | Feedback and grade |
| Han et al. (2025) | Image and story | Design and story |
| Hang and Man Ho (2025) | Text prompts | Flash card images |
| Ibanez et al. (2025) | Diagram drawing | Grade |
| Kunuku and Dehbozorgi (2025) | Video lectures | Quiz questions |
| U. Lee et al. (2024) | Art image and dialogue | Feedback |
| G. G. Lee and Zhai (2025) | Educational image | Text answers |
| J. Lee et al. (2025) | Text and image | Visualizations with textual hints |
| Marquez-Carpintero et al. (2025) | Image | Emotion and attention |
| Mittal et al. (2025) | Lecture audio, materials | Text answers |
| Nguyen and Hayward (2025) | Text and image | Feedback |
| Nguyen and Park (2025) | Assessment image | Feedback |
| Pang et al. (2026) | Online class image | Student engagement |
| Picard et al. (2025) | Class materials | Text answers |
| Rahmanian et al. (2025) | Diagram drawing | Grade |
| Sheng et al. (2025) | Handwriting trajectory | Problem solving type |
| Shu et al. (2025) | Notes | Quiz questions |
| Singh et al. (2023) | Text from textbooks | Matched images |
| Stamatakis et al. (2025) | Educational videos | Text questions |
| Su et al. (2025) | Writing pictures | Feedback |
| Tan et al. (2025) | Image | Story telling |
| Tapia-Mandiola and Araya (2025) | Worksheet picture | Color recognition, object count |
| Teotia et al. (2024) | Classroom image | Classroom emotion behavior |
| Tschope et al. (2025) | Performance video | Action category |
| Y. Wang et al. (2025) | Slide presentations | Lecture script |
| X. Wang et al. (2025) | Composition images | Grade |
| X. Wang et al. (2025) | Video frame | Persona creation |
| Xie et al. (2025) | Exam image | Grade |
| Zheng et al. (2025) | Artwork | Feedback and evaluation |
| Zhuang et al. (2025) | Picture and writing | Feedback |
| Benefit | Reference |
|---|---|
| Engagement support | Bossema et al. (2025); Marquez-Carpintero et al. (2025); Teotia et al. (2024); X. Wang et al. (2025) |
| Feedback and Q&A | Anderer et al. (2025); Busic et al. (2024); Fahmi and Bousmah (2025); J. Lee et al. (2025); U. Lee et al. (2024); Mittal et al. (2025); Nguyen and Park (2025); Picard et al. (2025); Su et al. (2025); Zhuang et al. (2025) |
| Multimodal knowledge acquisition | Anderer et al. (2024); J. Chen et al. (2024); Han et al. (2025); Hang and Man Ho (2025); Singh et al. (2023); Stamatakis et al. (2025); Tan et al. (2025); Y. Wang et al. (2025) |
| Benefit | Reference |
|---|---|
| Administrative efficiency | Bhaskaran and Pardos (2025); Mittal et al. (2025) |
| Assessment and feedback | Busic et al. (2024); Edwards et al. (2025); Ibanez et al. (2025); G. G. Lee and Zhai (2025); Rahmanian et al. (2025); Stamatakis et al. (2025); Tapia-Mandiola and Araya (2025); Tschope et al. (2025); X. Wang et al. (2025); Xie et al. (2025) |
| Classroom analytics | Abdelhadi et al. (2025); Cao et al. (2025); Pang et al. (2026) |
| Content authoring | J. Chen et al. (2024); Han et al. (2025); Hang and Man Ho (2025); Kunuku and Dehbozorgi (2025); Y. Wang et al. (2025) |
| Instructional support | Asseri et al. (2025); Dang et al. (2025); Fahmi and Bousmah (2025); Marquez-Carpintero et al. (2025); Nguyen and Hayward (2025); Sheng et al. (2025) |
| Reference | Dataset | Reproducibility |
|---|---|---|
| Bhaskaran and Pardos (2025) | 14 sample university transcripts | Code shared |
| J. Lee et al. (2025) | Questions from Scholastic Aptitude Test (SAT), math and geometry test | Code shared |
| Stamatakis et al. (2025) | 9280 videos curated from 2 public datasets | Code shared |
| Zheng et al. (2025) | 380 sessions from 5 art teachers | Code and dataset shared |
| Edwards et al. (2025) | 934 early design sketches | Dataset shared |
| Singh et al. (2023) | 35 textbooks from openstax | Dataset shared |
| Abdelhadi et al. (2025) | 10 scenarios | |
| Ambali Parambil et al. (2025) | 149 images | |
| Anderer et al. (2024) | 21 lectures | |
| Anderer et al. (2025) | 7 participants’ survey | |
| Asseri et al. (2025) | 75 images | |
| Bossema et al. (2025) | 18 learners | |
| Busic et al. (2024) | 10 design images | |
| Cao et al. (2025) | 6 months deployment | |
| J. Chen et al. (2024) | 10 images with 50 conversations | |
| Dang et al. (2025) | 26 students with 1317 recorded activities | |
| Han et al. (2025) | Experiments with 3 children | |
| Hang and Man Ho (2025) | 3 graphic design styles | |
| Ibanez et al. (2025) | Image submissions from 34 students | |
| Kunuku and Dehbozorgi (2025) | 15 MCQs | |
| U. Lee et al. (2024) | 1000 dialogue examples | |
| G. G. Lee and Zhai (2025) | 5 educational scenarios | |
| Marquez-Carpintero et al. (2025) | 14,425 images from 9 educational scenarios | |
| Mittal et al. (2025) | 28 lectures of 80 min each, 28 course notes documents, 430 student questions | |
| Nguyen and Hayward (2025) | 42 questions, an interview with 6 educators | |
| Nguyen and Park (2025) | 82 assessments | |
| Pang et al. (2026) | 3 publicly available datasets | |
| Picard et al. (2025) | 44 questions from 2 courses | |
| Rahmanian et al. (2025) | 40 diagrams | |
| Sheng et al. (2025) | 87,679 handwritings from 25 students | |
| Shu et al. (2025) | 66 MCQs | |
| Su et al. (2025) | 13 participants and 21 Chinese characters | |
| Tan et al. (2025) | 24 images | |
| Tapia-Mandiola and Araya (2025) | 5 worksheets | |
| Teotia et al. (2024) | 3 public emotion datasets | |
| Tschope et al. (2025) | 73 videos, 2 types of performance | |
| Y. Wang et al. (2025) | 20 presentations (320 slides) from 4 courses | |
| X. Wang et al. (2025) | 1511 images | |
| X. Wang et al. (2025) | 12 participants’ survey | |
| Xie et al. (2025) | 600 exam paper images, across 8 subjects and 6 question types | |
| Zhuang et al. (2025) | 5856 writing responses, 386 questionnaires |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the author. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Tian, J. Vision-Language Models in Teaching and Learning: A Systematic Literature Review. Educ. Sci. 2026, 16, 123. https://doi.org/10.3390/educsci16010123
Tian J. Vision-Language Models in Teaching and Learning: A Systematic Literature Review. Education Sciences. 2026; 16(1):123. https://doi.org/10.3390/educsci16010123
Chicago/Turabian StyleTian, Jing. 2026. "Vision-Language Models in Teaching and Learning: A Systematic Literature Review" Education Sciences 16, no. 1: 123. https://doi.org/10.3390/educsci16010123
APA StyleTian, J. (2026). Vision-Language Models in Teaching and Learning: A Systematic Literature Review. Education Sciences, 16(1), 123. https://doi.org/10.3390/educsci16010123

