Leveraging Multimodal Large Language Models to Analyse Student Exploration Behaviours in Educational Game Environments.

Background: Video data provides rich opportunities to examine student behaviour in game‐based learning environments, capturing not only observable actions but also subtle indicators of cognitive engagement. However, traditional video analysis is labor‐intensive, and current applications of AI to thi...

Descripción completa

Detalles Bibliográficos
Publicado en:Journal of Computer Assisted Learning Vol. 42; no. 4; pp. 1 - 17
Autores principales: Zhou, Yiqiu, Liu, Xiner, Zambrano, Andres Felipe, Paquette, Luc, Barany, Amanda, Ocumpaugh, Jaclyn, Baker, Ryan S., Ginger, Jeff
Formato: pictorial research tables/charts Journal Article
Publicado: Wiley-Blackwell Aug2026
Acceso en línea:Ver este registro en EBSCOhost
Descripción
Sumario:Background: Video data provides rich opportunities to examine student behaviour in game‐based learning environments, capturing not only observable actions but also subtle indicators of cognitive engagement. However, traditional video analysis is labor‐intensive, and current applications of AI to this task are limited by models trained on general‐purpose datasets, which often fail to capture the pedagogical meaning of student actions in authentic educational contexts. Objectives: This study explores how multimodal large language models (MLLMs), specifically LLaVA‐Video‐7B‐Qwen2, can support qualitative video analysis of student exploration behaviours in a Minecraft‐based STEM learning environment. Methods: We conducted an exploratory case study using screen recordings from a Minecraft‐based STEM learning environment. We tested multiple prompt strategies for guiding MLLM‐generated video descriptions and found that role‐assignment prompting performed most effectively. We evaluated model outputs using a mixed‐method framework that included quantitative scoring, GPT‐based judgement, and researcher validation. Results and Conclusions: Our findings show that MLLMs can reliably identify surface‐level behaviours, such as navigation patterns and object interactions, but struggle to infer the intent or goals behind student actions, leading to a significant rate of over‐interpretation (26.5% when explaining student strategies). The model's outputs are sensitive to prompt phrasing, underscoring the importance of prompt engineering. While current MLLMs show promise for streamlining parts of the video analysis workflow, their use in educational contexts requires structured oversight and careful interpretation to ensure reliability and relevance. Lay Summary: What is currently known about this topic? ○Video analysis is widely used in educational research but is time‐consuming and resource‐intensive.○Multimodal large language models (MLLMs) have shown potential in analysing video content but have mostly been evaluated on general tasks rather than educationally meaningful behaviours.What does this paper add? ○It demonstrates how an MLLM (LLaVA‐Video‐7B‐Qwen2) performs when analyzing gameplay screen captures from an educational context.○It introduces a structured approach to prompt design and evaluation for using MLLMs in educational video analysis.