Automatic Short‐Answer Grading in Sustainability Education: AI–Human Agreement.

Background: Sustainability education emphasises critical thinking and interdisciplinary understanding, making the assessment of students' learning outcomes complex. While Large Language Models (LLMs) have shown promise in educational assessment, their reliability in domains requiring contextual reas...

Descripción completa

Detalles Bibliográficos
Publicado en:Journal of Computer Assisted Learning Vol. 42; no. 1; pp. 1 - 16
Autores principales: Emirtekin, Emrah, Özarslan, Yasin
Formato: research tables/charts Journal Article
Publicado: Wiley-Blackwell Feb2026
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=191181612&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 191181612
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        02664909
        6M1
      jtl: Journal of Computer Assisted Learning
      issn: 02664909
      maglogo: Y
    pubinfo:
      dt: Feb2026
      vid: 42
      iid: 1
      pid: 480
      pub: Wiley-Blackwell
      place: Malden, Massachusetts
    artinfo:
      ui:
        191181612
        191181612
        191181612
        10.1002/jcal.70160
        191181612
      ppf: 1
      ppct: 15
      formats:
      tig:
        atl: Automatic Short‐Answer Grading in Sustainability Education: AI–Human Agreement.
      aug:
        au:
          Emirtekin, Emrah
          Özarslan, Yasin
        affil: Center for Distance Education Application and Research, Ege University, İzmir, Turkey
      sug:
        subj:
          Students, College
          Education, Interdisciplinary
          Environmental Sustainability Education
          Critical Thinking Evaluation
          Educational Measurement
          Artificial Intelligence Utilization
          Automation
          Reliability Evaluation
          Cognition Evaluation
          Human
          Validation Studies
          Consensus
          kappa Statistic
          Intraclass Correlation Coefficient
          Pearson's Correlation Coefficient
          Interrater Reliability
          Coefficient alpha
          Descriptive Statistics
          Confidence Intervals
          Construct Validity
          Data Analysis Software
      ab: Background: Sustainability education emphasises critical thinking and interdisciplinary understanding, making the assessment of students' learning outcomes complex. While Large Language Models (LLMs) have shown promise in educational assessment, their reliability in domains requiring contextual reasoning—such as sustainability—remains unclear. Objectives: This study aims to evaluate the agreement between human raters and several LLMs (GPT‐4o, Gemini 2.0 Flash, DeepSeek V3, LLaMA 3.3) in assessing short‐answer responses from a university‐level Sustainability course. It also investigates how this agreement varies across cognitive skill levels. Methods: A total of 232 short‐answer responses were evaluated using a rubric aligned with Bloom's Revised Taxonomy. Consensus scores from human raters were compared to LLM‐generated scores using multiple statistical measures, including Quadratic Weighted Kappa (QWK), Intraclass Correlation Coefficient (ICC), Pearson correlation, and distributional overlap. Results: Moderate agreement was found between LLMs and human raters in total scores (QWK: 0.585–0.640; r: 0.660–0.668; η̂$$ \hat{\eta} $$: 0.681–0.803). Inter‐rater reliability among humans was good to excellent (ICC: 0.667–0.800). Criterion‐level agreement declined as cognitive complexity increased, with notably low agreement on evaluating higher‐order skills. Conclusions: Overall, LLM–human agreement was moderate on total scores but declined at higher cognitive levels, indicating that LLMs are suitable for basic comprehension checks while human oversight remains necessary for complex reasoning. Practitioner Notes: What is already known about this topic ○LLMs are increasingly used in educational settings for grading and feedback.○Automatic Short‐Answer Grading (ASAG) has been widely explored in language and computer science education.○Assessing higher‐order cognitive skills remains a challenge for AI systems.What this paper adds ○This study provides empirical evidence on the performance of LLMs in evaluating short‐answer responses in sustainability education.○It highlights the gap in LLMs' reliability when assessing higher‐order skills, such as analysis and evaluation.○It shows substantial consistency among LLMs but divergence from human scoring in complex tasks.Implications for practice and/or policy ○LLMs can effectively support educators in grading foundational comprehension.○Human oversight remains critical, particularly when evaluating nuanced or interdisciplinary content.○Developing hybrid human‐AI assessment systems provides a practical framework for balancing the need for scalable assessment with the unwavering demand for educational validity.
      pubtype: Academic Journal
      doctype:
        research
        tables/charts
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N