Evaluation of a rule-based approach to automatic factual question generation using syntactic and semantic analysis.

We present a rule-based approach to automatic factual question generation implemented in the Adaptive Courseware and Natural Language Tutor, a natural language-based intelligent tutoring system. Since machine-generated questions are intended for adaptive teaching, learning and assessment, their accu...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 57; no. 4; pp. 1431 - 1462
Autores principales: Gašpar, Angelina, Grubišić, Ani, Šarić-Grgić, Ines
Formato: Artículo
Publicado: Springer Nature Dec2023
Materias:
Acceso en línea:Ver este registro en EBSCOhost
Descripción
Sumario:We present a rule-based approach to automatic factual question generation implemented in the Adaptive Courseware and Natural Language Tutor, a natural language-based intelligent tutoring system. Since machine-generated questions are intended for adaptive teaching, learning and assessment, their accuracy is of the utmost importance. However, the generation of high-quality questions is still challenging. The proposed approach relies on pre-processing techniques and syntactic and semantic feature extraction to transform declarative sentences and their segments into questions. The quality of questions, generated from domain specific texts, was evaluated by using mixed evaluation strategies: (1) human evaluation, (2) qualitative error analysis, (3) automatic evaluation, (4) human and automatic evaluation of machine-generated questions from paraphrases compared to a set of human-authored questions, (5) preliminary comparison to other approaches. The human evaluation involved two teachers of English as a foreign language who set up evaluation criteria (grammaticality, semantic accuracy, and answerability) and a group of 30 English language graduates. Student-generated questions were validated and used as reference questions for automatic evaluation based on similarity metrics (BLEU-4, METEOR, CHRF, NIST and ROUGE-L). Human and automatic evaluation results were satisfactory but improved significantly with the paraphrasing strategy. The preliminary comparison to other approaches showed that the proposed rule-based approach performed equally well despite its limitations.