Developing a corpus of plagiarised short answers.

Plagiarism is widely acknowledged to be a significant and increasing problem for higher education institutions (McCabe ; Judge ). A wide range of solutions, including several commercial systems, have been proposed to assist the educator in the task of identifying plagiarised work, or even to detect...

Full description

Bibliographic Details
Published in:Language Resources & Evaluation Vol. 45; no. 1; pp. 5 - 25
Main Authors: Clough, Paul, Stevenson, Mark
Format: Article
Published: Springer Nature Feb2011
Subjects:
Online Access:View this record in EBSCOhost
Description
Summary:Plagiarism is widely acknowledged to be a significant and increasing problem for higher education institutions (McCabe ; Judge ). A wide range of solutions, including several commercial systems, have been proposed to assist the educator in the task of identifying plagiarised work, or even to detect them automatically. Direct comparison of these systems is made difficult by the problems in obtaining genuine examples of plagiarised student work. We describe our initial experiences with constructing a corpus consisting of answers to short questions in which plagiarism has been simulated. This corpus is designed to represent types of plagiarism that are not included in existing corpora and will be a useful addition to the set of resources available for the evaluation of plagiarism detection systems.