Semi‐automatic coding of open‐ended text responses in large‐scale assessments.
Background: In the context of large‐scale educational assessments, the effort required to code open‐ended text responses is considerably more expensive and time‐consuming than the evaluation of multiple‐choice responses because it requires trained personnel and long manual coding sessions. Aim: Our...
| Publicado en: | Journal of Computer Assisted Learning Vol. 39; no. 3; pp. 841 - 855 |
|---|---|
| Autores principales: | , , |
| Formato: | equations & formulas research tables/charts Journal Article |
| Publicado: |
Wiley-Blackwell
Jun2023
|
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=163886535&site=ehost-live header: @attributes: shortDbName: ccm uiTerm: 163886535 longDbName: CINAHL Complete uiTag: AN controlInfo: bkinfo: dissinfo: jinfo: jid: 02664909 6M1 jtl: Journal of Computer Assisted Learning issn: 02664909 maglogo: Y pubinfo: dt: Jun2023 vid: 39 iid: 3 pid: 480 pub: Wiley-Blackwell place: Malden, Massachusetts artinfo: ui: 163886535 159017036 163886535 163886535 10.1111/jcal.12717 163886535 ppf: 841 ppct: 14 formats: tig: atl: Semi‐automatic coding of open‐ended text responses in large‐scale assessments. aug: au: Andersen, Nico Zehner, Fabian Goldhammer, Frank affil: DIPF Leibniz Institute for Research and Information in Education, Frankfurt am Main, Germany sug: subj: Coding, Computer-Assisted Methods Automation Educational Measurement Natural Language Processing Computer-Assisted Instruction Semantics Uncertainty Confidence Minimum Data Set Students, Foreign Student Performance Appraisal Simulations Workload Measurement Descriptive Statistics Human ab: Background: In the context of large‐scale educational assessments, the effort required to code open‐ended text responses is considerably more expensive and time‐consuming than the evaluation of multiple‐choice responses because it requires trained personnel and long manual coding sessions. Aim: Our semi‐supervised coding method eco (exploring coding assistant) dynamically supports human raters by automatically coding a subset of the responses. Method: We map normalized response texts into a semantic space and cluster response vectors based on their semantic similarity. Assuming that similar codes represent semantically similar responses, we propagate codes to responses in optimally homogeneous clusters. Cluster homogeneity is assessed by strategically querying informative responses and presenting them to a human rater. Following each manual coding, the method estimates the code distribution respecting a certainty interval and assumes a homogeneous distribution if certainty exceeds a predefined threshold. If a cluster is determined to certainly comprise homogeneous responses, all remaining responses are coded accordingly automatically. We evaluated the method in a simulation using different data sets. Results: With an average miscoding of about 3%, the method reduced the manual coding effort by an average of about 52%. Conclusion: Combining the advantages of automatic and manual coding produces considerable coding accuracy and reduces the required manual effort. Lay Description: We present a new (semi‐)automatic method for coding text responses.The method supports the human coders in their work and acts as an assistant. The exploring coding assistant eco groups all responses in a first unsupervised step and explore the clusters by continuously selecting responses for manual coding to gather supervised information.Based on the coding information, eco automatically codes similar responses and reduces the manual coding effort. pubtype: Academic Journal doctype: equations & formulas research tables/charts Journal Article ougenre: Article language: English refInfo: holdings: @attributes: islocal: N |
|---|