Statistical quality estimation for partially subjective classification tasks through crowdsourcing.
When constructing a large-scale data resource, the quality of artifacts has great significance, especially when they are generated by creators through crowdsourcing. A widely used approach is to estimate the quality of each artifact based on evaluations by reviewers. However, the commonly used vote-...
| Publicado en: | Language Resources & Evaluation Vol. 57; no. 1; pp. 31 - 57 |
|---|---|
| Autores principales: | , |
| Formato: | Artículo |
| Publicado: |
Springer Nature
Mar2023
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=162506767&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 162506767 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Mar2023 vid: 57 iid: 1 pid: 237 pub: Springer Nature artinfo: ui: 162506767 10.1007/s10579-022-09617-0 ppf: 31 ppct: 26 formats: fmt: – @attributes: type: T – @attributes: type: P size: 1.4MB tig: atl: Statistical quality estimation for partially subjective classification tasks through crowdsourcing. aug: au: Sato, Yoshinao Miyazawa, Kouki affil: Fairy Devices Inc., Tokyo, Japan su: Crowdsourcing Latent variables Classification Speech perception sug: subj: Crowdsourcing Latent variables Classification Speech perception keyword: Latent variable model Partially subjective task Quality estimation ab: When constructing a large-scale data resource, the quality of artifacts has great significance, especially when they are generated by creators through crowdsourcing. A widely used approach is to estimate the quality of each artifact based on evaluations by reviewers. However, the commonly used vote-counting method to aggregate reviewers' evaluations does not work effectively for partially subjective tasks. In such a task, a single correct answer cannot necessarily be defined. We propose a statistical quality estimation method for partially subjective classification tasks to infer the quality of artifacts considering the abilities and biases of creators and reviewers as latent variables. In our experiments, we use the partially subjective task of classifying speech into one of the following four attitudes: agreement, disagreement, stalling, and question. We collect a speech corpus through crowdsourcing and apply the proposed method to it. The results show that the proposed method estimates the quality of speech more effectively than vote aggregation, as measured by correlation with a fine-grained classification performed by experts. Furthermore, we compare the speech attitude classification performance of a neural network model on two subsets of our corpus extracted using the voting and proposed methods. The results indicate that we can effectively extract a consistent and high-quality subset of a corpus using the proposed method. This method facilitates the efficient collection of large-scale data resources for mutually exclusive classification, even if the task is partially subjective. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2023. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2023 holdings: @attributes: islocal: N |
|---|