An evaluation of classification models for question topic categorization.
We study the problem of question topic classification using a very large real-world Community Question Answering ( CQA) dataset from Yahoo! Answers. The dataset comprises 3.9 million questions and these questions are organized into more than 1,000 categories in a hierarchy. To the best knowledge, th...
| Publicado en: | Journal of the American Society for Information Science & Technology Vol. 63; no. 5; pp. 889 - 904 |
|---|---|
| Autores principales: | , , , , |
| Formato: | pictorial research tables/charts Journal Article |
| Publicado: |
Wiley-Blackwell
May2012
|
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=104560870&site=ehost-live header: @attributes: shortDbName: ccm uiTerm: 104560870 longDbName: CINAHL Complete uiTag: AN controlInfo: bkinfo: dissinfo: jinfo: jid: 15322882 IGD jtl: Journal of the American Society for Information Science & Technology issn: 15322882 maglogo: Y pubinfo: dt: May2012 vid: 63 iid: 5 pid: 480 pub: Wiley-Blackwell place: Malden, Massachusetts artinfo: ui: 104560870 74751477 10.1002/asi.22611 104560870 ppf: 889 ppct: 15 formats: tig: atl: An evaluation of classification models for question topic categorization. aug: au: Qu, Bo Cong, Gao Li, Cuiping Sun, Aixin Chen, Hong affil: asi22611-aff-0001 sug: subj: Classification Methods Information Services Online Services Human Internet Information Retrieval ab: We study the problem of question topic classification using a very large real-world Community Question Answering ( CQA) dataset from Yahoo! Answers. The dataset comprises 3.9 million questions and these questions are organized into more than 1,000 categories in a hierarchy. To the best knowledge, this is the first systematic evaluation of the performance of different classification methods on question topic classification as well as short texts. Specifically, we empirically evaluate the following in classifying questions into CQA categories: (a) the usefulness of n-gram features and bag-of-word features; (b) the performance of three standard classification algorithms (naive Bayes, maximum entropy, and support vector machines); (c) the performance of the state-of-the-art hierarchical classification algorithms; (d) the effect of training data size on performance; and (e) the effectiveness of the different components of CQA data, including subject, content, asker, and the best answer. The experimental results show what aspects are important for question topic classification in terms of both effectiveness and efficiency. We believe that the experimental findings from this study will be useful in real-world classification problems. pubtype: Academic Journal doctype: pictorial research tables/charts Journal Article ougenre: Article language: English refInfo: holdings: @attributes: islocal: N |
|---|