Mining a corpus of biographical texts using keywords.
Using statistically derived keywords to characterize texts has become an important research method for digital humanists and corpus linguists in areas such as literary analysis and the exploration of genre difference. Keywords—and the associated concepts of ‘keyness’ and ‘key-keyness’—have inspired...
| Published in: | Literary & Linguistic Computing Vol. 25; no. 1; pp. 23 - 36 |
|---|---|
| Main Author: | |
| Format: | Article |
| Published: |
Oxford University Press / USA
Apr2010
|
| Subjects: | |
| Online Access: | View this record in EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=53297471&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 53297471 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 02681145 BJ1 jtl: Literary & Linguistic Computing issn: 02681145 maglogo: N pubinfo: dt: Apr2010 vid: 25 iid: 1 pid: 622 pub: Oxford University Press / USA artinfo: ui: 53297471 10.1093/llc/fqp035 ppf: 23 ppct: 13 formats: fmt: @attributes: type: P size: 361KB tig: atl: Mining a corpus of biographical texts using keywords. aug: au: Conway, Mike affil: National Institute of Informatics, Japan su: Keywords Research methodology Humanities education Corpora Forums Adult education workshops sug: subj: Keywords Research methodology Humanities education Corpora Forums Adult education workshops ab: Using statistically derived keywords to characterize texts has become an important research method for digital humanists and corpus linguists in areas such as literary analysis and the exploration of genre difference. Keywords—and the associated concepts of ‘keyness’ and ‘key-keyness’—have inspired conferences and workshops, many and varied research papers, and are central to several modern corpus processing tools. In this article, we present evidence that (at least for the task of biographical sentence classification) frequent words characterize texts better than keywords or key-keywords. Using the naïve Bayes learning algorithm in conjunction with frequency-, keyword-, and key-keyword-based text representation to classify a corpus of biographical sentences, we discovered that the use of frequent words alone provided a classification accuracy better than either the keyword or key-keyword representations at a statistically significant level. This result suggests that (for the biographical sentence classification task at least) frequent words characterize texts better than keywords derived using more computationally intensive methods. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: © 2019 EADH: The European Association for Digital Humanities. item: Literary & Linguistic Computing holder: Oxford University Press / USA dt: @attributes: year: 2010 holdings: @attributes: islocal: N |
|---|