Mining a corpus of biographical texts using keywords.

Using statistically derived keywords to characterize texts has become an important research method for digital humanists and corpus linguists in areas such as literary analysis and the exploration of genre difference. Keywords—and the associated concepts of ‘keyness’ and ‘key-keyness’—have inspired...

Full description

Bibliographic Details
Published in:Literary & Linguistic Computing Vol. 25; no. 1; pp. 23 - 36
Main Author: Conway, Mike
Format: Article
Published: Oxford University Press / USA Apr2010
Subjects:
Online Access:View this record in EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=53297471&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 53297471
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        02681145
        BJ1
      jtl: Literary & Linguistic Computing
      issn: 02681145
      maglogo: N
    pubinfo:
      dt: Apr2010
      vid: 25
      iid: 1
      pid: 622
      pub: Oxford University Press / USA
    artinfo:
      ui:
        53297471
        10.1093/llc/fqp035
      ppf: 23
      ppct: 13
      formats:
        fmt:
          @attributes:
            type: P
            size: 361KB
      tig:
        atl: Mining a corpus of biographical texts using keywords.
      aug:
        au: Conway, Mike
        affil: National Institute of Informatics, Japan
      su:
        Keywords
        Research methodology
        Humanities education
        Corpora
        Forums
        Adult education workshops
      sug:
        subj:
          Keywords
          Research methodology
          Humanities education
          Corpora
          Forums
          Adult education workshops
      ab: Using statistically derived keywords to characterize texts has become an important research method for digital humanists and corpus linguists in areas such as literary analysis and the exploration of genre difference. Keywords—and the associated concepts of ‘keyness’ and ‘key-keyness’—have inspired conferences and workshops, many and varied research papers, and are central to several modern corpus processing tools. In this article, we present evidence that (at least for the task of biographical sentence classification) frequent words characterize texts better than keywords or key-keywords. Using the naïve Bayes learning algorithm in conjunction with frequency-, keyword-, and key-keyword-based text representation to classify a corpus of biographical sentences, we discovered that the use of frequent words alone provided a classification accuracy better than either the keyword or key-keyword representations at a statistically significant level. This result suggests that (for the biographical sentence classification task at least) frequent words characterize texts better than keywords derived using more computationally intensive methods.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: © 2019 EADH: The European Association for Digital Humanities.
      item: Literary & Linguistic Computing
      holder: Oxford University Press / USA
      dt:
        @attributes:
          year: 2010
    holdings:
      @attributes:
        islocal: N