Automatically Categorizing Written Texts by Author Gender.

The problem of automatically determining the gender of a document's author would appear to be a more subtle problem than those of categorization by topic or authorship attribution. Nevertheless, it is shown that automated text categorization techniques can exploit combinations of simple lexical and...

Descripción completa

Detalles Bibliográficos
Publicado en:Literary & Linguistic Computing Vol. 17; no. 4; pp. 401 - 412
Autores principales: Koppel, Moshe, Argamon, Shlomo, Shimoni, Anat Rachel
Formato: Artículo
Publicado: Oxford University Press / USA Nov2002
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=10037175&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 10037175
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        02681145
        BJ1
      jtl: Literary & Linguistic Computing
      issn: 02681145
      maglogo: N
    pubinfo:
      dt: Nov2002
      vid: 17
      iid: 4
      pid: 622
      pub: Oxford University Press / USA
    artinfo:
      ui:
        10037175
        10.1093/llc/17.4.401
      ppf: 401
      ppct: 11
      formats:
        fmt:
          @attributes:
            type: P
            size: 124KB
      tig:
        atl: Automatically Categorizing Written Texts by Author Gender.
      aug:
        au:
          Koppel, Moshe
          Argamon, Shlomo
          Shimoni, Anat Rachel
      su:
        Text processing (Computer science)
        Electronic data processing
      sug:
        subj:
          Text processing (Computer science)
          Electronic data processing
      ab: The problem of automatically determining the gender of a document's author would appear to be a more subtle problem than those of categorization by topic or authorship attribution. Nevertheless, it is shown that automated text categorization techniques can exploit combinations of simple lexical and syntactic features to infer the gender of the author of an unseen formal written document with approximately 80 per cent accuracy. The same techniques can be used to determine if a document is fiction or non-fiction with approximately 98 per cent accuracy.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: © 2019 EADH: The European Association for Digital Humanities.
      item: Literary & Linguistic Computing
      holder: Oxford University Press / USA
      dt:
        @attributes:
          year: 2002
    holdings:
      @attributes:
        islocal: N