Beyond lexical frequencies: using R for text analysis in the digital humanities.

This paper presents a combination of R packages—user contributed toolkits written in a common core programming language—to facilitate the humanistic investigation of digitised, text-based corpora. Our survey of text analysis packages includes those of our own creation (cleanNLP and fasttextM) as wel...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 53; no. 4; pp. 707 - 734
Autores principales: Arnold, Taylor, Ballier, Nicolas, Lissón, Paula, Tilton, Lauren
Formato: Artículo
Publicado: Springer Nature Dec2019
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=139882007&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 139882007
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Dec2019
      vid: 53
      iid: 4
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        139882007
        10.1007/s10579-019-09456-6
      ppf: 707
      ppct: 27
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 1.9MB
      tig:
        atl: Beyond lexical frequencies: using R for text analysis in the digital humanities.
      aug:
        au:
          Arnold, Taylor
          Ballier, Nicolas
          Lissón, Paula
          Tilton, Lauren
        affil:
          University of Richmond, Virginia, USA
          UFR Études Anglophones, Université Paris Diderot, Paris, France
          Department of Linguistics, Universität Potsdam, Potsdam, Germany
      su:
        Digital humanities
        Natural language processing
        Corpora
        Programming languages
        Prefabricated buildings
        Research teams
      sug:
        subj:
          Digital humanities
          Natural language processing
          Corpora
          Programming languages
          Prefabricated buildings
          Research teams
      keyword:
        R
        Text interoperability
        Text mining
      ab: This paper presents a combination of R packages—user contributed toolkits written in a common core programming language—to facilitate the humanistic investigation of digitised, text-based corpora. Our survey of text analysis packages includes those of our own creation (cleanNLP and fasttextM) as well as packages built by other research groups (stringi, readtext, hyphenatr, quanteda, and hunspell). By operating on generic object types, these packages unite research innovations in corpus linguistics, natural language processing, machine learning, statistics, and digital humanities. We begin by extrapolating on the theoretical benefits of R as an elaborate gluing language for bringing together several areas of expertise and compare it to linguistic concordancers and other tool-based approaches to text analysis in the digital humanities. We then showcase the practical benefits of an ecosystem by illustrating how R packages have been integrated into a digital humanities project. Throughout, the focus is on moving beyond the bag-of-words, lexical frequency model by incorporating linguistically-driven analyses in research.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2019. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2019
    holdings:
      @attributes:
        islocal: N