Beyond lexical frequencies: using R for text analysis in the digital humanities.
This paper presents a combination of R packages—user contributed toolkits written in a common core programming language—to facilitate the humanistic investigation of digitised, text-based corpora. Our survey of text analysis packages includes those of our own creation (cleanNLP and fasttextM) as wel...
| Publicado en: | Language Resources & Evaluation Vol. 53; no. 4; pp. 707 - 734 |
|---|---|
| Autores principales: | , , , |
| Formato: | Artículo |
| Publicado: |
Springer Nature
Dec2019
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=139882007&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 139882007 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Dec2019 vid: 53 iid: 4 pid: 237 pub: Springer Nature artinfo: ui: 139882007 10.1007/s10579-019-09456-6 ppf: 707 ppct: 27 formats: fmt: – @attributes: type: T – @attributes: type: P size: 1.9MB tig: atl: Beyond lexical frequencies: using R for text analysis in the digital humanities. aug: au: Arnold, Taylor Ballier, Nicolas Lissón, Paula Tilton, Lauren affil: University of Richmond, Virginia, USA UFR Études Anglophones, Université Paris Diderot, Paris, France Department of Linguistics, Universität Potsdam, Potsdam, Germany su: Digital humanities Natural language processing Corpora Programming languages Prefabricated buildings Research teams sug: subj: Digital humanities Natural language processing Corpora Programming languages Prefabricated buildings Research teams keyword: R Text interoperability Text mining ab: This paper presents a combination of R packages—user contributed toolkits written in a common core programming language—to facilitate the humanistic investigation of digitised, text-based corpora. Our survey of text analysis packages includes those of our own creation (cleanNLP and fasttextM) as well as packages built by other research groups (stringi, readtext, hyphenatr, quanteda, and hunspell). By operating on generic object types, these packages unite research innovations in corpus linguistics, natural language processing, machine learning, statistics, and digital humanities. We begin by extrapolating on the theoretical benefits of R as an elaborate gluing language for bringing together several areas of expertise and compare it to linguistic concordancers and other tool-based approaches to text analysis in the digital humanities. We then showcase the practical benefits of an ecosystem by illustrating how R packages have been integrated into a digital humanities project. Throughout, the focus is on moving beyond the bag-of-words, lexical frequency model by incorporating linguistically-driven analyses in research. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2019. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2019 holdings: @attributes: islocal: N |
|---|