In search of founding era registers: automatic modeling of registers from the corpus of Founding Era American English.
Registers are situationally defined text varieties, such as letters, essays, or news articles, that are considered to be one of the most important predictors of linguistic variation. Often historical databases of language lack register information, which could greatly enhance their usability (e.g. E...
| Publicado en: | Digital Scholarship in the Humanities Vol. 38; no. 4; pp. 1659 - 1678 |
|---|---|
| Autores principales: | , , |
| Formato: | Artículo |
| Publicado: |
Oxford University Press / USA
Dec2023
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=174444634&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 174444634 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 2055768X JEO9 jtl: Digital Scholarship in the Humanities issn: 2055768X maglogo: N pubinfo: dt: Dec2023 vid: 38 iid: 4 pid: 622 pub: Oxford University Press / USA artinfo: ui: 174444634 10.1093/llc/fqad049 ppf: 1659 ppct: 19 formats: fmt: – @attributes: type: T – @attributes: type: P size: 653KB tig: atl: In search of founding era registers: automatic modeling of registers from the corpus of Founding Era American English. aug: au: Repo, Liina Hashimoto, Brett Laippala, Veronika affil: School of Languages and Translation Studies, University of Turku , FI-20014 University of Turku , Finland Department of Linguistics, Brigham Young University , Provo, UT 84602, USA su: Deep learning Language models American English language Natural language processing Automatic identification English language sug: subj: Deep learning Language models American English language Natural language processing Automatic identification English language keyword: BERT historical natural language processing Late Modern English register text classification ab: Registers are situationally defined text varieties, such as letters, essays, or news articles, that are considered to be one of the most important predictors of linguistic variation. Often historical databases of language lack register information, which could greatly enhance their usability (e.g. Early English Books Online). This article examines register variation in Late Modern English and automatic register identification in historical corpora. We model register variation in the corpus of Founding Era American English (COFEA) and develop machine-learning methods for automatic register identification in COFEA. We also extract and analyze the most significant grammatical characteristics estimated by the classifier for the best-predicted registers and found that letters and journals in the 1700s were characterized by informational density. The chosen method enables us to learn more about registers in the Founding Era. We show that some registers can be reliably identified from COFEA, the best overall performance achieved by the deep learning model Bidirectional Encoder Representations from Transformers with an F1-score of 97 per cent. This suggests that deep learning models could be utilized in other studies concerned with historical language and its automatic classification. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: © 2019 EADH: The European Association for Digital Humanities. item: Digital Scholarship in the Humanities holder: Oxford University Press / USA dt: @attributes: year: 2023 holdings: @attributes: islocal: N |
|---|