Exploring the role of lexis and grammar for the stable identification of register in an unrestricted corpus of web documents.

The Internet offers great possibilities for many scientific disciplines that utilize text data. However, the potential of online data can be limited by the lack of information on the genre or register of the documents, as register—whether a text is, e.g., a news article or a recipe—is arguably the m...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 55; no. 3; pp. 757 - 789
Autores principales: Laippala, Veronika, Egbert, Jesse, Biber, Douglas, Kyröläinen, Aki-Juhani
Formato: Artículo
Publicado: Springer Nature Sep2021
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=151686297&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 151686297
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Sep2021
      vid: 55
      iid: 3
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        151686297
        10.1007/s10579-020-09519-z
      ppf: 757
      ppct: 32
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 1.7MB
      tig:
        atl: Exploring the role of lexis and grammar for the stable identification of register in an unrestricted corpus of web documents.
      aug:
        au:
          Laippala, Veronika
          Egbert, Jesse
          Biber, Douglas
          Kyröläinen, Aki-Juhani
        affil:
          University of Turku, Turku, Finland
          Northern Arizona University, Flagstaff, USA
          McMaster University & Brock University, Hamilton, ON, Canada
      su:
        Grammar
        Automatic identification
        Corpora
        Forecasting
        Linguists
      sug:
        subj:
          Grammar
          Automatic identification
          Corpora
          Forecasting
          Linguists
      keyword:
        Discriminative features
        Model stability
        Online data
        Online registers
        SVM
        Text classification
        Web genre identification
        Web genres
        Web-as-corpus
      ab: The Internet offers great possibilities for many scientific disciplines that utilize text data. However, the potential of online data can be limited by the lack of information on the genre or register of the documents, as register—whether a text is, e.g., a news article or a recipe—is arguably the most important predictor of linguistic variation (see Biber in Corpus Linguist Linguist Theory 8:9–37, 2012). Despite having received significant attention in recent years, the modeling of online registers has faced a number of challenges, and previous studies have presented contradictory results. In particular, these have concerned (1) the extent to which registers can be automatically identified in a large, unrestricted corpus of web documents and (2) the stability of the models, specifically the kinds of linguistic features that achieve the best performance while reflecting the registers instead of corpus idiosyncrasies. Furthermore, although the linguistic properties of registers vary importantly in a number of ways that may affect their modeling, this variation is often bypassed. In this article, we tackle these issues. We model online registers in the largest available corpus of online registers, the Corpus of Online Registers of English (CORE). Additionally, we evaluate the stability of the models towards corpus idiosyncrasies, analyze the role of different linguistic features in them, and examine how individual registers differ in these two aspects. We show that (1) competitive classification performance on a large-scale, unrestricted corpus can be achieved through a combination of lexico-grammatical features, (2) the inclusion of grammatical information improves the stability of the model, whereas many of the previously best-performing feature sets are less stable, and that (3) registers can be placed in a continuum based on the discriminative importance of lexis and grammar. These register-specific characteristics can explain the variation observed in previous studies concerning the automatic identification of online registers and the importance of different linguistic features for them. Thus, our results offer explanations for the jungle-likeness of online data and provide essential information on online registers for all studies using online data.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2021. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2021
    holdings:
      @attributes:
        islocal: N