Text complexity of open educational resources in Portuguese: mixing written and spoken registers in a multi-task approach.

This paper presents a study on text complexity of Open Educational Resources (OER) in Brazilian Portuguese. In a data analysis of the Brazilian Ministry of Education Integrated Platform (MEC-RED) carried out in September 2020, 86% of the resources on the platform did not have any grade level classif...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 56; no. 2; pp. 621 - 651
Autores principales: Gazzola, Murilo, Leal, Sidney, Pedroni, Breno, Theoto Rocha, Fábio, Pompéia, Sabine, Aluísio, Sandra
Formato: Artículo
Publicado: Springer Nature Jun2022
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=157410220&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 157410220
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Jun2022
      vid: 56
      iid: 2
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        157410220
        10.1007/s10579-021-09571-3
      ppf: 621
      ppct: 30
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 725KB
      tig:
        atl: Text complexity of open educational resources in Portuguese: mixing written and spoken registers in a multi-task approach.
      aug:
        au:
          Gazzola, Murilo
          Leal, Sidney
          Pedroni, Breno
          Theoto Rocha, Fábio
          Pompéia, Sabine
          Aluísio, Sandra
        affil:
          Institute of Mathematics and Computer Science, University of São Paulo, São Carlos, Brazil
          Department of Psychobiology, Universidade Federal de São Paulo, São Paulo, Brazil
      su:
        Open educational resources
        Natural language processing
        Linguistic complexity
        Language ability
        Grade levels
      sug:
        subj:
          Open educational resources
          Natural language processing
          Linguistic complexity
          Language ability
          Grade levels
      keyword:
        Multi-task learning
        Spontaneous speech
        Text complexity
        Transcribed narratives
      ab: This paper presents a study on text complexity of Open Educational Resources (OER) in Brazilian Portuguese. In a data analysis of the Brazilian Ministry of Education Integrated Platform (MEC-RED) carried out in September 2020, 86% of the resources on the platform did not have any grade level classification, making it difficult to find, use, and expand them. The text complexity task in the Natural Language Processing research area can be used to identify texts that have adequate linguistic complexity for specific grade levels, allowing to complete the stage of education metadata in MEC-RED. However, some types of MEC-RED's resources do not present any information about their stage of education, making it unfeasible to compile a balanced dataset of OER for training a text complexity predictor. This study is driven and enabled by a recently created corpus of transcribed spoken narratives produced by fourth graders to first graders of high school which were collected to evaluate the development of language abilities. A multi-task learning (MTL) approach via hard parameter sharing of hidden layers was adopted to train three models that share all parameters in their hidden layers. The main objective of this study was to explore the relationship between three text complexity tasks by jointly learning to predict text readability, using coarse and fine-grained datasets of written, spoken and domain texts (a small dataset of OER resources) to overcome the lack of grade classified resources in MEC-RED. Our MTL model with two auxiliary tasks presents a F-measure of 0.955, an improvement of 0.15 points over our previous results.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2022. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2022
    holdings:
      @attributes:
        islocal: N