Text complexity of open educational resources in Portuguese: mixing written and spoken registers in a multi-task approach.
This paper presents a study on text complexity of Open Educational Resources (OER) in Brazilian Portuguese. In a data analysis of the Brazilian Ministry of Education Integrated Platform (MEC-RED) carried out in September 2020, 86% of the resources on the platform did not have any grade level classif...
| Publicado en: | Language Resources & Evaluation Vol. 56; no. 2; pp. 621 - 651 |
|---|---|
| Autores principales: | , , , , , |
| Formato: | Artículo |
| Publicado: |
Springer Nature
Jun2022
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=157410220&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 157410220 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Jun2022 vid: 56 iid: 2 pid: 237 pub: Springer Nature artinfo: ui: 157410220 10.1007/s10579-021-09571-3 ppf: 621 ppct: 30 formats: fmt: – @attributes: type: T – @attributes: type: P size: 725KB tig: atl: Text complexity of open educational resources in Portuguese: mixing written and spoken registers in a multi-task approach. aug: au: Gazzola, Murilo Leal, Sidney Pedroni, Breno Theoto Rocha, Fábio Pompéia, Sabine Aluísio, Sandra affil: Institute of Mathematics and Computer Science, University of São Paulo, São Carlos, Brazil Department of Psychobiology, Universidade Federal de São Paulo, São Paulo, Brazil su: Open educational resources Natural language processing Linguistic complexity Language ability Grade levels sug: subj: Open educational resources Natural language processing Linguistic complexity Language ability Grade levels keyword: Multi-task learning Spontaneous speech Text complexity Transcribed narratives ab: This paper presents a study on text complexity of Open Educational Resources (OER) in Brazilian Portuguese. In a data analysis of the Brazilian Ministry of Education Integrated Platform (MEC-RED) carried out in September 2020, 86% of the resources on the platform did not have any grade level classification, making it difficult to find, use, and expand them. The text complexity task in the Natural Language Processing research area can be used to identify texts that have adequate linguistic complexity for specific grade levels, allowing to complete the stage of education metadata in MEC-RED. However, some types of MEC-RED's resources do not present any information about their stage of education, making it unfeasible to compile a balanced dataset of OER for training a text complexity predictor. This study is driven and enabled by a recently created corpus of transcribed spoken narratives produced by fourth graders to first graders of high school which were collected to evaluate the development of language abilities. A multi-task learning (MTL) approach via hard parameter sharing of hidden layers was adopted to train three models that share all parameters in their hidden layers. The main objective of this study was to explore the relationship between three text complexity tasks by jointly learning to predict text readability, using coarse and fine-grained datasets of written, spoken and domain texts (a small dataset of OER resources) to overcome the lack of grade classified resources in MEC-RED. Our MTL model with two auxiliary tasks presents a F-measure of 0.955, an improvement of 0.15 points over our previous results. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2022. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2022 holdings: @attributes: islocal: N |
|---|