Automatic readability assessment for sentences: neural, hybrid and large language models.
Automatic readability assessment (ARA) aims to determine the cognitive load of a reader to comprehend a given text. ARA research has been mostly conducted at the text level, with numerous studies showing strong performance of neural and hybrid models. ARA at the sentence level, however, has received...
| Publicado en: | Language Resources & Evaluation Vol. 59; no. 3; pp. 2265 - 2297 |
|---|---|
| Autores principales: | , , |
| Formato: | Artículo |
| Publicado: |
Springer Nature
Sep2025
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=186909059&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 186909059 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Sep2025 vid: 59 iid: 3 pid: 237 pub: Springer Nature artinfo: ui: 186909059 10.1007/s10579-024-09800-5 ppf: 2265 ppct: 32 formats: fmt: – @attributes: type: T – @attributes: type: P size: 1.5MB tig: atl: Automatic readability assessment for sentences: neural, hybrid and large language models. aug: au: Liu, Fengkai Jin, Tan Lee, John S. Y. affil: https://ror.org/03q8dnn23 Department of Linguistics and Translation, City University of Hong Kong, Hong Kong SAR, China https://ror.org/01kq0pv72 School of International Culture, South China Normal University, Guangzhou, Guangdong, China su: Natural language processing Language models Readability formulas Ensemble learning Artificial neural networks Semantics sug: subj: Natural language processing Language models Readability formulas Ensemble learning Artificial neural networks Semantics keyword: Automatic readability assessment Communication and Culture Linguistics Hybrid models Language Large language models Sentence readability assessment ab: Automatic readability assessment (ARA) aims to determine the cognitive load of a reader to comprehend a given text. ARA research has been mostly conducted at the text level, with numerous studies showing strong performance of neural and hybrid models. ARA at the sentence level, however, has received less attention, even though many applications in natural language processing (NLP) require assessment of the difficulty of individual sentences. This article compares the performance of neural models, hybrid models and large language models (LLMs) for sentence-level ARA, making three main contributions. First, we construct the first Chinese sentence-level ARA datasets, with nearly 70K sentences, to facilitate evaluation on Chinese. Second, we present the first experimental results on applying LLMs to sentence-level ARA. Finally, while previous work focused mostly on English data, we show that hybrid models outperform traditional classifiers, neural models, and LLMs in both English and Chinese data. The best hybrid model obtained state-of-the-art results on the Wall Street Journal dataset, surpassing the previous best result by almost 15% absolute. It also achieved competitive results on the CEFR-SP dataset. In detailed analyses, we identify the linguistic features that most significantly contributed to the performance of the hybrid model. We show that 10 linguistic features are correlated to readability across all datasets, and models trained on this reduced feature set achieve performance that rivals the full set. These results not only yield new insights into hybrid models for sentence-level ARA, but also set new benchmarks for future research in both English and Chinese. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2025. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2025 holdings: @attributes: islocal: N |
|---|