Vector space explorations of literary language.
Literary novels are said to distinguish themselves from other novels through conventions associated with literariness. We investigate the task of predicting the literariness of novels as perceived by readers, based on a large reader survey of contemporary Dutch novels. Previous research showed that...
| Publicado en: | Language Resources & Evaluation Vol. 53; no. 4; pp. 625 - 651 |
|---|---|
| Autores principales: | , , |
| Formato: | Artículo |
| Publicado: |
Springer Nature
Dec2019
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=139882004&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 139882004 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Dec2019 vid: 53 iid: 4 pid: 237 pub: Springer Nature artinfo: ui: 139882004 10.1007/s10579-018-09442-4 ppf: 625 ppct: 26 formats: fmt: – @attributes: type: T – @attributes: type: P size: 822KB tig: atl: Vector space explorations of literary language. aug: au: van Cranenburgh, Andreas van Dalen-Oskam, Karina van Zundert, Joris affil: Information Science, University of Groningen, Groningen, The Netherlands Huygens ING, Royal Netherlands Academy of Arts and Sciences, Amsterdam, The Netherlands Universiteit van Amsterdam, Amsterdam, The Netherlands su: Vector spaces Standard language Space exploration Child development deviations Machine learning Prediction models Literary research sug: subj: Vector spaces Standard language Space exploration Child development deviations Machine learning Prediction models Literary research keyword: Document embeddings Literariness Literature Topic models ab: Literary novels are said to distinguish themselves from other novels through conventions associated with literariness. We investigate the task of predicting the literariness of novels as perceived by readers, based on a large reader survey of contemporary Dutch novels. Previous research showed that ratings of literariness are predictable from texts to a substantial extent using machine learning, suggesting that it may be possible to explain the consensus among readers on which novels are literary as a consensus on the kind of writing style that characterizes literature. Although we have not yet collected human judgments to establish the influence of writing style directly (we use a survey with judgments based on the titles of novels), we can try to analyze the behavior of machine learning models on particular text fragments as a proxy for human judgments. In order to explore aspects of the texts associated with literariness, we divide the texts of the novels in chunks of 2–3 pages and create vector space representations using topic models (Latent Dirichlet Allocation) and neural document embeddings (Distributed Bag-of-Words Paragraph Vectors). We analyze the semantic complexity of the novels using distance measures, supporting the notion that literariness can be partly explained as a deviation from the norm. Furthermore, we build predictive models and identify specific keywords and stylistic markers related to literariness. While genre plays a role, we find that the greater part of factors affecting judgments of literariness are explicable in bag-of-words terms, even in short text fragments and among novels with higher literary ratings. The code and notebook used to produce the results in this paper are available at https://github.com/andreasvc/litvecspace. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2019. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2019 holdings: @attributes: islocal: N |
|---|