Identifying communicative functions in discourse with content types.
Texts are not monolithic entities but rather coherent collections of micro illocutionary acts which help to convey a unitary message of content and purpose. Identifying such text segments is challenging because they require a fine-grained level of analysis even within a single sentence. At the same...
| Publicado en: | Language Resources & Evaluation Vol. 56; no. 2; pp. 417 - 451 |
|---|---|
| Autores principales: | , , |
| Formato: | Artículo |
| Publicado: |
Springer Nature
Jun2022
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=157410217&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 157410217 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Jun2022 vid: 56 iid: 2 pid: 237 pub: Springer Nature artinfo: ui: 157410217 10.1007/s10579-021-09554-4 ppf: 417 ppct: 34 formats: fmt: – @attributes: type: T – @attributes: type: P size: 824KB tig: atl: Identifying communicative functions in discourse with content types. aug: au: Caselli, Tommaso Sprugnoli, Rachele Moretti, Giovanni affil: Rijksuniversiteit Groningen, Groningen, Netherlands Università Cattolica del Sacro Cuore, Milan, Italy su: Linguistic analysis Data mining Historical source material Discourse Corpora sug: subj: Linguistic analysis Data mining Historical source material Discourse Corpora keyword: Across genre Across time Content types Corpus annotation Neural networks ab: Texts are not monolithic entities but rather coherent collections of micro illocutionary acts which help to convey a unitary message of content and purpose. Identifying such text segments is challenging because they require a fine-grained level of analysis even within a single sentence. At the same time, accessing them facilitates the analysis of the communicative functions of a text as well as the identification of relevant information. We propose an empirical framework for modelling micro illocutionary acts at clause level, that we call content types, grounded on linguistic theories of text types, in particular on the framework proposed by Werlich in 1976. We make available a newly annotated corpus of 279 documents (for a total of more than 180,000 tokens) belonging to different genres and temporal periods, based on a dedicated annotation scheme. We obtain an average Cohen's kappa of 0.89 at token level. We achieve an average F1 score of 74.99% on the automatic classification of content types using a bi-LSTM model. Similar results are obtained on contemporary and historical documents, while performances on genres are more varied. This work promotes a discourse-oriented approach to information extraction and cross-fertilisation across disciplines through a computationally-aided linguistic analysis. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2022. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2022 holdings: @attributes: islocal: N |
|---|