Identifying communicative functions in discourse with content types.

Texts are not monolithic entities but rather coherent collections of micro illocutionary acts which help to convey a unitary message of content and purpose. Identifying such text segments is challenging because they require a fine-grained level of analysis even within a single sentence. At the same...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 56; no. 2; pp. 417 - 451
Autores principales: Caselli, Tommaso, Sprugnoli, Rachele, Moretti, Giovanni
Formato: Artículo
Publicado: Springer Nature Jun2022
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=157410217&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 157410217
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Jun2022
      vid: 56
      iid: 2
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        157410217
        10.1007/s10579-021-09554-4
      ppf: 417
      ppct: 34
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 824KB
      tig:
        atl: Identifying communicative functions in discourse with content types.
      aug:
        au:
          Caselli, Tommaso
          Sprugnoli, Rachele
          Moretti, Giovanni
        affil:
          Rijksuniversiteit Groningen, Groningen, Netherlands
          Università Cattolica del Sacro Cuore, Milan, Italy
      su:
        Linguistic analysis
        Data mining
        Historical source material
        Discourse
        Corpora
      sug:
        subj:
          Linguistic analysis
          Data mining
          Historical source material
          Discourse
          Corpora
      keyword:
        Across genre
        Across time
        Content types
        Corpus annotation
        Neural networks
      ab: Texts are not monolithic entities but rather coherent collections of micro illocutionary acts which help to convey a unitary message of content and purpose. Identifying such text segments is challenging because they require a fine-grained level of analysis even within a single sentence. At the same time, accessing them facilitates the analysis of the communicative functions of a text as well as the identification of relevant information. We propose an empirical framework for modelling micro illocutionary acts at clause level, that we call content types, grounded on linguistic theories of text types, in particular on the framework proposed by Werlich in 1976. We make available a newly annotated corpus of 279 documents (for a total of more than 180,000 tokens) belonging to different genres and temporal periods, based on a dedicated annotation scheme. We obtain an average Cohen's kappa of 0.89 at token level. We achieve an average F1 score of 74.99% on the automatic classification of content types using a bi-LSTM model. Similar results are obtained on contemporary and historical documents, while performances on genres are more varied. This work promotes a discourse-oriented approach to information extraction and cross-fertilisation across disciplines through a computationally-aided linguistic analysis.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2022. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2022
    holdings:
      @attributes:
        islocal: N