Construction of an aligned monolingual treebank for studying semantic similarity.

Modern paraphrase research would benefit from large corpora with detailed annotations. However, currently these corpora are still thin on the ground. In this paper, we describe the development of such a corpus for Dutch, which takes the form of a parallel monolingual treebank consisting of over 2 mi...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 48; no. 2; pp. 279 - 307
Autores principales: Marsi, Erwin, Krahmer, Emiel
Formato: Artículo
Publicado: Springer Nature Jun2014
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=95865803&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 95865803
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Jun2014
      vid: 48
      iid: 2
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        95865803
        10.1007/s10579-013-9252-1
      ppf: 279
      ppct: 28
      formats:
        fmt:
          @attributes:
            type: P
            size: 729KB
      tig:
        atl: Construction of an aligned monolingual treebank for studying semantic similarity.
      aug:
        au:
          Marsi, Erwin
          Krahmer, Emiel
        affil:
          Department of Computer and Information Science, Norwegian University of Science and Technology, Sem Sælands vei 7-9 7491 Trondheim Norway
          Tilburg Center for Cognition and Communication (TiCC), Tilburg University, 5000 LE Tilburg The Netherlands
      su:
        Paraphrase
        Corpora
        Annotations
        Monolingualism
        Dutch language
        Language & languages
      sug:
        subj:
          Paraphrase
          Corpora
          Annotations
          Monolingualism
          Dutch language
          Language & languages
      keyword:
        Alignment
        Comparable text
        Corpus
        Monolingual treebank
        Parallel text
        Semantic relations
        Semantic similarity
        Tree alignment
      ab: Modern paraphrase research would benefit from large corpora with detailed annotations. However, currently these corpora are still thin on the ground. In this paper, we describe the development of such a corpus for Dutch, which takes the form of a parallel monolingual treebank consisting of over 2 million tokens and covering various text genres, including both parallel and comparable text. This publicly available corpus is richly annotated with alignments between syntactic nodes, which are also classified using five different semantic similarity relations. A quarter of the corpus is manually annotated, and this informs the development of an automatic tree aligner used to annotate the remainder of the corpus. We argue that this corpus is the first of this size and kind, and offers great potential for paraphrasing research.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2014. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2014
    holdings:
      @attributes:
        islocal: N