In no uncertain terms: a dataset for monolingual and multilingual automatic term extraction from comparable corpora.

Automatic term extraction is a productive field of research within natural language processing, but it still faces significant obstacles regarding datasets and evaluation, which require manual term annotation. This is an arduous task, made even more difficult by the lack of a clear distinction betwe...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 54; no. 2; pp. 385 - 419
Autores principales: Rigouts Terryn, Ayla, Hoste, Véronique, Lefever, Els
Formato: Artículo
Publicado: Springer Nature Jun2020
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=143152354&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 143152354
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Jun2020
      vid: 54
      iid: 2
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        143152354
        10.1007/s10579-019-09453-9
      ppf: 385
      ppct: 34
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 840KB
      tig:
        atl: In no uncertain terms: a dataset for monolingual and multilingual automatic term extraction from comparable corpora.
      aug:
        au:
          Rigouts Terryn, Ayla
          Hoste, Véronique
          Lefever, Els
        affil: LT3 Language and Translation Technology Team, Department of Translation, Interpreting and Communication, Ghent University, Groot-Brittanniëlaan 45, 9000, Ghent, Belgium
      su:
        Natural language processing
        Corpora
        Terms & phrases
      sug:
        subj:
          Natural language processing
          Corpora
          Terms & phrases
      keyword:
        ATR
        Automatic term extraction
        Comparable corpora
        Term annotation
        Terminology
      ab: Automatic term extraction is a productive field of research within natural language processing, but it still faces significant obstacles regarding datasets and evaluation, which require manual term annotation. This is an arduous task, made even more difficult by the lack of a clear distinction between terms and general language, which results in low inter-annotator agreement. There is a large need for well-documented, manually validated datasets, especially in the rising field of multilingual term extraction from comparable corpora, which presents a unique new set of challenges. In this paper, a new approach is presented for both monolingual and multilingual term annotation in comparable corpora. The detailed guidelines with different term labels, the domain- and language-independent methodology and the large volumes annotated in three different languages and four different domains make this a rich resource. The resulting datasets are not just suited for evaluation purposes but can also serve as a general source of information about terms and even as training data for supervised methods. Moreover, the gold standard for multilingual term extraction from comparable corpora contains information about term variants and translation equivalents, which allows an in-depth, nuanced evaluation.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2020. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2020
    holdings:
      @attributes:
        islocal: N