NEREL: a Russian information extraction dataset with rich annotation for nested entities, relations, and wikidata entity links.

This paper describes NEREL—a Russian news dataset suited for three tasks: nested named entity recognition, relation extraction, and entity linking. Compared to flat entities, nested named entities provide a richer and more complete annotation while also increasing the coverage of relations annotatio...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 58; no. 2; pp. 547 - 584
Autores principales: Loukachevitch, Natalia, Artemova, Ekaterina, Batura, Tatiana, Braslavski, Pavel, Ivanov, Vladimir, Manandhar, Suresh, Pugachev, Alexander, Rozhkov, Igor, Shelmanov, Artem, Tutubalina, Elena, Yandutov, Alexey
Formato: Artículo
Publicado: Springer Nature Jun2024
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=178064680&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 178064680
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Jun2024
      vid: 58
      iid: 2
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        178064680
        10.1007/s10579-023-09674-z
      ppf: 547
      ppct: 37
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 1.5MB
      tig:
        atl: NEREL: a Russian information extraction dataset with rich annotation for nested entities, relations, and wikidata entity links.
      aug:
        au:
          Loukachevitch, Natalia
          Artemova, Ekaterina
          Batura, Tatiana
          Braslavski, Pavel
          Ivanov, Vladimir
          Manandhar, Suresh
          Pugachev, Alexander
          Rozhkov, Igor
          Shelmanov, Artem
          Tutubalina, Elena
          Yandutov, Alexey
        affil:
          https://ror.org/010pmpe69 Lomonosov Moscow State University, Moscow, Russia
          ISP RAS Research Center for Trusted Artificial Intelligence, Moscow, Russia
          HSE University, Moscow, Russia
          https://ror.org/04t2ss102 Novosibirsk State University, Novosibirsk, Russia
          https://ror.org/01dc8vg74 A.P. Ershov Institute of Informatics Systems, Novosibirsk, Russia
          https://ror.org/00hs7dr46 Ural Federal University, Yekaterinburg, Russia
          https://ror.org/02b7jh107 Innopolis University, Innopolis, Russia
          https://ror.org/021xhya68 Madan Bhandari University of Science and Technology, Chitlang, Nepal
          https://ror.org/014a87f14 AIRI, Moscow, Russia
          https://ror.org/0258gkt32 Mohamed bin Zayed University of Artificial Intelligence (MBZUAI), Abu Dhabi, UAE
          Sber AI, Moscow, Russia
      su:
        Data mining
        Annotations
        News websites
      sug:
        subj:
          Data mining
          Annotations
          News websites
      keyword:
        68T35
        68T50
        Entity linking
        Named entity recognition
        Nested entities
        Nested relations
        Relation extraction
      ab: This paper describes NEREL—a Russian news dataset suited for three tasks: nested named entity recognition, relation extraction, and entity linking. Compared to flat entities, nested named entities provide a richer and more complete annotation while also increasing the coverage of relations annotation and entity linking. Relations between nested named entities may cross entity boundaries to connect to shorter entities nested within longer ones, which makes it harder to detect such relations. NEREL is currently the largest Russian dataset annotated with entities and relations: it comprises 29 named entity types and 49 relation types. At the time of writing, the dataset contains 56 K named entities and 39 K relations annotated in 933 person-oriented news articles. NEREL is annotated with relations at three levels: (1) within nested named entities, (2) within sentences, and (3) with relations crossing sentence boundaries. We provide benchmark evaluation of current state-of-the-art methods in all three tasks. The dataset is freely available at https://github.com/nerel-ds/NEREL.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2024. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2024
    holdings:
      @attributes:
        islocal: N