CachacaNER: a dataset for named entity recognition in texts about the cachaça beverage.

Named Entity Recognition (NER) is the task of identifying and classifying tokens in texts corresponding to a set of pre-defined categories, such as names of people, organizations and locations. Datasets labeled for this task are essential for training supervised machine learning models. Although the...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 58; no. 4; pp. 1315 - 1334
Autores principales: Silva, Priscilla, Franco, Arthur, Santos, Thiago, Brito, Mozar, Pereira, Denilson
Formato: Artículo
Publicado: Springer Nature Dec2024
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=180627301&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 180627301
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Dec2024
      vid: 58
      iid: 4
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        180627301
        10.1007/s10579-023-09665-0
      ppf: 1315
      ppct: 19
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 1.7MB
      tig:
        atl: CachacaNER: a dataset for named entity recognition in texts about the cachaça beverage.
      aug:
        au:
          Silva, Priscilla
          Franco, Arthur
          Santos, Thiago
          Brito, Mozar
          Pereira, Denilson
        affil:
          https://ror.org/0122bmm03 Department of Computer Science, Federal University of Lavras, P.O. Box 3037, 37200-900, Lavras, MG, Brazil
          https://ror.org/0122bmm03 Department of Agroindustrial Management, Federal University of Lavras, P.O. Box 3037, 37200-900, Lavras, MG, Brazil
      su:
        Machine learning
        Supervised learning
        Portuguese language
        Text recognition
        English language
      sug:
        subj:
          Machine learning
          Supervised learning
          Portuguese language
          Text recognition
          English language
      keyword:
        Cachaça
        Dataset
        Labeled data
        Named entity recognition
        NER
      ab: Named Entity Recognition (NER) is the task of identifying and classifying tokens in texts corresponding to a set of pre-defined categories, such as names of people, organizations and locations. Datasets labeled for this task are essential for training supervised machine learning models. Although there are many datasets labeled with texts for English, in the Portuguese language they are scarcer. This work contributes to the creation and evaluation of a manually labeled dataset for the NER task, with texts in Brazilian Portuguese, in the specific domain of the beverage called Cachaça. This is a popular drink in Brazil, and of great economic importance. This is the first NER dataset in the beverage domain, and can be useful for other types of beverages with similar entity categories, such as wine and beer. We describe the process of data collection, creation of the dataset and its experimental evaluation. As a result, we created a dataset containing over 180,000 tokens labeled in 17 entity categories. The labeling obtained an agreement coefficient of 0.857 among the labelers, according to the Fleiss' Kappa metric, which is considered almost perfect. In our experimental evaluation, we obtained a micro-F1 value equal to 0.933 in the test set. The size of the dataset, as well as the result of its experimental evaluation, are comparable to other datasets in the Portuguese language, even though ours has a greater number of entity categories.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2024. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2024
    holdings:
      @attributes:
        islocal: N