UFLA-FORMS: an academic forms dataset for information extraction in the Portuguese language.

Information Extraction aims to analyze and extract relevant information in document samples. For visual documents, such as academic and commercial forms, key-value pair extraction is capable of extracting and grouping the requested information automatically. The state of the art presents multimodal...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 59; no. 3; pp. 2185 - 2212
Autores principales: Gonçalves Lima, Victor, Alves Pereira, Denilson
Formato: Artículo
Publicado: Springer Nature Sep2025
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=186909061&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 186909061
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Sep2025
      vid: 59
      iid: 3
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        186909061
        10.1007/s10579-024-09802-3
      ppf: 2185
      ppct: 27
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 2.1MB
      tig:
        atl: UFLA-FORMS: an academic forms dataset for information extraction in the Portuguese language.
      aug:
        au:
          Gonçalves Lima, Victor
          Alves Pereira, Denilson
        affil: https://ror.org/0122bmm03 Department of Computer Science, Federal University of Lavras, P.O. Box 3037, 37203-202, Lavras, MG, Brazil
      su:
        Portuguese language
        Data mining
        Data analysis
      sug:
        subj:
          Portuguese language
          Data mining
          Data analysis
      keyword:
        Dataset
        Document intelligence
        Information and Computing Sciences Artificial Intelligence and Image Processing
        Information extraction
        NLP
      ab: Information Extraction aims to analyze and extract relevant information in document samples. For visual documents, such as academic and commercial forms, key-value pair extraction is capable of extracting and grouping the requested information automatically. The state of the art presents multimodal models capable of analyzing document's image, text and layout. To study and use these models in downstream tasks, labeled datasets must be available for fine-tuning of models. However, publicly available data is scarce, even more so for the Portuguese language. In this context, this work aims to provide a dataset of academic forms, named UFLA-FORMS, for the task of Information Extraction. The dataset is composed of 200 manually labeled samples, containing 7710 entities with 4442 relationships pairs between them. The labeling obtained a Kappa coefficient of agreement of 0.918 for the labeling of entities and 0.909 for the relationships attributed between them. The dataset was experimentally evaluated through cross-validation with hyper-parameter search in Named Entity Recognition and Relation Extraction tasks, obtaining, respectively, an average F 1 of 0.921 and 0.846.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2025. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2025
    holdings:
      @attributes:
        islocal: N