UFLA-FORMS: an academic forms dataset for information extraction in the Portuguese language.
Information Extraction aims to analyze and extract relevant information in document samples. For visual documents, such as academic and commercial forms, key-value pair extraction is capable of extracting and grouping the requested information automatically. The state of the art presents multimodal...
| Publicado en: | Language Resources & Evaluation Vol. 59; no. 3; pp. 2185 - 2212 |
|---|---|
| Autores principales: | , |
| Formato: | Artículo |
| Publicado: |
Springer Nature
Sep2025
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=186909061&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 186909061 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Sep2025 vid: 59 iid: 3 pid: 237 pub: Springer Nature artinfo: ui: 186909061 10.1007/s10579-024-09802-3 ppf: 2185 ppct: 27 formats: fmt: – @attributes: type: T – @attributes: type: P size: 2.1MB tig: atl: UFLA-FORMS: an academic forms dataset for information extraction in the Portuguese language. aug: au: Gonçalves Lima, Victor Alves Pereira, Denilson affil: https://ror.org/0122bmm03 Department of Computer Science, Federal University of Lavras, P.O. Box 3037, 37203-202, Lavras, MG, Brazil su: Portuguese language Data mining Data analysis sug: subj: Portuguese language Data mining Data analysis keyword: Dataset Document intelligence Information and Computing Sciences Artificial Intelligence and Image Processing Information extraction NLP ab: Information Extraction aims to analyze and extract relevant information in document samples. For visual documents, such as academic and commercial forms, key-value pair extraction is capable of extracting and grouping the requested information automatically. The state of the art presents multimodal models capable of analyzing document's image, text and layout. To study and use these models in downstream tasks, labeled datasets must be available for fine-tuning of models. However, publicly available data is scarce, even more so for the Portuguese language. In this context, this work aims to provide a dataset of academic forms, named UFLA-FORMS, for the task of Information Extraction. The dataset is composed of 200 manually labeled samples, containing 7710 entities with 4442 relationships pairs between them. The labeling obtained a Kappa coefficient of agreement of 0.918 for the labeling of entities and 0.909 for the relationships attributed between them. The dataset was experimentally evaluated through cross-validation with hyper-parameter search in Named Entity Recognition and Relation Extraction tasks, obtaining, respectively, an average F 1 of 0.921 and 0.846. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2025. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2025 holdings: @attributes: islocal: N |
|---|