A Spanish dataset for reproducible benchmarked offline handwriting recognition.
In this paper, a public dataset for Offline Handwriting Recognition, along with an appropriate evaluation method to provide benchmark indicators at sentence level, is presented. This dataset, called SPA-Sentences, consists of offline handwritten Spanish sentences extracted from 1617 forms produced b...
| Publicado en: | Language Resources & Evaluation Vol. 56; no. 3; pp. 1009 - 1023 |
|---|---|
| Autores principales: | , |
| Formato: | Artículo |
| Publicado: |
Springer Nature
Sep2022
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=158609444&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 158609444 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Sep2022 vid: 56 iid: 3 pid: 237 pub: Springer Nature artinfo: ui: 158609444 10.1007/s10579-022-09587-3 ppf: 1009 ppct: 14 formats: fmt: – @attributes: type: T – @attributes: type: P size: 508KB tig: atl: A Spanish dataset for reproducible benchmarked offline handwriting recognition. aug: au: España-Boquera, Salvador Castro-Bleda, Maria Jose affil: VRAIN Valencian Research Institute for Artificial Intelligence, Universitat Politècnica de València, Valencia, Spain su: International Association of Machinists & Aerospace Workers Spanish language Long-term memory Handwriting sug: subj: International Association of Machinists & Aerospace Workers Spanish language Long-term memory Handwriting keyword: Benchmarking Connectionist temporal classification (CTC) Convolutional neural networks (CNN) Datasets Deep learning Evaluation Experimental reproducibility Handwriting recognition Long short term memory (LSTM) networks Offline handwriting recognition Spanish resources ab: In this paper, a public dataset for Offline Handwriting Recognition, along with an appropriate evaluation method to provide benchmark indicators at sentence level, is presented. This dataset, called SPA-Sentences, consists of offline handwritten Spanish sentences extracted from 1617 forms produced by the same number of writers. A total of 13,691 sentences comprising around 100,000 word instances out of a vocabulary of 3288 words occur in the collection. Careful attention has been paid to make the baseline experiments both reproducible and competitive. To this end, experiments are based on state-of-the-art recognition techniques combining convolutional blocks with one-dimensional Bidirectional Long Short Term Memory (LSTM) networks using Connectionist Temporal Classification (CTC) decoding. The scripts with the entire experimental setting have been made available. The SPA-Sentences dataset and its baseline evaluation are freely available for research purposes via the institutional University repository. We expect the research community to include this corpus, as is usually done with English IAM and French RIMES datasets, in their battery of experiments when reporting novel handwriting recognition techniques. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2022. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2022 holdings: @attributes: islocal: N |
|---|