Multi-modal page stream segmentation with convolutional neural networks.
In recent years, (retro-)digitizing paper-based files became a major undertaking for private and public archives as well as an important task in electronic mailroom applications. As first steps, the workflow usually involves batch scanning and optical character recognition (OCR) of documents. In the...
| Publicado en: | Language Resources & Evaluation Vol. 55; no. 1; pp. 127 - 151 |
|---|---|
| Autores principales: | , |
| Formato: | Artículo |
| Publicado: |
Springer Nature
Mar2021
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=149616765&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 149616765 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Mar2021 vid: 55 iid: 1 pid: 237 pub: Springer Nature artinfo: ui: 149616765 10.1007/s10579-019-09476-2 ppf: 127 ppct: 24 formats: fmt: @attributes: type: P size: 1.2MB tig: atl: Multi-modal page stream segmentation with convolutional neural networks. aug: au: Wiedemann, Gregor Heyer, Gerhard affil: Department of Computer Science, Hamburg University, Vogt-Kölln-Str. 30, 22527, Hamburg, Germany Department of Computer Science, Leipzig University, Augustusplatz 9, 04109, Leipzig, Germany su: Convolutional neural networks Optical character recognition sug: subj: Convolutional neural networks Optical character recognition keyword: Convolutional neural nets Digital mailroom Document flow segmentation Page stream segmentation Text classification ab: In recent years, (retro-)digitizing paper-based files became a major undertaking for private and public archives as well as an important task in electronic mailroom applications. As first steps, the workflow usually involves batch scanning and optical character recognition (OCR) of documents. In the case of multi-page documents, the preservation of document contexts is a major requirement. To facilitate workflows involving very large amounts of paper scans, page stream segmentation (PSS) is the task to automatically separate a stream of scanned images into coherent multi-page documents. In a digitization project together with a German federal archive, we developed a novel approach for PSS based on convolutional neural networks (CNN). As a first project, we combine visual information from scanned images with semantic information from OCR-ed texts for this task. The multi-modal combination of features in a single classification architecture allows for major improvements towards optimal document separation. Further to multimodality, our PSS approach profits from transfer-learning and sequential page modeling. We achieve accuracy up to 95% on multi-page documents on our in-house dataset and up to 93% on a publicly available dataset. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2021. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2021 holdings: @attributes: islocal: N |
|---|