Downs and Acrosses: Textual Markup on a Stroke Level.
Textual encoding is one of the main focuses of Humanities Computing. However, existing encoding schemes and initiatives focus on `text' from the character level upwards, and are of little use to scholars, such as papyrologists and palaeographers, who study the constituent strokes of individual chara...
| Publicado en: | Literary & Linguistic Computing Vol. 19; no. 3; pp. 397 - 415 |
|---|---|
| Autores principales: | , |
| Formato: | Artículo |
| Publicado: |
Oxford University Press / USA
Sep2004
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=14385816&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 14385816 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 02681145 BJ1 jtl: Literary & Linguistic Computing issn: 02681145 maglogo: N pubinfo: dt: Sep2004 vid: 19 iid: 3 pid: 622 pub: Oxford University Press / USA artinfo: ui: 14385816 10.1093/llc/19.3.397 ppf: 397 ppct: 18 formats: fmt: @attributes: type: P size: 1.1MB tig: atl: Downs and Acrosses: Textual Markup on a Stroke Level. aug: au: Terras, Melissas Robertson, Paul affil: University College London, UK Massachusetts Institute of Technology, USA su: Document markup languages Text processing (Computer science) XML (Extensible Markup Language) Computational linguistics Word processing Paleography sug: subj: Document markup languages Text processing (Computer science) XML (Extensible Markup Language) Computational linguistics Word processing Paleography ab: Textual encoding is one of the main focuses of Humanities Computing. However, existing encoding schemes and initiatives focus on `text' from the character level upwards, and are of little use to scholars, such as papyrologists and palaeographers, who study the constituent strokes of individual characters. This paper discusses the development of a markup system used to annotate a corpus of images of Roman texts, resulting in an XML representation of each character on a stroke by stroke basis. The XML data generated allows further interrogation of the palaeographic data, increasing the knowledge available regarding the palaeography of the documentation produced by the Roman Army. Additionally, the corpus was used to train an Artificial Intelligence system to effectively 'read' in stroke data of unknown text and output possible, reliable, interpretations of that text: the next step in aiding historians in the reading of ancient texts. The development and implementation of the markup scheme is introduced, the results of our initial encoding effort are presented, and it is demonstrated that textual markup on a stroke level can extend the remit of marked-up digital texts in the humanities. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: © 2019 EADH: The European Association for Digital Humanities. item: Literary & Linguistic Computing holder: Oxford University Press / USA dt: @attributes: year: 2004 holdings: @attributes: islocal: N |
|---|