Downs and Acrosses: Textual Markup on a Stroke Level.

Textual encoding is one of the main focuses of Humanities Computing. However, existing encoding schemes and initiatives focus on `text' from the character level upwards, and are of little use to scholars, such as papyrologists and palaeographers, who study the constituent strokes of individual chara...

Descripción completa

Detalles Bibliográficos
Publicado en:Literary & Linguistic Computing Vol. 19; no. 3; pp. 397 - 415
Autores principales: Terras, Melissas, Robertson, Paul
Formato: Artículo
Publicado: Oxford University Press / USA Sep2004
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=14385816&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 14385816
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        02681145
        BJ1
      jtl: Literary & Linguistic Computing
      issn: 02681145
      maglogo: N
    pubinfo:
      dt: Sep2004
      vid: 19
      iid: 3
      pid: 622
      pub: Oxford University Press / USA
    artinfo:
      ui:
        14385816
        10.1093/llc/19.3.397
      ppf: 397
      ppct: 18
      formats:
        fmt:
          @attributes:
            type: P
            size: 1.1MB
      tig:
        atl: Downs and Acrosses: Textual Markup on a Stroke Level.
      aug:
        au:
          Terras, Melissas
          Robertson, Paul
        affil:
          University College London, UK
          Massachusetts Institute of Technology, USA
      su:
        Document markup languages
        Text processing (Computer science)
        XML (Extensible Markup Language)
        Computational linguistics
        Word processing
        Paleography
      sug:
        subj:
          Document markup languages
          Text processing (Computer science)
          XML (Extensible Markup Language)
          Computational linguistics
          Word processing
          Paleography
      ab: Textual encoding is one of the main focuses of Humanities Computing. However, existing encoding schemes and initiatives focus on `text' from the character level upwards, and are of little use to scholars, such as papyrologists and palaeographers, who study the constituent strokes of individual characters. This paper discusses the development of a markup system used to annotate a corpus of images of Roman texts, resulting in an XML representation of each character on a stroke by stroke basis. The XML data generated allows further interrogation of the palaeographic data, increasing the knowledge available regarding the palaeography of the documentation produced by the Roman Army. Additionally, the corpus was used to train an Artificial Intelligence system to effectively 'read' in stroke data of unknown text and output possible, reliable, interpretations of that text: the next step in aiding historians in the reading of ancient texts. The development and implementation of the markup scheme is introduced, the results of our initial encoding effort are presented, and it is demonstrated that textual markup on a stroke level can extend the remit of marked-up digital texts in the humanities.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: © 2019 EADH: The European Association for Digital Humanities.
      item: Literary & Linguistic Computing
      holder: Oxford University Press / USA
      dt:
        @attributes:
          year: 2004
    holdings:
      @attributes:
        islocal: N