From legacy encodings to Unicode: the graphical and logical principles in the scripts of South Asia.

Much electronic text in the languages of South Asia has been published on the Internet. However, while Unicode has emerged as the favoured encoding system of corpus and computational linguists, most South Asian language data on the web uses one of a wide range of non-standard legacy encodings. This...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 41; no. 1; pp. 1 - 26
Autor principal: Hardie, Andrew
Formato: Artículo
Publicado: Springer Nature Feb2007
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=27190391&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 27190391
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Feb2007
      vid: 41
      iid: 1
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        27190391
        10.1007/s10579-006-9003-7
      ppf: 1
      ppct: 25
      formats:
        fmt:
          @attributes:
            type: P
            size: 203KB
      tig:
        atl: From legacy encodings to Unicode: the graphical and logical principles in the scripts of South Asia.
      aug:
        au: Hardie, Andrew
        affil: Department of Linguistics and English Language , University of Lancaster , Lancaster LA1 4YT UK
      su:
        SGML (Document markup language)
        Encoding
        Ethnology
        Internet service providers
        Unicode (Computer character set)
        South Asia
      sug:
        subj:
          South Asia
          SGML (Document markup language)
          Encoding
          Ethnology
          Internet service providers
          Unicode (Computer character set)
      keyword:
        Conjunct consonant
        Conversion
        Devanagari
        Font
        Legacy text
        South Asian languages/scripts
        Unicode
        Virama
        Vowel diacritic
      ab: Much electronic text in the languages of South Asia has been published on the Internet. However, while Unicode has emerged as the favoured encoding system of corpus and computational linguists, most South Asian language data on the web uses one of a wide range of non-standard legacy encodings. This paper describes the difficulties inherent in converting text in these encodings to Unicode. Among the various legacy encodings for South Asian scripts, the most problematic are 8-bit fonts based on graphical principles (as opposed to the logical principles of Unicode). Graphical fonts typically encode several features in ways highly incompatible with Unicode. For instance, half-form glyphs used to construct conjunct consonants are typically separate code points in 8-bit fonts; in Unicode they are represented by the full consonant followed by virama. There are many more such cases. The solution described here is an approach to text conversion based on mapping rules. A small number of generalised rules (plus the capacity for more specialised rules) captures the behaviour of each character in a font, building up a conversion algorithm for that encoding. This system is embedded in a font-mapping program, outputting CES-compliant SGML Unicode. This program, a generalised text-conversion tool, has been employed extensively in corpus-building for South Asian languages.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2007. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2007
    holdings:
      @attributes:
        islocal: N