PROCESSING NATURAL MALAY TEXTS: A DATA-DRIVEN APPROACH.

This research represents the first attempt to produce a working system for the automatic processing of texts of Bahasa Melayu 'Malay'. At the heart of the system is an integrated relational lexical database called MALEX, which draws on the experience of working on English and other languages, but wh...

Descripción completa

Detalles Bibliográficos
Publicado en:TRAMES: A Journal of the Humanities & Social Sciences Vol. 14; no. 1; pp. 90 - 104
Autor principal: Don, Zuraidah Mohd
Formato: Artículo
Publicado: Teaduste Akadeemia Kirjastus 2010
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ssf&AN=48679420&site=ehost-live
header:
  @attributes:
    shortDbName: ssf
    uiTerm: 48679420
    longDbName: Social Sciences Full Text (H.W. Wilson)
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        14060922
        DT1
      jtl: TRAMES: A Journal of the Humanities & Social Sciences
      issn: 14060922
      maglogo: N
    pubinfo:
      dt: 2010
      vid: 14
      iid: 1
      pid: 11112
      pub: Teaduste Akadeemia Kirjastus
    artinfo:
      ui:
        48679420
        10.3176/tr.2010.1.06
      ppf: 90
      ppct: 14
      formats:
        fmt:
          @attributes:
            type: P
            size: 104KB
      tig:
        atl: PROCESSING NATURAL MALAY TEXTS: A DATA-DRIVEN APPROACH.
      aug:
        au: Don, Zuraidah Mohd
        affil: University of Malaya
      su:
        Written communication
        Linguistics
        Malay language
        Text processing (Computer science)
        Databases
        Corpora
        Lexicon
      sug:
        subj:
          Written communication
          Linguistics
          Malay language
          Text processing (Computer science)
          Databases
          Corpora
          Lexicon
      keyword:
        corpus
        lexicon
        Malay
        part of speech
        text
        corpus
        lexicon
        Malay
        part of speech
        text
      ab: This research represents the first attempt to produce a working system for the automatic processing of texts of Bahasa Melayu 'Malay'. At the heart of the system is an integrated relational lexical database called MALEX, which draws on the experience of working on English and other languages, but which is specifically tailored to the conditions of Malay. The development of the database is from the beginning entirely data driven, and is based on the analysis of a corpus of naturally produced Malay texts. In designing procedures which access the database, properties of the text are consistently and rigorously distinguished from properties of the lexicon and of the grammar. The system is currently used to provide information for a range of applications, for grammatical tagging, stemming and lemmatisation, parsing, and for generating phonological representations. It is hoped and intended that the design features of MALEX will be transferable, and provide a model for the development of working systems for other Asian languages.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: N
    holdings:
      @attributes:
        islocal: N