Urdu part of speech tagging using conditional random fields.

Part of speech (POS) tagging, the assignment of syntactic categories for words in running text, is significant to natural language processing as a preliminary task in applications such as speech processing, information extraction, and others. Urdu language processing presents a challenge due to the...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 53; no. 3; pp. 331 - 363
Autores principales: Khan, Wahab, Daud, Ali, Nasir, Jamal Abdul, Amjad, Tehmina, Arafat, Sachi, Aljohani, Naif, Alotaibi, Fahd S.
Formato: Artículo
Publicado: Springer Nature Sep2019
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=138297593&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 138297593
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Sep2019
      vid: 53
      iid: 3
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        138297593
        10.1007/s10579-018-9439-6
      ppf: 331
      ppct: 32
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 1.4MB
      tig:
        atl: Urdu part of speech tagging using conditional random fields.
      aug:
        au:
          Khan, Wahab
          Daud, Ali
          Nasir, Jamal Abdul
          Amjad, Tehmina
          Arafat, Sachi
          Aljohani, Naif
          Alotaibi, Fahd S.
        affil:
          Department of Computer Science and Software Engineering, IIU, 44000, Islamabad, Pakistan
          Faculty of Computing and Information Technology, King Abdulaziz University, Jeddah, Saudi Arabia
      su:
        Conditional random fields
        Parts of speech
        Natural language processing
        Support vector machines
        Data mining
      sug:
        subj:
          Conditional random fields
          Parts of speech
          Natural language processing
          Support vector machines
          Data mining
      keyword:
        Conditional random field (CRF)
        Part of speech (POS)
        Support vector machine (SVM)
        Urdu
      ab: Part of speech (POS) tagging, the assignment of syntactic categories for words in running text, is significant to natural language processing as a preliminary task in applications such as speech processing, information extraction, and others. Urdu language processing presents a challenge due to the dual behaviour of various Urdu POS tags in differing situations (morphosyntactic ambiguity). This paper addresses this challenge by developing a novel tagging approach using linear-chain conditional random fields (CRF). Our work is the first instance of a CRF approach for Urdu POS tagging. The proposed model employs a strong, stable and balanced language-independent as well as language dependent feature set. The language-dependent feature considered includes part-of-speech tag of the previous word and suffix of the current word while the language-independent features includes the 'context words window'. Our approach was evaluated against support vector machine techniques for Urdu POS—considered as state of the art—on two benchmark datasets. The results show our CRF approach to improve upon the F-measure of prior attempts by 8.3–8.5%.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2019. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2019
    holdings:
      @attributes:
        islocal: N