A word sense disambiguation corpus for Urdu.
The aim of word sense disambiguation (WSD) is to correctly identify the meaning of a word in context. All natural languages exhibit word sense ambiguities and these are often hard to resolve automatically. Consequently WSD is considered an important problem in natural language processing (NLP). Stan...
| Publicado en: | Language Resources & Evaluation Vol. 53; no. 3; pp. 397 - 419 |
|---|---|
| Autores principales: | , , , |
| Formato: | Artículo |
| Publicado: |
Springer Nature
Sep2019
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=138297592&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 138297592 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Sep2019 vid: 53 iid: 3 pid: 237 pub: Springer Nature artinfo: ui: 138297592 10.1007/s10579-018-9438-7 ppf: 397 ppct: 22 formats: fmt: – @attributes: type: T – @attributes: type: P size: 1.5MB tig: atl: A word sense disambiguation corpus for Urdu. aug: au: Saeed, Ali Nawab, Rao Muhammad Adeel Stevenson, Mark Rayson, Paul affil: COMSATS University Islamabad, Lahore, Pakistan University of Sheffield, Sheffield, UK Lancaster University, Lancaster, UK su: Natural language processing Natural languages Semantics Bag design Corpora Language research sug: subj: Natural language processing Natural languages Semantics Bag design Corpora Language research keyword: Lexical sample task Sense tagged Urdu corpus Word sense disambiguation ab: The aim of word sense disambiguation (WSD) is to correctly identify the meaning of a word in context. All natural languages exhibit word sense ambiguities and these are often hard to resolve automatically. Consequently WSD is considered an important problem in natural language processing (NLP). Standard evaluation resources are needed to develop, evaluate and compare WSD methods. A range of initiatives have lead to the development of benchmark WSD corpora for a wide range of languages from various language families. However, there is a lack of benchmark WSD corpora for South Asian languages including Urdu, despite there being over 300 million Urdu speakers and a large amounts of Urdu digital text available online. To address that gap, this study describes a novel benchmark corpus for the Urdu Lexical Sample WSD task. This corpus contains 50 target words (30 nouns, 11 adjectives, and 9 verbs). A standard, manually crafted dictionary called Urdu Lughat is used as a sense inventory. Four baseline WSD approaches were applied to the corpus. The results show that the best performance was obtained using a simple Bag of Words approach. To encourage NLP research on the Urdu language the corpus is freely available to the research community. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2019. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2019 holdings: @attributes: islocal: N |
|---|