A web-based Bengali news corpus for named entity recognition.
The rapid development of language resources and tools using machine learning techniques for less computerized languages requires appropriately tagged corpus. A tagged Bengali news corpus has been developed from the web archive of a widely read Bengali newspaper. A web crawler retrieves the web pages...
| Publicado en: | Language Resources & Evaluation Vol. 42; no. 2; pp. 173 - 183 |
|---|---|
| Autores principales: | , |
| Formato: | Artículo |
| Publicado: |
Springer Nature
May2008
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=33379637&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 33379637 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: May2008 vid: 42 iid: 2 pid: 237 pub: Springer Nature artinfo: ui: 33379637 10.1007/s10579-008-9064-x ppf: 173 ppct: 10 formats: fmt: @attributes: type: P size: 205KB tig: atl: A web-based Bengali news corpus for named entity recognition. aug: au: Ekbal, Asif Bandyopadhyay, Sivaji affil: Department of Computer Science and Engineering , Jadavpur University , Kolkata 700032 India su: Web development Language digital resources Machine learning HTML (Document markup language) News gathering Equipment & supplies sug: subj: Web development Language digital resources Machine learning HTML (Document markup language) News gathering Equipment & supplies keyword: Named entity Named entity recognition News corpus Web as corpus Web-based tagged Bengali news corpus ab: The rapid development of language resources and tools using machine learning techniques for less computerized languages requires appropriately tagged corpus. A tagged Bengali news corpus has been developed from the web archive of a widely read Bengali newspaper. A web crawler retrieves the web pages in Hyper Text Markup Language (HTML) format from the news archive. At present, the corpus contains approximately 34 million wordforms. Named Entity Recognition (NER) systems based on pattern based shallow parsing with or without using linguistic knowledge have been developed using a part of this corpus. The NER system that uses linguistic knowledge has performed better yielding highest F-Score values of 75.40%, 72.30%, 71.37%, and 70.13% for person, location, organization, and miscellaneous names, respectively. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2008. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2008 holdings: @attributes: islocal: N |
|---|