A web-based Bengali news corpus for named entity recognition.

The rapid development of language resources and tools using machine learning techniques for less computerized languages requires appropriately tagged corpus. A tagged Bengali news corpus has been developed from the web archive of a widely read Bengali newspaper. A web crawler retrieves the web pages...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 42; no. 2; pp. 173 - 183
Autores principales: Ekbal, Asif, Bandyopadhyay, Sivaji
Formato: Artículo
Publicado: Springer Nature May2008
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=33379637&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 33379637
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: May2008
      vid: 42
      iid: 2
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        33379637
        10.1007/s10579-008-9064-x
      ppf: 173
      ppct: 10
      formats:
        fmt:
          @attributes:
            type: P
            size: 205KB
      tig:
        atl: A web-based Bengali news corpus for named entity recognition.
      aug:
        au:
          Ekbal, Asif
          Bandyopadhyay, Sivaji
        affil: Department of Computer Science and Engineering , Jadavpur University , Kolkata 700032 India
      su:
        Web development
        Language digital resources
        Machine learning
        HTML (Document markup language)
        News gathering
        Equipment & supplies
      sug:
        subj:
          Web development
          Language digital resources
          Machine learning
          HTML (Document markup language)
          News gathering
          Equipment & supplies
      keyword:
        Named entity
        Named entity recognition
        News corpus
        Web as corpus
        Web-based tagged Bengali news corpus
      ab: The rapid development of language resources and tools using machine learning techniques for less computerized languages requires appropriately tagged corpus. A tagged Bengali news corpus has been developed from the web archive of a widely read Bengali newspaper. A web crawler retrieves the web pages in Hyper Text Markup Language (HTML) format from the news archive. At present, the corpus contains approximately 34 million wordforms. Named Entity Recognition (NER) systems based on pattern based shallow parsing with or without using linguistic knowledge have been developed using a part of this corpus. The NER system that uses linguistic knowledge has performed better yielding highest F-Score values of 75.40%, 72.30%, 71.37%, and 70.13% for person, location, organization, and miscellaneous names, respectively.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2008. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2008
    holdings:
      @attributes:
        islocal: N