A Statistical Learning Approach to Improving the Accuracy of Chinese Word Segmentation.

In Chinese, there is no delimiter separating successive words in a sentence. Chinese word segmentation, which is a process of identifying word boundaries in text, is an essential step for Chinese language processing There are different word segmentation algorithms. However, because of the irregulari...

Descripción completa

Detalles Bibliográficos
Publicado en:Literary & Linguistic Computing Vol. 11; no. 2; pp. 87 - 93
Autores principales: Leung, Chi-Hong, Kan, Wing-Kay
Formato: Artículo
Publicado: Oxford University Press / USA Jun1996
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=79232905&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 79232905
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        02681145
        BJ1
      jtl: Literary & Linguistic Computing
      issn: 02681145
      maglogo: N
    pubinfo:
      dt: Jun1996
      vid: 11
      iid: 2
      pid: 622
      pub: Oxford University Press / USA
    artinfo:
      ui:
        79232905
        10.1093/llc/11.2.87
      ppf: 87
      ppct: 6
      formats:
        fmt:
          @attributes:
            type: P
            size: 559KB
      tig:
        atl: A Statistical Learning Approach to Improving the Accuracy of Chinese Word Segmentation.
      aug:
        au:
          Leung, Chi-Hong
          Kan, Wing-Kay
        affil: Department of Computer Science and Engineering, The Chinese University of Hong Kong
      sug:
      ab: In Chinese, there is no delimiter separating successive words in a sentence. Chinese word segmentation, which is a process of identifying word boundaries in text, is an essential step for Chinese language processing There are different word segmentation algorithms. However, because of the irregularities of syntactic and semantic features in Chinese, it is difficult to obtain word segmentation accuracy of 100%. To solve this problem, a statistical learning approach to improving the accuracy of Chinese word segmentation is proposed. Based on statistical correlations between incorrect segmented strings and their contexts, a number of rules governing the modification from incorrect segmented strings to correct ones are constructed. These rules can be applied to word segmentation results obtained by an automatic word segmentation algorithm. They can modify the word segmentation results and make them more accurate. Experimental results have shown that this approach is a practical method to assist in the process of automatic Chinese word segmentation. The rules constructed in the experiment are studied and the limitation of this approach is discussed.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: © 2019 EADH: The European Association for Digital Humanities.
      item: Literary & Linguistic Computing
      holder: Oxford University Press / USA
      dt:
        @attributes:
          year: 1996
    holdings:
      @attributes:
        islocal: N