A Statistical Learning Approach to Improving the Accuracy of Chinese Word Segmentation.
In Chinese, there is no delimiter separating successive words in a sentence. Chinese word segmentation, which is a process of identifying word boundaries in text, is an essential step for Chinese language processing There are different word segmentation algorithms. However, because of the irregulari...
| Publicado en: | Literary & Linguistic Computing Vol. 11; no. 2; pp. 87 - 93 |
|---|---|
| Autores principales: | , |
| Formato: | Artículo |
| Publicado: |
Oxford University Press / USA
Jun1996
|
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=79232905&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 79232905 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 02681145 BJ1 jtl: Literary & Linguistic Computing issn: 02681145 maglogo: N pubinfo: dt: Jun1996 vid: 11 iid: 2 pid: 622 pub: Oxford University Press / USA artinfo: ui: 79232905 10.1093/llc/11.2.87 ppf: 87 ppct: 6 formats: fmt: @attributes: type: P size: 559KB tig: atl: A Statistical Learning Approach to Improving the Accuracy of Chinese Word Segmentation. aug: au: Leung, Chi-Hong Kan, Wing-Kay affil: Department of Computer Science and Engineering, The Chinese University of Hong Kong sug: ab: In Chinese, there is no delimiter separating successive words in a sentence. Chinese word segmentation, which is a process of identifying word boundaries in text, is an essential step for Chinese language processing There are different word segmentation algorithms. However, because of the irregularities of syntactic and semantic features in Chinese, it is difficult to obtain word segmentation accuracy of 100%. To solve this problem, a statistical learning approach to improving the accuracy of Chinese word segmentation is proposed. Based on statistical correlations between incorrect segmented strings and their contexts, a number of rules governing the modification from incorrect segmented strings to correct ones are constructed. These rules can be applied to word segmentation results obtained by an automatic word segmentation algorithm. They can modify the word segmentation results and make them more accurate. Experimental results have shown that this approach is a practical method to assist in the process of automatic Chinese word segmentation. The rules constructed in the experiment are studied and the limitation of this approach is discussed. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: © 2019 EADH: The European Association for Digital Humanities. item: Literary & Linguistic Computing holder: Oxford University Press / USA dt: @attributes: year: 1996 holdings: @attributes: islocal: N |
|---|