Word Alignment in English–Chinese Parallel Corpora.
Word alignment in bilingual or multilingual parallel corpora has been a challenging issue for natural language engineering. An efficient algorithm for automatically aligning word translation equivalents across different languages will be of use for a number of practical applications such as multilin...
| Publicado en: | Literary & Linguistic Computing Vol. 17; no. 2; pp. 207 - 231 |
|---|---|
| Autor principal: | |
| Formato: | Artículo |
| Publicado: |
Oxford University Press / USA
Jun2002
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=44441552&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 44441552 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 02681145 BJ1 jtl: Literary & Linguistic Computing issn: 02681145 maglogo: N pubinfo: dt: Jun2002 vid: 17 iid: 2 pid: 622 pub: Oxford University Press / USA artinfo: ui: 44441552 10.1093/llc/17.2.207 ppf: 207 ppct: 24 formats: fmt: @attributes: type: P size: 395KB tig: atl: Word Alignment in English–Chinese Parallel Corpora. aug: au: Piao, Scott Songlin affil: Department of Linguistics and Modern English Language, Lancaster University, UK su: Algorithms Sentences (Grammar) Bilingualism Loanwords Language policy Translating & interpreting Language arts Corpora Linguistic analysis sug: subj: Algorithms Sentences (Grammar) Bilingualism Loanwords Language policy Translating & interpreting Language arts Corpora Linguistic analysis ab: Word alignment in bilingual or multilingual parallel corpora has been a challenging issue for natural language engineering. An efficient algorithm for automatically aligning word translation equivalents across different languages will be of use for a number of practical applications such as multilingual lexical construction, machine translation, etc. This paper presents a hybrid algorithm for English–Chinese word alignment, which incorporates co‐occurrence association measures, word distribution distances, English word lemmatization, and part‐of‐speech information. Eleven co‐occurrence association coefficients and eight distance measures of word distribution are explored to compare their efficiency for word alignment. The paper also describes an experiment in which the algorithm is evaluated on sentence‐aligned English–Chinese parallel corpora. In the experiment, the algorithm produced encouraging success rates on two test corpora, with the highest success rate of 89.37 per cent. It provides a practical tool for extracting word translation equivalents from English–Chinese parallel corpora. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: © 2019 EADH: The European Association for Digital Humanities. item: Literary & Linguistic Computing holder: Oxford University Press / USA dt: @attributes: year: 2002 holdings: @attributes: islocal: N |
|---|