Efficient corpus development for lexicography: building the New Corpus for Ireland.
In a 12-month project we have developed a new, register-diverse, 55-million-word bilingual corpus—the New Corpus for Ireland (NCI)—to support the creation of a new English-to-Irish dictionary. The paper describes the strategies we employed, and the solutions to problems encountered. We believe we ha...
| Publicado en: | Language Resources & Evaluation Vol. 40; no. 2; pp. 127 - 153 |
|---|---|
| Autores principales: | , , |
| Formato: | Artículo |
| Publicado: |
Springer Nature
May2006
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=24151661&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 24151661 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: May2006 vid: 40 iid: 2 pid: 237 pub: Springer Nature artinfo: ui: 24151661 10.1007/s10579-006-9011-7 ppf: 127 ppct: 26 formats: fmt: @attributes: type: P size: 380KB tig: atl: Efficient corpus development for lexicography: building the New Corpus for Ireland. aug: au: Kilgarriff, Adam Rundell, Michael Dhonnchadha, Elaine Uí affil: Lexicography MasterClass Ltd, Brighton, UK. Trinity College, Dublin, Ireland. su: Corpora Lexicography Computational linguistics Natural language processing Encyclopedias & dictionaries Language & languages sug: subj: Corpora Lexicography Computational linguistics Natural language processing Encyclopedias & dictionaries Language & languages keyword: Corpus linguistics Dictionaries Gaelic Hiberno-English Irish Language technology ab: In a 12-month project we have developed a new, register-diverse, 55-million-word bilingual corpus—the New Corpus for Ireland (NCI)—to support the creation of a new English-to-Irish dictionary. The paper describes the strategies we employed, and the solutions to problems encountered. We believe we have a good model for corpus creation for lexicography, and others may find it useful as a blueprint. The corpus has two parts, one Irish, the other Hiberno-English (English as spoken in Ireland). We describe its design, collection and encoding. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2006. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2006 holdings: @attributes: islocal: N |
|---|