Natural language processing and early-modern dirty data: applying IBM Languageware to the 1641 depositions.
This article provides an account of the steps involved in adapting IBM's Languageware natural language processing software to a large corpus of highly non-standard 17th century documents. It examines the challenges encountered as part of this process, and outlines the approach adopted to provide a r...
| Publicado en: | Literary & Linguistic Computing Vol. 27; no. 1; pp. 39 - 55 |
|---|---|
| Autores principales: | , |
| Formato: | Artículo |
| Publicado: |
Oxford University Press / USA
Apr2012
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| Sumario: | This article provides an account of the steps involved in adapting IBM's Languageware natural language processing software to a large corpus of highly non-standard 17th century documents. It examines the challenges encountered as part of this process, and outlines the approach adopted to provide a robust and reusable tool for the linguistic analysis of early modern source texts. |
|---|