Constructing specialised corpora through analysing domain representativeness of websites.
The role of the Web for text corpus construction is becoming increasingly significant. However, the contribution of the Web is largely confined to building a general virtual corpus or low quality specialised corpora. In this paper, we introduce a new technique called SPARTAN for constructing special...
| Publicado en: | Language Resources & Evaluation Vol. 45; no. 2; pp. 209 - 242 |
|---|---|
| Autores principales: | , , |
| Formato: | Artículo |
| Publicado: |
Springer Nature
May2011
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=60133437&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 60133437 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: May2011 vid: 45 iid: 2 pid: 237 pub: Springer Nature artinfo: ui: 60133437 10.1007/s10579-011-9141-4 ppf: 209 ppct: 33 formats: fmt: @attributes: type: P size: 472KB tig: atl: Constructing specialised corpora through analysing domain representativeness of websites. aug: au: Wong, Wilson Liu, Wei Bennamoun, Mohammed affil: School of Computer Science and Software Engineering, The University of Western Australia, Crawley 6009 Australia su: Rankings of websites World Wide Web Virtual communities Search engines Corpora Computational linguistics Electronic information resource searching sug: subj: Rankings of websites World Wide Web Virtual communities Search engines Corpora Computational linguistics Electronic information resource searching keyword: Boilerplate removal Corpus construction Specialised corpus Term recognition Virtual corpus Web-derived corpus Website ranking ab: The role of the Web for text corpus construction is becoming increasingly significant. However, the contribution of the Web is largely confined to building a general virtual corpus or low quality specialised corpora. In this paper, we introduce a new technique called SPARTAN for constructing specialised corpora from the Web by systematically analysing website contents. Our evaluations show that the corpora constructed using our technique are independent of the search engines employed. In particular, SPARTAN-derived corpora outperform all corpora based on existing techniques for the task of term recognition. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2011. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2011 holdings: @attributes: islocal: N |
|---|