Comparing web-crawled and traditional corpora.
Using a multi-dimensional (MD) analysis of register variability, the study compares two corpora of Czech: Koditex, a "traditional" corpus carefully designed using various sources with rich metadata, and Araneum Bohemicum Maximum, a web-crawled corpus with an opportunistic composition representative...
| Published in: | Language Resources & Evaluation Vol. 54; no. 3; pp. 713 - 746 |
|---|---|
| Main Authors: | , , , , , , |
| Format: | Article |
| Published: |
Springer Nature
Sep2020
|
| Subjects: | |
| Online Access: | View this record in EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=144950773&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 144950773 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Sep2020 vid: 54 iid: 3 pid: 237 pub: Springer Nature artinfo: ui: 144950773 10.1007/s10579-020-09487-4 ppf: 713 ppct: 33 formats: fmt: – @attributes: type: T – @attributes: type: P size: 816KB tig: atl: Comparing web-crawled and traditional corpora. aug: au: Cvrček, Václav Komrsková, Zuzana Lukeš, David Poukarová, Petra Řehořková, Anna Zasina, Adrian Jan Benko, Vladimír affil: Institute of the Czech National Corpus, Faculty of Arts, Charles University, Prague, Czech Republic Ľ. Štúr Institute of Linguistics, Slovak Academy of Sciences, Bratislava, Slovakia UNESCO Chair in Plurilingual and Multicultural Communication, Comenius University in Bratislava, Bratislava, Slovakia su: Meta Platforms Inc. Corpora User-generated content Metadata sug: subj: Meta Platforms Inc. Corpora User-generated content Metadata keyword: Crawling Czech Multi-dimensional analysis Register Variation Web corpus ab: Using a multi-dimensional (MD) analysis of register variability, the study compares two corpora of Czech: Koditex, a "traditional" corpus carefully designed using various sources with rich metadata, and Araneum Bohemicum Maximum, a web-crawled corpus with an opportunistic composition representative of the "searchable" web. Both types of corpora are projected onto the space induced by the MD model, with the main objective being to find out whether they overlap in the linguistic variation they cover, or whether one introduces some specific variation which cannot be found in the other. We also document a crucial methodological point which has broader relevance for MD analyses in general, namely that texts have to be of similar lengths in order for their scores on the dimensions to be comparable. Results indicate that some traditional text categories, such as journalism or non-fiction, are characterized by language phenomena which are equally well covered by web-crawled data, though of course traditional corpora keep their edge in terms of the richness of the accompanying metadata. But overall, the range of variation in Koditex is broader as it contains texts which have no adequate substitute (i.e. texts with a comparable set of linguistic characteristics, regardless of their extratextual label) in data acquired through general-purpose web-crawling techniques. These include informal conversations, private correspondence, some types of fiction, but also user-generated content (comments on Facebook, forums etc.). pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2020. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2020 holdings: @attributes: islocal: N |
|---|