The Janes project: language resources and tools for Slovene user generated content.
The paper presents the results of the Janes project, which aimed to develop language resources and tools for Slovene user generated content. The paper first describes the 200 million word Janes corpus, containing tweets, forum posts, news comments, user and talk pages from Wikipedia, and blogs and b...
| Publicado en: | Language Resources & Evaluation Vol. 54; no. 1; pp. 223 - 247 |
|---|---|
| Autores principales: | , , |
| Formato: | Artículo |
| Publicado: |
Springer Nature
Mar2020
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=142203865&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 142203865 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Mar2020 vid: 54 iid: 1 pid: 237 pub: Springer Nature artinfo: ui: 142203865 10.1007/s10579-018-9425-z ppf: 223 ppct: 24 formats: fmt: – @attributes: type: T – @attributes: type: P size: 532KB tig: atl: The Janes project: language resources and tools for Slovene user generated content. aug: au: Fišer, Darja Ljubešić, Nikola Erjavec, Tomaž affil: Department of Translation, Faculty of Arts, University of Ljubljana, Aškerčeva cesta 2, 1000, Ljubljana, Slovenia Department of Knowledge Technologies, Jožef Stefan Institute, Jamova cesta 39, 1000, Ljubljana, Slovenia Department of Information and Communication Sciences, Faculty of Humanities and Social Sciences, University of Zagreb, Ivana Lučića 3, 10000, Zagreb, Croatia su: Wikipedia Online comments User-generated content Language & languages sug: subj: Wikipedia Online comments User-generated content Language & languages keyword: Corpora Manually annotated datasets Slovene language Text normalisation User generated content ab: The paper presents the results of the Janes project, which aimed to develop language resources and tools for Slovene user generated content. The paper first describes the 200 million word Janes corpus, containing tweets, forum posts, news comments, user and talk pages from Wikipedia, and blogs and blog comments, where each text is accompanied by rich metadata. The developed processing tools for Slovene user generated content are presented next, which include a tokeniser, word-normaliser, part-of-speech tagger and lemmatiser, and a named entity recogniser. A set of manually annotated datasets was also produced, both for tool training as well as for linguistic research. The developed resources and tools are made publicly available under Creative Commons licences in the repository of the CLARIN.SI research infrastructure and on GitHub, while the corpora are also available through the CLARIN.SI concordancers. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2020. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2020 holdings: @attributes: islocal: N |
|---|