The Janes project: language resources and tools for Slovene user generated content.

The paper presents the results of the Janes project, which aimed to develop language resources and tools for Slovene user generated content. The paper first describes the 200 million word Janes corpus, containing tweets, forum posts, news comments, user and talk pages from Wikipedia, and blogs and b...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 54; no. 1; pp. 223 - 247
Autores principales: Fišer, Darja, Ljubešić, Nikola, Erjavec, Tomaž
Formato: Artículo
Publicado: Springer Nature Mar2020
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=142203865&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 142203865
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Mar2020
      vid: 54
      iid: 1
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        142203865
        10.1007/s10579-018-9425-z
      ppf: 223
      ppct: 24
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 532KB
      tig:
        atl: The Janes project: language resources and tools for Slovene user generated content.
      aug:
        au:
          Fišer, Darja
          Ljubešić, Nikola
          Erjavec, Tomaž
        affil:
          Department of Translation, Faculty of Arts, University of Ljubljana, Aškerčeva cesta 2, 1000, Ljubljana, Slovenia
          Department of Knowledge Technologies, Jožef Stefan Institute, Jamova cesta 39, 1000, Ljubljana, Slovenia
          Department of Information and Communication Sciences, Faculty of Humanities and Social Sciences, University of Zagreb, Ivana Lučića 3, 10000, Zagreb, Croatia
      su:
        Wikipedia
        Online comments
        User-generated content
        Language & languages
      sug:
        subj:
          Wikipedia
          Online comments
          User-generated content
          Language & languages
      keyword:
        Corpora
        Manually annotated datasets
        Slovene language
        Text normalisation
        User generated content
      ab: The paper presents the results of the Janes project, which aimed to develop language resources and tools for Slovene user generated content. The paper first describes the 200 million word Janes corpus, containing tweets, forum posts, news comments, user and talk pages from Wikipedia, and blogs and blog comments, where each text is accompanied by rich metadata. The developed processing tools for Slovene user generated content are presented next, which include a tokeniser, word-normaliser, part-of-speech tagger and lemmatiser, and a named entity recogniser. A set of manually annotated datasets was also produced, both for tool training as well as for linguistic research. The developed resources and tools are made publicly available under Creative Commons licences in the repository of the CLARIN.SI research infrastructure and on GitHub, while the corpora are also available through the CLARIN.SI concordancers.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2020. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2020
    holdings:
      @attributes:
        islocal: N