Microblog language identification: overcoming the limitations of short, unedited and idiomatic text.

Multilingual posts can potentially affect the outcomes of content analysis on microblog platforms. To this end, language identification can provide a monolingual set of content for analysis. We find the unedited and idiomatic language of microblogs to be challenging for state-of-the-art language ide...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 47; no. 1; pp. 195 - 216
Autores principales: Carter, Simon, Weerkamp, Wouter, Tsagkias, Manos
Formato: Artículo
Publicado: Springer Nature Mar2013
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=85873232&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 85873232
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Mar2013
      vid: 47
      iid: 1
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        85873232
        10.1007/s10579-012-9195-y
      ppf: 195
      ppct: 21
      formats:
        fmt:
          @attributes:
            type: P
            size: 428KB
      tig:
        atl: Microblog language identification: overcoming the limitations of short, unedited and idiomatic text.
      aug:
        au:
          Carter, Simon
          Weerkamp, Wouter
          Tsagkias, Manos
        affil: ISLA, University of Amsterdam, Science Park 904 1098 XH Amsterdam The Netherlands
      su:
        Linguistic identity
        Microblogs
        Instant messaging
        Online social networks
        Content analysis
      sug:
        subj:
          Linguistic identity
          Microblogs
          Instant messaging
          Online social networks
          Content analysis
      keyword:
        Language identification
        Text classification
      ab: Multilingual posts can potentially affect the outcomes of content analysis on microblog platforms. To this end, language identification can provide a monolingual set of content for analysis. We find the unedited and idiomatic language of microblogs to be challenging for state-of-the-art language identification methods. To account for this, we identify five microblog characteristics that can help in language identification: the language profile of the blogger (blogger), the content of an attached hyperlink (link), the language profile of other users mentioned (mention) in the post, the language profile of a tag (tag), and the language of the original post (conversation), if the post we examine is a reply. Further, we present methods that combine these priors in a post-dependent and post-independent way. We present test results on 1,000 posts from five languages (Dutch, English, French, German, and Spanish), which show that our priors improve accuracy by 5 % over a domain specific baseline, and show that post-dependent combination of the priors achieves the best performance. When suitable training data does not exist, our methods still outperform a domain unspecific baseline. We conclude with an examination of the language distribution of a million tweets, along with temporal analysis, the usage of twitter features across languages, and a correlation study between classifications made and geo-location and language metadata fields.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2013. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2013
    holdings:
      @attributes:
        islocal: N