Microblog language identification: overcoming the limitations of short, unedited and idiomatic text.
Multilingual posts can potentially affect the outcomes of content analysis on microblog platforms. To this end, language identification can provide a monolingual set of content for analysis. We find the unedited and idiomatic language of microblogs to be challenging for state-of-the-art language ide...
| Publicado en: | Language Resources & Evaluation Vol. 47; no. 1; pp. 195 - 216 |
|---|---|
| Autores principales: | , , |
| Formato: | Artículo |
| Publicado: |
Springer Nature
Mar2013
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=85873232&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 85873232 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Mar2013 vid: 47 iid: 1 pid: 237 pub: Springer Nature artinfo: ui: 85873232 10.1007/s10579-012-9195-y ppf: 195 ppct: 21 formats: fmt: @attributes: type: P size: 428KB tig: atl: Microblog language identification: overcoming the limitations of short, unedited and idiomatic text. aug: au: Carter, Simon Weerkamp, Wouter Tsagkias, Manos affil: ISLA, University of Amsterdam, Science Park 904 1098 XH Amsterdam The Netherlands su: Linguistic identity Microblogs Instant messaging Online social networks Content analysis sug: subj: Linguistic identity Microblogs Instant messaging Online social networks Content analysis keyword: Language identification Text classification ab: Multilingual posts can potentially affect the outcomes of content analysis on microblog platforms. To this end, language identification can provide a monolingual set of content for analysis. We find the unedited and idiomatic language of microblogs to be challenging for state-of-the-art language identification methods. To account for this, we identify five microblog characteristics that can help in language identification: the language profile of the blogger (blogger), the content of an attached hyperlink (link), the language profile of other users mentioned (mention) in the post, the language profile of a tag (tag), and the language of the original post (conversation), if the post we examine is a reply. Further, we present methods that combine these priors in a post-dependent and post-independent way. We present test results on 1,000 posts from five languages (Dutch, English, French, German, and Spanish), which show that our priors improve accuracy by 5 % over a domain specific baseline, and show that post-dependent combination of the priors achieves the best performance. When suitable training data does not exist, our methods still outperform a domain unspecific baseline. We conclude with an examination of the language distribution of a million tweets, along with temporal analysis, the usage of twitter features across languages, and a correlation study between classifications made and geo-location and language metadata fields. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2013. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2013 holdings: @attributes: islocal: N |
|---|