TweetLID: a benchmark for tweet language identification.
Language identification, as the task of determining the language a given text is written in, has progressed substantially in recent decades. However, three main issues remain still unresolved: (1) distinction of similar languages, (2) detection of multilingualism in a single document, and (3) identi...
| Publicado en: | Language Resources & Evaluation Vol. 50; no. 4; pp. 729 - 767 |
|---|---|
| Autores principales: | , , , , , , , |
| Formato: | Artículo |
| Publicado: |
Springer Nature
Dec2016
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=119384314&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 119384314 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Dec2016 vid: 50 iid: 4 pid: 237 pub: Springer Nature artinfo: ui: 119384314 10.1007/s10579-015-9317-4 ppf: 729 ppct: 38 formats: fmt: @attributes: type: P size: 677KB tig: atl: TweetLID: a benchmark for tweet language identification. aug: au: Zubiaga, Arkaitz Vicente, Iñaki Gamallo, Pablo Pichel, José Alegria, Iñaki Aranberri, Nora Ezeiza, Aitzol Fresno, Víctor affil: University of Warwick , Coventry UK Elhuyar , Usurbil Spain USC , Santiago de Compostela Spain imaxin|software , Santiago de Compostela Spain University of the Basque Country , Donostia-San Sebastián Spain UNED , Madrid Spain su: Language identification (Computational linguistics) Natural language processing Multilingualism Language & languages Microblogs sug: subj: Language identification (Computational linguistics) Natural language processing Multilingualism Language & languages Microblogs keyword: Language identification Short texts Similar languages Tweets ab: Language identification, as the task of determining the language a given text is written in, has progressed substantially in recent decades. However, three main issues remain still unresolved: (1) distinction of similar languages, (2) detection of multilingualism in a single document, and (3) identifying the language of short texts. In this paper, we describe our work on the development of a benchmark to encourage further research in these three directions, set forth an evaluation framework suitable for the task, and make a dataset of annotated tweets publicly available for research purposes. We also describe the shared task we organized to validate and assess the evaluation framework and dataset with systems submitted by seven different participants, and analyze the performance of these systems. The evaluation of the results submitted by the participants of the shared task helped us shed some light on the shortcomings of state-of-the-art language identification systems, and gives insight into the extent to which the brevity, multilingualism, and language similarity found in texts exacerbate the performance of language identifiers. Our dataset with nearly 35,000 tweets and the evaluation framework provide researchers and practitioners with suitable resources to further study the aforementioned issues on language identification within a common setting that enables to compare results with one another. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2016. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2016 holdings: @attributes: islocal: N |
|---|