TweetLID: a benchmark for tweet language identification.

Language identification, as the task of determining the language a given text is written in, has progressed substantially in recent decades. However, three main issues remain still unresolved: (1) distinction of similar languages, (2) detection of multilingualism in a single document, and (3) identi...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 50; no. 4; pp. 729 - 767
Autores principales: Zubiaga, Arkaitz, Vicente, Iñaki, Gamallo, Pablo, Pichel, José, Alegria, Iñaki, Aranberri, Nora, Ezeiza, Aitzol, Fresno, Víctor
Formato: Artículo
Publicado: Springer Nature Dec2016
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=119384314&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 119384314
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Dec2016
      vid: 50
      iid: 4
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        119384314
        10.1007/s10579-015-9317-4
      ppf: 729
      ppct: 38
      formats:
        fmt:
          @attributes:
            type: P
            size: 677KB
      tig:
        atl: TweetLID: a benchmark for tweet language identification.
      aug:
        au:
          Zubiaga, Arkaitz
          Vicente, Iñaki
          Gamallo, Pablo
          Pichel, José
          Alegria, Iñaki
          Aranberri, Nora
          Ezeiza, Aitzol
          Fresno, Víctor
        affil:
          University of Warwick , Coventry UK
          Elhuyar , Usurbil Spain
          USC , Santiago de Compostela Spain
          imaxin|software , Santiago de Compostela Spain
          University of the Basque Country , Donostia-San Sebastián Spain
          UNED , Madrid Spain
      su:
        Language identification (Computational linguistics)
        Natural language processing
        Multilingualism
        Language & languages
        Microblogs
      sug:
        subj:
          Language identification (Computational linguistics)
          Natural language processing
          Multilingualism
          Language & languages
          Microblogs
      keyword:
        Language identification
        Short texts
        Similar languages
        Tweets
      ab: Language identification, as the task of determining the language a given text is written in, has progressed substantially in recent decades. However, three main issues remain still unresolved: (1) distinction of similar languages, (2) detection of multilingualism in a single document, and (3) identifying the language of short texts. In this paper, we describe our work on the development of a benchmark to encourage further research in these three directions, set forth an evaluation framework suitable for the task, and make a dataset of annotated tweets publicly available for research purposes. We also describe the shared task we organized to validate and assess the evaluation framework and dataset with systems submitted by seven different participants, and analyze the performance of these systems. The evaluation of the results submitted by the participants of the shared task helped us shed some light on the shortcomings of state-of-the-art language identification systems, and gives insight into the extent to which the brevity, multilingualism, and language similarity found in texts exacerbate the performance of language identifiers. Our dataset with nearly 35,000 tweets and the evaluation framework provide researchers and practitioners with suitable resources to further study the aforementioned issues on language identification within a common setting that enables to compare results with one another.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2016. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2016
    holdings:
      @attributes:
        islocal: N