PinLID: a dataset for Pinglish language identiftcation based on code-mixing sentence on unstructured resources.

Language identification is a major task in natural language processing. It serves as an initial and effective stage in critical tasks such as information extraction, sentiment analysis, and question answering. Most research on language identification has focused on monolingual contexts, performing p...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 59; no. 3; pp. 3215 - 3242
Autores principales: Ghafouri, Arash, Naderi, Hasan, Firouzmandi, Mahdi
Formato: Conference Paper/Materials
Publicado: Springer Nature Sep2025
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=186909046&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 186909046
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Sep2025
      vid: 59
      iid: 3
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        186909046
        10.1007/s10579-024-09783-3
      ppf: 3215
      ppct: 27
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 1.8MB
      tig:
        atl: PinLID: a dataset for Pinglish language identiftcation based on code-mixing sentence on unstructured resources.
      aug:
        au:
          Ghafouri, Arash
          Naderi, Hasan
          Firouzmandi, Mahdi
        affil: https://ror.org/01jw2p796 Department of Computer Engineering, Iran University of Science and Technology, Tehran, Iran
      su:
        Language identification (Computational linguistics)
        Natural language processing
        Persian language
        Acquisition of data
        Code switching (Linguistics)
        Machine learning
        Social media
      sug:
        subj:
          Language identification (Computational linguistics)
          Natural language processing
          Persian language
          Acquisition of data
          Code switching (Linguistics)
          Machine learning
          Social media
      keyword:
        Code-mixed language
        Communication and Culture Linguistics Information and Computing Sciences Artificial Intelligence and Image Processing Psychology and Cognitive Sciences Cognitive Sciences
        Language
        Language identification
        Persian-English text
        Twitter
      ab: Language identification is a major task in natural language processing. It serves as an initial and effective stage in critical tasks such as information extraction, sentiment analysis, and question answering. Most research on language identification has focused on monolingual contexts, performing poorly with texts containing code-mixing. Identifying the language in social media texts, such as those on Twitter, poses challenges due to high levels of code-mixing. Consequently, creating an accurate language identification tool for code-mixed texts is essential for intelligent systems that rely on natural language processing, such as advanced search engines and question-answering systems. Recently, significant research has been conducted in non-Persian languages in this field. However, no substantial efforts have been made to recognize languages in code-mixed Persian texts. In this paper, we introduce a dataset called PinLID, collected from tweets with Persian-English code-mixing, labeled at both the sentence and token levels using a supervised learning approach to language identification. We evaluated the dataset using various machine learning classification algorithms, including the classical SVM method, the multilingual BERT language model, XLM-RoBERTa, ParsBERT, AriaBERT, and PersianLLaMA: Persian Large Language Model. The testing yielded results as high as 99.59% F1 score at both the sentence and token levels in the test data.
      pubtype: Academic Journal
      doctype: Conference Paper/Materials
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2025. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2025
    holdings:
      @attributes:
        islocal: N