PinLID: a dataset for Pinglish language identiftcation based on code-mixing sentence on unstructured resources.
Language identification is a major task in natural language processing. It serves as an initial and effective stage in critical tasks such as information extraction, sentiment analysis, and question answering. Most research on language identification has focused on monolingual contexts, performing p...
| Publicado en: | Language Resources & Evaluation Vol. 59; no. 3; pp. 3215 - 3242 |
|---|---|
| Autores principales: | , , |
| Formato: | Conference Paper/Materials |
| Publicado: |
Springer Nature
Sep2025
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=186909046&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 186909046 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Sep2025 vid: 59 iid: 3 pid: 237 pub: Springer Nature artinfo: ui: 186909046 10.1007/s10579-024-09783-3 ppf: 3215 ppct: 27 formats: fmt: – @attributes: type: T – @attributes: type: P size: 1.8MB tig: atl: PinLID: a dataset for Pinglish language identiftcation based on code-mixing sentence on unstructured resources. aug: au: Ghafouri, Arash Naderi, Hasan Firouzmandi, Mahdi affil: https://ror.org/01jw2p796 Department of Computer Engineering, Iran University of Science and Technology, Tehran, Iran su: Language identification (Computational linguistics) Natural language processing Persian language Acquisition of data Code switching (Linguistics) Machine learning Social media sug: subj: Language identification (Computational linguistics) Natural language processing Persian language Acquisition of data Code switching (Linguistics) Machine learning Social media keyword: Code-mixed language Communication and Culture Linguistics Information and Computing Sciences Artificial Intelligence and Image Processing Psychology and Cognitive Sciences Cognitive Sciences Language Language identification Persian-English text ab: Language identification is a major task in natural language processing. It serves as an initial and effective stage in critical tasks such as information extraction, sentiment analysis, and question answering. Most research on language identification has focused on monolingual contexts, performing poorly with texts containing code-mixing. Identifying the language in social media texts, such as those on Twitter, poses challenges due to high levels of code-mixing. Consequently, creating an accurate language identification tool for code-mixed texts is essential for intelligent systems that rely on natural language processing, such as advanced search engines and question-answering systems. Recently, significant research has been conducted in non-Persian languages in this field. However, no substantial efforts have been made to recognize languages in code-mixed Persian texts. In this paper, we introduce a dataset called PinLID, collected from tweets with Persian-English code-mixing, labeled at both the sentence and token levels using a supervised learning approach to language identification. We evaluated the dataset using various machine learning classification algorithms, including the classical SVM method, the multilingual BERT language model, XLM-RoBERTa, ParsBERT, AriaBERT, and PersianLLaMA: Persian Large Language Model. The testing yielded results as high as 99.59% F1 score at both the sentence and token levels in the test data. pubtype: Academic Journal doctype: Conference Paper/Materials src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2025. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2025 holdings: @attributes: islocal: N |
|---|