Sentence boundary detection of various forms of Tunisian Arabic.
Sentence boundary detection (SBD) is an essential step for a very large number of natural language processing applications such as parsing, information retrieval, automatic summarization, machine translation, etc. In this paper, we tackle the problem of SBD of dialectal Arabic, especially for the Tu...
| Publicado en: | Language Resources & Evaluation Vol. 56; no. 1; pp. 357 - 386 |
|---|---|
| Autores principales: | , , , |
| Formato: | Artículo |
| Publicado: |
Springer Nature
Mar2022
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=155080011&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 155080011 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Mar2022 vid: 56 iid: 1 pid: 237 pub: Springer Nature artinfo: ui: 155080011 10.1007/s10579-021-09538-4 ppf: 357 ppct: 29 formats: fmt: – @attributes: type: T – @attributes: type: P size: 906KB tig: atl: Sentence boundary detection of various forms of Tunisian Arabic. aug: au: Mekki, Asma Zribi, Inès Ellouze, Mariem Belguith, Lamia Hadrich affil: ANLP Research Group, MIRACL, University of Sfax, Sfax, Tunisia su: Natural language processing Machine translating Support vector machines Machine learning Random fields sug: subj: Natural language processing Machine translating Support vector machines Machine learning Random fields keyword: CRF Deep learning Dialectal Arabic Sentence boundary detection SVM Tunisian Arabic ab: Sentence boundary detection (SBD) is an essential step for a very large number of natural language processing applications such as parsing, information retrieval, automatic summarization, machine translation, etc. In this paper, we tackle the problem of SBD of dialectal Arabic, especially for the Tunisian dialect. We compare the efficiency of three learning algorithms: Deep Neuronal Networks (DNN), Support Vector Machines (SVM) and Conditional Random Fields (CRF) to detect the boundaries of sentences written in different types of dialect. The best model achieved an F-measure of 84.37% using CRF which is a popular formalism for structured prediction in NLP and it has been widely applied in text segmentation. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2022. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2022 holdings: @attributes: islocal: N |
|---|