Sentence boundary detection of various forms of Tunisian Arabic.

Sentence boundary detection (SBD) is an essential step for a very large number of natural language processing applications such as parsing, information retrieval, automatic summarization, machine translation, etc. In this paper, we tackle the problem of SBD of dialectal Arabic, especially for the Tu...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 56; no. 1; pp. 357 - 386
Autores principales: Mekki, Asma, Zribi, Inès, Ellouze, Mariem, Belguith, Lamia Hadrich
Formato: Artículo
Publicado: Springer Nature Mar2022
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=155080011&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 155080011
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Mar2022
      vid: 56
      iid: 1
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        155080011
        10.1007/s10579-021-09538-4
      ppf: 357
      ppct: 29
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 906KB
      tig:
        atl: Sentence boundary detection of various forms of Tunisian Arabic.
      aug:
        au:
          Mekki, Asma
          Zribi, Inès
          Ellouze, Mariem
          Belguith, Lamia Hadrich
        affil: ANLP Research Group, MIRACL, University of Sfax, Sfax, Tunisia
      su:
        Natural language processing
        Machine translating
        Support vector machines
        Machine learning
        Random fields
      sug:
        subj:
          Natural language processing
          Machine translating
          Support vector machines
          Machine learning
          Random fields
      keyword:
        CRF
        Deep learning
        Dialectal Arabic
        Sentence boundary detection
        SVM
        Tunisian Arabic
      ab: Sentence boundary detection (SBD) is an essential step for a very large number of natural language processing applications such as parsing, information retrieval, automatic summarization, machine translation, etc. In this paper, we tackle the problem of SBD of dialectal Arabic, especially for the Tunisian dialect. We compare the efficiency of three learning algorithms: Deep Neuronal Networks (DNN), Support Vector Machines (SVM) and Conditional Random Fields (CRF) to detect the boundaries of sentences written in different types of dialect. The best model achieved an F-measure of 84.37% using CRF which is a popular formalism for structured prediction in NLP and it has been widely applied in text segmentation.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2022. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2022
    holdings:
      @attributes:
        islocal: N