Develop corpora and methods for cross-lingual text reuse detection for English Urdu language pair at lexical, syntactical, and phrasal levels.

In recent years, Cross-Lingual Text Reuse Detection (CLTRD) has attracted the attention of the research community because large digital repositories and efficient Machine Translation systems are readily and freely available, which makes it easier to reuse text across the languages and very difficult...

Full description

Bibliographic Details
Published in:Language Resources & Evaluation Vol. 56; no. 4; pp. 1103 - 1131
Main Authors: Muneer, Iqra, Nawab, Rao Muhammad Adeel
Format: Article
Published: Springer Nature Dec2022
Subjects:
Online Access:View this record in EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=159972001&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 159972001
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Dec2022
      vid: 56
      iid: 4
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        159972001
        10.1007/s10579-022-09613-4
      ppf: 1103
      ppct: 28
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 885KB
      tig:
        atl: Develop corpora and methods for cross-lingual text reuse detection for English Urdu language pair at lexical, syntactical, and phrasal levels.
      aug:
        au:
          Muneer, Iqra
          Nawab, Rao Muhammad Adeel
        affil:
          COMSATS University Islamabad, Lahore Campus, Lahore, Pakistan
          University of Engineering & Technology Lahore, Narowal Campus, Pakistan
      su:
        Urdu language
        Corpora
        English language
        Machine translating
        Scientific community
      sug:
        subj:
          Urdu language
          Corpora
          English language
          Machine translating
          Scientific community
      keyword:
        Cross-lingual semantic tagger
        Cross-lingual sentence transformer
        Cross-lingual text reuse
        Cross-lingual word embedding
        English-Urdu language pair
        Lexical
        Phrasal
        Syntactical
      ab: In recent years, Cross-Lingual Text Reuse Detection (CLTRD) has attracted the attention of the research community because large digital repositories and efficient Machine Translation systems are readily and freely available, which makes it easier to reuse text across the languages and very difficult to detect it. In the previous studies, the problem of CLTRD for the English-Urdu language pair has been explored at the sentence/passage and document level, and benchmark corpora and methods have been developed. However, there is a lack of benchmark corpora and methods for the CLTRD for the English-Urdu language pair at the lexical, syntactical, and phrasal levels. To fulfill this research gap, this study presents three large benchmark corpora for detecting the Cross-Lingual Text Reuse (CLTR) at three levels of rewrite (Wholly Derived (WD), Partially Derived (PD), and Non Derived (ND)). The CLEU-Lex, CLEU-Syn and CLEU-Phr corpora contain 66,485 (WD = 22,236, PD = 20,315 and ND = 23,934), 60,267 (WD = 20,007, PD = 16,979 and ND = 23,281) and 60,106 (WD = 23,862, PD = 15,878 and ND = 20,366) CLTR pairs respectively. As a secondary major contribution, we have applied the Cross-Lingual Word Embedding (CLWE), Cross-Lingual Semantic Tagger (CLST), and Cross-Lingual Sentence Transformer (CLSTR) based methods on our three proposed corpora for the CLTRD. Our extensive experimentation showed that for the binary classification task, the best results on the CLEU-Lex corpus were obtained using the cross-lingual sentence transformer ( F 1 = 0.80). For the CLEU-Syn and CLEU-Phr corpora, the best results were obtained using the cross-lingual sentence transformer and a combination of the CLWE, CLST and CLSTR methods ( F 1 = 0.92 on CLEU-Syn and F 1 = 0.94 on CLEU-Phr). For the ternary classification task, the best results on the CLEU-Lex corpus were obtained using the cross-lingual sentence transformer method ( F 1 = 0.69). For the CLEU-Syn corpus, the best results were obtained using a combination of the CLWE, CLST, and CLSTR methods ( F 1 = 0.82). For the CLEU-Phr corpus the best results were obtained using cross-lingual sentence transformer and combination of CLWE, CLST, and CLSTR methods ( F 1 = 0.78). To foster and promote research in Urdu (a low-resourced language) all the three proposed corpora are free and publicly available for research purposes.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2022. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2022
    holdings:
      @attributes:
        islocal: N