COUNTER: corpus of Urdu news text reuse.

Text reuse is the act of borrowing text from existing documents to create new texts. Freely available and easily accessible large online repositories are not only making reuse of text more common in society but also harder to detect. A major hindrance in the development and evaluation of existing/ne...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 51; no. 3; pp. 777 - 804
Autores principales: Sharjeel, Muhammad, Nawab, Rao, Rayson, Paul
Formato: Artículo
Publicado: Springer Nature Sep2017
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=124484776&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 124484776
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Sep2017
      vid: 51
      iid: 3
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        124484776
        10.1007/s10579-016-9367-2
      ppf: 777
      ppct: 27
      formats:
        fmt:
          @attributes:
            type: P
            size: 1.5MB
      tig:
        atl: COUNTER: corpus of Urdu news text reuse.
      aug:
        au:
          Sharjeel, Muhammad
          Nawab, Rao
          Rayson, Paul
        affil:
          Department of Computer Science , COMSATS Institute of Information Technology , Lahore Pakistan
          School of Computing and Communications , Lancaster University , Bailrigg UK
      su:
        Corpora
        Institutional repositories
        Asian languages
        Urdu language
        Digital libraries
      sug:
        subj:
          Corpora
          Institutional repositories
          Asian languages
          Urdu language
          Digital libraries
      keyword:
        Corpus generation
        Mono-lingual text reuse
        Urdu news corpus
        Urdu text reuse detection
      ab: Text reuse is the act of borrowing text from existing documents to create new texts. Freely available and easily accessible large online repositories are not only making reuse of text more common in society but also harder to detect. A major hindrance in the development and evaluation of existing/new mono-lingual text reuse detection methods, especially for South Asian languages, is the unavailability of standardized benchmark corpora. Amongst other things, a gold standard corpus enables researchers to directly compare existing state-of-the-art methods. In our study, we address this gap by developing a benchmark corpus for one of the widely spoken but under resourced languages i.e. Urdu. The COrpus of Urdu News TExt Reuse (COUNTER) corpus contains 1200 documents with real examples of text reuse from the field of journalism. It has been manually annotated at document level with three levels of reuse: wholly derived, partially derived and non derived. We also apply a number of similarity estimation methods on our corpus to show how it can be used for the development, evaluation and comparison of text reuse detection systems for the Urdu language. The corpus is a vital resource for the development and evaluation of text reuse detection systems in general and specifically for Urdu language.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2017. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2017
    holdings:
      @attributes:
        islocal: N