COUNTER: corpus of Urdu news text reuse.
Text reuse is the act of borrowing text from existing documents to create new texts. Freely available and easily accessible large online repositories are not only making reuse of text more common in society but also harder to detect. A major hindrance in the development and evaluation of existing/ne...
| Publicado en: | Language Resources & Evaluation Vol. 51; no. 3; pp. 777 - 804 |
|---|---|
| Autores principales: | , , |
| Formato: | Artículo |
| Publicado: |
Springer Nature
Sep2017
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=124484776&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 124484776 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Sep2017 vid: 51 iid: 3 pid: 237 pub: Springer Nature artinfo: ui: 124484776 10.1007/s10579-016-9367-2 ppf: 777 ppct: 27 formats: fmt: @attributes: type: P size: 1.5MB tig: atl: COUNTER: corpus of Urdu news text reuse. aug: au: Sharjeel, Muhammad Nawab, Rao Rayson, Paul affil: Department of Computer Science , COMSATS Institute of Information Technology , Lahore Pakistan School of Computing and Communications , Lancaster University , Bailrigg UK su: Corpora Institutional repositories Asian languages Urdu language Digital libraries sug: subj: Corpora Institutional repositories Asian languages Urdu language Digital libraries keyword: Corpus generation Mono-lingual text reuse Urdu news corpus Urdu text reuse detection ab: Text reuse is the act of borrowing text from existing documents to create new texts. Freely available and easily accessible large online repositories are not only making reuse of text more common in society but also harder to detect. A major hindrance in the development and evaluation of existing/new mono-lingual text reuse detection methods, especially for South Asian languages, is the unavailability of standardized benchmark corpora. Amongst other things, a gold standard corpus enables researchers to directly compare existing state-of-the-art methods. In our study, we address this gap by developing a benchmark corpus for one of the widely spoken but under resourced languages i.e. Urdu. The COrpus of Urdu News TExt Reuse (COUNTER) corpus contains 1200 documents with real examples of text reuse from the field of journalism. It has been manually annotated at document level with three levels of reuse: wholly derived, partially derived and non derived. We also apply a number of similarity estimation methods on our corpus to show how it can be used for the development, evaluation and comparison of text reuse detection systems for the Urdu language. The corpus is a vital resource for the development and evaluation of text reuse detection systems in general and specifically for Urdu language. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2017. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2017 holdings: @attributes: islocal: N |
|---|