ViMs: a high-quality Vietnamese dataset for abstractive multi-document summarization.
Automatic text summarization is important in this era due to the exponential growth of documents available on the Internet. In the Vietnamese language, VietnameseMDS is the only publicly available dataset for this task. Although the dataset has 199 clusters, there are only three documents in each cl...
| Publicado en: | Language Resources & Evaluation Vol. 54; no. 4; pp. 893 - 921 |
|---|---|
| Autores principales: | , , , , , |
| Formato: | Artículo |
| Publicado: |
Springer Nature
Dec2020
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=146751852&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 146751852 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Dec2020 vid: 54 iid: 4 pid: 237 pub: Springer Nature artinfo: ui: 146751852 10.1007/s10579-020-09495-4 ppf: 893 ppct: 28 formats: fmt: @attributes: type: P size: 1.9MB tig: atl: ViMs: a high-quality Vietnamese dataset for abstractive multi-document summarization. aug: au: Tran, Nhi-Thao Nghiem, Minh-Quoc Nguyen, Nhung T. H. Nguyen, Ngan Luu-Thuy Van Chi, Nam Dinh, Dien affil: Faculty of Information Technology, HCMC University of Science, Ho Chi Minh City, Vietnam Faculty of Computer Science, HCMC University of Information Technology, Ho Chi Minh City, Vietnam su: Document clustering Exponential functions Vietnamese people Employee reviews Scientific community sug: subj: Document clustering Exponential functions Vietnamese people Employee reviews Scientific community keyword: Abstractive summarization Automatic summarization Multi-document summarization Vietnamese dataset ab: Automatic text summarization is important in this era due to the exponential growth of documents available on the Internet. In the Vietnamese language, VietnameseMDS is the only publicly available dataset for this task. Although the dataset has 199 clusters, there are only three documents in each cluster, which is small compared to typical datasets in English. This motivates us to construct ViMs—a big and high-quality Vietnamese dataset for abstractive multi-document summarization. To that end, we recruited 29 annotators and enhanced MDSWriter—an open-source annotation tool, to support the annotators in creating gold standard summaries. As a result, ViMs has 600 summaries corresponding to 300 clusters of 1,945 documents. We have verified the reliability of our dataset by using a variety of metrics including conventional Cohen's κ , relaxed Cohen's κ —a new metric that we propose to make it more suitable for abstractive summarization, and ROUGE scores. A relaxed κ score of 0.55 indicate that ViMs could attain moderate agreement between annotators. Meanwhile, ROUGE scores are 0.729 of ROUGE-1, 0.507 of ROUGE-2 and 0.524 of ROUGE-SU4. We have further evaluated ViMs by using three different summarization systems: TextRank, CFVi and MUSEEC. Their performances are 0.628, 0.711 and 0.732 of ROUGE-1, respectively. These results show that the ViMs dataset is suitable for both training and evaluating multi-document summarization systems. We have made the dataset and evaluation results of this work publicly available for research community. It is noted that unlike previous work that only published the final summarization dataset, we also publish intermediate annotation results, which can be used in other NLP problems such as sentence classification. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2020. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2020 holdings: @attributes: islocal: N |
|---|