ViMs: a high-quality Vietnamese dataset for abstractive multi-document summarization.

Automatic text summarization is important in this era due to the exponential growth of documents available on the Internet. In the Vietnamese language, VietnameseMDS is the only publicly available dataset for this task. Although the dataset has 199 clusters, there are only three documents in each cl...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 54; no. 4; pp. 893 - 921
Autores principales: Tran, Nhi-Thao, Nghiem, Minh-Quoc, Nguyen, Nhung T. H., Nguyen, Ngan Luu-Thuy, Van Chi, Nam, Dinh, Dien
Formato: Artículo
Publicado: Springer Nature Dec2020
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=146751852&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 146751852
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Dec2020
      vid: 54
      iid: 4
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        146751852
        10.1007/s10579-020-09495-4
      ppf: 893
      ppct: 28
      formats:
        fmt:
          @attributes:
            type: P
            size: 1.9MB
      tig:
        atl: ViMs: a high-quality Vietnamese dataset for abstractive multi-document summarization.
      aug:
        au:
          Tran, Nhi-Thao
          Nghiem, Minh-Quoc
          Nguyen, Nhung T. H.
          Nguyen, Ngan Luu-Thuy
          Van Chi, Nam
          Dinh, Dien
        affil:
          Faculty of Information Technology, HCMC University of Science, Ho Chi Minh City, Vietnam
          Faculty of Computer Science, HCMC University of Information Technology, Ho Chi Minh City, Vietnam
      su:
        Document clustering
        Exponential functions
        Vietnamese people
        Employee reviews
        Scientific community
      sug:
        subj:
          Document clustering
          Exponential functions
          Vietnamese people
          Employee reviews
          Scientific community
      keyword:
        Abstractive summarization
        Automatic summarization
        Multi-document summarization
        Vietnamese dataset
      ab: Automatic text summarization is important in this era due to the exponential growth of documents available on the Internet. In the Vietnamese language, VietnameseMDS is the only publicly available dataset for this task. Although the dataset has 199 clusters, there are only three documents in each cluster, which is small compared to typical datasets in English. This motivates us to construct ViMs—a big and high-quality Vietnamese dataset for abstractive multi-document summarization. To that end, we recruited 29 annotators and enhanced MDSWriter—an open-source annotation tool, to support the annotators in creating gold standard summaries. As a result, ViMs has 600 summaries corresponding to 300 clusters of 1,945 documents. We have verified the reliability of our dataset by using a variety of metrics including conventional Cohen's κ , relaxed Cohen's κ —a new metric that we propose to make it more suitable for abstractive summarization, and ROUGE scores. A relaxed κ score of 0.55 indicate that ViMs could attain moderate agreement between annotators. Meanwhile, ROUGE scores are 0.729 of ROUGE-1, 0.507 of ROUGE-2 and 0.524 of ROUGE-SU4. We have further evaluated ViMs by using three different summarization systems: TextRank, CFVi and MUSEEC. Their performances are 0.628, 0.711 and 0.732 of ROUGE-1, respectively. These results show that the ViMs dataset is suitable for both training and evaluating multi-document summarization systems. We have made the dataset and evaluation results of this work publicly available for research community. It is noted that unlike previous work that only published the final summarization dataset, we also publish intermediate annotation results, which can be used in other NLP problems such as sentence classification.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2020. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2020
    holdings:
      @attributes:
        islocal: N