Mining an English-Chinese parallel Dataset of Financial News.

Parallel text datasets are a valuable for educational purposes, machine translation, and cross-language information retrieval, but few are domain-oriented. We have created a Chinese–English parallel dataset in the domain of finance technology, using the Financial Times website, from which we grabbed...

Descripción completa

Detalles Bibliográficos
Publicado en:Journal of Open Humanities Data Vol. 8; no. 1; pp. 1 - 13
Autores principales: TURENNE, NICOLAS, ZIWEI CHEN, GUITAO FAN, JIANLONG LI, YIWEN LI, SIYUAN WANG, JIAQI ZHOU
Formato: Artículo
Publicado: Ubiquity Press 2022
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=162571059&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 162571059
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        2059481X
        MPQW
      jtl: Journal of Open Humanities Data
      issn: 2059481X
      maglogo: N
    pubinfo:
      dt: 2022
      vid: 8
      iid: 1
      pid: 83901
      pub: Ubiquity Press
    artinfo:
      ui:
        162571059
        10.5334/johd.62
      ppf: 1
      ppct: 12
      formats:
      tig:
        atl: Mining an English-Chinese parallel Dataset of Financial News.
      aug:
        au:
          TURENNE, NICOLAS
          ZIWEI CHEN
          GUITAO FAN
          JIANLONG LI
          YIWEN LI
          SIYUAN WANG
          JIAQI ZHOU
        affil: BNU-HKBU United International College, UIC, Division of Science and Technology, Zhuhai Guangdong,
      su:
        Parallel algorithms
        Data analysis
        Text mining
        Machine translating
        Bilingualism
      sug:
        subj:
          Parallel algorithms
          Data analysis
          Text mining
          Machine translating
          Bilingualism
      keyword:
        classification
        clustering
        English-Chinese
        patterns
        text mining
      ab: Parallel text datasets are a valuable for educational purposes, machine translation, and cross-language information retrieval, but few are domain-oriented. We have created a Chinese–English parallel dataset in the domain of finance technology, using the Financial Times website, from which we grabbed 60,473 news items from between 2007 and 2021. This dataset is a bilingual Chinese–English parallel dataset of news in the domain of finance. It is open access in its original state without transformation, and has been made not for machine translation as has been used, but for intelligent mining, in which we conducted many experiments using up-to-date text mining techniques: clustering (topic modeling, community detection, k-means), topic prediction (naive Bayes, SVM, LSTM, Bert), and pattern discovery (dictionary based, time series). We present the usage of these techniques as a framework for other studies, not only as an application but with an interpretation.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      dt:
        @attributes:
          year: 2022
    holdings:
      @attributes:
        islocal: N