Mining an English-Chinese parallel Dataset of Financial News.
Parallel text datasets are a valuable for educational purposes, machine translation, and cross-language information retrieval, but few are domain-oriented. We have created a Chinese–English parallel dataset in the domain of finance technology, using the Financial Times website, from which we grabbed...
| Publicado en: | Journal of Open Humanities Data Vol. 8; no. 1; pp. 1 - 13 |
|---|---|
| Autores principales: | , , , , , , |
| Formato: | Artículo |
| Publicado: |
Ubiquity Press
2022
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=162571059&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 162571059 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 2059481X MPQW jtl: Journal of Open Humanities Data issn: 2059481X maglogo: N pubinfo: dt: 2022 vid: 8 iid: 1 pid: 83901 pub: Ubiquity Press artinfo: ui: 162571059 10.5334/johd.62 ppf: 1 ppct: 12 formats: tig: atl: Mining an English-Chinese parallel Dataset of Financial News. aug: au: TURENNE, NICOLAS ZIWEI CHEN GUITAO FAN JIANLONG LI YIWEN LI SIYUAN WANG JIAQI ZHOU affil: BNU-HKBU United International College, UIC, Division of Science and Technology, Zhuhai Guangdong, su: Parallel algorithms Data analysis Text mining Machine translating Bilingualism sug: subj: Parallel algorithms Data analysis Text mining Machine translating Bilingualism keyword: classification clustering English-Chinese patterns text mining ab: Parallel text datasets are a valuable for educational purposes, machine translation, and cross-language information retrieval, but few are domain-oriented. We have created a Chinese–English parallel dataset in the domain of finance technology, using the Financial Times website, from which we grabbed 60,473 news items from between 2007 and 2021. This dataset is a bilingual Chinese–English parallel dataset of news in the domain of finance. It is open access in its original state without transformation, and has been made not for machine translation as has been used, but for intelligent mining, in which we conducted many experiments using up-to-date text mining techniques: clustering (topic modeling, community detection, k-means), topic prediction (naive Bayes, SVM, LSTM, Bert), and pattern discovery (dictionary based, time series). We present the usage of these techniques as a framework for other studies, not only as an application but with an interpretation. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y dt: @attributes: year: 2022 holdings: @attributes: islocal: N |
|---|