The impact of corpus domain on word representation: a study on Persian word embeddings.
Word embedding, has been a great success story for natural language processing in recent years. The main purpose of this approach is providing a vector representation of words based on neural network language modeling. Using a large training corpus, the model most learns from co-occurrences of words...
| Published in: | Language Resources & Evaluation Vol. 52; no. 4; pp. 997 - 1020 |
|---|---|
| Main Authors: | , |
| Format: | Article |
| Published: |
Springer Nature
Dec2018
|
| Subjects: | |
| Online Access: | View this record in EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=132695160&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 132695160 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Dec2018 vid: 52 iid: 4 pid: 237 pub: Springer Nature artinfo: ui: 132695160 10.1007/s10579-018-9419-x ppf: 997 ppct: 23 formats: fmt: – @attributes: type: T – @attributes: type: P size: 1.4MB tig: atl: The impact of corpus domain on word representation: a study on Persian word embeddings. aug: au: Hadifar, Amir Momtazi, Saeedeh affil: Computer Engineering and Information Technology Department, Amirkabir University of Technology, Hafez Avenue, Tehran, Iran su: Natural language processing Artificial neural networks Persian language Semantic computing Computational mathematics sug: subj: Natural language processing Artificial neural networks Persian language Semantic computing Computational mathematics keyword: Distributional semantic models Persian Word embedding Word2vec ab: Word embedding, has been a great success story for natural language processing in recent years. The main purpose of this approach is providing a vector representation of words based on neural network language modeling. Using a large training corpus, the model most learns from co-occurrences of words, namely Skip-gram model, and capture semantic features of words. Moreover, adding the recently introduced character embedding model to the objective function, the model can also focus on morphological features of words. In this paper, we study the impact of training corpus on the results of word embedding and show how the genre of training data affects the type of information captured by word embedding models. We perform our experiments on the Persian language. In line of our experiments, providing two well-known evaluation datasets for Persian, namely Google semantic/syntactic analogy and Wordsim353, is also part of the contribution of this paper. The experiments include computation of word embedding from various public Persian corpora with different genres and sizes while considering comprehensive lexical and semantic comparison between them. We identify words whose usages differ between these datasets resulted totally different vector representation which ends to significant impact on different domains in which the results vary up to 9% on Google analogy and up to 6% on Wordsim353. The resulted word embedding for each of the individual corpora as well as their combinations will be publicly available for any further research based on word embedding for Persian. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2018. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2018 holdings: @attributes: islocal: N |
|---|