The impact of corpus domain on word representation: a study on Persian word embeddings.

Word embedding, has been a great success story for natural language processing in recent years. The main purpose of this approach is providing a vector representation of words based on neural network language modeling. Using a large training corpus, the model most learns from co-occurrences of words...

Full description

Bibliographic Details
Published in:Language Resources & Evaluation Vol. 52; no. 4; pp. 997 - 1020
Main Authors: Hadifar, Amir, Momtazi, Saeedeh
Format: Article
Published: Springer Nature Dec2018
Subjects:
Online Access:View this record in EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=132695160&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 132695160
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Dec2018
      vid: 52
      iid: 4
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        132695160
        10.1007/s10579-018-9419-x
      ppf: 997
      ppct: 23
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 1.4MB
      tig:
        atl: The impact of corpus domain on word representation: a study on Persian word embeddings.
      aug:
        au:
          Hadifar, Amir
          Momtazi, Saeedeh
        affil: Computer Engineering and Information Technology Department, Amirkabir University of Technology, Hafez Avenue, Tehran, Iran
      su:
        Natural language processing
        Artificial neural networks
        Persian language
        Semantic computing
        Computational mathematics
      sug:
        subj:
          Natural language processing
          Artificial neural networks
          Persian language
          Semantic computing
          Computational mathematics
      keyword:
        Distributional semantic models
        Persian
        Word embedding
        Word2vec
      ab: Word embedding, has been a great success story for natural language processing in recent years. The main purpose of this approach is providing a vector representation of words based on neural network language modeling. Using a large training corpus, the model most learns from co-occurrences of words, namely Skip-gram model, and capture semantic features of words. Moreover, adding the recently introduced character embedding model to the objective function, the model can also focus on morphological features of words. In this paper, we study the impact of training corpus on the results of word embedding and show how the genre of training data affects the type of information captured by word embedding models. We perform our experiments on the Persian language. In line of our experiments, providing two well-known evaluation datasets for Persian, namely Google semantic/syntactic analogy and Wordsim353, is also part of the contribution of this paper. The experiments include computation of word embedding from various public Persian corpora with different genres and sizes while considering comprehensive lexical and semantic comparison between them. We identify words whose usages differ between these datasets resulted totally different vector representation which ends to significant impact on different domains in which the results vary up to 9% on Google analogy and up to 6% on Wordsim353. The resulted word embedding for each of the individual corpora as well as their combinations will be publicly available for any further research based on word embedding for Persian.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2018. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2018
    holdings:
      @attributes:
        islocal: N