A paper-text perspective.

Purpose In the era of Big Data, network digital resources are growing rapidly, especially the short-text resources, such as tweets, comments, messages and so on, are showing a vigorous vitality. This study aims to compare the categories discriminative capacity (CDC) of Chinese language fragments wit...

Descripción completa

Detalles Bibliográficos
Publicado en:Electronic Library Vol. 35; no. 4; pp. 689 - 709
Autores principales: Wang, Hao, Deng, Sanhong
Formato: equations & formulas research tables/charts Journal Article
Publicado: Emerald Publishing Limited 2017
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=125679398&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 125679398
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        02640473
        2OJ
      jtl: Electronic Library
      issn: 02640473
      maglogo: N
    pubinfo:
      dt: 2017
      vid: 35
      iid: 4
      pid: 465
      pub: Emerald Publishing Limited
    artinfo:
      ui:
        125679398
        125679398
        125679398
        10.1108/EL-09-2016-0192
        125679398
      ppf: 689
      ppct: 20
      formats:
      tig:
        atl: A paper-text perspective.
      aug:
        au:
          Wang, Hao
          Deng, Sanhong
        affil: School of Information Management, Nanjing University, Nanjing, China
      sug:
        subj:
          Data Analytics Methods
          Language China
          Classification
          Serial Publications China
          Human
          China
          Experimental Studies
          P-Value
          Bibliometrics
      ab: Purpose In the era of Big Data, network digital resources are growing rapidly, especially the short-text resources, such as tweets, comments, messages and so on, are showing a vigorous vitality. This study aims to compare the categories discriminative capacity (CDC) of Chinese language fragments with different granularities and to explore and verify feasibility, rationality and effectiveness of the low-granularity feature, such as Chinese characters in Chinese short-text classification (CSTC).Design/methodology/approach This study takes discipline classification of journal articles from CSSCI as a simulation environment. On the basis of sorting out the distribution rules of classification features with various granularities, including keywords, terms and characters, the classification effects accessed by the SVM algorithm are comprehensively compared and evaluated from three angles of using the same experiment samples, testing before and after feature optimization, and introducing external data.Findings The granularity of a classification feature has an important impact on CSTC. In general, the larger the granularity is, the better the classification result is, and vice versa. However, a low-granularity feature is also feasible, and its CDC could be improved by reasonable weight setting, even exceeding a high-granularity feature if synthetically considering classification precision, computational complexity and text coverage.Originality/value This is the first study to propose that Chinese characters are more suitable as descriptive features in CSTC than terms and keywords and to demonstrate that CDC of Chinese character features could be strengthened by mixing frequency and position as weight.
      pubtype: Academic Journal
      doctype:
        equations & formulas
        research
        tables/charts
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N