Simple-random-sampling-based multiclass text classification algorithm.

Multiclass text classification (MTC) is a challenging issue and the corresponding MTC algorithms can be used in many applications. The space-time overhead of the algorithms must be concerned about the era of big data. Through the investigation of the token frequency distribution in a Chinese web doc...

Descripción completa

Detalles Bibliográficos
Publicado en:Scientific World Journal pp. 517498 - 517499
Autores principales: Liu, Wuying, Wang, Lin, Yi, Mianzhu
Formato: Journal Article
Publicado: Wiley-Blackwell 2014
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=103820617&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 103820617
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        1537744X
        1BX5
      jtl: Scientific World Journal
      issn: 1537744X
      maglogo: N
    pubinfo:
      dt: 2014
      pid: 480
      pub: Wiley-Blackwell
      place: Malden, Massachusetts
    artinfo:
      ui:
        103820617
        NLM24778587
        2012565564
        10.1155/2014/517498
        NLM24778587
        PMC3977423
        103820617
      ppf: 517498
      ppct: 1
      formats:
      tig:
        atl: Simple-random-sampling-based multiclass text classification algorithm.
      aug:
        au:
          Liu, Wuying
          Wang, Lin
          Yi, Mianzhu
        affil: Department of Language Engineering, PLA University of Foreign Languages, Luoyang, Henan 471003, China ; College of Computer, National University of Defense Technology, Changsha, Hunan 410073, China.
      sug:
        subj:
          Algorithms
          Models, Theoretical
      ab: Multiclass text classification (MTC) is a challenging issue and the corresponding MTC algorithms can be used in many applications. The space-time overhead of the algorithms must be concerned about the era of big data. Through the investigation of the token frequency distribution in a Chinese web document collection, this paper reexamines the power law and proposes a simple-random-sampling-based MTC (SRSMTC) algorithm. Supported by a token level memory to store labeled documents, the SRSMTC algorithm uses a text retrieval approach to solve text classification problems. The experimental results on the TanCorp data set show that SRSMTC algorithm can achieve the state-of-the-art performance at greatly reduced space-time requirements.
      pubtype: Academic Journal
      doctype: Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N