A document is known by the company it keeps: neighborhood consensus for short text categorization.

During the last decades the Web has become the greatest repository of digital information. In order to organize all this information, several text categorization methods have been developed, achieving accurate results in most cases and in very different domains. Due to the recent usage of Internet a...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 47; no. 1; pp. 127 - 150
Autores principales: Ramírez-de-la-Rosa, Gabriela, Montes-y-Gómez, Manuel, Solorio, Thamar, Villaseñor-Pineda, Luis
Formato: Artículo
Publicado: Springer Nature Mar2013
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=85873225&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 85873225
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Mar2013
      vid: 47
      iid: 1
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        85873225
        10.1007/s10579-012-9192-1
      ppf: 127
      ppct: 23
      formats:
        fmt:
          @attributes:
            type: P
            size: 818KB
      tig:
        atl: A document is known by the company it keeps: neighborhood consensus for short text categorization.
      aug:
        au:
          Ramírez-de-la-Rosa, Gabriela
          Montes-y-Gómez, Manuel
          Solorio, Thamar
          Villaseñor-Pineda, Luis
        affil:
          Department of Computer and Information Sciences, University of Alabama at Birmingham, Birmingham USA
          Department of Computational Sciences, National Institute for Astrophysics, Optics and Electronics, Puebla Mexico
      su:
        Prototype (Linguistics)
        World Wide Web
        Electronic information resources
        Categorization (Linguistics)
        Internet users
      sug:
        subj:
          Prototype (Linguistics)
          World Wide Web
          Electronic information resources
          Categorization (Linguistics)
          Internet users
      keyword:
        News titles
        Prototype-based classification
        Short text categorization
        Unlabeled information
      ab: During the last decades the Web has become the greatest repository of digital information. In order to organize all this information, several text categorization methods have been developed, achieving accurate results in most cases and in very different domains. Due to the recent usage of Internet as communication media, short texts such as news, tweets, blogs, and product reviews are more common every day. In this context, there are two main challenges; on the one hand, the length of these documents is short, and therefore, the word frequencies are not informative enough, making text categorization even more difficult than usual. On the other hand, topics are changing constantly at a fast rate, causing the lack of adequate amounts of training data. In order to deal with these two problems we consider a text classification method that is supported on the idea that similar documents may belong to the same category. Mainly, we propose a neighborhood consensus classification method that classifies documents by considering their own information as well as information about the category assigned to other similar documents from the same target collection. In particular, the short texts we used in our evaluation are news titles with an average of 8 words. Experimental results are encouraging; they indicate that leveraging information from similar documents helped to improve classification accuracy and that the proposed method is especially useful when labeled training resources are limited.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2013. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2013
    holdings:
      @attributes:
        islocal: N