Classifying unlabeled short texts using a fuzzy declarative approach.

Web 2.0 provides user-friendly tools that allow persons to create and publish content online. User generated content often takes the form of short texts (e.g., blog posts, news feeds, snippets, etc). This has motivated an increasing interest on the analysis of short texts and, specifically, on their...

Full description

Bibliographic Details
Published in:Language Resources & Evaluation Vol. 47; no. 1; pp. 151 - 179
Main Authors: Romero, Francisco, Julián-Iranzo, Pascual, Soto, Andrés, Ferreira-Satler, Mateus, Gallardo-Casero, Juan
Format: Article
Published: Springer Nature Mar2013
Subjects:
Online Access:View this record in EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=85873227&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 85873227
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Mar2013
      vid: 47
      iid: 1
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        85873227
        10.1007/s10579-012-9203-2
      ppf: 151
      ppct: 28
      formats:
        fmt:
          @attributes:
            type: P
            size: 836KB
      tig:
        atl: Classifying unlabeled short texts using a fuzzy declarative approach.
      aug:
        au:
          Romero, Francisco
          Julián-Iranzo, Pascual
          Soto, Andrés
          Ferreira-Satler, Mateus
          Gallardo-Casero, Juan
        affil:
          Department of Information Technologies and Systems, University of Castilla La Mancha, Paseo de la Universidad, 4 13071 Ciudad Real Spain
          Department of Computer Science, Universidad Autònoma del Carmen, Ciudad del Carmen CP 24160 Campeche Mèxico
      su:
        Categorization (Linguistics)
        Linguistic analysis
        Ontologies (Information retrieval)
        Thesauri
        Subject headings
      sug:
        subj:
          Categorization (Linguistics)
          Linguistic analysis
          Ontologies (Information retrieval)
          Thesauri
          Subject headings
      keyword:
        Ontologies
        Text categorization
        Unlabeled short texts
      ab: Web 2.0 provides user-friendly tools that allow persons to create and publish content online. User generated content often takes the form of short texts (e.g., blog posts, news feeds, snippets, etc). This has motivated an increasing interest on the analysis of short texts and, specifically, on their categorisation. Text categorisation is the task of classifying documents into a certain number of predefined categories. Traditional text classification techniques are mainly based on word frequency statistical analysis and have been proved inadequate for the classification of short texts where word occurrence is too small. On the other hand, the classic approach to text categorization is based on a learning process that requires a large number of labeled training texts to achieve an accurate performance. However labeled documents might not be available, when unlabeled documents can be easily collected. This paper presents an approach to text categorisation which does not need a pre-classified set of training documents. The proposed method only requires the category names as user input. Each one of these categories is defined by means of an ontology of terms modelled by a set of what we call proximity equations. Hence, our method is not category occurrence frequency based, but highly depends on the definition of that category and how the text fits that definition. Therefore, the proposed approach is an appropriate method for short text classification where the frequency of occurrence of a category is very small or even zero. Another feature of our method is that the classification process is based on the ability of an extension of the standard Prolog language, named , for flexible matching and knowledge representation. This declarative approach provides a text classifier which is quick and easy to build, and a classification process which is easy for the user to understand. The results of experiments showed that the proposed method achieved a reasonably useful performance.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2013. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2013
    holdings:
      @attributes:
        islocal: N