Using semantic roles to improve text classification in the requirements domain.

Engineering activities often produce considerable documentation as a by-product of the development process. Due to their complexity, technical analysts can benefit from text processing techniques able to identify concepts of interest and analyze deficiencies of the documents in an automated fashion....

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 52; no. 3; pp. 801 - 838
Autores principales: Rago, Alejandro, Diaz-Pace, J. Andres, Marcos, Claudia
Formato: Artículo
Publicado: Springer Nature Sep2018
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=131216685&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 131216685
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Sep2018
      vid: 52
      iid: 3
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        131216685
        10.1007/s10579-017-9406-7
      ppf: 801
      ppct: 37
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 2.9MB
      tig:
        atl: Using semantic roles to improve text classification in the requirements domain.
      aug:
        au:
          Rago, Alejandro
          Diaz-Pace, J. Andres
          Marcos, Claudia
        affil:
          ISISTAN Research Institute, UNICEN University, Tandil, Argentina
          CONICET, Buenos Aires, Argentina
          CIC, Buenos Aires, Argentina
      su:
        Machine learning
        Requirements engineering
        Discourse analysis
        Knowledge representation (Information theory)
        Wikipedia
      sug:
        subj:
          Machine learning
          Requirements engineering
          Discourse analysis
          Knowledge representation (Information theory)
          Wikipedia
      keyword:
        Knowledge representation
        Natural language processing
        Semantic enrichment
        Text classification
        Use case specification
      ab: Engineering activities often produce considerable documentation as a by-product of the development process. Due to their complexity, technical analysts can benefit from text processing techniques able to identify concepts of interest and analyze deficiencies of the documents in an automated fashion. In practice, text sentences from the documentation are usually transformed to a vector space model, which is suitable for traditional machine learning classifiers. However, such transformations suffer from problems of synonyms and ambiguity that cause classification mistakes. For alleviating these problems, there has been a growing interest in the semantic enrichment of text. Unfortunately, using general-purpose thesaurus and encyclopedias to enrich technical documents belonging to a given domain (e.g. requirements engineering) often introduces noise and does not improve classification. In this work, we aim at boosting text classification by exploiting information about semantic roles. We have explored this approach when building a multi-label classifier for identifying special concepts, called domain actions, in textual software requirements. After evaluating various combinations of semantic roles and text classification algorithms, we found that this kind of semantically-enriched data leads to improvements of up to 18% in both precision and recall, when compared to non-enriched data. Our enrichment strategy based on semantic roles also allowed classifiers to reach acceptable accuracy levels with small training sets. Moreover, semantic roles outperformed Wikipedia- and WordNET-based enrichments, which failed to boost requirements classification with several techniques. These results drove the development of two requirements tools, which we successfully applied in the processing of textual use cases.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2018. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2018
    holdings:
      @attributes:
        islocal: N