Linguistic analysis of datasets for semantic textual similarity.

Semantic Textual Similarity (STS), which measures the equivalence of meanings between two textual segments, is an important and useful task in Natural Language Processing. In this article, we have analyzed the datasets provided by the Semantic Evaluation (SemEval) 2012–2014 campaigns for this task i...

Descripción completa

Detalles Bibliográficos
Publicado en:Digital Scholarship in the Humanities Vol. 35; no. 2; pp. 471 - 485
Autores principales: Wang, Chunlin, Castellón, Irene, Comelles, Elisabet
Formato: Artículo
Publicado: Oxford University Press / USA Jun2020
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=144382990&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 144382990
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        2055768X
        JEO9
      jtl: Digital Scholarship in the Humanities
      issn: 2055768X
      maglogo: N
    pubinfo:
      dt: Jun2020
      vid: 35
      iid: 2
      pid: 622
      pub: Oxford University Press / USA
    artinfo:
      ui:
        144382990
        10.1093/llc/fqy076
      ppf: 471
      ppct: 14
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 185KB
      tig:
        atl: Linguistic analysis of datasets for semantic textual similarity.
      aug:
        au:
          Wang, Chunlin
          Castellón, Irene
          Comelles, Elisabet
        affil:
          Artificial Solutions Iberia S.L. , Barcelona
          Departamento de Filología Catalana y Lingüística General , Universidad de Barcelona, Gran Via de les Corts Catalanes, Barcelona
          Departamento de Lenguas y Literaturas Modernas y de Estudios Ingleses , Universidad de Barcelona, Gran Via de les Corts Catalanes, Barcelona
      su:
        Linguistic analysis
        Industrialized building
        Semantics
        Vocabulary
        Corpora
      sug:
        subj:
          Linguistic analysis
          Industrialized building
          Semantics
          Vocabulary
          Corpora
      ab: Semantic Textual Similarity (STS), which measures the equivalence of meanings between two textual segments, is an important and useful task in Natural Language Processing. In this article, we have analyzed the datasets provided by the Semantic Evaluation (SemEval) 2012–2014 campaigns for this task in order to find out appropriate linguistic features for each dataset, taking into account the influence that linguistic features at different levels (e.g. syntactic constituents and lexical semantics) might have on the sentence similarity. Results indicate that a linguistic feature may have a different effect on different corpus due to the great difference in sentence structure and vocabulary between datasets. Thus, we conclude that the selection of linguistic features according to the genre of the text might be a good strategy for obtaining better results in the STS task. This analysis could be a useful reference for measuring system building and linguistic feature tuning.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: © 2019 EADH: The European Association for Digital Humanities.
      item: Digital Scholarship in the Humanities
      holder: Oxford University Press / USA
      dt:
        @attributes:
          year: 2020
    holdings:
      @attributes:
        islocal: N