Linguistic analysis of datasets for semantic textual similarity.
Semantic Textual Similarity (STS), which measures the equivalence of meanings between two textual segments, is an important and useful task in Natural Language Processing. In this article, we have analyzed the datasets provided by the Semantic Evaluation (SemEval) 2012–2014 campaigns for this task i...
| Publicado en: | Digital Scholarship in the Humanities Vol. 35; no. 2; pp. 471 - 485 |
|---|---|
| Autores principales: | , , |
| Formato: | Artículo |
| Publicado: |
Oxford University Press / USA
Jun2020
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=144382990&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 144382990 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 2055768X JEO9 jtl: Digital Scholarship in the Humanities issn: 2055768X maglogo: N pubinfo: dt: Jun2020 vid: 35 iid: 2 pid: 622 pub: Oxford University Press / USA artinfo: ui: 144382990 10.1093/llc/fqy076 ppf: 471 ppct: 14 formats: fmt: – @attributes: type: T – @attributes: type: P size: 185KB tig: atl: Linguistic analysis of datasets for semantic textual similarity. aug: au: Wang, Chunlin Castellón, Irene Comelles, Elisabet affil: Artificial Solutions Iberia S.L. , Barcelona Departamento de Filología Catalana y Lingüística General , Universidad de Barcelona, Gran Via de les Corts Catalanes, Barcelona Departamento de Lenguas y Literaturas Modernas y de Estudios Ingleses , Universidad de Barcelona, Gran Via de les Corts Catalanes, Barcelona su: Linguistic analysis Industrialized building Semantics Vocabulary Corpora sug: subj: Linguistic analysis Industrialized building Semantics Vocabulary Corpora ab: Semantic Textual Similarity (STS), which measures the equivalence of meanings between two textual segments, is an important and useful task in Natural Language Processing. In this article, we have analyzed the datasets provided by the Semantic Evaluation (SemEval) 2012–2014 campaigns for this task in order to find out appropriate linguistic features for each dataset, taking into account the influence that linguistic features at different levels (e.g. syntactic constituents and lexical semantics) might have on the sentence similarity. Results indicate that a linguistic feature may have a different effect on different corpus due to the great difference in sentence structure and vocabulary between datasets. Thus, we conclude that the selection of linguistic features according to the genre of the text might be a good strategy for obtaining better results in the STS task. This analysis could be a useful reference for measuring system building and linguistic feature tuning. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: © 2019 EADH: The European Association for Digital Humanities. item: Digital Scholarship in the Humanities holder: Oxford University Press / USA dt: @attributes: year: 2020 holdings: @attributes: islocal: N |
|---|