Diversité sociale et sémantique : représentation socio-sémantique d'un corpus scientifique, le cas du corpus ACL Anthology.

We propose a new method to extract multiword expressions from scientific papers. Our approach is made of two major steps: a first list of candidates is extracted based on a score using frequency and specificity information. This list is then filtered based on the status of the term in the abstract o...

Descripción completa

Detalles Bibliográficos
Publicado en:Nouvelles Perspectives en Sciences Sociales Vol. 11; no. 2; pp. 145 - 180
Autores principales: OMODEI, ELISA, GUO, YUFAN, COINTET, JEAN-PHILIPPE, POIBEAU, THIERRY
Formato: Artículo
Publicado: Editions Prise de parole nov2015
Acceso en línea:Ver este registro en EBSCOhost
Descripción
Sumario:We propose a new method to extract multiword expressions from scientific papers. Our approach is made of two major steps: a first list of candidates is extracted based on a score using frequency and specificity information. This list is then filtered based on the status of the term in the abstract of the scientific papers under investigation. These abstracts are annotated using a text zoning analyser. The terms are then classified in different categories according to the text zoning analysis: we make a difference between terms appearing in the method section of the abstract vs terms appearing in other zones. This method is applied to the ACL Anthology collection, containing the papers published by the ACL between 1980 and 2008. We show that the technique we use allows us to model interesting facts concerning the evolution of the domain and of the methods used in computational linguistics.