Diversité sociale et sémantique : représentation socio-sémantique d'un corpus scientifique, le cas du corpus ACL Anthology.
We propose a new method to extract multiword expressions from scientific papers. Our approach is made of two major steps: a first list of candidates is extracted based on a score using frequency and specificity information. This list is then filtered based on the status of the term in the abstract o...
| Publicado en: | Nouvelles Perspectives en Sciences Sociales Vol. 11; no. 2; pp. 145 - 180 |
|---|---|
| Autores principales: | , , , |
| Formato: | Artículo |
| Publicado: |
Editions Prise de parole
nov2015
|
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=ssf&AN=113571425&site=ehost-live header: @attributes: shortDbName: ssf uiTerm: 113571425 longDbName: Social Sciences Full Text (H.W. Wilson) uiTag: AN controlInfo: bkinfo: jinfo: jid: 17128307 84PR jtl: Nouvelles Perspectives en Sciences Sociales issn: 17128307 maglogo: N pubinfo: dt: nov2015 vid: 11 iid: 2 pid: 48324 pub: Editions Prise de parole artinfo: ui: 113571425 10.7202/1035935ar ppf: 145 ppct: 35 formats: fmt: @attributes: type: P size: 3.6MB tig: atl: Diversité sociale et sémantique : représentation socio-sémantique d'un corpus scientifique, le cas du corpus ACL Anthology. aug: au: OMODEI, ELISA GUO, YUFAN COINTET, JEAN-PHILIPPE POIBEAU, THIERRY affil: Department of Mathematics and Computer Science Rovira i Virgili University, Tarragone IBM Research, Almaden, San Jose Institut des Systèmes Complexes de Paris-Île de France (ISC-PIF) Laboratoire Lattice, UMR 89094, CNRS, ENS et Université Paris 3 Sorbonne Nouvelle keyword: Discourse Term Extraction ACL Anthology analyse discursive Corpus extraction de termes text zoning Discourse Term Extraction ACL Anthology analyse discursive Corpus extraction de termes text zoning ab: We propose a new method to extract multiword expressions from scientific papers. Our approach is made of two major steps: a first list of candidates is extracted based on a score using frequency and specificity information. This list is then filtered based on the status of the term in the abstract of the scientific papers under investigation. These abstracts are annotated using a text zoning analyser. The terms are then classified in different categories according to the text zoning analysis: we make a difference between terms appearing in the method section of the abstract vs terms appearing in other zones. This method is applied to the ACL Anthology collection, containing the papers published by the ACL between 1980 and 2008. We show that the technique we use allows us to model interesting facts concerning the evolution of the domain and of the methods used in computational linguistics. pubtype: Academic Journal doctype: Article src: R language: French refInfo: copyright: @attributes: flag: N holdings: @attributes: islocal: N |
|---|