Discourse lexicon induction for multiple languages and its use for gender profiling.
We propose a novel way to create categorized discourse lexicons for multiple languages. We combine information from the Penn Discourse Treebank with statistical machine translation techniques on the Europarl corpus. Using gender profiling as an application, we evaluate our approach by comparing it w...
| Published in: | Digital Scholarship in the Humanities Vol. 34; no. 1; pp. 208 - 221 |
|---|---|
| Main Authors: | , |
| Format: | Article |
| Published: |
Oxford University Press / USA
Apr2019
|
| Subjects: | |
| Online Access: | View this record in EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=135432200&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 135432200 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 2055768X JEO9 jtl: Digital Scholarship in the Humanities issn: 2055768X maglogo: N pubinfo: dt: Apr2019 vid: 34 iid: 1 pid: 622 pub: Oxford University Press / USA artinfo: ui: 135432200 10.1093/llc/fqy025 ppf: 208 ppct: 13 formats: fmt: – @attributes: type: T – @attributes: type: P size: 204KB tig: atl: Discourse lexicon induction for multiple languages and its use for gender profiling. aug: au: Verhoeven, Ben Daelemans, Walter affil: CLiPS Research Center, University of Antwerp, Belgium su: Discourse Lexicon English language Dutch language German language Discourse analysis Lexicology sug: subj: Discourse Lexicon English language Dutch language German language Discourse analysis Lexicology ab: We propose a novel way to create categorized discourse lexicons for multiple languages. We combine information from the Penn Discourse Treebank with statistical machine translation techniques on the Europarl corpus. Using gender profiling as an application, we evaluate our approach by comparing it with an approach using features from a knowledge-based lexicon and with an Rhetorical structure theory (RST) discourse parser. Our experiments are performed on corpora for three languages (English, Dutch, and German) in two genres (news and blogs). We include a feature analysis in which we look for (in)consistencies of discourse features related to male and female authors between the different experimental settings. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: © 2019 EADH: The European Association for Digital Humanities. item: Digital Scholarship in the Humanities holder: Oxford University Press / USA dt: @attributes: year: 2019 holdings: @attributes: islocal: N |
|---|