Identifying translationese at the word and sub-word level.

We use text classification to distinguish automatically between original and translated texts in Hebrew, a morphologically complex language. To this end, we design several linguistically informed feature sets that capture word-level and sub-word-level (in particular, morphological) properties of Heb...

Full description

Bibliographic Details
Published in:Digital Scholarship in the Humanities Vol. 31; no. 1; pp. 30 - 55
Main Authors: Avner, Ehud Alexander, Ordan, Noam, Wintner, Shuly
Format: Article
Published: Oxford University Press / USA 4/1/2016
Subjects:
Online Access:View this record in EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=114160245&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 114160245
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        2055768X
        JEO9
      jtl: Digital Scholarship in the Humanities
      issn: 2055768X
      maglogo: N
    pubinfo:
      dt: 4/1/2016
      vid: 31
      iid: 1
      pid: 622
      pub: Oxford University Press / USA
    artinfo:
      ui:
        114160245
        10.1093/llc/fqu047
      ppf: 30
      ppct: 25
      formats:
        fmt:
          @attributes:
            type: P
            size: 1.7MB
      tig:
        atl: Identifying translationese at the word and sub-word level.
      aug:
        au:
          Avner, Ehud Alexander
          Ordan, Noam
          Wintner, Shuly
        affil:
          Universität Potsdam, Germany
          Universität des Saarlandes, Germany
          University of Haifa, Israel
      su:
        Hebrew language
        Biblical language & style
        Jewish languages
        Morphology (Grammar)
        Humanities research
      sug:
        subj:
          Hebrew language
          Biblical language & style
          Jewish languages
          Morphology (Grammar)
          Humanities research
      ab: We use text classification to distinguish automatically between original and translated texts in Hebrew, a morphologically complex language. To this end, we design several linguistically informed feature sets that capture word-level and sub-word-level (in particular, morphological) properties of Hebrew. Such features are abstract enough to allow for the development of accurate, robust classifiers, and they also lend themselves to linguistic interpretation. Careful evaluation shows that some of the classifiers we define are, indeed, highly accurate, and scale up nicely to domains that they were not trained on. In addition, analysis of the best features provides insight into the morphological properties of translated texts.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: © 2019 EADH: The European Association for Digital Humanities.
      item: Digital Scholarship in the Humanities
      holder: Oxford University Press / USA
      dt:
        @attributes:
          year: 2016
    holdings:
      @attributes:
        islocal: N