El corpus paral·lel del Diari Oficial de la Generalitat de Catalunya: compilació, anàlisi i exemples d'ús.

In this paper the process of compilation of the parallel corpus from the Official Diary of the Catalan Government (DOGC) is presented. It describes the downloading process, the tools and processes for the treatment and linguistic analysis. The final result is a big parallel corpus that is freely ava...

Descripción completa

Detalles Bibliográficos
Publicado en:Zeitschrift für Katalanistik Vol. 30; pp. 269 - 292
Autor principal: Oliver, Antoni
Formato: Artículo
Publicado: Zeitschrift fur Katalanistik 2017
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=125047337&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 125047337
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        09322221
        1WQB
      jtl: Zeitschrift für Katalanistik
      issn: 09322221
      maglogo: N
    pubinfo:
      dt: 2017
      vid: 30
      pid: 27971
      pub: Zeitschrift fur Katalanistik
    artinfo:
      ui:
        125047337
        10.46586/zfk.2017.269-291
      ppf: 269
      ppct: 23
      formats:
        fmt:
          @attributes:
            type: P
            size: 2.1MB
      tig:
        atl: El corpus paral·lel del Diari Oficial de la Generalitat de Catalunya: compilació, anàlisi i exemples d'ús.
      aug:
        au: Oliver, Antoni
        affil: Universitat Oberta de Catalunya, Estudis d'Arts i Humanitats, Avda. Tibidabo, 39-43, E-08035 Barcelona
      sug:
      keyword:
        Natural Language Processing
        Parallel corpus
        statistical machine translation
        terminology extraction
        translation memory
      ab: In this paper the process of compilation of the parallel corpus from the Official Diary of the Catalan Government (DOGC) is presented. It describes the downloading process, the tools and processes for the treatment and linguistic analysis. The final result is a big parallel corpus that is freely available in several formats and with several annotation levels. This corpus is a very valuable resource for different applications. As example, three possible fields of application are described: as a translation memory to be used in a Computer-Assisted Translation tool; for terminology extraction and query and for training statistical machine translation systems.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: Catalan
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Copyright of Zeitschrift für Katalanistik is the property of Zeitschrift fur Katalanistik and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use.
      item: Zeitschrift für Katalanistik
      holder: Zeitschrift fur Katalanistik
      dt:
        @attributes:
          year: 2017
    holdings:
      @attributes:
        islocal: N