ANNIS: A new architecture for generic corpus query and visualization.

This article is concerned with the data structures, properties of query languages, and visualization facilities required for the generic representation of richly annotated, heterogeneous linguistic corpora. We propose that above and beyond a general graph-based data model, which is becoming increasi...

Descripción completa

Detalles Bibliográficos
Publicado en:Digital Scholarship in the Humanities Vol. 31; no. 1; pp. 118 - 140
Autores principales: Krause, Thomas, Zeldes, Amir
Formato: Artículo
Publicado: Oxford University Press / USA 4/1/2016
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=114160250&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 114160250
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        2055768X
        JEO9
      jtl: Digital Scholarship in the Humanities
      issn: 2055768X
      maglogo: N
    pubinfo:
      dt: 4/1/2016
      vid: 31
      iid: 1
      pid: 622
      pub: Oxford University Press / USA
    artinfo:
      ui:
        114160250
        10.1093/llc/fqu057
      ppf: 118
      ppct: 22
      formats:
        fmt:
          @attributes:
            type: P
            size: 2MB
      tig:
        atl: ANNIS: A new architecture for generic corpus query and visualization.
      aug:
        au:
          Krause, Thomas
          Zeldes, Amir
        affil:
          Humboldt-Universität zu Berlin, Germany
          Georgetown University, USA
      su:
        Query languages (Computer science)
        Genericalness (Linguistics)
        Linguistics research
        Data visualization
        Digital humanities
      sug:
        subj:
          Query languages (Computer science)
          Genericalness (Linguistics)
          Linguistics research
          Data visualization
          Digital humanities
      ab: This article is concerned with the data structures, properties of query languages, and visualization facilities required for the generic representation of richly annotated, heterogeneous linguistic corpora. We propose that above and beyond a general graph-based data model, which is becoming increasingly popular in many complex annotation formats, a well-defined concept of multiple, potentially conflicting segmentation layers must be introduced to deal with different sources and applications of corpus data flexibly. We also propose a generic solution for specialized corpus visualizations in a Web interface using annotation-triggered style sheets, which leverage the power of modern browsers and CSS for multiple and highly customizable views of primary data. We offer an implementation and evaluation of our architecture in ANNIS3, an open-source browser-based architecture for corpus search and visualization. We present three case studies to test the coverage of the system, encompassing core linguistic and digital humanities use-cases including richly annotated newspaper treebanks, multilingual diplomatic and normalized manuscript materials edited in TEI, and analysis of multimodal recordings of spoken language.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: © 2019 EADH: The European Association for Digital Humanities.
      item: Digital Scholarship in the Humanities
      holder: Oxford University Press / USA
      dt:
        @attributes:
          year: 2016
    holdings:
      @attributes:
        islocal: N