Unleashing the killer corpus: experiences in creating the multi-everything AMI Meeting Corpus.

The AMI Meeting Corpus contains 100 h of meetings captured using many synchronized recording devices, and is designed to support work in speech and video processing, language engineering, corpus linguistics, and organizational psychology. It has been transcribed orthographically, with annotated subs...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 41; no. 2; pp. 181 - 191
Autor principal: Carletta, Jean
Formato: Artículo
Publicado: Springer Nature May2007
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=27362876&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 27362876
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: May2007
      vid: 41
      iid: 2
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        27362876
        10.1007/s10579-007-9040-x
      ppf: 181
      ppct: 10
      formats:
        fmt:
          @attributes:
            type: P
            size: 285KB
      tig:
        atl: Unleashing the killer corpus: experiences in creating the multi-everything AMI Meeting Corpus.
      aug:
        au: Carletta, Jean
        affil: University of Edinburgh , Edinburgh EH8 9LW UK
      su:
        Meetings
        Annotations
        Recording instruments
        Speech processing systems
        Discourse
        XML (Extensible Markup Language)
      sug:
        subj:
          Meetings
          Annotations
          Recording instruments
          Speech processing systems
          Discourse
          XML (Extensible Markup Language)
      keyword:
        Annotated corpora
        Discourse annotation
      ab: The AMI Meeting Corpus contains 100 h of meetings captured using many synchronized recording devices, and is designed to support work in speech and video processing, language engineering, corpus linguistics, and organizational psychology. It has been transcribed orthographically, with annotated subsets for everything from named entities, dialogue acts, and summaries to simple gaze and head movement. In this written version of an LREC conference keynote address, I describe the data and how it was created. If this is “killer” data, that presupposes a platform that it will “sell”; in this case, that is the NITE XML Toolkit, which allows a distributed set of users to create, store, browse, and search annotations for the same base data that are both time-aligned against signal and related to each other structurally.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2007. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2007
    holdings:
      @attributes:
        islocal: N