Unleashing the killer corpus: experiences in creating the multi-everything AMI Meeting Corpus.
The AMI Meeting Corpus contains 100 h of meetings captured using many synchronized recording devices, and is designed to support work in speech and video processing, language engineering, corpus linguistics, and organizational psychology. It has been transcribed orthographically, with annotated subs...
| Publicado en: | Language Resources & Evaluation Vol. 41; no. 2; pp. 181 - 191 |
|---|---|
| Autor principal: | |
| Formato: | Artículo |
| Publicado: |
Springer Nature
May2007
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=27362876&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 27362876 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: May2007 vid: 41 iid: 2 pid: 237 pub: Springer Nature artinfo: ui: 27362876 10.1007/s10579-007-9040-x ppf: 181 ppct: 10 formats: fmt: @attributes: type: P size: 285KB tig: atl: Unleashing the killer corpus: experiences in creating the multi-everything AMI Meeting Corpus. aug: au: Carletta, Jean affil: University of Edinburgh , Edinburgh EH8 9LW UK su: Meetings Annotations Recording instruments Speech processing systems Discourse XML (Extensible Markup Language) sug: subj: Meetings Annotations Recording instruments Speech processing systems Discourse XML (Extensible Markup Language) keyword: Annotated corpora Discourse annotation ab: The AMI Meeting Corpus contains 100 h of meetings captured using many synchronized recording devices, and is designed to support work in speech and video processing, language engineering, corpus linguistics, and organizational psychology. It has been transcribed orthographically, with annotated subsets for everything from named entities, dialogue acts, and summaries to simple gaze and head movement. In this written version of an LREC conference keynote address, I describe the data and how it was created. If this is “killer” data, that presupposes a platform that it will “sell”; in this case, that is the NITE XML Toolkit, which allows a distributed set of users to create, store, browse, and search annotations for the same base data that are both time-aligned against signal and related to each other structurally. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2007. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2007 holdings: @attributes: islocal: N |
|---|