The Spoken Wikipedia Corpus collection: Harvesting, alignment and an application to hyperlistening.

Spoken corpora are important for speech research, but are expensive to create and do not necessarily reflect (read or spontaneous) speech 'in the wild'. We report on our conversion of the preexisting and freely available Spoken Wikipedia into a speech resource. The Spoken Wikipedia project unites vo...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 53; no. 2; pp. 303 - 330
Autores principales: Baumann, Timo, Köhn, Arne, Hennig, Felix
Formato: Artículo
Publicado: Springer Nature Jun2019
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=136891169&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 136891169
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Jun2019
      vid: 53
      iid: 2
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        136891169
        10.1007/s10579-017-9410-y
      ppf: 303
      ppct: 27
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 1.5MB
      tig:
        atl: The Spoken Wikipedia Corpus collection: Harvesting, alignment and an application to hyperlistening.
      aug:
        au:
          Baumann, Timo
          Köhn, Arne
          Hennig, Felix
        affil:
          School of Computer Science, Language Technology Institute, Carnegie Mellon University, 5000 Forbes Ave, 15213, Pittsburgh, PA, USA
          FB Informatik, Natural Language Systems group, Universität Hamburg, Vogt-Kölln-Straße 30, 22527, Hamburg, Germany
      su:
        Wikipedia
        Speech processing software
        Corpora
        Hypertext systems
        Annotations
      sug:
        subj:
          Wikipedia
          Speech processing software
          Corpora
          Hypertext systems
          Annotations
      keyword:
        Annotation
        Eyes-free speech access
        Found data
        Robust text–speech alignment
        Speech corpus
        Spoken hypertext
      ab: Spoken corpora are important for speech research, but are expensive to create and do not necessarily reflect (read or spontaneous) speech 'in the wild'. We report on our conversion of the preexisting and freely available Spoken Wikipedia into a speech resource. The Spoken Wikipedia project unites volunteer readers of Wikipedia articles. There are initiatives to create and sustain Spoken Wikipedia versions in many languages and hence the available data grows over time. Thousands of spoken articles are available to users who prefer a spoken over the written version. We turn these semi-structured collections into structured and time-aligned corpora, keeping the exact correspondence with the original hypertext as well as all available metadata. Thus, we make the Spoken Wikipedia accessible for sustainable research. We present our open-source software pipeline that downloads, extracts, normalizes and text–speech aligns the Spoken Wikipedia. Additional language versions can be exploited by adapting configuration files or extending the software if necessary for language peculiarities. We also present and analyze the resulting corpora for German, English, and Dutch, which presently total 1005 h and grow at an estimated 87 h per year. The corpora, together with our software, are available via http://islrn.org/resources/684-927-624-257-3/. As a prototype usage of the time-aligned corpus, we describe an experiment about the preferred modalities for interacting with information-rich read-out hypertext. We find alignments to help improve user experience and factual information access by enabling targeted interaction.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2019. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2019
    holdings:
      @attributes:
        islocal: N