Developing ODIN: A Multilingual Repository of Annotated Language Data for Hundreds of the World's Languages.

In this article, we review the process of building ODIN, the Online Database of Interlinear Text (http://odin.linguistlist.org) a multilingual repository of linguistically analyzed language data. ODIN is built from interlinear text that has been harvested from scholarly linguistic documents posted o...

Descripción completa

Detalles Bibliográficos
Publicado en:Literary & Linguistic Computing Vol. 25; no. 3; pp. 303 - 320
Autores principales: Lewis, William D., Xia, Fei
Formato: Artículo
Publicado: Oxford University Press / USA Sep2010
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=53375036&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 53375036
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        02681145
        BJ1
      jtl: Literary & Linguistic Computing
      issn: 02681145
      maglogo: N
    pubinfo:
      dt: Sep2010
      vid: 25
      iid: 3
      pid: 622
      pub: Oxford University Press / USA
    artinfo:
      ui:
        53375036
        10.1093/llc/fqq006
      ppf: 303
      ppct: 17
      formats:
        fmt:
          @attributes:
            type: P
            size: 433KB
      tig:
        atl: Developing ODIN: A Multilingual Repository of Annotated Language Data for Hundreds of the World's Languages.
      aug:
        au:
          Lewis, William D.
          Xia, Fei
        affil:
          Microsoft Research, U.S.A.
          Microsoft Research, One Microsoft Way, 99\1637, Redmond, WA 98052, U.S.A.
          Department of Linguistics, University of Washington, U.S.A.
      su:
        Foreign language education
        Multilingualism
        Online databases
        Computational linguistics
        Extraction (Linguistics)
        Linguists
        Annotations
        Parsing (Computer grammar)
      sug:
        subj:
          Foreign language education
          Multilingualism
          Online databases
          Computational linguistics
          Extraction (Linguistics)
          Linguists
          Annotations
          Parsing (Computer grammar)
      ab: In this article, we review the process of building ODIN, the Online Database of Interlinear Text (http://odin.linguistlist.org) a multilingual repository of linguistically analyzed language data. ODIN is built from interlinear text that has been harvested from scholarly linguistic documents posted on the web. At the time of this writing, ODIN holds nearly 190,000 instances of interlinear text representing annotated language data for more than 1,000 languages (representing data from >10% of the world's languages). ODIN's charter has been to make these data available to linguists and other language researchers via search, providing the facility to find instances of language data and related resources (i.e. the documents from which data were extracted) by language name, language family, and even annotations used to markup the data (e.g. NOM, ACC, ERG, PST, 3SG). Further, we have sought to enrich the data we have collected and extract ‘knowledge’ from the enriched content. To enrich the data, we use a variety of statistical tagging and parsing methods applied in the English translations. An enhanced search facility allows users to find data across languages for a variety of syntactic constructions and constituent orders, facilitating unprecedented automated and online discovery of language data.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: © 2019 EADH: The European Association for Digital Humanities.
      item: Literary & Linguistic Computing
      holder: Oxford University Press / USA
      dt:
        @attributes:
          year: 2010
    holdings:
      @attributes:
        islocal: N