Discovery of Language Resources on the Web: Information Extraction from Heterogeneous Documents.

The present article is concerned with the problem of automatic database population via information extraction (IE) from web pages obtained from heterogeneous sources, such as those retrieved by a domain crawler. Specifically, we address the task of filling single multi-field templates from individua...

Descripción completa

Detalles Bibliográficos
Publicado en:Literary & Linguistic Computing Vol. 22; no. 3; pp. 329 - 344
Autores principales: Pekar, Viktor, Evans, Richard
Formato: Artículo
Publicado: Oxford University Press / USA Sep2007
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=26863970&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 26863970
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        02681145
        BJ1
      jtl: Literary & Linguistic Computing
      issn: 02681145
      maglogo: N
    pubinfo:
      dt: Sep2007
      vid: 22
      iid: 3
      pid: 622
      pub: Oxford University Press / USA
    artinfo:
      ui:
        26863970
        10.1093/llc/fqm010
      ppf: 329
      ppct: 15
      formats:
        fmt:
          @attributes:
            type: P
            size: 237KB
      tig:
        atl: Discovery of Language Resources on the Web: Information Extraction from Heterogeneous Documents.
      aug:
        au:
          Pekar, Viktor
          Evans, Richard
        affil: School of Humanities, Languages, and Social Sciences, University of Wolverhampton, Stafford Street, Wolverhampton, WV1 1SB, UK
      su:
        Data mining
        Natural language processing
        Databases
        Websites
        Electronic information resources
        Electronic data processing
      sug:
        subj:
          Data mining
          Natural language processing
          Databases
          Websites
          Electronic information resources
          Electronic data processing
      ab: The present article is concerned with the problem of automatic database population via information extraction (IE) from web pages obtained from heterogeneous sources, such as those retrieved by a domain crawler. Specifically, we address the task of filling single multi-field templates from individual documents, a common scenario that involves free-format documents with the same communicative goal such as job adverts, CVs, or meeting/seminar announcements. We discuss challenges that arise in this scenario and propose solutions to them at different levels of the processing of web page content. Our main focus is on the issue of information extraction, which we address with a two-step machine learning approach that first aims to determine segments of a page that are likely to contain relevant facts and then delimits specific natural language expressions with which to fill template fields. We also present a range of techniques for the enrichment of web pages with semantic annotations, such as recognition of named entities, domain terminology and coreference resolution, and examine their effect on the information extraction method. We evaluate the developed IE system on the task of automatically populating a database with information on language resources available on the web.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: © 2019 EADH: The European Association for Digital Humanities.
      item: Literary & Linguistic Computing
      holder: Oxford University Press / USA
      dt:
        @attributes:
          year: 2007
    holdings:
      @attributes:
        islocal: N