Discovery of Language Resources on the Web: Information Extraction from Heterogeneous Documents.
The present article is concerned with the problem of automatic database population via information extraction (IE) from web pages obtained from heterogeneous sources, such as those retrieved by a domain crawler. Specifically, we address the task of filling single multi-field templates from individua...
| Publicado en: | Literary & Linguistic Computing Vol. 22; no. 3; pp. 329 - 344 |
|---|---|
| Autores principales: | , |
| Formato: | Artículo |
| Publicado: |
Oxford University Press / USA
Sep2007
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=26863970&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 26863970 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 02681145 BJ1 jtl: Literary & Linguistic Computing issn: 02681145 maglogo: N pubinfo: dt: Sep2007 vid: 22 iid: 3 pid: 622 pub: Oxford University Press / USA artinfo: ui: 26863970 10.1093/llc/fqm010 ppf: 329 ppct: 15 formats: fmt: @attributes: type: P size: 237KB tig: atl: Discovery of Language Resources on the Web: Information Extraction from Heterogeneous Documents. aug: au: Pekar, Viktor Evans, Richard affil: School of Humanities, Languages, and Social Sciences, University of Wolverhampton, Stafford Street, Wolverhampton, WV1 1SB, UK su: Data mining Natural language processing Databases Websites Electronic information resources Electronic data processing sug: subj: Data mining Natural language processing Databases Websites Electronic information resources Electronic data processing ab: The present article is concerned with the problem of automatic database population via information extraction (IE) from web pages obtained from heterogeneous sources, such as those retrieved by a domain crawler. Specifically, we address the task of filling single multi-field templates from individual documents, a common scenario that involves free-format documents with the same communicative goal such as job adverts, CVs, or meeting/seminar announcements. We discuss challenges that arise in this scenario and propose solutions to them at different levels of the processing of web page content. Our main focus is on the issue of information extraction, which we address with a two-step machine learning approach that first aims to determine segments of a page that are likely to contain relevant facts and then delimits specific natural language expressions with which to fill template fields. We also present a range of techniques for the enrichment of web pages with semantic annotations, such as recognition of named entities, domain terminology and coreference resolution, and examine their effect on the information extraction method. We evaluate the developed IE system on the task of automatically populating a database with information on language resources available on the web. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: © 2019 EADH: The European Association for Digital Humanities. item: Literary & Linguistic Computing holder: Oxford University Press / USA dt: @attributes: year: 2007 holdings: @attributes: islocal: N |
|---|