DeepDive: Declarative Knowledge Base Construction.
The dark data extraction or knowledge base construction (KBC) problem is to populate a relational database with information from unstructured data sources, such as emails, webpages, and PDFs. KBC is a long-standing problem in industry and research that encompasses problems of data extraction, cleani...
| Publicado en: | Communications of the ACM Vol. 60; no. 5; pp. 93 - 103 |
|---|---|
| Autores principales: | , , , , , , , |
| Formato: | Artículo |
| Publicado: |
Association for Computing Machinery
May2017
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=122701130&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 122701130 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 00010782 ACM jtl: Communications of the ACM issn: 00010782 maglogo: N pubinfo: dt: May2017 vid: 60 iid: 5 pid: 68 pub: Association for Computing Machinery artinfo: ui: 122701130 10.1145/3060586 ppf: 93 ppct: 10 formats: tig: atl: DeepDive: Declarative Knowledge Base Construction. aug: au: Ce Zhang Ré, Christopher Cafarella, Michael De Sa, Christopher Ratner, Alex Jaeho Shin Feiran Wang Sen Wu affil: ETH Zurich, Zurich, Switzerland Computer Science Department, Stanford University, Stanford, CA Lattice Data, Inc., Palo Alto, CA su: Knowledge base Data extraction Machine learning Algorithm research Database design sug: subj: Knowledge base Data extraction Machine learning Algorithm research Database design ab: The dark data extraction or knowledge base construction (KBC) problem is to populate a relational database with information from unstructured data sources, such as emails, webpages, and PDFs. KBC is a long-standing problem in industry and research that encompasses problems of data extraction, cleaning, and integration. We describe DeepDive, a system that combines database and machine learning ideas to help to develop KBC systems. The key idea in DeepDive is to frame traditional extract--transform--load (ETL) style data management problems as a single large statistical inference task that is declaratively defined by the user. DeepDive leverages the effectiveness and efficiency of statistical inference and machine learning for difficult extraction tasks, whereas not requiring users to directly write any probabilistic inference algorithms. Instead, domain experts interact with DeepDive by defining features or rules about the domain. DeepDive has been successfully applied to domains such as pharmacogenomics, paleobiology, and antihuman trafficking enforcement, achieving human-caliber quality at machine-caliber scale. We present the applications, abstractions, and techniques used in DeepDive to accelerate the construction of such dark data extraction systems. pubtype: Periodical doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y dt: @attributes: year: 2017 holdings: @attributes: islocal: N |
|---|