Automatic Extraction of Figures from Scientific Publications in High-Energy Physics.
Plots and figures play an important role in the process of understanding a scientific publication, providing overviews of large amounts of data or ideas that are difficult to intuitively present using only the text. State-of-the-art digital libraries, which serve as gateways to knowledge encoded in...
| Publicado en: | Information Technology & Libraries Vol. 32; no. 4; pp. 25 - 53 |
|---|---|
| Autores principales: | , , |
| Formato: | algorithm equations & formulas research tables/charts Journal Article |
| Publicado: |
American Library Association
Dec2013
|
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=93378683&site=ehost-live header: @attributes: shortDbName: ccm uiTerm: 93378683 longDbName: CINAHL Complete uiTag: AN controlInfo: bkinfo: dissinfo: jinfo: jid: 07309295 ITL jtl: Information Technology & Libraries issn: 07309295 maglogo: N pubinfo: dt: Dec2013 vid: 32 iid: 4 pid: 55 pub: American Library Association place: Chicago, Illinois artinfo: ui: 93378683 104131346 104131346 10.6017/ital.v32i4.3670 93378683 ppf: 25 ppct: 28 formats: fmt: @attributes: type: P tig: atl: Automatic Extraction of Figures from Scientific Publications in High-Energy Physics. aug: au: Praczyk, Piotr Adam Nogueras-Iso, Javier Mele, Salvatore affil: PhD student, Universidad de Zaragoza, Spain sug: subj: Physics Graphics Access to Information Electronic Journals Information Retrieval Libraries, Electronic Algorithms Metadata Human Funding Source ab: Plots and figures play an important role in the process of understanding a scientific publication, providing overviews of large amounts of data or ideas that are difficult to intuitively present using only the text. State-of-the-art digital libraries, which serve as gateways to knowledge encoded in scholarly writings, do not yet take full advantage of the graphical content of documents. Enabling machines to automatically unlock the meaning of scientific illustrations would allow immense improvements in the way scientists work and the way knowledge is processed. In this paper, we present a novel solution for the initial problem of processing graphical content, obtaining figures from scholarly publications stored in PDF. Our method relies on vector properties of documents and, as such, does not introduce additional errors, unlike methods based on raster image processing. Emphasis has been placed on correctly processing documents in high-energy physics. The described approach distinguishes different classes of objects appearing in PDF documents and uses spatial clustering techniques to group objects into larger logical entities. Many heuristics allow the rejection of incorrect figure candidates and the extraction of different types of metadata. pubtype: Academic Journal doctype: algorithm equations & formulas research tables/charts Journal Article ougenre: Article language: English refInfo: holdings: @attributes: islocal: N |
|---|