Automatic Extraction of Figures from Scientific Publications in High-Energy Physics.

Plots and figures play an important role in the process of understanding a scientific publication, providing overviews of large amounts of data or ideas that are difficult to intuitively present using only the text. State-of-the-art digital libraries, which serve as gateways to knowledge encoded in...

Descripción completa

Detalles Bibliográficos
Publicado en:Information Technology & Libraries Vol. 32; no. 4; pp. 25 - 53
Autores principales: Praczyk, Piotr Adam, Nogueras-Iso, Javier, Mele, Salvatore
Formato: algorithm equations & formulas research tables/charts Journal Article
Publicado: American Library Association Dec2013
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=93378683&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 93378683
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        07309295
        ITL
      jtl: Information Technology & Libraries
      issn: 07309295
      maglogo: N
    pubinfo:
      dt: Dec2013
      vid: 32
      iid: 4
      pid: 55
      pub: American Library Association
      place: Chicago, Illinois
    artinfo:
      ui:
        93378683
        104131346
        104131346
        10.6017/ital.v32i4.3670
        93378683
      ppf: 25
      ppct: 28
      formats:
        fmt:
          @attributes:
            type: P
      tig:
        atl: Automatic Extraction of Figures from Scientific Publications in High-Energy Physics.
      aug:
        au:
          Praczyk, Piotr Adam
          Nogueras-Iso, Javier
          Mele, Salvatore
        affil: PhD student, Universidad de Zaragoza, Spain
      sug:
        subj:
          Physics
          Graphics
          Access to Information
          Electronic Journals
          Information Retrieval
          Libraries, Electronic
          Algorithms
          Metadata
          Human
          Funding Source
      ab: Plots and figures play an important role in the process of understanding a scientific publication, providing overviews of large amounts of data or ideas that are difficult to intuitively present using only the text. State-of-the-art digital libraries, which serve as gateways to knowledge encoded in scholarly writings, do not yet take full advantage of the graphical content of documents. Enabling machines to automatically unlock the meaning of scientific illustrations would allow immense improvements in the way scientists work and the way knowledge is processed. In this paper, we present a novel solution for the initial problem of processing graphical content, obtaining figures from scholarly publications stored in PDF. Our method relies on vector properties of documents and, as such, does not introduce additional errors, unlike methods based on raster image processing. Emphasis has been placed on correctly processing documents in high-energy physics. The described approach distinguishes different classes of objects appearing in PDF documents and uses spatial clustering techniques to group objects into larger logical entities. Many heuristics allow the rejection of incorrect figure candidates and the extraction of different types of metadata.
      pubtype: Academic Journal
      doctype:
        algorithm
        equations & formulas
        research
        tables/charts
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N