A flexible template generation and matching method with applications for publication reference metadata extraction.

Conventional rule‐based approaches use exact template matching to capture linguistic information and necessarily need to enumerate all variations. We propose a novel flexible template generation and matching scheme called the principle‐based approach (PBA) based on sequence alignment, and employ it...

Full description

Bibliographic Details
Published in:Journal of the Association for Information Science & Technology Vol. 72; no. 1; pp. 32 - 46
Main Authors: Yang, Ting‐Hao, Hsieh, Yu‐Lun, Liu, Shih‐Hung, Chang, Yung‐Chun, Hsu, Wen‐Lian
Format: equations & formulas research tables/charts Journal Article
Published: Wiley-Blackwell Jan2021
Online Access:View this record in EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=147599138&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 147599138
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        23301635
        H6JN
      jtl: Journal of the Association for Information Science & Technology
      issn: 23301635
      maglogo: N
    pubinfo:
      dt: Jan2021
      vid: 72
      iid: 1
      pid: 480
      pub: Wiley-Blackwell
      place: Malden, Massachusetts
    artinfo:
      ui:
        147599138
        145805013
        147599138
        147599138
        10.1002/asi.24391
        147599138
      ppf: 32
      ppct: 14
      formats:
      tig:
        atl: A flexible template generation and matching method with applications for publication reference metadata extraction.
      aug:
        au:
          Yang, Ting‐Hao
          Hsieh, Yu‐Lun
          Liu, Shih‐Hung
          Chang, Yung‐Chun
          Hsu, Wen‐Lian
        affil: Department of Computer Science, National Tsing Hua University, Hsinchu, Taiwan
      sug:
        subj:
          Metadata
          Information Management
          Automation
          Human
          Algorithms
          Deep Learning
          Machine Learning
          Paired T-Tests
          Data Analysis Software
          Funding Source
      ab: Conventional rule‐based approaches use exact template matching to capture linguistic information and necessarily need to enumerate all variations. We propose a novel flexible template generation and matching scheme called the principle‐based approach (PBA) based on sequence alignment, and employ it for reference metadata extraction (RME) to demonstrate its effectiveness. The main contributions of this research are threefold. First, we propose an automatic template generation that can capture prominent patterns using the dominating set algorithm. Second, we devise an alignment‐based template‐matching technique that uses a logistic regression model, which makes it more general and flexible than pure rule‐based approaches. Last, we apply PBA to RME on extensive cross‐domain corpora and demonstrate its robustness and generality. Experiments reveal that the same set of templates produced by the PBA framework not only deliver consistent performance on various unseen domains, but also surpass hand‐crafted knowledge (templates). We use four independent journal style test sets and one conference style test set in the experiments. When compared to renowned machine learning methods, such as conditional random fields (CRF), as well as recent deep learning methods (i.e., bi‐directional long short‐term memory with a CRF layer, Bi‐LSTM‐CRF), PBA has the best performance for all datasets.
      pubtype: Academic Journal
      doctype:
        equations & formulas
        research
        tables/charts
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N