A flexible template generation and matching method with applications for publication reference metadata extraction.
Conventional rule‐based approaches use exact template matching to capture linguistic information and necessarily need to enumerate all variations. We propose a novel flexible template generation and matching scheme called the principle‐based approach (PBA) based on sequence alignment, and employ it...
| Published in: | Journal of the Association for Information Science & Technology Vol. 72; no. 1; pp. 32 - 46 |
|---|---|
| Main Authors: | , , , , |
| Format: | equations & formulas research tables/charts Journal Article |
| Published: |
Wiley-Blackwell
Jan2021
|
| Online Access: | View this record in EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=147599138&site=ehost-live header: @attributes: shortDbName: ccm uiTerm: 147599138 longDbName: CINAHL Complete uiTag: AN controlInfo: bkinfo: dissinfo: jinfo: jid: 23301635 H6JN jtl: Journal of the Association for Information Science & Technology issn: 23301635 maglogo: N pubinfo: dt: Jan2021 vid: 72 iid: 1 pid: 480 pub: Wiley-Blackwell place: Malden, Massachusetts artinfo: ui: 147599138 145805013 147599138 147599138 10.1002/asi.24391 147599138 ppf: 32 ppct: 14 formats: tig: atl: A flexible template generation and matching method with applications for publication reference metadata extraction. aug: au: Yang, Ting‐Hao Hsieh, Yu‐Lun Liu, Shih‐Hung Chang, Yung‐Chun Hsu, Wen‐Lian affil: Department of Computer Science, National Tsing Hua University, Hsinchu, Taiwan sug: subj: Metadata Information Management Automation Human Algorithms Deep Learning Machine Learning Paired T-Tests Data Analysis Software Funding Source ab: Conventional rule‐based approaches use exact template matching to capture linguistic information and necessarily need to enumerate all variations. We propose a novel flexible template generation and matching scheme called the principle‐based approach (PBA) based on sequence alignment, and employ it for reference metadata extraction (RME) to demonstrate its effectiveness. The main contributions of this research are threefold. First, we propose an automatic template generation that can capture prominent patterns using the dominating set algorithm. Second, we devise an alignment‐based template‐matching technique that uses a logistic regression model, which makes it more general and flexible than pure rule‐based approaches. Last, we apply PBA to RME on extensive cross‐domain corpora and demonstrate its robustness and generality. Experiments reveal that the same set of templates produced by the PBA framework not only deliver consistent performance on various unseen domains, but also surpass hand‐crafted knowledge (templates). We use four independent journal style test sets and one conference style test set in the experiments. When compared to renowned machine learning methods, such as conditional random fields (CRF), as well as recent deep learning methods (i.e., bi‐directional long short‐term memory with a CRF layer, Bi‐LSTM‐CRF), PBA has the best performance for all datasets. pubtype: Academic Journal doctype: equations & formulas research tables/charts Journal Article ougenre: Article language: English refInfo: holdings: @attributes: islocal: N |
|---|