A random walk on an ontology: Using thesaurus structure for automatic subject indexing...[corrected] [published erratum appears in J AM SOC INF SCI TECHNOL 2013; 64(8):1757]
Relationships between terms and features are an essential component of thesauri, ontologies, and a range of controlled vocabularies. In this article, we describe ways to identify important concepts in documents using the relationships in a thesaurus or other vocabulary structures. We introduce a met...
| Publicado en: | Journal of the American Society for Information Science & Technology Vol. 64; no. 7; pp. 1330 - 1345 |
|---|---|
| Autores principales: | , |
| Formato: | algorithm equations & formulas research tables/charts Journal Article |
| Publicado: |
Wiley-Blackwell
Jul2013
|
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=104175458&site=ehost-live header: @attributes: shortDbName: ccm uiTerm: 104175458 longDbName: CINAHL Complete uiTag: AN controlInfo: bkinfo: dissinfo: jinfo: jid: 15322882 IGD jtl: Journal of the American Society for Information Science & Technology issn: 15322882 maglogo: Y pubinfo: dt: Jul2013 vid: 64 iid: 7 pid: 480 pub: Wiley-Blackwell place: Malden, Massachusetts artinfo: ui: 104175458 87947629 10.1002/asi.22853 104175458 ppf: 1330 ppct: 15 formats: tig: atl: A random walk on an ontology: Using thesaurus structure for automatic subject indexing...[corrected] [published erratum appears in J AM SOC INF SCI TECHNOL 2013; 64(8):1757] aug: au: Willis, Craig Losee, Robert M. affil: Graduate School of Library and Information Science, University of Illinois at Urbana-Champaign sug: subj: Abstracting and Indexing Methods Algorithms Subject Headings Agriculture Funding Source Human National Library of Medicine (U.S.) Probability ab: Relationships between terms and features are an essential component of thesauri, ontologies, and a range of controlled vocabularies. In this article, we describe ways to identify important concepts in documents using the relationships in a thesaurus or other vocabulary structures. We introduce a methodology for the analysis and modeling of the indexing process based on a weighted random walk algorithm. The primary goal of this research is the analysis of the contribution of thesaurus structure to the indexing process. The resulting models are evaluated in the context of automatic subject indexing using four collections of documents pre-indexed with 4 different thesauri ( AGROVOC [UN Food and Agriculture Organization], high-energy physics taxonomy [ HEP], National Agricultural Library Thesaurus [ NALT], and medical subject headings [ MeSH]). We also introduce a thesaurus-centric matching algorithm intended to improve the quality of candidate concepts. In all cases, the weighted random walk improves automatic indexing performance over matching alone with an increase in average precision ( AP) of 9% for HEP, 11% for MeSH, 35% for NALT, and 37% for AGROVOC. The results of the analysis support our hypothesis that subject indexing is in part a browsing process, and that using the vocabulary and its structure in a thesaurus contributes to the indexing process. The amount that the vocabulary structure contributes was found to differ among the 4 thesauri, possibly due to the vocabulary used in the corresponding thesauri and the structural relationships between the terms. Each of the thesauri and the manual indexing associated with it is characterized using the methods developed here. pubtype: Academic Journal doctype: algorithm equations & formulas research tables/charts Journal Article ougenre: Article language: English refInfo: holdings: @attributes: islocal: N |
|---|