Text Categorization from category name in an industry-motivated scenario.
In this work we suggest a novel Text Categorization (TC) scenario, motivated by an ad-hoc industrial need to assign documents to a set of predefined categories, while labeled training data for the categories is not available. The scenario is applicable in many industrial settings and is interesting...
| Publicado en: | Language Resources & Evaluation Vol. 49; no. 2; pp. 227 - 262 |
|---|---|
| Autores principales: | , , |
| Formato: | Artículo |
| Publicado: |
Springer Nature
Jun2015
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=102481851&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 102481851 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Jun2015 vid: 49 iid: 2 pid: 237 pub: Springer Nature artinfo: ui: 102481851 10.1007/s10579-015-9298-3 ppf: 227 ppct: 35 formats: fmt: @attributes: type: P size: 1.4MB tig: atl: Text Categorization from category name in an industry-motivated scenario. aug: au: Liebeskind, Chaya Kotlerman, Lili Dagan, Ido affil: Bar Ilan University, 5290002 Ramat Gan Israel su: Categorization (Linguistics) Semantics Hypertext systems Ad hoc computer networks Social media sug: subj: Categorization (Linguistics) Semantics Hypertext systems Ad hoc computer networks Social media keyword: Name-based Text Categorization Natural language processing Semantic similarity ab: In this work we suggest a novel Text Categorization (TC) scenario, motivated by an ad-hoc industrial need to assign documents to a set of predefined categories, while labeled training data for the categories is not available. The scenario is applicable in many industrial settings and is interesting from the academic perspective. We present a new dataset geared for the main characteristics of the scenario, and utilize it to investigate the name-based TC approach, which uses the category names as its only input and does not require training data. We evaluate and analyze the performance of state-of-the-art methods for this dataset to identify the shortcomings of these methods for our scenario, and suggest ways for overcoming these shortcomings. We utilize statistical correlation measured over a target corpus for improving the state-of-the-art, and offer a different classification scheme based on the characteristics of the setting. We evaluate our improvements and adaptations and show superior performance of our suggested method. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2015. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2015 holdings: @attributes: islocal: N |
|---|