An approach to enhance topic modeling by using paratext and nonnegative matrix factorizations.
Given the growing expansion in the development and use of computational methods in humanities research, it is necessary to propose methodologies that properly explore the questions posed by different disciplines, considering the locality of both data and the process behind its generation. In the pre...
| Publicado en: | Digital Scholarship in the Humanities Vol. 38; no. 1; pp. 87 - 99 |
|---|---|
| Autores principales: | , , |
| Formato: | Artículo |
| Publicado: |
Oxford University Press / USA
Apr2023
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=162941105&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 162941105 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 2055768X JEO9 jtl: Digital Scholarship in the Humanities issn: 2055768X maglogo: N pubinfo: dt: Apr2023 vid: 38 iid: 1 pid: 622 pub: Oxford University Press / USA artinfo: ui: 162941105 10.1093/llc/fqac043 ppf: 87 ppct: 12 formats: fmt: – @attributes: type: T – @attributes: type: P size: 331KB tig: atl: An approach to enhance topic modeling by using paratext and nonnegative matrix factorizations. aug: au: Flores-Garrido, Marisol García-Velázquez, Luis Miguel López-Vázquez, Julieta Arisbe affil: Escuela Nacional de Estudios Superiores Unidad Morelia, Universidad Nacional Autónoma de México , Mexico su: Matrix decomposition Nonnegative matrices Paratext Corpora Electronic data processing sug: subj: Matrix decomposition Nonnegative matrices Paratext Corpora Electronic data processing ab: Given the growing expansion in the development and use of computational methods in humanities research, it is necessary to propose methodologies that properly explore the questions posed by different disciplines, considering the locality of both data and the process behind its generation. In the present work, we explore the problem of automatically identifying the main topics in collections of Nahua discourses known as huehuetlahtollis. Each document in the collections is introduced through an extended title, and it is a natural question if enhancing the role of title terms during the unsupervised learning process could enrich results. Aiming at explainability, we consider a model based on nonnegative matrix factorizations (NMF). An overview of the historical process behind the composition of the explored corpora suggests that titles reflect the point of view of the collection's compiler in manners that justify viewing the paratext as a supplementary source on the material. Therefore, we propose a bi-objective NMF scheme that appropriately reflects the a priori knowledge on the corpus, linking and combining the information of titles and content to improve the accuracy in identifying topic groups and relevant terms within a corpus. By comparing three different schemes against the labels assigned by an expert, we show that our model better reflects the nature of data, translating into higher accuracy. Finally, we present some insights on the studied corpora derived from our analysis of identified relevant terms. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: © 2019 EADH: The European Association for Digital Humanities. item: Digital Scholarship in the Humanities holder: Oxford University Press / USA dt: @attributes: year: 2023 holdings: @attributes: islocal: N |
|---|