An approach to enhance topic modeling by using paratext and nonnegative matrix factorizations.

Given the growing expansion in the development and use of computational methods in humanities research, it is necessary to propose methodologies that properly explore the questions posed by different disciplines, considering the locality of both data and the process behind its generation. In the pre...

Descripción completa

Detalles Bibliográficos
Publicado en:Digital Scholarship in the Humanities Vol. 38; no. 1; pp. 87 - 99
Autores principales: Flores-Garrido, Marisol, García-Velázquez, Luis Miguel, López-Vázquez, Julieta Arisbe
Formato: Artículo
Publicado: Oxford University Press / USA Apr2023
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=162941105&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 162941105
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        2055768X
        JEO9
      jtl: Digital Scholarship in the Humanities
      issn: 2055768X
      maglogo: N
    pubinfo:
      dt: Apr2023
      vid: 38
      iid: 1
      pid: 622
      pub: Oxford University Press / USA
    artinfo:
      ui:
        162941105
        10.1093/llc/fqac043
      ppf: 87
      ppct: 12
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 331KB
      tig:
        atl: An approach to enhance topic modeling by using paratext and nonnegative matrix factorizations.
      aug:
        au:
          Flores-Garrido, Marisol
          García-Velázquez, Luis Miguel
          López-Vázquez, Julieta Arisbe
        affil: Escuela Nacional de Estudios Superiores Unidad Morelia, Universidad Nacional Autónoma de México , Mexico
      su:
        Matrix decomposition
        Nonnegative matrices
        Paratext
        Corpora
        Electronic data processing
      sug:
        subj:
          Matrix decomposition
          Nonnegative matrices
          Paratext
          Corpora
          Electronic data processing
      ab: Given the growing expansion in the development and use of computational methods in humanities research, it is necessary to propose methodologies that properly explore the questions posed by different disciplines, considering the locality of both data and the process behind its generation. In the present work, we explore the problem of automatically identifying the main topics in collections of Nahua discourses known as huehuetlahtollis. Each document in the collections is introduced through an extended title, and it is a natural question if enhancing the role of title terms during the unsupervised learning process could enrich results. Aiming at explainability, we consider a model based on nonnegative matrix factorizations (NMF). An overview of the historical process behind the composition of the explored corpora suggests that titles reflect the point of view of the collection's compiler in manners that justify viewing the paratext as a supplementary source on the material. Therefore, we propose a bi-objective NMF scheme that appropriately reflects the a priori knowledge on the corpus, linking and combining the information of titles and content to improve the accuracy in identifying topic groups and relevant terms within a corpus. By comparing three different schemes against the labels assigned by an expert, we show that our model better reflects the nature of data, translating into higher accuracy. Finally, we present some insights on the studied corpora derived from our analysis of identified relevant terms.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: © 2019 EADH: The European Association for Digital Humanities.
      item: Digital Scholarship in the Humanities
      holder: Oxford University Press / USA
      dt:
        @attributes:
          year: 2023
    holdings:
      @attributes:
        islocal: N