Reliability of large language models as a tool for knowledge extraction from biographical dictionaries: the case of the Polish Biographical Dictionary.

Large language models are tools with great potential for text processing. This study aims to assess the reliability of the models' results in extracting structured knowledge from unstructured textual sources, particularly biographies from the Polish Biographical Dictionary. The task of the model was...

Descripción completa

Detalles Bibliográficos
Publicado en:Digital Scholarship in the Humanities Vol. 40; no. 2; pp. 538 - 549
Autores principales: Jaskulski, Piotr, Latos, Tomasz, Ryńca, Mariusz, Zapała, Adam
Formato: Artículo
Publicado: Oxford University Press / USA Jun2025
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=186085058&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 186085058
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        2055768X
        JEO9
      jtl: Digital Scholarship in the Humanities
      issn: 2055768X
      maglogo: N
    pubinfo:
      dt: Jun2025
      vid: 40
      iid: 2
      pid: 622
      pub: Oxford University Press / USA
    artinfo:
      ui:
        186085058
        10.1093/llc/fqaf014
      ppf: 538
      ppct: 11
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 724KB
      tig:
        atl: Reliability of large language models as a tool for knowledge extraction from biographical dictionaries: the case of the Polish Biographical Dictionary.
      aug:
        au:
          Jaskulski, Piotr
          Latos, Tomasz
          Ryńca, Mariusz
          Zapała, Adam
        affil: Tadeusz Manteuffel Institute of History of the Polish Academy of Sciences, Warsaw, 00-272, Poland
      su:
        Language models
        Interment
        Data mining
        Birthplaces
        Family relations
      sug:
        subj:
          Language models
          Interment
          Data mining
          Birthplaces
          Family relations
      keyword:
        biographies
        information extraction
        large language models
      ab: Large language models are tools with great potential for text processing. This study aims to assess the reliability of the models' results in extracting structured knowledge from unstructured textual sources, particularly biographies from the Polish Biographical Dictionary. The task of the model was to extract information about the individuals, such as date and place of birth, death and burial, family relationships, important people, related settlements and institutions as well as occupied positions. The test was conducted on a sample of 250 biographies. The texts were written in Polish from the 1930s onwards and described the lives of individuals from various historical periods. The results show that the large language model (LLM) is very effective in identifying basic personal data, important family relationships, occupations, or offices held by the characters. Weaker results were obtained when attempting to find institutions and places associated with the protagonists. The outcome of the test suggests that LLMs can efficiently assist in digitizing and structuring historical biographical data and offer a promising tool for improving historical knowledge bases and speeding up the work compared to manual extraction of information.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: © 2019 EADH: The European Association for Digital Humanities.
      item: Digital Scholarship in the Humanities
      holder: Oxford University Press / USA
      dt:
        @attributes:
          year: 2025
    holdings:
      @attributes:
        islocal: N