In Pursuit of the Trivial.
We compare the performance of state-of-the-art Large Language Models on a recently released benchmarking set for automated question answering for Icelandic and compare it with performance on questions from an Icelandic trivia game. We find that the models perform worse for questions on Icelandic sub...
| Publicado en: | Digital Humanities in the Nordic & Baltic Countries Publications (DHNB Publications) Vol. 7; no. 2; pp. 1 - 9 |
|---|---|
| Autores principales: | , |
| Formato: | Artículo |
| Publicado: |
University of Oslo
2025
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=190250195&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 190250195 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 27041441 NUGX jtl: Digital Humanities in the Nordic & Baltic Countries Publications (DHNB Publications) issn: 27041441 maglogo: N pubinfo: dt: 2025 vid: 7 iid: 2 pid: 58757 pub: University of Oslo artinfo: ui: 190250195 10.5617/dhnbpub.12302 ppf: 1 ppct: 8 formats: tig: atl: In Pursuit of the Trivial. aug: au: Steingrímsson, Steinþór Ármannsson, Bjarki affil: The Árni Magnússon Institute for Icelandic Studies su: Trivia contests Question answering systems Language models Benchmark problems (Computer science) Computer performance Culture Cognitive ability sug: subj: Trivia contests Question answering systems Language models Benchmark problems (Computer science) Computer performance Culture Cognitive ability keyword: Automatic Question Answering Icelandic Large Language Models ab: We compare the performance of state-of-the-art Large Language Models on a recently released benchmarking set for automated question answering for Icelandic and compare it with performance on questions from an Icelandic trivia game. We find that the models perform worse for questions on Icelandic subjects, specifically Icelandic culture, but somewhat surprisingly do better on a trivia game for people than on the benchmark set meant for language models, built around data that the model has seen during training. We also call into question some aspects of the benchmarking set and discuss what playing trivia games can tell us - if anything - about the capabilities of these models. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y dt: @attributes: year: 2025 holdings: @attributes: islocal: N |
|---|