A comparative evaluation for question answering over Greek texts by using machine translation and BERT.
Although there are numerous and effective BERT models for question answering (QA) over plain texts in English, it is not the same for other languages, such as Greek. Since it can be time-consuming to train a new BERT model for a given language, we present a generic methodology for multilingual QA by...
| Publicado en: | Language Resources & Evaluation Vol. 59; no. 2; pp. 931 - 958 |
|---|---|
| Autores principales: | , , , |
| Formato: | Artículo |
| Publicado: |
Springer Nature
Jun2025
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=185240040&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 185240040 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Jun2025 vid: 59 iid: 2 pid: 237 pub: Springer Nature artinfo: ui: 185240040 10.1007/s10579-024-09745-9 ppf: 931 ppct: 27 formats: fmt: – @attributes: type: T – @attributes: type: P size: 1.5MB tig: atl: A comparative evaluation for question answering over Greek texts by using machine translation and BERT. aug: au: Mountantonakis, Michalis Mertzanis, Loukas Bastakis, Michalis Tzitzikas, Yannis affil: https://ror.org/02tf48g55 Institute of Computer Science, FORTH, Heraklion, Greece https://ror.org/00dr28g20 Department of Computer Science, University of Crete, Heraklion, Greece su: Language models Machine translating Language & languages Programming languages Greek language sug: subj: Language models Machine translating Language & languages Programming languages Greek language keyword: BERT Communication and Culture Linguistics Greek annotated test set Language Machine translation Question answering Short sentence similarity ab: Although there are numerous and effective BERT models for question answering (QA) over plain texts in English, it is not the same for other languages, such as Greek. Since it can be time-consuming to train a new BERT model for a given language, we present a generic methodology for multilingual QA by combining at runtime existing machine translation (MT) models and BERT QA models pretrained in English, and we perform a comparative evaluation for Greek language. Particularly, we propose a pipeline that (a) exploits widely used MT libraries for translating a question and a context from a source language to the English language, (b) extracts the answer from the translated English context through popular BERT models (pretrained in English corpus), (c) translates the answer back to the source language, and (d) evaluates the answer through semantic similarity metrics based on sentence embeddings, such as Bi-Encoder and BERTScore. For evaluating our system, we use 21 models, whereas we have created a test set with 20 texts and 200 questions and we have manually labelled 4200 answers. These resources can be reused for several tasks including QA and sentence similarity. Moreover, we use the existing multilingual test set XQuAD, with 240 texts and 1190 questions in Greek language. We focus on both the effectiveness and efficiency, through manually and machine labelled results. The results of the evaluation show that the proposed approach can be an efficient and effective alternative option to multilingual BERT. In particular, although the multilingual BERT QA model provides the highest scores for both human and automatic evaluation, all the models combining MT and BERT QA models are faster and some of them achieve quite similar scores. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2025. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2025 holdings: @attributes: islocal: N |
|---|