Prompting encoder models for zero-shot classification: a cross-domain study in Italian.
Addressing the challenge of limited annotated data in specialized fields and low-resource languages is crucial for the effective use of language models (LMs). While most large language models (LLMs) are trained on general-purpose English corpora, there is a notable gap in models specifically tailore...
| Publicado en: | Language Resources & Evaluation Vol. 59; no. 4; pp. 3659 - 3698 |
|---|---|
| Autores principales: | , , , , , |
| Formato: | Artículo |
| Publicado: |
Springer Nature
Dec2025
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=189912025&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 189912025 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Dec2025 vid: 59 iid: 4 pid: 237 pub: Springer Nature artinfo: ui: 189912025 10.1007/s10579-025-09853-0 ppf: 3659 ppct: 39 formats: fmt: – @attributes: type: T – @attributes: type: P size: 1.1MB tig: atl: Prompting encoder models for zero-shot classification: a cross-domain study in Italian. aug: au: Auriemma, Serena Miliani, Martina Madeddu, Mauro Bondielli, Alessandro Passaro, Lucia Lenci, Alessandro affil: https://ror.org/03ad39j10 CoLing Lab, Department of Philology, Literature and Linguistics, University of Pisa, 36 Santa Maria Street, 56126, Pisa, Italy https://ror.org/03ad39j10 Department of Computer Science, University of Pisa, 3 Largo Bruno Pontecorvo, 56127, Pisa, Italy su: Italian language Domain specificity Jargon (Terminology) Machine learning Natural language processing Feature extraction Language models Classification sug: subj: Italian language Domain specificity Jargon (Terminology) Machine learning Natural language processing Feature extraction Language models Classification keyword: Communication and Culture Linguistics Domain-adapted model Encoder Language Legal Prompting Public administration Zero-shot classification ab: Addressing the challenge of limited annotated data in specialized fields and low-resource languages is crucial for the effective use of language models (LMs). While most large language models (LLMs) are trained on general-purpose English corpora, there is a notable gap in models specifically tailored for Italian, particularly for technical and bureaucratic jargon. This paper explores the feasibility of employing smaller, domain-specific encoder LMs alongside prompting techniques to enhance performance in these specialized contexts. Our study concentrates on the Italian bureaucratic and legal language, experimenting with both general-purpose and further pre-trained encoder-only models. We evaluated the models on downstream tasks such as document classification and entity typing and conducted intrinsic evaluations using Pseudo-log-likelihood. The results indicate that while further pre-trained models may show diminished robustness in general knowledge, they exhibit superior adaptability for domain-specific tasks, even in a zero-shot setting. Furthermore, the application of calibration techniques and in-domain verbalizers significantly enhances the efficacy of encoder models. These domain-specialized models prove to be particularly advantageous in scenarios where in-domain resources or expertise are scarce. In conclusion, our findings offer new insights into the use of Italian models in specialized contexts, which may have a significant impact on both research and industrial applications in the digital transformation era. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2025. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2025 holdings: @attributes: islocal: N |
|---|