GPT models for text annotation: An empirical exploration in public policy research.

Text annotation, the practice of labeling text following a predetermined scheme, is essential to qualitative public policy research. Despite its importance, annotating large qualitative data faces challenges of high labor and time costs. Recent developments in large language models (LLMs), specifica...

Descripción completa

Detalles Bibliográficos
Publicado en:Policy Studies Journal Vol. 54; no. 1; pp. 1 - 18
Autores principales: Churchill, Alexander, Pichika, Shamitha, Xu, Chengxin, Liu, Ying
Formato: Artículo
Publicado: Wiley-Blackwell Feb2026
Materias:
Acceso en línea:Ver este registro en EBSCOhost
Descripción
Sumario:Text annotation, the practice of labeling text following a predetermined scheme, is essential to qualitative public policy research. Despite its importance, annotating large qualitative data faces challenges of high labor and time costs. Recent developments in large language models (LLMs), specifically models with generative pretrained transformers (GPTs), show a potential approach that may alleviate the burden of manual text annotation. In this report, we first introduce a small sample pretest strategy for researchers to decide whether to use Open AI's GPT models for text annotation. In addition, we test if GPT models can substitute human coders by comparing the results of two GPT models with different prompting strategies against human annotation. Using email messages collected from a national corresponding experiment in the US nursing home market as an example, on average, we demonstrate 86.25% percentage agreement between GPT and human annotations. We also show that GPT models possess context‐based limitations. Our report ends with reflections and suggestions for readers who are interested in using GPT models for text annotation.