Structured LLM Augmentation for Clinical Information Extraction...20th World Congress on Medical and Health Informatics, Aug 09 - 13, 2025, Taipei, Taiwan.

Information extraction tasks, such as Named Entity Recognition (NER) and Relation Extraction (RE), are essential for advancing clinical research and applications. However, these tasks are hindered by the scarcity of labeled clinical documents due to privacy concerns and high annotation costs. This s...

Descripción completa

Detalles Bibliográficos
Publicado en:Studies in Health Technology & Informatics Vol. 329; pp. 971 - 977
Autores principales: Ying Wei, Qi Li, Pillai, Jay
Formato: proceedings research tables/charts Journal Article
Publicado: Sage Publications Inc. 2025
Acceso en línea:Ver este registro en EBSCOhost
Descripción
Sumario:Information extraction tasks, such as Named Entity Recognition (NER) and Relation Extraction (RE), are essential for advancing clinical research and applications. However, these tasks are hindered by the scarcity of labeled clinical documents due to privacy concerns and high annotation costs. This study introduces a novel framework combining Large Language Models (LLMs) for data augmentation with an adapted BERT model for clinical information extraction. The framework encodes entity and relational information within clinical note segments, enabling LLMs to generate diverse and contextually accurate augmentations while preserving structural integrity. Augmented data is used to train a segmentationbased BERT model, overcoming sequence length limitations and integrating global context via BiLSTM. Evaluations on public and proprietary datasets demonstrate significant performance improvements, highlighting the approach's potential to address data scarcity in clinical information extraction tasks.