Tackling Challenges in Large Language Model–Based Data Extraction via Context Engineering: A Commentary on Jansen et al. (2025).

Systematic reviews, particularly meta-analyses, involve crucial yet labor-intensive and error-prone stages of data extraction. Recent advances in large language models (LLMs) have unlocked new avenues for automating this process, potentially enhancing both efficiency and reliability. Recently, Janse...

Full description

Bibliographic Details
Published in:Psychological Bulletin Vol. 152; no. 4; pp. 404 - 420
Main Authors: Lu, Junsong, Wang, X. T. (XiaoTian)
Format: Article
Published: American Psychological Association Apr2026
Subjects:
Online Access:View this record in EBSCOhost
Description
Summary:Systematic reviews, particularly meta-analyses, involve crucial yet labor-intensive and error-prone stages of data extraction. Recent advances in large language models (LLMs) have unlocked new avenues for automating this process, potentially enhancing both efficiency and reliability. Recently, Jansen et al. (2025) systematically evaluated the accuracy and error patterns of LLM-assisted data extraction across 22 reviews published in Psychological Bulletin. Their findings indicated that while achieving acceptable-to-good accuracy for some variables describing study characteristics, LLMs struggled with numerical variables, especially those related to effect sizes. In this commentary, we discuss the current challenges of automated data extraction and potential pathways to improve the work reported in Jansen et al.'s study. We situate our discussion within the framework of context engineering, aiming to refine the information provided to LLMs through dynamic optimization strategies tailored to specific tasks. We identify five key challenges that reflect either LLMs' unique patterns or standard practices in research synthesis: parsing semistructured data, understanding long contexts, performing arithmetic induction, engaging in complex reasoning, and ensuring the reproducibility of coding protocols. We then outline potential solutions inspired by context engineering implementations such as retrieval-augmented generation and tool-integrated reasoning. For illustration, we present four examples: extracting semistructured data via optical character recognition, reliably computing effect sizes through function calls, performing adaptive retrieval with LLM-based agents, and iteratively improving outputs through self-refinement. We conclude by calling for future research in automated data extraction to advance beyond simple instruction-following paradigms toward more reliable forms of context engineering. Public Significance Statement: Conducting systematic reviews and meta-analyses is becoming increasingly time-consuming and costly as the volume of scientific literature continues to grow exponentially. This commentary highlights that—although large language models hold promises for automating data extraction—they still face significant limitations, including difficulties in handling numerical information and performing tasks that require complex reasoning. We identify key challenges and outline directions for potential solutions, such as improved document processing and the integration of external computational tools, to enhance the accuracy, transparency, and reliability of artificial intelligence–assisted data extraction; to reduce costs; and to accelerate the synthesis of scientific evidence.