LLM-Based Medical Studies: Systematic Review Search Strategy & Evidence Tiering
The rapid evolution of large language models (LLMs) is transforming many fields and clinical medicine is no exception. A recent systematic review, published as a preprint and currently undergoing peer review, details a rigorous approach to evaluating the growing body of research on LLMs in healthcare. The study, conducted between January 2022 and September 2025, leveraged both traditional systematic review methods and an innovative, LLM-assisted framework to analyze thousands of studies. This work highlights the increasing need for standardized reporting and evaluation as LLMs grow more integrated into clinical practice.
Searching for Evidence in a Flood of Data
Researchers faced a significant challenge: a rapidly expanding volume of publications exploring LLMs like GPT, ChatGPT, LLaMA, Claude, Gemini, and Bard. To address this, they searched three major databases – PubMed, Embase, and Scopus – using carefully constructed search terms. These terms combined general descriptors of LLMs with specific model names, focusing on original research articles, conference papers, preprints, and letters related to health. The search deliberately excluded reviews, meta-analyses, surveys, and commentaries to focus on primary evidence. The specific query strings used for each database are available in the study’s supporting information, demonstrating a commitment to transparency, and reproducibility. You can find the PubMed query string here, and details on the Scopus and Embase queries within the preprint itself.
A Tiered Approach to Evaluating Evidence
Recognizing the varying levels of rigor in studies evaluating LLMs, the researchers developed a system for categorizing evidence. This system defines four tiers:
- Tier S: Real-world, prospective evaluations of deployed systems in live clinical environments, using randomized, controlled, and ideally blinded studies.
- Tier I: Retrospective or prospective evaluations on real, never-before-seen clinical data.
- Tier II: Evaluations using simulated clinical situations, open-ended questions, or subjective patient ratings.
- Tier III: Assessments based on board exams, multiple-choice questions, or case studies with clear-cut answers.
This tiered system allows for a nuanced understanding of the strength of evidence supporting different LLM applications. Tier S represents the most robust evidence, while Tier III provides the least insight into real-world clinical performance.
LLMs Assisting the Review Process
The sheer volume of studies necessitated an innovative approach. Researchers implemented an LLM-assisted screening pipeline using GPT-5, a powerful language model, to initially filter studies for inclusion. GPT-5 was instructed to identify studies that reported original evaluations of LLMs on clinical tasks, excluding those using non-LLM models or applying LLMs to non-clinical contexts. To validate this automated process, a blinded manual review of 500 randomly chosen studies was conducted. This comparison allowed researchers to assess the accuracy and reliability of the LLM-based screening. The precise inclusion and exclusion criteria used by both humans and the LLM are publicly available on a GitHub repository, further enhancing transparency.
Beyond initial screening, GPT-5 was also used to “tier” included studies according to the framework described above. The model was prompted to assign each study to one of the four tiers, using a standardized prompt also available on the GitHub repository. This automated tiering process significantly accelerated the review process, allowing researchers to focus on higher-level analysis.
Validating the AI’s Performance
While LLMs offer significant potential for streamlining systematic reviews, it’s crucial to validate their performance. The researchers rigorously compared the LLM’s screening and tiering decisions to those made by human experts. For inclusion screening, they calculated Cohen’s κ, a statistical measure of agreement, between human screeners and between humans and the LLM. Similar comparisons were made for tiering. These validation steps are essential for ensuring the reliability and trustworthiness of the LLM-assisted review process.
Findings and Implications
The systematic review ultimately included 4,609 studies. The researchers also extracted metadata from each study, including the specific LLM evaluated, the clinical specialty involved, and whether the LLM outperformed humans. This data extraction was also assisted by GPT-5, with manual review to ensure accuracy and consistency. The study’s findings provide a comprehensive overview of the current state of LLM research in clinical medicine, identifying key trends and gaps in the literature.
The use of LLMs to assist in systematic reviews is a relatively new development, and this study contributes to the growing body of evidence supporting its feasibility and potential benefits. However, it’s important to note that the researchers did not conduct a formal risk of bias assessment for individual studies, as the review’s primary focus was on mapping the landscape of LLM research rather than evaluating treatment effects. Instead, they relied on the evidence tier framework to capture aspects of methodological rigor.
The Future of Evidence Synthesis
This work builds on the broader effort to improve the transparency and rigor of systematic reviews. The PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) guidelines, designed to enhance the reporting of systematic reviews, are central to this effort. PRISMA provides authors with guidance on how to clearly and completely report their methods and findings. New extensions like PRISMA-trAIce are emerging to specifically address the challenges of incorporating AI into the systematic review process, as detailed in a recent publication in JMIR AI. The development of PRISMA-DFLLM, which combines PRISMA guidelines with domain-specific finetuned LLMs, represents another step towards leveraging the power of AI to accelerate evidence synthesis.
As LLMs continue to evolve, and as more research emerges on their applications in clinical medicine, standardized reporting and rigorous evaluation will be crucial. The framework outlined in this systematic review, combined with ongoing efforts to refine guidelines and develop new tools, will help ensure that LLMs are used responsibly and effectively to improve patient care. The next steps involve continued monitoring of the literature, refinement of the evidence tiering system, and further validation of LLM-assisted review methods. Researchers are also exploring the potential for “living systematic reviews,” which are continuously updated as new evidence becomes available, with the aid of LLMs.