The short answer
You can use AI in a systematic review if it does not compromise the rigour of the review, a human stays responsible for every judgement, and you report the use fully. That is the shared position of Cochrane, the Campbell Collaboration, JBI and the Collaboration for Environmental Evidence, published together in November 2025.
In practice AI works best as a second screener, a prioritisation aid and an extraction assistant whose output you check. It is not yet reliable enough to make inclusion decisions on its own.
What RAISE and the joint position statement say
RAISE (Responsible use of AI in evidence SynthEsis) is a set of recommendations for people who carry out reviews, build tools for them, or choose and use those tools. In November 2025 Cochrane, Campbell, JBI and CEE published a joint position statement based on it in their journals at the same time.
The key points: evidence synthesists remain ultimately responsible for their review. AI may be used if it will not compromise methodological rigour or integrity. Human oversight is required. Any AI that makes or suggests judgements, such as eligibility, risk of bias or data extraction, must be fully and transparently reported. Cochrane described the statement as a pivotal moment for the evidence synthesis community.
AI for screening: what the evidence shows
Screening is the most studied use. Results vary with the model, the prompt and the topic, so treat any single number with care.
- Machine-learning classifiers are well established. Cochrane’s RCT classifier retrieved 99.5% of the randomised trials included in Cochrane reviews while letting reviewers skip a large share of irrelevant records (Thomas and colleagues, 2021).
- Active-learning tools such as ASReview rank records so that likely includes come first. Published stopping rules, such as the SAFE procedure, help you decide when it is reasonable to stop screening.
- Large language models as screeners show high sensitivity in some studies (Tran and colleagues, 2024, reported 94.6% to 99.8% for GPT-3.5 under a sensitive setting) and lower, more variable performance in others. Khraisha and colleagues (2024) urged substantial caution for title and abstract screening.
- A 2025 scoping review in the Journal of Clinical Epidemiology summed up LLMs in systematic reviews as "on the rise, but not yet ready for use".
Data extraction and search strategies
For data extraction, Gartlehner and colleagues (2024) found that an LLM extracted 160 data elements from ten trials with about 96% accuracy, but also made errors, including one invented value. That is good enough to speed up extraction and not good enough to skip checking.
For search strategies, Wang and colleagues (SIGIR 2023) found that ChatGPT can write Boolean queries with high precision but lower recall, and that the same prompt produces different queries on different runs. Use AI to brainstorm terms, then build and test the strategy yourself or with a librarian.
How to report AI use
PRISMA 2020 already asks you to report how automation tools were used in study selection, including whether any records were excluded by a machine alone. Records marked as ineligible by automation tools have their own box in the flow diagram.
The joint position statement asks you to report, for each AI use: the tool, version and dates; what it was used for and which parts of the review it affected; why it was justified and how outputs were validated; any interests in the tool; and its limitations. Help with spelling and grammar usually does not need declaring.
- Example: "Titles and abstracts were screened by one reviewer, with [tool, version] as a second screener. All records the tool proposed to exclude were checked by a human reviewer, and disagreements were resolved by discussion. Prompts are provided in Supplementary File 2."
A responsible AI workflow for your review
- 1. Write your protocol first, including how you will use AI and how you will check it. The joint statement includes a template for declaring AI use in a protocol.
- 2. Build the search with a librarian or a tested strategy. Use AI only to brainstorm synonyms.
- 3. Screen with AI as a second screener or prioritisation aid, and keep a human decision on every record.
- 4. Validate on a sample: compare AI and human decisions on a few hundred records and report agreement.
- 5. Let AI draft extraction tables, then check every value against the paper.
- 6. Report everything in your methods section and PRISMA flow diagram.
How Kahubi supports this workflow
Kahubi’s systematic review flow follows the same principle: the AI suggests include or exclude with a short reason based on your criteria, and you decide. Both are logged, so you can report agreement, and the PRISMA numbers come straight from the record. When you ask extraction questions, the agent reads the included papers in full and cites each one, so you can check every value.
Sources
- Flemyng E, Noel-Storr A, Macura B, Gartlehner G, Thomas J, Meerpohl JJ, et al. (2025). Position statement on artificial intelligence use in evidence synthesis across Cochrane, the Campbell Collaboration, JBI and the Collaboration for Environmental Evidence. Cochrane Database of Systematic Reviews, ED000178.
- RAISE: Responsible use of AI in evidence SynthEsis. Open Science Framework.
- Cochrane (2025). Setting standards for responsible AI use in evidence synthesis.
- Cochrane (2025). How Cochrane is advancing responsible AI in evidence synthesis.
- PRISMA 2020 item 8: selection process. EQUATOR Network.
- Thomas J, et al. (2021). Machine learning reduced workload with minimal risk of missing studies: development and evaluation of a randomized controlled trial classifier for Cochrane Reviews. Journal of Clinical Epidemiology 133:140-151.
- van de Schoot R, et al. (2021). An open source machine learning framework for efficient and transparent systematic reviews. Nature Machine Intelligence 3:125-133.
- Boetje J, van de Schoot R (2024). The SAFE procedure: a practical stopping heuristic for active learning-based screening in systematic reviews. Systematic Reviews 13:81.
- Tran VT, et al. (2024). Sensitivity and specificity of using GPT-3.5 Turbo models for title and abstract screening in systematic reviews. Annals of Internal Medicine 177:791-799.
- Khraisha Q, et al. (2024). Can large language models replace humans in systematic reviews? Research Synthesis Methods 15(4):616-626.
- Lieberum JL, et al. (2025). Large language models for conducting systematic reviews: on the rise, but not yet ready for use. Journal of Clinical Epidemiology 181:111746.
- Gartlehner G, et al. (2024). Data extraction for evidence synthesis using a large language model: a proof-of-concept study. Research Synthesis Methods 15(4):576-589.
- Wang S, Scells H, Koopman B, Zuccon G (2023). Can ChatGPT write a good Boolean query for systematic review literature search? SIGIR 2023.
Last updated 2026-10-09.