The short answer
Use AI tools to get started faster, find studies you would otherwise miss, and read and organise what you find. Do not use them as the only search for a systematic review. Independent evaluations show they find a smaller share of relevant studies than a well-designed database search, and their results can change between runs.
Used alongside a traditional search, they are genuinely useful. Several evaluations found that AI tools surfaced relevant studies the original searches had missed.
The main types of AI literature tools
Most tools fall into four groups, and many products now combine several of them.
| Type | What it does | Examples |
|---|---|---|
| Deep research agents | Run a multi-step web search and write a long report with sources | OpenAI deep research, Gemini Deep Research, Perplexity |
| Academic AI search | Search scholarly indexes with natural-language questions and summarise papers | Elicit, Consensus, Undermind, SciSpace, Scite |
| Citation-network tools | Map papers that cite or resemble your seed papers | ResearchRabbit, Litmaps, Connected Papers |
| Research workspaces | Combine search with a library, reading, screening and writing | Kahubi, Paperguide, Anara |
Kahubi publishes this guide.
Where the papers come from
Most academic AI tools search open scholarly indexes, mainly Semantic Scholar and OpenAlex, rather than subscription databases such as Scopus, Web of Science or MEDLINE. These indexes are large, with hundreds of millions of records. For paywalled papers, though, the AI usually sees only the title and abstract. Librarian Aaron Tay notes, for example, that Consensus reads paywalled papers only at abstract level.
That matters in two ways: the AI cannot check details that appear only in the full text, and coverage of some fields and languages is thinner.
What independent evaluations show
Librarians and information specialists have tested these tools against established searches. The findings are consistent.
- Lau and Golder (2025, Cochrane Evidence Synthesis and Methods) compared Elicit with the searches behind four published reviews. Elicit found on average about 40% of the included studies, against about 95% for the original searches. Its precision was far higher, and it found some studies the originals missed. Their conclusion: not sensitive enough to replace traditional searching, but useful as an extra source.
- Bernard and colleagues (2025, BMC Medical Research Methodology) ran the same Elicit search three times and got 246, 169 and 172 results. They recommend AI tools as complements, not replacements.
- Tomczyk and colleagues (2024, Journal of Librarianship and Information Science) found Scopus and Web of Science more accurate than Elicit and SciSpace, while the AI tools returned more unique results.
- Parisi and Sutton (2024, Journal of EAHIL), two health information specialists, concluded that for systematic literature searching the limitations of ChatGPT outweigh its strengths, and that humans must stay in the loop.
Reproducibility, and how "agentic" the tools really are
A good literature search can be repeated by someone else. AI search often cannot. Aaron Tay tested the AI assistants inside Scopus and Web of Science in 2025 and found that the generated query changed between identical runs, in one tool about every other run. Plain Boolean search remains the most reproducible approach.
He also examined the "deep research" features of academic tools and found they are mostly fixed workflows, in which the model makes decisions only at set points, rather than fully autonomous agents. That is not a flaw in itself, but marketing can suggest more flexible reasoning than the tools deliver.
How to use AI tools well in a literature review
- Use AI search to explore a topic, find key papers and sharpen your question before you design the formal search.
- For a systematic or scoping review, run a documented database search as your main search and report AI tools as a supplementary source.
- Save the exact question, tool, date and results of every AI search, because you may not get the same results again.
- Read the papers, not only the summaries. AI summaries of real papers can still be wrong.
- Verify every reference before you cite it.
- Talk to your librarian. Many university libraries now publish guidance on AI search tools.
Where Kahubi fits
Kahubi searches OpenAlex and Semantic Scholar from a conversation, imports the papers you choose into your library with open-access PDFs, and lets you ask questions across the full texts with citations. For a systematic review, the dedicated flow keeps search strings, screening decisions and PRISMA numbers together. You can also import results from Scopus, Web of Science or PubMed as RIS or BibTeX, so the AI works alongside your documented database search rather than instead of it.
Sources
- Lau O, Golder S (2025). Comparison of Elicit AI and traditional literature searching in evidence syntheses using four case studies. Cochrane Evidence Synthesis and Methods 3(6):e70050.
- Bernard N, et al. (2025). Using artificial intelligence for systematic review: the example of Elicit. BMC Medical Research Methodology 25:75.
- Tomczyk P, Brüggemann P, Mergner N, Petrescu M (2024). Are AI tools better than traditional tools in literature searching? Evidence from e-commerce research. Journal of Librarianship and Information Science 58(1):135-145.
- Parisi V, Sutton A (2024). The role of ChatGPT in developing systematic literature searches: an evidence summary. Journal of EAHIL 20(2):30-34.
- Tay A (2025). Deep research, shallow agency: what academic deep research can and can’t do. Musings about librarianship.
- Tay A (2025). The reproducibility and interpretability of academic AI search engines. Musings about librarianship.
- Tay A (2025). A 2025 deep dive of Consensus. Musings about librarianship.
- University of British Columbia Library. AI search tools used in literature reviews and comprehensive searching (evidence list).
Last updated 2026-10-09.