AI tools for literature reviews: what they can and cannot do

AI search tools can find relevant papers in minutes, and some surface studies a keyword search misses. Published evaluations also show where they fall short. Here is an honest map of the tools, what librarians found when they tested them, and how to use them well.

Updated 2026-10-09

The short answer

Use AI tools to get started faster, find studies you would otherwise miss, and read and organise what you find. Do not use them as the only search for a systematic review. Independent evaluations show they find a smaller share of relevant studies than a well-designed database search, and their results can change between runs.

Used alongside a traditional search, they are genuinely useful. Several evaluations found that AI tools surfaced relevant studies the original searches had missed.

The main types of AI literature tools

Most tools fall into four groups, and many products now combine several of them.

TypeWhat it doesExamples
Deep research agentsRun a multi-step web search and write a long report with sourcesOpenAI deep research, Gemini Deep Research, Perplexity
Academic AI searchSearch scholarly indexes with natural-language questions and summarise papersElicit, Consensus, Undermind, SciSpace, Scite
Citation-network toolsMap papers that cite or resemble your seed papersResearchRabbit, Litmaps, Connected Papers
Research workspacesCombine search with a library, reading, screening and writingKahubi, Paperguide, Anara

Kahubi publishes this guide.

Where the papers come from

Most academic AI tools search open scholarly indexes, mainly Semantic Scholar and OpenAlex, rather than subscription databases such as Scopus, Web of Science or MEDLINE. These indexes are large, with hundreds of millions of records. For paywalled papers, though, the AI usually sees only the title and abstract. Librarian Aaron Tay notes, for example, that Consensus reads paywalled papers only at abstract level.

That matters in two ways: the AI cannot check details that appear only in the full text, and coverage of some fields and languages is thinner.

What independent evaluations show

Librarians and information specialists have tested these tools against established searches. The findings are consistent.

  • Lau and Golder (2025, Cochrane Evidence Synthesis and Methods) compared Elicit with the searches behind four published reviews. Elicit found on average about 40% of the included studies, against about 95% for the original searches. Its precision was far higher, and it found some studies the originals missed. Their conclusion: not sensitive enough to replace traditional searching, but useful as an extra source.
  • Bernard and colleagues (2025, BMC Medical Research Methodology) ran the same Elicit search three times and got 246, 169 and 172 results. They recommend AI tools as complements, not replacements.
  • Tomczyk and colleagues (2024, Journal of Librarianship and Information Science) found Scopus and Web of Science more accurate than Elicit and SciSpace, while the AI tools returned more unique results.
  • Parisi and Sutton (2024, Journal of EAHIL), two health information specialists, concluded that for systematic literature searching the limitations of ChatGPT outweigh its strengths, and that humans must stay in the loop.

Reproducibility, and how "agentic" the tools really are

A good literature search can be repeated by someone else. AI search often cannot. Aaron Tay tested the AI assistants inside Scopus and Web of Science in 2025 and found that the generated query changed between identical runs, in one tool about every other run. Plain Boolean search remains the most reproducible approach.

He also examined the "deep research" features of academic tools and found they are mostly fixed workflows, in which the model makes decisions only at set points, rather than fully autonomous agents. That is not a flaw in itself, but marketing can suggest more flexible reasoning than the tools deliver.

How to use AI tools well in a literature review

  • Use AI search to explore a topic, find key papers and sharpen your question before you design the formal search.
  • For a systematic or scoping review, run a documented database search as your main search and report AI tools as a supplementary source.
  • Save the exact question, tool, date and results of every AI search, because you may not get the same results again.
  • Read the papers, not only the summaries. AI summaries of real papers can still be wrong.
  • Verify every reference before you cite it.
  • Talk to your librarian. Many university libraries now publish guidance on AI search tools.

Where Kahubi fits

Kahubi searches OpenAlex and Semantic Scholar from a conversation, imports the papers you choose into your library with open-access PDFs, and lets you ask questions across the full texts with citations. For a systematic review, the dedicated flow keeps search strings, screening decisions and PRISMA numbers together. You can also import results from Scopus, Web of Science or PubMed as RIS or BibTeX, so the AI works alongside your documented database search rather than instead of it.

Sources

  1. Lau O, Golder S (2025). Comparison of Elicit AI and traditional literature searching in evidence syntheses using four case studies. Cochrane Evidence Synthesis and Methods 3(6):e70050.
  2. Bernard N, et al. (2025). Using artificial intelligence for systematic review: the example of Elicit. BMC Medical Research Methodology 25:75.
  3. Tomczyk P, Brüggemann P, Mergner N, Petrescu M (2024). Are AI tools better than traditional tools in literature searching? Evidence from e-commerce research. Journal of Librarianship and Information Science 58(1):135-145.
  4. Parisi V, Sutton A (2024). The role of ChatGPT in developing systematic literature searches: an evidence summary. Journal of EAHIL 20(2):30-34.
  5. Tay A (2025). Deep research, shallow agency: what academic deep research can and can’t do. Musings about librarianship.
  6. Tay A (2025). The reproducibility and interpretability of academic AI search engines. Musings about librarianship.
  7. Tay A (2025). A 2025 deep dive of Consensus. Musings about librarianship.
  8. University of British Columbia Library. AI search tools used in literature reviews and comprehensive searching (evidence list).

Last updated 2026-10-09.

Frequently asked questions

What is the best AI tool for a literature review?
It depends on the stage. Academic AI search tools such as Elicit, Consensus and Undermind are good for discovery, citation maps such as ResearchRabbit and Litmaps for exploring connections, and research workspaces such as Kahubi for reading, screening and writing in one place. For a systematic review, combine any of them with a documented database search.
Can AI replace database searching for a systematic review?
Not yet. A 2025 evaluation in Cochrane Evidence Synthesis and Methods found that Elicit identified about 40% of the studies included in four reviews, against about 95% for the original searches. AI tools work best as a supplementary search.
Are AI literature search results reproducible?
Often not fully. Published studies and librarian tests have found that identical AI searches can return different results. Record the tool, date, question and results of each search.
Do AI research tools read the full text of papers?
Usually only for open-access papers or publisher partners. For paywalled papers most tools see the title and abstract. Uploading the PDFs you have access to into a tool that reads your own library gives better answers.

Related

Search, read and write in one place

Find papers in OpenAlex and Semantic Scholar, ask questions across the full texts and cite them in your draft. Free plan included.