All free tools
Free · no signup

Which statistical test should I use?

Answer a few questions about your design — what you're comparing, how many groups, what kind of data — and get the standard test for your situation, linked to a free calculator where we have one.

Let the AI do this — sign up freeUpload your data, get the write-up back.

What do you want to do?

When to use this

Use this at two moments: while designing a study (the planned analysis shapes the design — a repeated-measures version of your experiment may need half the participants), and when you are sitting on collected data unsure what the conventional analysis is. The tree covers the workhorse decisions: group comparisons (t-tests, ANOVA and their rank-based cousins), associations (Pearson, Spearman, chi-square), prediction (linear and logistic regression), and reliability (kappa, alpha, ICC) — the tests behind the vast majority of published quantitative results.

How the chooser decides

  • The measurement level of your outcome (continuous, ordinal, categorical) — the biggest fork in the road.
  • Design: how many groups, and whether measurements are independent or paired/repeated.
  • Distribution: rank-based alternatives are offered where normality is doubtful.
  • It recommends the conventional default — specific fields may have their own standards.

Common mistakes

  • Choosing the test after peeking at which one gives p < .05.
  • Treating a single Likert item as continuous data.
  • Running independent-samples tests on paired data (or vice versa) — the pairing changes everything.
  • Stopping at the p-value: whatever test you run, report the effect size and confidence interval too.

Frequently asked questions

How do I know if my data are "roughly normal"?
Plot a histogram of each group (or of the paired differences) and look for gross skew or extreme outliers — eyeballing beats formal tests here. Shapiro–Wilk and friends are overpowered at large n (flagging trivial deviations exactly when they matter least) and underpowered at small n. With n ≳ 30 per group, the t-family is robust to moderate non-normality anyway.
What if my outcome is a Likert item?
A single Likert item (1–5) is ordinal: prefer rank-based tests (Mann–Whitney, Spearman). A SCORE formed by summing or averaging several items is conventionally treated as continuous, so t-tests and ANOVAs are standard — this distinction resolves most Likert arguments.
The chooser recommends a test with no calculator — now what?
Some recommendations (repeated-measures ANOVA, Kruskal–Wallis, regression, ICC) go beyond what a paste-in web calculator handles well. Inside Kahubi you can upload your CSV and the AI runs these directly on your dataset, with assumption checks and a written interpretation.
Can I just run several tests and pick the significant one?
No — that is p-hacking. Each additional test inflates the false-positive rate, and choosing after seeing the results invalidates the p-values. Choose the test from the design (ideally preregistered), run it once, and report it whatever it says.

Related free tools

Not sure the generic answer fits your study?

Inside Kahubi, the AI looks at your actual dataset — variables, distributions, design — recommends the analysis, runs it, and writes up the results.

Kahubi AI
Hi! Ask me anything you’d normally google mid-study — study design, statistics, literature, writing.

Free accounts include AI chat, literature search, and every tool on this site.

Let the AI run the test it recommends

Upload your CSV in Kahubi and the agent picks the right analysis, checks assumptions, runs it, and drafts the results section — statistics reported correctly, in your style.