All free tools
Free · no signup

Effect size benchmarks — what counts as small, medium, large?

The conventional cutoffs for every common effect size and reliability statistic, on one page — with conversions between metrics and the caveats the one-line answers leave out.

Let the AI do this — sign up freeUpload your data, get the write-up back.
Group difference and association measures (Cohen, 1988)
StatisticUsed forSmallMediumLarge
Cohen's d / Hedges' gDifference between two means0.200.500.80
Correlation rLinear association.10.30.50
η² (eta squared)ANOVA variance explained.01.06.14
Cohen's fANOVA (power analysis)0.100.250.40
Odds ratioBinary outcomes (2×2)1.683.476.71
Cramér's V (df* = 1)Contingency tables.10.30.50
Regression variance explained.02.13.26

Cohen offered these reluctantly, "for use when no better basis is available." A d of 0.2 can be hugely consequential (aspirin and heart attacks) and a d of 0.8 trivial — context beats convention.

Agreement and reliability
StatisticRangeConventional reading
Cohen's kappa (κ)≤ .20 / .21–.40 / .41–.60 / .61–.80 / .81–1.0Slight / fair / moderate / substantial / almost perfect (Landis & Koch, 1977)
Cronbach's alpha (α)≥ .70 / ≥ .80 / ≥ .90Acceptable for research / good / needed for clinical decisions (but > .95 suggests redundant items)
ICC< .50 / .50–.75 / .75–.90 / > .90Poor / moderate / good / excellent (Koo & Li, 2016)
Krippendorff's alpha≥ .667 / ≥ .80Tentative conclusions / firm conclusions (Krippendorff, 2004)
KR-20 (dichotomous items)same as Cronbach’s αSame benchmarks as alpha
Converting between effect size metrics
From → toFormulaExample
d → r (equal groups)r = d / √(d² + 4)d = 0.50 → r = .24
r → dd = 2r / √(1 − r²)r = .30 → d = 0.63
d → odds ratioOR ≈ exp(1.81 × d)d = 0.50 → OR ≈ 2.5
η² → ff = √(η² / (1 − η²))η² = .06 → f = 0.25
r → R²R² = r²r = .50 → R² = .25
d → % non-overlap (U3)U3 = Φ(d)d = 0.50 → 69% of group 1 above group 2 mean

Conversions assume roughly normal outcomes and (for d ↔ r) similar group sizes — treat them as good approximations for synthesis, not identities.

Using benchmarks responsibly

Benchmarks answer "is this effect big for a generic study?", but reviewers increasingly ask the better question: "is it big for THIS literature?" Field-specific norms differ by a factor of two or more — in social psychology the median published effect is around r = .21, in some areas of education d = 0.20 from a cheap intervention is remarkable. When you can, anchor interpretation to effects your readers know.

Also distinguish statistical from practical significance in both directions: with n = 10,000 a meaningless d = 0.05 is "significant", while an important d = 0.6 in a pilot of 15 is not. Report the effect size with its confidence interval and let the magnitude — not the star count — carry the interpretation.

Frequently asked questions

Is a Cronbach's alpha of .65 acceptable?
Sometimes. The .70 cutoff (Nunnally, 1978) was proposed for early-stage research instruments. Short scales (3–4 items) mechanically produce lower alphas, so .65 on a 3-item scale may be fine — report mean inter-item correlation (.15–.50 is the target) alongside. For decisions about individuals (clinical, hiring), demand .90+.
What is a good effect size for a power analysis?
Not the benchmark — the smallest effect that would matter in your context (the SESOI), or a meta-analytic estimate from your literature shrunk for publication bias. Powering for "medium because Cohen said .50" is the most common design error in the social sciences.
Why does partial eta squared differ from eta squared?
In multi-factor ANOVAs, η²p removes the other effects from the denominator, so it is larger than η² and NOT comparable to the .01/.06/.14 benchmarks (which were set for one-way designs). Software defaults to η²p — label which one you report.

Related free tools

Have a question the table can't answer?

The AI inside Kahubi works with your actual data and library — it runs analyses, checks assumptions, and writes the results up in your style.

Kahubi AI
Hi! Ask me anything you’d normally google mid-study — study design, statistics, literature, writing.

Free accounts include AI chat, literature search, and every tool on this site.

Stop copying numbers between tools

Inside Kahubi, the AI agent runs this analysis directly on your uploaded dataset — then writes the results section in your own writing style, with the statistics reported correctly.