Effect size benchmarks — what counts as small, medium, large?
The conventional cutoffs for every common effect size and reliability statistic, on one page — with conversions between metrics and the caveats the one-line answers leave out.
| Statistic | Used for | Small | Medium | Large |
|---|---|---|---|---|
| Cohen's d / Hedges' g | Difference between two means | 0.20 | 0.50 | 0.80 |
| Correlation r | Linear association | .10 | .30 | .50 |
| η² (eta squared) | ANOVA variance explained | .01 | .06 | .14 |
| Cohen's f | ANOVA (power analysis) | 0.10 | 0.25 | 0.40 |
| Odds ratio | Binary outcomes (2×2) | 1.68 | 3.47 | 6.71 |
| Cramér's V (df* = 1) | Contingency tables | .10 | .30 | .50 |
| R² | Regression variance explained | .02 | .13 | .26 |
Cohen offered these reluctantly, "for use when no better basis is available." A d of 0.2 can be hugely consequential (aspirin and heart attacks) and a d of 0.8 trivial — context beats convention.
| Statistic | Range | Conventional reading |
|---|---|---|
| Cohen's kappa (κ) | ≤ .20 / .21–.40 / .41–.60 / .61–.80 / .81–1.0 | Slight / fair / moderate / substantial / almost perfect (Landis & Koch, 1977) |
| Cronbach's alpha (α) | ≥ .70 / ≥ .80 / ≥ .90 | Acceptable for research / good / needed for clinical decisions (but > .95 suggests redundant items) |
| ICC | < .50 / .50–.75 / .75–.90 / > .90 | Poor / moderate / good / excellent (Koo & Li, 2016) |
| Krippendorff's alpha | ≥ .667 / ≥ .80 | Tentative conclusions / firm conclusions (Krippendorff, 2004) |
| KR-20 (dichotomous items) | same as Cronbach’s α | Same benchmarks as alpha |
| From → to | Formula | Example |
|---|---|---|
| d → r (equal groups) | r = d / √(d² + 4) | d = 0.50 → r = .24 |
| r → d | d = 2r / √(1 − r²) | r = .30 → d = 0.63 |
| d → odds ratio | OR ≈ exp(1.81 × d) | d = 0.50 → OR ≈ 2.5 |
| η² → f | f = √(η² / (1 − η²)) | η² = .06 → f = 0.25 |
| r → R² | R² = r² | r = .50 → R² = .25 |
| d → % non-overlap (U3) | U3 = Φ(d) | d = 0.50 → 69% of group 1 above group 2 mean |
Conversions assume roughly normal outcomes and (for d ↔ r) similar group sizes — treat them as good approximations for synthesis, not identities.
Using benchmarks responsibly
Benchmarks answer "is this effect big for a generic study?", but reviewers increasingly ask the better question: "is it big for THIS literature?" Field-specific norms differ by a factor of two or more — in social psychology the median published effect is around r = .21, in some areas of education d = 0.20 from a cheap intervention is remarkable. When you can, anchor interpretation to effects your readers know.
Also distinguish statistical from practical significance in both directions: with n = 10,000 a meaningless d = 0.05 is "significant", while an important d = 0.6 in a pilot of 15 is not. Report the effect size with its confidence interval and let the magnitude — not the star count — carry the interpretation.
Frequently asked questions
Is a Cronbach's alpha of .65 acceptable?
What is a good effect size for a power analysis?
Why does partial eta squared differ from eta squared?
Related free tools
Have a question the table can't answer?
The AI inside Kahubi works with your actual data and library — it runs analyses, checks assumptions, and writes the results up in your style.
Free accounts include AI chat, literature search, and every tool on this site.
Stop copying numbers between tools
Inside Kahubi, the AI agent runs this analysis directly on your uploaded dataset — then writes the results section in your own writing style, with the statistics reported correctly.