Cohen's kappa calculator (inter-rater reliability)
Paste the codes assigned by two raters — one item per line, any labels — and get κ with a confidence interval, raw agreement, and a copy-ready reliability sentence for your methods section.
Or skip the manual work
In the app, you chat with your data — the AI runs the analysis and writes it up.
Free accounts include AI chat, data analysis, and every tool on this site.
When to use this
Use this whenever two people independently categorize the same items and you need to show the coding is reliable: qualitative coding of interview excerpts against a codebook, include/exclude decisions in systematic review screening, diagnostic classifications, or content analysis of documents. Raw percent agreement overstates reliability because two raters using the same categories will agree part of the time by pure chance — kappa reports agreement beyond that chance level, which is what reviewers ask for.
Key assumptions
- Exactly two raters coding the same items independently, with a fixed set of categories.
- Items are in the same order in both boxes — line 7 of Rater 1 and line 7 of Rater 2 are the same item.
- Labels are matched exactly (case-sensitive): "Include" and "include" count as different categories.
Common mistakes
- Reporting only percent agreement — 80% agreement can be κ = .30 when one category dominates.
- Ignoring the kappa paradox: with very skewed category use, κ can be low despite high agreement. Report both.
- Using Cohen’s kappa for three or more raters — that calls for Fleiss’ kappa or Krippendorff’s alpha.
Frequently asked questions
What is an acceptable kappa?
My agreement is high but kappa is low — why?
Does kappa handle ordered categories or partial credit?
Related free tools
Stop copying numbers between tools
Inside Kahubi, the AI agent runs this analysis directly on your uploaded dataset — then writes the results section in your own writing style, with the statistics reported correctly.