All free tools
Free · no signup

Cohen's kappa calculator (inter-rater reliability)

Paste the codes assigned by two raters — one item per line, any labels — and get κ with a confidence interval, raw agreement, and a copy-ready reliability sentence for your methods section.

Let the AI do this — sign up freeUpload your data, get the write-up back.

Or skip the manual work

In the app, you chat with your data — the AI runs the analysis and writes it up.

Kahubi AILive in the app
Typing numbers in by hand works — but inside Kahubi you just upload your dataset and ask. I check the assumptions, run the right test, and hand you an APA-ready results section. Try me.

Free accounts include AI chat, data analysis, and every tool on this site.

Your numbers

One item per line. Labels can be anything: category names, numbers, include/exclude.

Same items in the same order as Rater 1.

Runs in your browser — data never leaves this page.

When to use this

Use this whenever two people independently categorize the same items and you need to show the coding is reliable: qualitative coding of interview excerpts against a codebook, include/exclude decisions in systematic review screening, diagnostic classifications, or content analysis of documents. Raw percent agreement overstates reliability because two raters using the same categories will agree part of the time by pure chance — kappa reports agreement beyond that chance level, which is what reviewers ask for.

Key assumptions

  • Exactly two raters coding the same items independently, with a fixed set of categories.
  • Items are in the same order in both boxes — line 7 of Rater 1 and line 7 of Rater 2 are the same item.
  • Labels are matched exactly (case-sensitive): "Include" and "include" count as different categories.

Common mistakes

  • Reporting only percent agreement — 80% agreement can be κ = .30 when one category dominates.
  • Ignoring the kappa paradox: with very skewed category use, κ can be low despite high agreement. Report both.
  • Using Cohen’s kappa for three or more raters — that calls for Fleiss’ kappa or Krippendorff’s alpha.

Frequently asked questions

What is an acceptable kappa?
The Landis & Koch conventions: ≤ .20 slight, .21–.40 fair, .41–.60 moderate, .61–.80 substantial, .81–1.00 almost perfect. Most journals treat κ ≥ .61 as acceptable for content/qualitative coding and κ ≥ .80 as the bar in clinical measurement. Report the CI, not just the point estimate.
My agreement is high but kappa is low — why?
The kappa paradox. If 90% of items fall in one category, chance agreement is already very high, leaving kappa little room. It is real behavior of the statistic, not an error. Report percent agreement alongside κ, and consider prevalence-adjusted kappa (PABAK) as a sensitivity check.
Does kappa handle ordered categories or partial credit?
This calculator computes unweighted kappa: near-misses count as full disagreement. For ordinal codes (e.g. severity ratings 1–5) use weighted kappa, which credits close ratings. For more than two raters use Fleiss’ kappa or Krippendorff’s alpha.

Related free tools

Stop copying numbers between tools

Inside Kahubi, the AI agent runs this analysis directly on your uploaded dataset — then writes the results section in your own writing style, with the statistics reported correctly.