Part of the Maths & Statistics suite · 235 calculators

Cohen's Kappa Calculator

How much two raters agree beyond chance — the statistic that shows why raw agreement percentages are almost always misleading.

κ = (observed agreement − expected agreement) ÷ (1 − expected agreement).

Results update as you type
Results
Cohen's kappa
0.4
Observed agreement
Agreement expected by chance
Strength of agreement
Standard error
95% CI — lower
95% CI — upper
Cases rated
Reviewed September 2026. Pure mathematics: the result does not depend on where you are. Terminology follows the national curriculum (maths, brackets, decimal point).
No account required · Google Analytics off unless allowedCalculator arithmetic runs in your browserResults update as you type
All calculations run 100% in your browser. The calculator code does not submit your figures to GlobalCalc to obtain a result.
About cohen's kappa

How the cohen's kappa calculator works

κ = (observed agreement − expected agreement) ÷ (1 − expected agreement). It rescales agreement so that chance is 0 and perfection is 1.

The correction matters enormously when one category dominates. Two raters who both label 95% of cases "negative" will agree about 90% of the time by luck alone, so 92% raw agreement is barely better than guessing — and kappa says so.

Formula: κ = (pₒ − pₑ) / (1 − pₑ)

Worked examples

InputsCohen's kappaNote
Fifty cases, moderate agreement0.4κ = 0.40 despite 70% raw agreement
Perfect agreement1κ = 1
A rare category0.36842194% agreement, κ only 0.36

Frequently asked questions

Why not just report percentage agreement?

Because chance agreement is often very high. Two raters labelling a rare condition will agree most of the time whatever they do.

What is a good kappa?

Landis and Koch's labels — slight, fair, moderate, substantial, almost perfect at 0.2, 0.4, 0.6 and 0.8 — are conventional but arbitrary. Most fields want at least 0.6.

Can kappa be negative?

Yes, when agreement is worse than chance. It is rare and usually means the raters understood the categories differently.

What is the kappa paradox?

With a very unbalanced category split, kappa can be low despite near-total agreement, because chance agreement is nearly total too. Report both numbers.

What about more than two raters?

Fleiss' kappa generalises it. For ordered categories, weighted kappa credits near-misses.

Where these figures come from

Last checked: September 2026. Formulas are fixed by mathematics and do not change with tax years or regulations.