← Back to blog
Study DesignAI Evaluation

How Do You Measure Inter-Rater Reliability?

3 min read

Pick the statistic that matches your data: Cohen's kappa for two raters assigning categories, weighted kappa if the categories are ordered, Fleiss' kappa or Krippendorff's alpha for more raters, and an intraclass correlation (ICC) for numeric scores. Whatever you pick, don't report raw percent agreement or a Pearson correlation on their own. Both can look excellent when the raters barely agree at all.

Why percent agreement and correlation mislead

Say two raters label 100 customer comments as "toxic" or "not toxic," and each calls about 90% of them not toxic. Even if they labeled at random, they would agree about 82% of the time, just because both keep landing on the common answer. So when they report 88% agreement, it sounds great, but it is only a little better than a coin weighted toward "no." Kappa corrects for that: here it comes out around 0.33, which is a far more honest summary.

Correlation fails in a different way. If one rater scores every essay exactly two points higher than the other, their correlation is a perfect 1.0, yet they never give the same score once. Correlation measures whether raters rank things alike, not whether they agree. If the scores feed a cutoff, such as pass/fail or flag/don't flag, that difference matters.

Match the statistic to the ratings

  • Two raters, unordered categories: Cohen's kappa. The standard choice for things like coding interview transcripts or labeling model outputs as correct or incorrect.
  • Ordered categories (a 1-5 rubric, mild/moderate/severe): weighted kappa, so a 4-versus-5 disagreement counts less than a 1-versus-5.
  • Three or more raters, or not everyone rated everything: Fleiss' kappa when every item has the same number of ratings, and Krippendorff's alpha when coverage is uneven. Alpha also handles nominal, ordinal, and interval data, which makes it a good default for messy real-world labeling.
  • Continuous scores: an ICC, but say which one. Pick absolute agreement if the actual values matter and consistency if only the ranking does. Use single-measure if one rater will score each case later, and average-measure if you will always average several raters. These choices can move the number a lot, and a reviewer who sees "ICC = 0.85" with no type given will ask.

Three mistakes that undermine the number

  • Reporting kappa without the base rate. When one category is rare, kappa can be low even though raters agree on almost every case. Report how common each category is and how often raters agreed within each one, so readers can see where the disagreement actually lies.
  • Skipping the confidence interval. With 50 double-rated items, a kappa of 0.6 could plausibly be anywhere from about 0.4 to 0.8. That range runs from "shaky" to "solid." Rate enough items to narrow it, and report it.
  • Measuring reliability during training. Raters who discuss cases as they go will agree beautifully. Calibrate first, then have them rate a fresh sample independently, and compute reliability on that.

Why AI teams should care

If your LLM eval relies on human labels as ground truth, inter-rater reliability sets a ceiling on everything after it. When two trained annotators only reach a kappa of 0.5 with each other, a model that "disagrees with the gold label" 25% of the time may be doing as well as a person can. The useful comparison is model-versus-human agreement set against human-versus-human agreement on the same items, measured with the same statistic. The same logic applies to LLM-as-judge setups: check the judge against humans before you trust it to replace them.

If you're designing a coding scheme, a rating rubric, or a human-labeling protocol and want the reliability numbers to hold up to reviewers, that's part of the measurement and design work we do.

Related reading

Get new posts by email

One email whenever we publish something new. No spam, unsubscribe anytime.

Check your inbox — click the confirmation link to finish subscribing.

{{ subscribeErrorMsg }}