Pick the statistic that matches your data: Cohen's kappa for two raters assigning categories, weighted kappa if the categories are ordered, Fleiss' kappa or Krippendorff's alpha for more raters, and an intraclass correlation (ICC) for numeric scores. Whatever you pick, don't report raw percent agreement or a Pearson correlation on their own. Both can look excellent when the raters barely agree at all.
Say two raters label 100 customer comments as "toxic" or "not toxic," and each calls about 90% of them not toxic. Even if they labeled at random, they would agree about 82% of the time, just because both keep landing on the common answer. So when they report 88% agreement, it sounds great, but it is only a little better than a coin weighted toward "no." Kappa corrects for that: here it comes out around 0.33, which is a far more honest summary.
Correlation fails in a different way. If one rater scores every essay exactly two points higher than the other, their correlation is a perfect 1.0, yet they never give the same score once. Correlation measures whether raters rank things alike, not whether they agree. If the scores feed a cutoff, such as pass/fail or flag/don't flag, that difference matters.
If your LLM eval relies on human labels as ground truth, inter-rater reliability sets a ceiling on everything after it. When two trained annotators only reach a kappa of 0.5 with each other, a model that "disagrees with the gold label" 25% of the time may be doing as well as a person can. The useful comparison is model-versus-human agreement set against human-versus-human agreement on the same items, measured with the same statistic. The same logic applies to LLM-as-judge setups: check the judge against humans before you trust it to replace them.
If you're designing a coding scheme, a rating rubric, or a human-labeling protocol and want the reliability numbers to hold up to reviewers, that's part of the measurement and design work we do.
One email whenever we publish something new. No spam, unsubscribe anytime.
{{ subscribeErrorMsg }}