Suppose several raters each score the same set of things: essays, patients, video clips. You want to know whether the scores can be trusted. If two raters watch the same clip, will they land on the same number? If you swapped in a different rater, would the result hold up? That is a question about interrater reliability, and an intraclass correlation coefficient (ICC) is the number that answers it.
An ICC runs from 0 to 1 and reports the share of the variation in scores that reflects real differences between the things being rated, rather than disagreement or noise between raters. Near 1, almost all the spread in scores is genuine subject-to-subject difference and the raters barely disagree, and the ratings are highly reliable. Near 0, the raters are effectively adding noise, and a score tells you more about who happened to rate it than about the subject.
intraclass estimates that number by fitting a
mixed model, rather than from the classical ANOVA
mean-squares formulas older tools use. The model separates subject
variation from rater variation as variance components. The two
approaches agree on clean, balanced data, but the model-based one also
handles missing ratings, multiple designs, and honest confidence
intervals. See Engines for the
machinery. This article walks the whole pipeline on a small example:
fit → estimate → interpret. (New to any of the terms as
they come up? The Glossary defines
each one in a sentence.)
We use the classic Shrout & Fleiss (1979) example, shipped with
the package as ratings: 6 subjects each rated by the same 4
raters, one rating per cell. intraclass wants
long format: one rating per row, with columns for the
subject, the rater, and the score.
Call icc() with the data and the three columns
(unquoted). With the defaults, a single two-way random
fit reports every defined formulation: absolute agreement and
consistency, each for a single rater and for the average of your raters.
Those four are ICC(A,1), ICC(A,k),
ICC(C,1), and ICC(C,k). The next sections
unpack the labels. They are grouped by error definition in the printout.
We set a seed so the confidence interval is
reproducible.
fit <- icc(ratings, score, subject, rater, seed = 2024)
fit
#> ── Intraclass correlation: two-way random, absolute agreement & consistency ────
#> Subjects: 6 | Raters: 4 (random) | Observations: 24 of 24 cells (complete)
#> Engine: glmmTMB (REML) | CI: 95% montecarlo (10000 draws)
#>
#> index estimate 95% CI
#> Absolute agreement
#> ICC(A,1) 0.290 [0.053, 0.714]
#> ICC(A,k) 0.620 [0.183, 0.909]
#> Consistency
#> ICC(C,1) 0.715 [0.334, 0.924]
#> ICC(C,k) 0.909 [0.667, 0.980]
#>
#> Variance components: subject 2.556, rater 5.244, residual 1.019
#> Shrout & Fleiss equivalent: ICC(A,1) = ICC(2,1), ICC(A,k) = ICC(2,k)tidy() returns one row per coefficient, with the
estimate and confidence interval. glance() returns a
one-row model summary, including the variance components.
tidy(fit)
#> # A tibble: 4 × 11
#> term occasions type level sf_index estimate std.error conf.low conf.high
#> <chr> <dbl> <chr> <chr> <chr> <dbl> <dbl> <dbl> <dbl>
#> 1 ICC(A,1) NA agree… <NA> ICC(2,1) 0.290 0.180 0.0531 0.714
#> 2 ICC(A,k) NA agree… <NA> ICC(2,k) 0.620 0.201 0.183 0.909
#> 3 ICC(C,1) NA consi… <NA> <NA> 0.715 0.155 0.334 0.924
#> 4 ICC(C,k) NA consi… <NA> <NA> 0.909 0.0811 0.667 0.980
#> # ℹ 2 more variables: conf.level <dbl>, method <chr>
glance(fit)
#> # A tibble: 1 × 24
#> n_subjects n_raters n_clusters n_obs n_cells balanced raters replicates
#> <int> <int> <int> <int> <int> <lgl> <chr> <lgl>
#> 1 6 4 NA 24 24 TRUE random FALSE
#> # ℹ 16 more variables: multilevel <lgl>, ml_design <chr>, k_eff <dbl>,
#> # k_c_eff <dbl>, var_cluster <dbl>, var_subject <dbl>, var_rater <dbl>,
#> # var_cluster_rater <dbl>, var_subject_rater <dbl>, var_residual <dbl>,
#> # n_o <int>, engine <chr>, ci_method <chr>, conf.level <dbl>, rhat <dbl>,
#> # ess_bulk <dbl>summary() is the other reading of the same fit. It
reprints the report above, then appends a short interpretive
note for each error definition present, plus one about what a
single rating per cell cannot separate. Nothing is recomputed: the notes
are read off the design, so they track whatever you fitted:
summary(fit)
#> ── Intraclass correlation: two-way random, absolute agreement & consistency ────
#> Subjects: 6 | Raters: 4 (random) | Observations: 24 of 24 cells (complete)
#> Engine: glmmTMB (REML) | CI: 95% montecarlo (10000 draws)
#>
#> index estimate 95% CI
#> Absolute agreement
#> ICC(A,1) 0.290 [0.053, 0.714]
#> ICC(A,k) 0.620 [0.183, 0.909]
#> Consistency
#> ICC(C,1) 0.715 [0.334, 0.924]
#> ICC(C,k) 0.909 [0.667, 0.980]
#>
#> Variance components: subject 2.556, rater 5.244, residual 1.019
#> Shrout & Fleiss equivalent: ICC(A,1) = ICC(2,1), ICC(A,k) = ICC(2,k)
#>
#> Absolute agreement counts the rater main effect (systematic differences in rater level) as error.
#> Consistency ignores the rater main effect (systematic differences in rater level); only relative standing counts.
#> A single rating per cell confounds the subject-by-rater interaction with
#> residual error.Single vs. the average of several raters.
ICC(A,1) = 0.29 is the reliability of a single
rater: how much you can trust one person’s score. ICC(A,k)
= 0.62 is the reliability of the mean of all 4 raters.
Averaging cancels out independent rater noise, so the mean is always
more reliable than one rater alone (this is the
Spearman–Brown relationship: more raters, higher
reliability, with diminishing returns). Report ICC(A,k)
when the averaged score is what you will actually use, and
ICC(A,1) when downstream users see one rater’s
judgment.
Absolute agreement vs. rank order. These are absolute-agreement coefficients: if one rater scores consistently higher than another, that systematic gap counts as error. Here the raters differ sharply in average level, which is why the agreement ICC is low. That is a signal that the rating procedure has a level problem worth fixing. (If you only care that raters rank subjects the same way, ask for consistency instead, described below.)
There is no universal cutoff, but two widely cited rules of thumb give a rough vocabulary. Koo & Li (2016) propose:
| ICC | Label |
|---|---|
| < 0.50 | poor |
| 0.50–0.75 | moderate |
| 0.75–0.90 | good |
| > 0.90 | excellent |
An older scheme, Cicchetti (1994), draws its lines a little differently (< 0.40 poor, 0.40–0.59 fair, 0.60–0.74 good, 0.75–1.00 excellent). Treat these as conventions, not laws. The bar that matters depends on the stakes of your decision, and on which ICC you are reading. An agreement coefficient and a consistency coefficient on the same data are not comparable to the same cutoff.
Most important, and this is Koo & Li’s own recommendation:
judge the confidence interval, not just the point
estimate. Here the four-rater mean ICC(A,k) = 0.62
reads as “moderate,” but its 95% interval runs from 0.18 to 0.91, from
“poor” all the way to “excellent.” With only six subjects, the data
simply cannot pin the reliability down to one band. That is why
icc() never returns a point estimate without an interval,
and never labels a result for you: the honest summary is the whole
interval.
The interval icc() reports is a Monte-Carlo
interval, and it is boundary-aware.
In plain terms: the default interval does not rely on a textbook formula
that misbehaves when a variance is near zero. That is a common
situation: raters who barely differ push the rater variance to its zero
boundary. Instead icc() simulates many plausible parameter
sets from the fitted model, and reads the interval off the resulting
spread of ICCs. This keeps the interval well-behaved right at that
boundary. The Interval
methods article covers how it works and the alternatives (a
parametric bootstrap, and Bayesian credible intervals).
Suppose a constant per-rater offset is acceptable: you care only that
raters rank subjects the same way, not that they land on the
same number. Then ask for consistency with
type = "consistency". It drops the rater main effect from
the error, so it is never smaller than the agreement coefficient:
icc(ratings, score, subject, rater, type = "consistency", seed = 2024)
#> ── Intraclass correlation: two-way random, consistency ─────────────────────────
#> Subjects: 6 | Raters: 4 (random) | Observations: 24 of 24 cells (complete)
#> Engine: glmmTMB (REML) | CI: 95% montecarlo (10000 draws)
#>
#> index estimate 95% CI
#> ICC(C,1) 0.715 [0.334, 0.924]
#> ICC(C,k) 0.909 [0.667, 0.980]
#>
#> Variance components: subject 2.556, rater 5.244, residual 1.019The gap between the two (here ICC(C,1) = 0.71
vs. ICC(A,1) = 0.29) is a direct read-out of how much
systematic rater-level difference is present.
By default raters are treated as random: a sample
you want to generalize beyond, so that your reliability claim covers
raters you did not use. If your raters are the entire population of
interest you can pass raters = "fixed", the classic
ICC(3,1). But icc() will warn: random is the
recommended default for interrater reliability, and on balanced data the
number is identical anyway.
Several choices define which ICC is correct for your study.
Are the raters crossed, so that the same set judges everyone, or
interchangeable (model = "oneway")? Absolute agreement or
consistency? A single rater or the average? Fixed or random raters? A
complete or an incomplete design? The Choosing an ICC article
(vignette("choosing-an-icc")) walks that decision step by
step.
Cicchetti, D. V. (1994). Guidelines, criteria, and rules of thumb for evaluating normed and standardized assessment instruments in psychology. Psychological Assessment, 6(4), 284–290.
Koo, T. K., & Li, M. Y. (2016). A guideline of selecting and reporting intraclass correlation coefficients for reliability research. Journal of Chiropractic Medicine, 15(2), 155–163.
Shrout, P. E., & Fleiss, J. L. (1979). Intraclass correlations: Uses in assessing rater reliability. Psychological Bulletin, 86(2), 420–428.