---
title: "Expert-Panel Content Validation: Relevance, Essentiality, and Congruence"
output: rmarkdown::html_vignette
vignette: >
  %\VignetteIndexEntry{Expert-Panel Content Validation: Relevance, Essentiality, and Congruence}
  %\VignetteEngine{knitr::rmarkdown}
  %\VignetteEncoding{UTF-8}
---

```{r setup, include=FALSE}
knitr::opts_chunk$set(collapse = TRUE, comment = "#>")
library(contentvalidR)
```

## Why expert-panel methods need separate modes

Expert-panel content validation is not one statistical task. A relevance rating,
an essential/not-essential judgment, and an item-objective congruence judgment
ask experts different questions and therefore support different indices.
`expert_validity()` uses an explicit `mode` so those designs are not treated as
interchangeable.

The three modes are:

- **relevance**: Aiken's V plus CVI and modified kappa, with a panel-level
  agreement coefficient;
- **essentiality**: Lawshe's CVR with exact binomial inference; and
- **congruence**: Rovinelli-Hambleton IOC.

All three are quantitative complements to qualitative expert comments, construct
coverage, comprehensibility review, and other parts of the content-validity
argument.

## Relevance: Aiken V, score intervals, and CVI

Suppose six experts rate item relevance from 1 (not relevant) to 4 (highly
relevant):

```{r relevance}
R <- matrix(
  c(4,4,4,4,4,4,
    4,4,4,3,4,4,
    4,3,4,4,3,4,
    3,3,4,3,2,3),
  nrow = 6,
  dimnames = list(NULL, paste0("Item", 1:4))
)

fit <- expert_validity(R, mode = "relevance", lo = 1, hi = 4, seed = 1)
fit
summary(fit)
```

Aiken's V rescales the bounded expert ratings to the 0-1 interval. The default
confidence interval is the score interval proposed by Penfield and Giacobbi
(2004), rather than a simulation-dependent bootstrap interval. Bootstrap
intervals remain available through the low-level function:

```{r aiken-bootstrap}
aikens_v(R, lo = 1, hi = 4, ci = "bootstrap", B = 200, seed = 1)
```

For CVI, the workflow dichotomizes ratings at `relevance_cut`. On a 1-4 scale
the default is 3, so ratings of 3 or 4 count as relevant. Declare a different
threshold if the study protocol used one.

The workflow reports common panel-size I-CVI guidelines (1.00 for panels of
3-5 experts and .78 for 6 or more) as **review aids**. They are not presented as
universal proof that an item is or is not content valid. Modified kappa provides
a chance-corrected complement to I-CVI.

At the scale level, S-CVI/Ave and S-CVI/UA are reported together. S-CVI/Ave is
generally less brittle than universal agreement, but both should be interpreted
alongside the distribution of item-level evidence.

## Panel-level agreement

I-CVI and modified kappa describe one item at a time. Relevance mode also
reports how consistently the panel rated the whole item set, as one coefficient
with a bootstrap interval:

```{r agreement}
fit$scale_summary[, c("agreement", "agreement_low", "agreement_high")]
fit$details$agreement
```

The default coefficient is Krippendorff's alpha. It accepts any number of
experts and missing ratings, and Zapf et al. (2016) recommend it when ratings
are ordinal or incomplete, which describes most expert panels. It is a general
reliability coefficient (Hayes & Krippendorff, 2007) rather than one developed
for content validity; no publication applying it specifically to
content-validity panels was found.

Choose the measurement level that matches the rating scale. Relevance ratings
are treated as ordinal by default. `agreement_level = "interval"` treats the
distances between scale points as equal, and `"nominal"` treats every
disagreement as equally serious:

```{r agreement-level}
expert_validity(R, mode = "relevance", lo = 1, hi = 4,
                agreement_level = "interval", agreement_B = 0)$scale_summary$agreement
```

### Why alpha can be low when experts agree

Alpha compares the disagreement within items with the disagreement expected if
the same ratings were scattered across items at random. When a panel rates
nearly every item 4, very little disagreement is expected by chance, so a few 3s
pull alpha down even though most rating pairs are identical. Feinstein and
Cicchetti (1990) described the same pattern for kappa. The output reports the
share of identical rating pairs next to alpha so the two can be read together.
A low alpha alongside a high share of identical pairs is not by itself evidence
of a poor panel.

### Gwet's AC1

Gwet's (2008) AC1 was designed to stay high in that situation, and it is
available with `agreement = "ac1"`. It is never the default. Vach and Gerke
(2023) show that AC1 rises as ratings concentrate in one category even when
agreement does not change, and that it can be above zero when experts rate
independently. Its output always repeats that critique. In relevance mode, AC1
is computed on the relevant/not-relevant decision at `relevance_cut`:

```{r agreement-ac1}
ac1_fit <- expert_validity(R, mode = "relevance", lo = 1, hi = 4,
                           agreement = "ac1", agreement_B = 0)
ac1_fit$details$agreement
```

### The interval

The interval resamples items with all of their ratings intact, the procedure
Zapf et al. (2016) evaluated; they found that Krippendorff's original bootstrap,
which ignores dependence between raters, reached only about 60% coverage. The
interval varies slightly between runs, so set `seed` to make it reproducible, or
set `agreement_B = 0` to skip it. `panel_agreement()` runs the same analysis on
any rater-by-item matrix.

## Essentiality: Lawshe CVR with exact critical values

Lawshe's task asks experts whether an item is essential. With twelve experts:

```{r essentiality}
expert_validity(c(10, 8, 6), mode = "essentiality", N = 12)
```

`cvr()` derives the smallest essential count whose one-sided binomial upper-tail
probability is no greater than `alpha`. This makes the panel-size dependency
explicit and follows the exact-probability logic revisited by Ayre and Scally
(2014).

Judge-by-item binary data can be supplied directly:

```{r essential-matrix}
E <- cbind(
  Item1 = c(1,1,1,1,1,1,1,1),
  Item2 = c(1,1,1,1,1,0,0,0)
)
expert_validity(E, mode = "essentiality")
```

A failure to clear the exact criterion is labeled `Review`, not automatic
deletion. Expert rationales and domain coverage matter when deciding whether an
item should be rewritten, retained for breadth, or removed.

## Congruence: item-objective alignment

IOC uses expert ratings of -1, 0, and +1 for item-objective congruence. A target
mapping lets the workflow compare intended and competing objectives:

```{r congruence}
d <- expand.grid(
  item = c("I1", "I2"),
  judge = 1:4,
  objective = c("A", "B")
)
d$target_objective <- ifelse(d$item == "I1", "A", "B")
d$score <- ifelse(d$objective == d$target_objective, 1, -1)

expert_validity(d, mode = "congruence")
```

The workflow reports target IOC, the strongest competitor, and their margin.
This is a diagnostic comparison, not a manufactured significance test. If no
target mapping is provided, all IOC cells are returned descriptively.

## Missing ratings

Missing data are never silently ignored by default. Set `na.rm = TRUE` only when
itemwise/cellwise deletion matches the study protocol. Effective expert counts
and missing counts are then reported so downstream interpretation uses the
actual panel size.

## Plotting expert evidence

Each mode uses a plot matched to the expert task rather than forcing unlike
indices into one generic chart.

```{r expert-plots, fig.width=7, fig.height=4}
plot(expert_validity(R, mode = "relevance", lo = 1, hi = 4))
plot(expert_validity(c(10, 8, 6), mode = "essentiality", N = 12))
plot(expert_validity(d, mode = "congruence"))
```

Relevance mode displays Aiken's V with its score interval and overlays I-CVI as
a separate marker. Essentiality mode displays observed CVR against the exact
panel-specific critical CVR. Congruence mode connects target IOC to the strongest
competitor so the alignment margin is visually explicit. These displays are
diagnostic summaries; they do not create new validity thresholds.

## Reporting

A concise methods/results description should identify:

1. who the experts were and why they were qualified;
2. the exact task and response scale;
3. the index and inference/CI procedure used, and for relevance ratings the
   agreement coefficient and its measurement level;
4. the panel size, including item-specific missingness;
5. quantitative item and scale evidence; and
6. how expert comments, construct coverage, and comprehensibility informed the
   final item decisions.

A content-validity coefficient is evidence about a defined expert task. It is
not, by itself, a complete validity argument.

## References

Aiken, L. R. (1980). Content validity and reliability of single items or
questionnaires. *Educational and Psychological Measurement, 40*(4), 955-959. https://doi.org/10.1177/001316448004000419

Lawshe, C. H. (1975). A quantitative approach to content validity. *Personnel
Psychology, 28*(4), 563-575. https://doi.org/10.1111/j.1744-6570.1975.tb01393.x

Ayre, C., & Scally, A. J. (2014). Critical values for Lawshe's content validity
ratio: Revisiting the original methods of calculation. *Measurement and
Evaluation in Counseling and Development, 47*(1), 79-86.
https://doi.org/10.1177/0748175613513808

Feinstein, A. R., & Cicchetti, D. V. (1990). High agreement but low kappa: I.
The problems of two paradoxes. *Journal of Clinical Epidemiology, 43*(6),
543-549.

Gwet, K. L. (2008). Computing inter-rater reliability and its variance in the
presence of high agreement. *British Journal of Mathematical and Statistical
Psychology, 61*(1), 29-48. https://doi.org/10.1348/000711006X126600

Hayes, A. F., & Krippendorff, K. (2007). Answering the call for a standard
reliability measure for coding data. *Communication Methods and Measures, 1*(1),
77-89. https://doi.org/10.1080/19312450709336664

Krippendorff, K. (2011). *Computing Krippendorff's alpha-reliability.* Annenberg
School for Communication, University of Pennsylvania.

Penfield, R. D., & Giacobbi, P. R., Jr. (2004). Applying a score confidence
interval to Aiken's item content-relevance index. *Measurement in Physical
Education and Exercise Science, 8*(4), 213-225.
https://doi.org/10.1207/S15327841MPEE0804_3

Polit, D. F., Beck, C. T., & Owen, S. V. (2007). Is the CVI an acceptable
indicator of content validity? Appraisal and recommendations. *Research in
Nursing & Health, 30*(4), 459-467. https://doi.org/10.1002/nur.20199

Rovinelli, R. J., & Hambleton, R. K. (1977). On the use of content specialists
in the assessment of criterion-referenced test item validity. *Dutch Journal of
Educational Research, 2*, 49-60.

Turner, R. C., & Carlson, L. (2003). Indexes of item-objective congruence for
multidimensional items. *International Journal of Testing, 3*(2), 163-171.
https://doi.org/10.1207/S15327574IJT0302_5

Vach, W., & Gerke, O. (2023). Gwet's AC1 is not a substitute for Cohen's kappa:
A comparison of basic properties. *MethodsX, 10*, 102212.

Zapf, A., Castell, S., Morawietz, L., & Karch, A. (2016). Measuring inter-rater
reliability for nominal data: Which coefficients and confidence intervals are
appropriate? *BMC Medical Research Methodology, 16*, 93.
https://doi.org/10.1186/s12874-016-0200-9
