---
title: "How many evaluators, prompts, and runs?"
output:
  rmarkdown::html_vignette:
    toc: true
vignette: >
  %\VignetteIndexEntry{How many evaluators, prompts, and runs?}
  %\VignetteEngine{knitr::rmarkdown}
  %\VignetteEncoding{UTF-8}
---

```{r setup, include = FALSE, purl = FALSE}
knitr::opts_chunk$set(collapse = TRUE, comment = "#>")
options(width = 100)
```

A complete **Gtheory4LLM** workflow for the reliability of mean continuous
subjective-quality judgments.

> **All data below are synthetic.** This tutorial uses no manuscript results,
> source-study data, API calls, or real LLM outputs. Its results demonstrate the
> software workflow; they do not establish an adequate evaluator or prompt count
> for a real task.

```{r library}
# This vignette is also installed as a runnable script:
#   source(system.file("doc", "LLM-workflow.R", package = "Gtheory4LLM"))
library(Gtheory4LLM)
```

## 1. Define the quantity and the sampling design

The quantity is an item's mean continuous quality score over evaluators,
prompts, and repeated runs. We generate a complete panel of 24 items, four
evaluators, three prompts, and two runs: 576 measurements. The model assumes
exchangeable random effects for those populations and a Gaussian score model.
The scores are simulated continuous numbers, not numeric codes assigned to
ordered categories.

| Facet | Meaning in this example |
|---|---|
| Item | The object of measurement. Stable differences between items form the universe-score variance. |
| Evaluator and prompt | Both have main effects and direct item interactions. Different items can respond differently to these facets. |
| Run | Scoped to an evaluator/prompt pair. A run group labels all items. The run source is a common shift across items; remaining item-specific run variation belongs to the residual in this simulation. |

```{r design}
set.seed(20260912)
llm_data <- expand.grid(item = factor(seq_len(24)),
  evaluator = factor(paste0("model_", seq_len(4))),
  prompt = factor(paste0("prompt_", seq_len(3))),
  run = factor(seq_len(2)))

llm_design <- gt_design("item", c("evaluator", "prompt", "run"),
  crossed = c("evaluator", "prompt"),
  nested = list(run = c("evaluator", "prompt")),
  item_interactions = "additive", instrument_interactions = "additive")
llm_design$terms_requested
```

This declares six random sources: item, evaluator, prompt, item-by-evaluator,
item-by-prompt, and evaluator-by-prompt-by-run. The run group is parent-scoped
even though the coded panel is complete. Item-by-run and additional
evaluator-by-prompt effects are not silently added. Omitting them is an
assumption of this known simulation; a real study needs its own justification.
The exact Gaussian engine currently requires a complete coded Cartesian panel
with one row per full cell.

```{r simulate}
source_sd <- c(item = 1.2, evaluator = 0.4, prompt = 0.3,
  "item:evaluator" = 0.5, "item:prompt" = 0.4,
  "evaluator:prompt:run" = 0.25)
llm_data$quality <- 0
for (source in names(source_sd)) {
  members <- strsplit(source, ":", fixed = TRUE)[[1L]]
  group <- do.call(interaction, c(llm_data[members], list(drop = TRUE)))
  effects <- rnorm(nlevels(group), sd = source_sd[[source]])
  llm_data$quality <- llm_data$quality + effects[as.integer(group)]
}
llm_data$quality <- llm_data$quality + rnorm(nrow(llm_data), sd = 0.6)
```

## 2. Inspect feasibility before fitting

`gt_preflight()` reports observed counts, source dimensions, free parameters,
configured size limits, and the supported reliability scale, without building a
model matrix or running an optimizer.

```{r preflight}
preflight <- gt_preflight(llm_data, "quality", llm_design)
print(preflight)
preflight$sources
stopifnot(preflight$fitting_feasible)
```

Passing this report means that the listed structural and resource checks pass.
It does not prove identification, reliable estimation, numerical acceptance, or
approximation adequacy. Read failed checks before changing a resource limit or
simplifying a scientific model.

The panel audit counts missing cells and cells above or below the declared
replication. It retains only a bounded set of examples, using the observed
coded levels; it cannot discover a level that never appeared in the data.

```{r panel-audit, fig.width=7, fig.height=4}
preflight$panel_audit$summary
plot(preflight, type = "cells")
incomplete <- gt_preflight(llm_data[-1, ], "quality", llm_design)
incomplete$panel_audit$missing_cells
```

These diagnostics do not fill cells or remove records. Investigate unexpected
codes or missing measurements before deciding how the study should be analysed.

## 3. Fit and inspect acceptance

```{r fit}
fit <- gt_fit(llm_data, "quality", llm_design,
  control = gt_control(gaussian = list(optimizer = "CSOLNP",
    extra_tries = 2L, retry_seed = 415L, threads = 1L)))
print(summary(fit))
diagnostics <- gt_diagnostics(fit)
print(diagnostics[c("numerically_accepted", "selected_attempt",
                    "acceptance_failures", "approximation_adequacy")])
stopifnot(isTRUE(diagnostics$numerically_accepted))
gt_components(fit)
```

The Gaussian default estimator is REML. The trial budget allows the initial fit
plus two additional attempts; it is not resampling. The chosen optimizer remains
fixed across attempts. Inspect numerical acceptance and source-boundary
diagnostics before interpreting the components. A completed optimizer alone is
insufficient. If an installation does not support the selected optimizer or a
fit is rejected, address that explicitly.

The summary table also carries an asymptotic standard error for each source
variance, obtained from the numerically differentiated restricted-likelihood
Hessian. A component resting on a variance boundary is flagged, because Wald
theory does not apply there.

```{r variances}
summary(fit)$variances
```

## 4. Compute relative and absolute reliability

```{r reliability}
reliability <- gt_reliability(fit)
reliability
reliability$interpretation
```

`Erho2` is the relative coefficient G: common evaluator, prompt, and run shifts
do not change relative item comparisons. `Phi` is the absolute coefficient and
also includes those shifts as error. Item-dependent interactions contribute to
both error definitions. Here both coefficients refer to observed Gaussian-score
averages under the declared design. They measure consistency under that model,
not accuracy against human reference labels. A coefficient of 0.80 means that
80% of the variance in item means is universe-score variance under this model;
it is not 80% labeling accuracy.

The interval is a delta-method Wald interval on the logit scale, propagating the
fitted parameter covariance matrix through the source covariances. It describes
estimation uncertainty in those covariances under the declared model. It is not
a guarantee about a future sample of evaluators or prompts.

```{r reliability-figure, fig.width=8, fig.height=4.5}
plot(reliability, target = 0.80,
  main = "Reliability of the declared mean quality score")
reliability_table <- as.data.frame(reliability)
```

The dashed reference is an illustrative target. Missing intervals are shown as
unavailable, never as a claim of perfect precision. The table records the scale,
uncertainty status and any conditional inference, alongside the estimates.

## 5. Compare measurement budgets

```{r dstudy}
grid <- expand.grid(evaluator = c(2L, 4L, 6L),
  prompt = c(2L, 3L, 4L), run = c(1L, 2L))
dstudy <- gt_dstudy(fit, grid)
planning <- as.data.frame(dstudy)
# Prefixed allocation columns avoid collisions with outcome/estimate names.
# These readable aliases are just for this tutorial's known facet names.
planning$evaluator <- planning$allocation_evaluator
planning$prompt <- planning$allocation_prompt
planning$run <- planning$allocation_run
planning$measurements_per_item <- planning$measurements_per_object
planning <- planning[order(planning$measurements_per_item, -planning$Phi), ]
rownames(planning) <- NULL
print(planning)
```

Runs are counted within each evaluator/prompt pair, so the budget per item is
evaluators x prompts x runs. A row using more levels than observed extrapolates
the fitted variance model to the declared exchangeable population.

```{r threshold}
screen <- gt_dstudy_target(dstudy, target = 0.80, coefficient = "Phi",
  outcome = "quality")
screen[screen$meets_target %in% TRUE, ]
screen[screen$fewest_measurements %in% TRUE, ]
```

A threshold of 0.80 is illustrative; choose and justify your own criterion. Read
the interval columns alongside the point projection: several allocations here
have overlapping intervals, so the ordering of nearby rows is not established by
these data.

The screening helper keeps every supplied candidate, including those that miss
the target or have an unavailable estimate. It marks all ties for the fewest
measurements among qualifying candidates. It performs no search outside this
grid and does not guarantee that a future panel will meet the target.

```{r dstudy-figure, fig.width=7, fig.height=5}
plot(dstudy, coefficient = "Phi", outcome = "quality", target = 0.80,
  main = "Candidate measurement allocations")
```

Different allocations can use the same total number of measurements; they are
separate points rather than one assumed smooth curve. The figure identifies
allocations extending beyond the observed facet counts. Such projections keep
the fitted source distributions fixed.

Both result tables contain ordinary scalar columns, suitable for CSV output:

```{r export, eval=FALSE, purl=FALSE}
write.csv(reliability_table, "reliability.csv", row.names = FALSE)
write.csv(as.data.frame(dstudy), "decision-study.csv", row.names = FALSE)
pdf("decision-study.pdf", width = 7, height = 5)
plot(dstudy, coefficient = "Phi", outcome = "quality", target = 0.80)
dev.off()
```

## 6. Save a portable analysis report

A report collects the design, fitted model context, diagnostic flags, coefficient
and decision-study tables, and figures. It does not rerun optimization. Supplied
coefficient results are checked against projections from the fit; supplied
preflight information is checked for compatible aggregate specifications, which
does not prove that it came from identical observations.

```{r analysis-report}
analysis_report <- gt_report(fit, preflight = preflight,
  reliability = reliability, dstudy = dstudy)
print(analysis_report)
```

```{r save-analysis-report, eval=FALSE, purl=FALSE}
gt_export_report(analysis_report, "analysis-report.html")
```

The HTML file contains its figures and styling and can be opened offline. An
existing destination is refused unless replacement is explicitly requested.
The report omits observations, row examples and group-level identifiers, while
keeping variable and source names and aggregate results. This is not an
anonymization guarantee. Inspect study-specific names before sharing it.
Retained fitting provenance and the report-generation environment are separate;
missing fitting provenance is reported as unavailable.

## 7. Fixed facets

A facet is *random* when the study generalizes to a population of its levels,
and *fixed* when the universe of generalization is exactly the levels used. LLM
evaluation may target exactly a chosen temperature or prompt set, in which case
those facets are fixed. Selecting levels deliberately does not by itself decide
the universe of generalization.

Declare them with `fixed`. Following the mixed model of Brennan (2001), the
object-by-fixed-facet variance is averaged over that facet's levels and added to
universe-score variance, and a source built only from fixed facets shifts every
item equally and leaves the model.

```{r fixed}
fixed_prompt <- gt_reliability(fit, fixed = "prompt")
fixed_prompt
fixed_prompt$source_roles
```

Treating the prompt set as fixed raises both coefficients, because
item-by-prompt variation is now part of what the study is trying to measure
rather than error. A fixed facet's count cannot be changed, and a decision study
may not project over it.

`fixed` acts at the reliability stage: it changes how the already-fitted
components are aggregated into a coefficient. It does not change the G study,
so it cannot repair a source whose variance was misspecified during fitting.
Observed seed disagreement can change across temperatures even when latent
variance is constant, because category probabilities can change. If diagnostics
raise doubts about pooling temperatures, fitting within one temperature is a
useful sensitivity analysis for a conditional estimand; declaring a facet fixed
is a separate decision about generalization.

## 8. Match other outcomes to their observation model

| Outcome | Current scalar reliability support |
|---|---|
| Gaussian | Observed continuous-score averages. |
| Binary or ordinal | Explicit `scale = "latent"` after an accepted fit. This describes latent-response averages, not majority votes, label proportions, or observed ordinal-score averages. |
| Unordered categorical | No implemented default scalar G/Phi. A scientific score or category-probability estimand is needed. |

Declare category order and references explicitly. Do not recode binary, ordinal,
or nominal labels as the continuous quality score used above. Discrete fitting
uses a bounded dense Laplace engine with Gaussian latent random effects;
`gt_preflight()` reports its current limits, and changing the link does not
remove the random-effects distribution assumption. That engine reports point
estimates only: it computes no standard errors. Analytic reliability and D
studies require complete balanced coded panels even when discrete fitting admits
missing whole cells.

At one observation per object-by-facets cell, a discrete fit cannot identify the
full-cell source that the default design requests. Declare the same design
without it:

```{r discrete-design}
gt_design("item", "rater", full_cell = FALSE)$terms_requested
```

A complete panel can still contain little observed outcome variation. This
small, deliberately imbalanced example has one positive response. It is a data
profile, not an estimation example or a recommended design.

```{r outcome-information, fig.width=7, fig.height=4}
information_data <- expand.grid(item = seq_len(6), rater = seq_len(3))
information_data$label <- as.integer(information_data$item == 1 &
  information_data$rater == 1)
information_design <- gt_design("item", "rater", random = ~ item + rater)
information <- gt_preflight(information_data, "label", information_design,
  gt_family("binary"), max_examples = 0)
information$outcome_profile$label$categories
information$outcome_profile$label$by_variable
plot(information, type = "outcomes", outcome = "label")
```

The heatmap describes the fraction of groups in which each category appears.
A group with one observation necessarily has no observed variation. These
summaries supply no minimum-count threshold, independence claim or guarantee of
parameter recovery. Declared categories absent from the entire panel are shown
with zero counts and block fitting; the preflight report does not merge them or
silently change their order. No finite panel can reveal a category never
observed and never declared.

## Where to read next

Installed help: `?gt_design`, `?gt_family`, `?gt_preflight`, `?gt_control`,
`?gt_fit`, `?gt_diagnostics`, `?gt_reliability`, `?gt_dstudy`, and `?gt_report`.

Brennan, R. L. (2001). *Generalizability Theory*. Springer.
<doi:10.1007/978-1-4757-3456-0>
