Gtheory4LLM provides exact balanced Gaussian generalizability-theory models and a bounded experimental Laplace engine for small discrete G-theory models used in LLM annotation and evaluation studies. Specify which evaluator, prompt, and repeated-run sources matter for your study, fit their variation, and compare projected reliability across alternative numbers of evaluators and prompts.
This is a research beta. Passing the package’s numerical acceptance checks does not establish parameter recovery, interval coverage, Laplace approximation quality, or a scientifically sufficient number of evaluators. Read docs/LIMITATIONS.md before quoting a coefficient, and docs/VALIDATION_SCOPE.md for what has actually been validated.
Source version: 0.5.0. For versioned archives, manuals and publication status, see the repository manifest and GitHub releases. These repository records are excluded from the package archive; this source version does not assert that a corresponding release has been published.
R 4.5 or later and OpenMx are required. Nothing in this package’s own
R code needs R 4.5; the floor is inherited from OpenMx 2.22.11, which
calls a C entry point introduced in R 4.5.0 while declaring only
R (>= 3.5.0) itself. Declaring it here turns a later,
unexplained compilation failure into a clear refusal. Older pairings
such as R 4.3 with OpenMx 2.21.11 have been reported to pass these
checks but are not part of the validated gate; to use one, install from
source with a relaxed floor, or source("load_functions.R"),
which imposes no version requirement.
The current CRAN release, which can be older than this source version, installs with its dependencies:
install.packages("Gtheory4LLM")A specific release can also be installed from its versioned archive, for example the v0.4.1 archive:
install.packages("OpenMx")
install.packages(
"https://github.com/Veronica0206/Gtheory4LLM/releases/download/v0.4.1/Gtheory4LLM_0.4.1.tar.gz",
repos = NULL, type = "source"
)
library(Gtheory4LLM)The repository keeps checksummed release files in artifacts/.
A 0.4.0 bundle, Gtheory4LLM_0.4.0.tar.gz, was prepared and
tagged but never published; 0.4.1 supersedes it. To install this source
version, run R CMD build ., then
R CMD INSTALL Gtheory4LLM_0.5.0.tar.gz, the archive it
creates. For the exact checked release archive, use the versioned GitHub
release and verify its accompanying manifest. Building the vignette
needs knitr, rmarkdown, and pandoc; using the installed package does
not. CRAN availability is separate from GitHub availability. The version
CRAN carries is recorded in the repository development
status.
The installed tutorial creates synthetic continuous judgments for 24 items, 4 evaluators, 3 prompts, and 2 runs within each evaluator/prompt pair. It completes design inspection, fitting, diagnostics, G/Phi calculation, and an 18-allocation decision study. It requires no API keys or external data.
source(system.file("doc", "LLM-workflow.R",
package = "Gtheory4LLM", mustWork = TRUE))The same workflow is installed as a vignette:
browseVignettes("Gtheory4LLM"), or
vignette("LLM-workflow", package = "Gtheory4LLM"). Its
source is vignettes/LLM-workflow.Rmd. The
main steps are:
# llm_data and llm_design are created by the tutorial above.
gt_preflight(llm_data, "quality", llm_design)
summary(fit)
gt_diagnostics(fit)$numerically_accepted
gt_reliability(fit)
gt_dstudy(fit, expand.grid(evaluator = c(2, 4, 6),
prompt = c(2, 3, 4), run = c(1, 2)))G describes relative comparisons between items. Phi additionally includes absolute shifts across the declared measurement conditions. Decision studies average over specified random facet populations and hold fitted source covariances fixed. More evaluators or prompts are meaningful projections only if those populations and covariance assumptions remain appropriate. These are point estimates carrying an interval. For example, G = 0.80 describes the universe-score share of variance for relative item comparisons under the specified model; it is not 80% labeling accuracy or agreement with a human reference. Phi additionally penalizes absolute panel shifts and is no larger than G under this model.
Gaussian fits report asymptotic standard errors and delta-method intervals where the likelihood curvature and boundary rules permit them. They use the restricted (REML) or profile (ML) likelihood Hessian:
summary(fit)$variances # variance, std_error, boundary flag
gt_reliability(fit)$per_trait # Erho2, Erho2_se, Erho2_lower, Erho2_upper, ...
logLik(fit); AIC(fit); gt_component_vcov(fit)These are asymptotic Wald quantities conditional on the declared model and allocation. A source resting on a variance boundary has no Wald standard error, and the discrete Laplace engine computes no observed information at all. What those intervals do and do not cover is stated once in docs/LIMITATIONS.md.
A facet is random when the study generalizes to a population of its
levels, and fixed when the universe is exactly the levels used. A chosen
temperature or prompt set may be fixed when inference is restricted to
those conditions; the choice follows the intended generalization, not
the facet name. Declare it with fixed, which applies the
mixed model of Brennan (2001): the object-by-fixed interaction is
averaged over that facet’s levels and joins universe-score variance, and
a source built only from fixed facets leaves the model.
gt_reliability(fit, fixed = "temp")A fixed facet’s count cannot be changed, and a decision study may not project over it.
as.data.frame() exports reliability or D-study results
with allocation, coefficient scale, uncertainty status and extrapolation
recorded alongside each estimate. gt_dstudy_target()
identifies supplied allocations meeting an estimated target and
preserves ties for the fewest measurements among them. It does not
optimize beyond the supplied grid or guarantee future reliability.
reliability <- gt_reliability(fit)
as.data.frame(reliability)
plot(reliability, target = 0.80)Preflight plots show cell completeness and replication, random-source
dimensions, or category representation across design groups. Outcome
profiles include category counts and groups with no observed variation;
they do not certify adequate information. D-study plots show
coefficients against measurements per object, with optional target lines
and extrapolation markers. Figures use base R graphics, so standard
PNG/PDF devices can save them. The installed LLM-workflow
vignette demonstrates these views and CSV export; the
visualization guide explains their interpretation.
A portable report combines retained model context, diagnostics, result tables and figures without rerunning optimization:
report <- gt_report(fit, reliability = reliability)
# Writes a standalone offline HTML file; existing files are protected.
# gt_export_report(report, "analysis-report.html")Reports exclude observations and group-level identifiers, while retaining variable/source names and aggregate results. Review those names before sharing. Unavailable fitting provenance stays unavailable; the report-generation session is recorded separately. See the reporting guide.
Study planning separates three quantities: distinct pilot items provide estimation information, annotations per item define the final score, and items per request define shared context and request overhead. Increasing one does not substitute for increasing another. These workflows are included in source version 0.5.0; their statistical boundaries are described below.
The study-planning vignette runs synthetic examples of all three additions. Its monetary amounts are illustrative user inputs.
gt_batch_reliability() and
gt_batch_dstudy() project absolute item-score error,
explicitly named item differences and weighted aggregate error for
supported Gaussian shared-call fits. A call shift cancels from a
same-batch difference but remains in an absolute score. Targets
condition on the recorded fixed item grouping; all instrumentation
facets remain random. These are point projections, with no sampling
intervals, regrouping or batch-size optimization.gt_plan() evaluates a supplied bounded grid using
dated, user-supplied costs and constraints. It retains every candidate,
minimum-cost ties, the cost versus reliability Pareto set, and changes
in error contributions from a declared baseline. It screens a point
estimate or, separately, an available lower confidence bound. Missing
bounds stay unavailable. It does not choose evaluator identities or
change a facet from random to fixed to improve a result.gt_simulate() and gt_pilot_plan() draw and
refit balanced Gaussian pilot panels under a fitted or explicit
parameter scenario. Pilot item/facet counts may vary while
target_counts holds the eventual annotation protocol
constant. The planner records every attempt, refusal and unavailable
interval, with Monte Carlo errors and conditional interval-width
summaries. This evaluates model-based precision, not calibrated coverage
or guaranteed study adequacy.For example, after fitting an accepted model without a shared-call source:
plan <- gt_plan(fit,
candidates = list(evaluator = 2:4, prompt = 1:3),
cost = function(d) data.frame(annotation = 0.01 * d$total_annotations,
setup = 2 * d$allocation_evaluator),
currency = "USD", cost_date = "2026-10-09", study_items = 1000,
target = 0.80, coefficient = "Phi")
subset(as.data.frame(plan), minimum_cost %in% TRUE)The cost planner uses ordinary reliability and therefore refuses
modelled Call fits. Use the separate batch projection
functions for their declared targets. Supplying request costs for an
independent-item model does not account for shared-call error, and
adding human-review costs does not itself model a benefit from review.
Pilot simulation currently supports Gaussian models without any batch
declaration. See the function manuals for the remaining scope and
limits.
A new shared-call fit retains compact item-to-batch labels even with
retain = list(data = FALSE): the layout is necessary for
named comparisons. Batch projection objects contain those labels and
supplied target weights. gt_report(fit) and its HTML export
exclude the labels; share a projection object only when its identifiers
are appropriate to disclose. Older fits lacking the layout must be
refitted rather than having their grouping guessed.
Facet names do not impose a hierarchy. Declare any nonempty subset of
the instrumentation facets to interact with the object, give other
facets explicit nesting parents, and choose higher-order interactions
separately. With four facets, all 15 nonempty crossed subsets are
supported by the design interface. An explicit random-source formula is
also available. gt_preflight() reports the requested and
resolved sources, dimensions, replication, resource limits, and
downstream operations before optimization; it does not certify
statistical identification or successful fitting.
| Outcome | Fitting | Supported G/Phi and D studies |
|---|---|---|
| Gaussian | Exact balanced ML/REML, univariate or joint multivariate | Observed-score coefficients and explicitly weighted composites |
| Binary | Bernoulli probit/logit, univariate or joint | Explicitly requested latent-response coefficients |
| Ordinal | Cumulative probit/logit, univariate or joint | Explicitly requested latent-response coefficients |
| Unordered categorical | Reference-category multinomial softmax, univariate or joint | No scalar coefficient implemented |
Discrete observations follow their chosen response distributions; their latent random effects remain Gaussian. Gaussian assumptions therefore remain part of the model. Latent binary/ordinal reliability concerns an averaged latent response, not observed proportions, majority votes, or category agreement. Joint Gaussian-discrete models are not implemented.
Traditional method-of-moments (MoM) G theory already accommodates random effects and crossed/nested designs. The package’s contribution is the common workflow for explicit source selection, joint covariance models, different response families, and supported allocation comparisons. It does not claim that random effects are novel or that likelihood is uniformly superior to MoM. See Brennan (2001) and Jiang et al. (2020) for G theory and earlier multivariate likelihood implementations.
At one observation per object-by-facets cell, a discrete fit cannot identify the object-by-all-facets source that the default design requests, and that source is never dropped silently. Declare the same design without it:
binary_design <- gt_design("item", "rater", full_cell = FALSE)full_cell = FALSE removes exactly that one term and
leaves every other requested source unchanged;
random = ~ item + rater writes the reduced model out in
full. Either way the omission is an explicit modeling choice.
A first discrete fit, start to finish — the guard above, preflight, binary and ordinal fits, diagnostics, latent reliability and a decision study — is installed as a runnable script:
source(system.file("examples", "discrete-first.R", package = "Gtheory4LLM"))system.file("examples", "discrete-boundary.R", package = "Gtheory4LLM")
shows the companion case: a constructed panel whose zero object variance
is a legitimate optimum rather than evidence about reliability.
gt_control() exposes named starting values, reproducible
additional fitting attempts, an optimizer held fixed across attempts,
the discrete resource limits, and which optional components a fit keeps.
Discrete fits retain rejected candidates and their reasons.
summary(fit) reports acceptance, selected attempt, boundary
status, and approximation status without printing model internals or
observations; detailed records remain available.
Two limits decide whether a model can be fitted at all. The Gaussian
engine requires a complete balanced coded panel, as do
gt_reliability() and gt_dstudy(). The dense
discrete backend is bounded by row, dimension, parameter, and
working-memory limits, and refuses before allocating rather than after
allocation fails. gt_preflight() reports both before any
optimization.
Every limitation is listed once, in docs/LIMITATIONS.md — what the designs, the Gaussian engine and the discrete engine do and do not support, which combinations are not implemented, and what “supported” means here. Read it before quoting a coefficient. What has actually been checked, and what a future study still has to establish, is in docs/VALIDATION_SCOPE.md; the initial statistical pilots preserve their coverage, approximation and recovery results, including boundaries and unavailable comparisons.
The real-data workflow audits the native outcome types and full designs of all three panels, then demonstrates how to interpret unsupported requests and numerical failures.
gt_example() lists eight outcome sets from three
real LLM annotation panels: hate speech, mental health,
and drug reviews. Each contains 21,600 measurements of 100 items crossed
with four evaluators, three prompts, six temperatures, and three seeds.
These are alternative codings of three panels, not eight independent
datasets. The package includes design identifiers and modeled
annotations from the public OSF deposit. The
original source CSV checksums match that deposit. Raw texts, original
corpus reference labels, and API metadata are omitted from the package
tables.
The six discrete native outcome sets each contain 21,600 rows and exceed the dense discrete engine limits by more than an order of magnitude; raising the limits does not make the dense engine practical at that size. They are data resources, not full-panel discrete fitting examples. Use a scientifically justified small design for this backend; do not remove item interactions merely to obtain an accepted fit. The two mental-health 7L/3L sets remain Gaussian working scores under both coding options.
Temperature-specific seed agreement is a useful descriptive diagnostic. Lower agreement at higher temperatures may reflect changes in category probabilities, latent variances, or both; agreement alone does not identify which changed.
Use fixed = "temp" when the coefficient is intended to
average over exactly the observed temperatures. This changes the
reliability estimand after fitting; it cannot repair misspecification in
the fitted variance model. If pooling across temperatures is
questionable, a within-temperature analysis is one sensitivity analysis,
with its own conditional estimand. It is not the only possible model and
does not generalize across temperatures automatically. See
help("gtheory_datasets").
Other native codings preserve binary, ordinal, or unordered
categorical outcomes; coding = "manuscript" reproduces the
seven historical Gaussian working-score codings. Mental-health 7L/3L
values are working scores, not established clinical severity scales. The
installed manual documents variables, category mappings, preprocessing,
source corpora, and the corresponding annotation studies. In particular,
the original mental-health preprocessing could default omitted items
within a parsed batch to NORMAL; these defaults cannot be distinguished
in the bundled tables. Interpret the labels with that limitation in
mind:
gt_example()
x <- gt_example("hate_speech")
help("gtheory_datasets", package = "Gtheory4LLM")
help("gtheory_hate_speech", package = "Gtheory4LLM")
help("gtheory_mental_health", package = "Gtheory4LLM")
help("gtheory_drug_review", package = "Gtheory4LLM")The task bibliography is
installed as
system.file("DATASET_REFERENCES.bib", package = "Gtheory4LLM").
Use citation("Gtheory4LLM") for the software citation.
Loading a table does not establish that every design or full joint
discrete fit is feasible.
A single release entrypoint checks sources and the distributed archive separately:
python3 scripts/run_validation.py --scope all --as-cranSee validation instructions for the locked environment, minimum-R and Windows/macOS checks, source-only checks, and retained evidence. The numerical suite includes independent dense Gaussian likelihoods, binary/ordinal comparisons with lme4 and ordinal, likelihood/derivative identities, joint covariance estimation, and coefficient aggregation. These are specific numerical checks, not a substitute for simulation-based statistical validation. A configured workflow is not evidence of a passing run.
Report installation failures, unclear design declarations, and numerical problems through GitHub issues. Please include the package/R versions, a small synthetic example, the declared family/design, and relevant diagnostics. Remove private observations and credentials before posting. Maintainer: Jin Liu, Veronica.Liu0206@gmail.com.
Package code is licensed under GPL-3. The real annotation tables retain the CC BY 4.0 license stated in the public deposit’s data codebook; see data attribution and license. Synthetic tutorial and test examples are covered by the code license. Public release files are limited to the package, its documented examples, tests, and software release artifacts. Unpublished manuscripts and research archives are excluded.
Each document has one job, so that nothing has to be kept true in two places.
| Document | What it covers |
|---|---|
| This README | Install, a complete worked workflow, and what the package is for |
| vignettes/LLM-workflow.Rmd | The end-to-end reliability tutorial, installed and runnable |
| vignettes/study-planning.Rmd | Fixed-layout batch targets, monetary allocation search and Gaussian pilot precision |
| docs/LIMITATIONS.md | Everything the package does not do, listed once |
| docs/VALIDATION_SCOPE.md | What has been checked, how, and what that does not establish |
| docs/REAL_DATA_WORKFLOW.md | The bundled panels, their native outcomes, and their resource ceilings |
| docs/DEVELOPMENT_STATUS.md | The current release state and what is being worked on |
| docs/ROADMAP.md | Planned work beyond this release, in the order it is planned |
| docs/DISCRETE_BACKEND_CONTRACT.md | Private dense response, matrix and mode interfaces |
| NEWS.md | Version history |
| docs/CONTRIBUTING.md, docs/RELEASE_CHECKLIST.md, docs/REPOSITORY_POLICY.md, SECURITY.md | Working on the package itself |
See the contribution guide for setup and statistical change requirements, and the release checklist for version, tag, archive, and manual correspondence. Software checks and scientific validation are reported separately. The development status tracks completed work and remaining extensions.