How many evaluators, prompts, and runs?

A complete Gtheory4LLM workflow for the reliability of mean continuous subjective-quality judgments.

All data below are synthetic. This tutorial uses no manuscript results, source-study data, API calls, or real LLM outputs. Its results demonstrate the software workflow; they do not establish an adequate evaluator or prompt count for a real task.

# This vignette is also installed as a runnable script:
#   source(system.file("doc", "LLM-workflow.R", package = "Gtheory4LLM"))
library(Gtheory4LLM)

1. Define the quantity and the sampling design

The quantity is an item’s mean continuous quality score over evaluators, prompts, and repeated runs. We generate a complete panel of 24 items, four evaluators, three prompts, and two runs: 576 measurements. The model assumes exchangeable random effects for those populations and a Gaussian score model. The scores are simulated continuous numbers, not numeric codes assigned to ordered categories.

Facet Meaning in this example
Item The object of measurement. Stable differences between items form the universe-score variance.
Evaluator and prompt Both have main effects and direct item interactions. Different items can respond differently to these facets.
Run Scoped to an evaluator/prompt pair. A run group labels all items. The run source is a common shift across items; remaining item-specific run variation belongs to the residual in this simulation.
set.seed(20260912)
llm_data <- expand.grid(item = factor(seq_len(24)),
  evaluator = factor(paste0("model_", seq_len(4))),
  prompt = factor(paste0("prompt_", seq_len(3))),
  run = factor(seq_len(2)))

llm_design <- gt_design("item", c("evaluator", "prompt", "run"),
  crossed = c("evaluator", "prompt"),
  nested = list(run = c("evaluator", "prompt")),
  item_interactions = "additive", instrument_interactions = "additive")
llm_design$terms_requested
#> [1] "evaluator"            "item"                 "prompt"               "item:evaluator"      
#> [5] "item:prompt"          "evaluator:prompt:run"

This declares six random sources: item, evaluator, prompt, item-by-evaluator, item-by-prompt, and evaluator-by-prompt-by-run. The run group is parent-scoped even though the coded panel is complete. Item-by-run and additional evaluator-by-prompt effects are not silently added. Omitting them is an assumption of this known simulation; a real study needs its own justification. The exact Gaussian engine currently requires a complete coded Cartesian panel with one row per full cell.

source_sd <- c(item = 1.2, evaluator = 0.4, prompt = 0.3,
  "item:evaluator" = 0.5, "item:prompt" = 0.4,
  "evaluator:prompt:run" = 0.25)
llm_data$quality <- 0
for (source in names(source_sd)) {
  members <- strsplit(source, ":", fixed = TRUE)[[1L]]
  group <- do.call(interaction, c(llm_data[members], list(drop = TRUE)))
  effects <- rnorm(nlevels(group), sd = source_sd[[source]])
  llm_data$quality <- llm_data$quality + effects[as.integer(group)]
}
llm_data$quality <- llm_data$quality + rnorm(nrow(llm_data), sd = 0.6)

2. Inspect feasibility before fitting

gt_preflight() reports observed counts, source dimensions, free parameters, configured size limits, and the supported reliability scale, without building a model matrix or running an optimizer.

preflight <- gt_preflight(llm_data, "quality", llm_design)
print(preflight)
#> G-theory preflight: 576 rows | exact_balanced_gaussian 
#> Observed levels: item=24, evaluator=4, prompt=3, run=2 
#> Retained random sources: 6  | Random dimensions: 223 | Model parameters: 8 
#> Structural/resource checks: PASS 
#> Panel cells (observed coded levels):
#>            status cell_count_exact
#>          complete              576
#>  under_replicated                0
#>   over_replicated                0
#>           missing                0
#> Audit examples returned: missing_cells=0, replication_issues=0, row_examples=0 | Per-table limit: 10 
#> Outcome profiles: 1 | See outcome_profile for category coverage, numeric summaries, and descriptive group variation.
#> Supported reliability scale: observed 
#> Reliability of observed Gaussian-score averages. 
#> Preflight checks structure and configured size limits. It does not establish identification, approximation adequacy, fit acceptance, or scientific validity.
preflight$sources
#>                 source observed_groups predictor_dimensions random_dimension covariance_parameters
#> 1            evaluator               4                    1                4                     1
#> 2                 item              24                    1               24                     1
#> 3               prompt               3                    1                3                     1
#> 4       item:evaluator              96                    1               96                     1
#> 5          item:prompt              72                    1               72                     1
#> 6 evaluator:prompt:run              24                    1               24                     1
stopifnot(preflight$fitting_feasible)

Passing this report means that the listed structural and resource checks pass. It does not prove identification, reliable estimation, numerical acceptance, or approximation adequacy. Read failed checks before changing a resource limit or simplifying a scientific model.

The panel audit counts missing cells and cells above or below the declared replication. It retains only a bounded set of examples, using the observed coded levels; it cannot discover a level that never appeared in the data.

preflight$panel_audit$summary
#>             status cell_count cell_count_exact
#> 1         complete        576              576
#> 2 under_replicated          0                0
#> 3  over_replicated          0                0
#> 4          missing          0                0
plot(preflight, type = "cells")

incomplete <- gt_preflight(llm_data[-1, ], "quality", llm_design)
incomplete$panel_audit$missing_cells
#>   item evaluator   prompt run
#> 1    1   model_1 prompt_1   1

These diagnostics do not fill cells or remove records. Investigate unexpected codes or missing measurements before deciding how the study should be analysed.

3. Fit and inspect acceptance

fit <- gt_fit(llm_data, "quality", llm_design,
  control = gt_control(gaussian = list(optimizer = "CSOLNP",
    extra_tries = 2L, retry_seed = 415L, threads = 1L)))
print(summary(fit))
#> G-theory fit summary: 1 outcome(s), 576 measurement rows
#> Outcomes: quality 
#> Families: gaussian | Estimator: REML 
#> Random sources: 6 | Instrumentation facets: 3 
#> Optimizer: CSOLNP | Attempts: 1 | Selected attempt: 1 
#> Optimizer completed: TRUE | Numerically accepted: TRUE 
#> Likelihood approximation: exact_balanced_gaussian_likelihood 
#> Boundary or nearly singular sources: none recorded 
#> -2 log likelihood: 1411 
#> 
#> Source variances (Gaussian observed-score covariance):
#>                source   trait variance std_error at_boundary
#>             evaluator quality  0.16915   0.15491       FALSE
#>                  item quality  1.95564   0.60959       FALSE
#>                prompt quality  0.06368   0.07587       FALSE
#>        item:evaluator quality  0.24121   0.05231       FALSE
#>           item:prompt quality  0.10283   0.03178       FALSE
#>  evaluator:prompt:run quality  0.04622   0.02085       FALSE
#>              Residual quality  0.38950   0.02707       FALSE
#> 
#> Fixed location or contrast estimates:
#> quality 
#> -0.4553 
#> 
#> Notes:
#> - Nesting parents define instrumentation groups shared across objects; nested sibling interactions are not automatically included.
#> - No requested source was dropped for this observation family.
#> 
#> Use gt_components(fit) for full covariance matrices and gt_diagnostics(fit) for details.
diagnostics <- gt_diagnostics(fit)
print(diagnostics[c("numerically_accepted", "selected_attempt",
                    "acceptance_failures", "approximation_adequacy")])
#> $numerically_accepted
#> [1] TRUE
#> 
#> $selected_attempt
#> [1] "1"
#> 
#> $acceptance_failures
#> NULL
#> 
#> $approximation_adequacy
#> [1] "exact_balanced_gaussian_likelihood"
stopifnot(isTRUE(diagnostics$numerically_accepted))
gt_components(fit)
#> $evaluator
#>           quality
#> quality 0.1691472
#> 
#> $item
#>         quality
#> quality 1.95564
#> 
#> $prompt
#>            quality
#> quality 0.06368045
#> 
#> $`item:evaluator`
#>           quality
#> quality 0.2412094
#> 
#> $`item:prompt`
#>           quality
#> quality 0.1028346
#> 
#> $`evaluator:prompt:run`
#>            quality
#> quality 0.04621689
#> 
#> $Residual
#>           quality
#> quality 0.3895031
#> 
#> attr(,"scale")
#> [1] "Gaussian observed-score covariance"
#> attr(,"interpretation")
#> [1] "Cross-category contrasts within one categorical outcome are not separate measured outcomes."

The Gaussian default estimator is REML. The trial budget allows the initial fit plus two additional attempts; it is not resampling. The chosen optimizer remains fixed across attempts. Inspect numerical acceptance and source-boundary diagnostics before interpreting the components. A completed optimizer alone is insufficient. If an installation does not support the selected optimizer or a fit is rejected, address that explicitly.

The summary table also carries an asymptotic standard error for each source variance, obtained from the numerically differentiated restricted-likelihood Hessian. A component resting on a variance boundary is flagged, because Wald theory does not apply there.

summary(fit)$variances
#>                 source   trait   variance  std_error at_boundary
#> 1            evaluator quality 0.16914721 0.15491014       FALSE
#> 2                 item quality 1.95564029 0.60958554       FALSE
#> 3               prompt quality 0.06368045 0.07587063       FALSE
#> 4       item:evaluator quality 0.24120936 0.05231458       FALSE
#> 5          item:prompt quality 0.10283459 0.03177643       FALSE
#> 6 evaluator:prompt:run quality 0.04621689 0.02084544       FALSE
#> 7             Residual quality 0.38950314 0.02707234       FALSE

4. Compute relative and absolute reliability

reliability <- gt_reliability(fit)
reliability
#> G-theory reliability | observed scale
#> Facet counts: evaluator=4, prompt=3, run=2 
#>  outcome  Erho2 Erho2_se Erho2_lower Erho2_upper    Phi  Phi_se Phi_lower Phi_upper
#>  quality 0.9464  0.01778      0.8988      0.9723 0.9173 0.03181    0.8299    0.9619
#> Erho2: relative comparisons; Phi: absolute decisions.
#> 95% intervals: delta method on the logit scale from the fitted parameter covariance.
#> Full covariance matrices and fit diagnostics remain in the returned object.
reliability$interpretation
#> [1] "Gaussian observed-score coefficient"

Erho2 is the relative coefficient G: common evaluator, prompt, and run shifts do not change relative item comparisons. Phi is the absolute coefficient and also includes those shifts as error. Item-dependent interactions contribute to both error definitions. Here both coefficients refer to observed Gaussian-score averages under the declared design. They measure consistency under that model, not accuracy against human reference labels. A coefficient of 0.80 means that 80% of the variance in item means is universe-score variance under this model; it is not 80% labeling accuracy.

The interval is a delta-method Wald interval on the logit scale, propagating the fitted parameter covariance matrix through the source covariances. It describes estimation uncertainty in those covariances under the declared model. It is not a guarantee about a future sample of evaluators or prompts.

plot(reliability, target = 0.80,
  main = "Reliability of the declared mean quality score")

reliability_table <- as.data.frame(reliability)

The dashed reference is an illustrative target. Missing intervals are shown as unavailable, never as a claim of perfect precision. The table records the scale, uncertainty status and any conditional inference, alongside the estimates.

5. Compare measurement budgets

grid <- expand.grid(evaluator = c(2L, 4L, 6L),
  prompt = c(2L, 3L, 4L), run = c(1L, 2L))
dstudy <- gt_dstudy(fit, grid)
planning <- as.data.frame(dstudy)
# Prefixed allocation columns avoid collisions with outcome/estimate names.
# These readable aliases are just for this tutorial's known facet names.
planning$evaluator <- planning$allocation_evaluator
planning$prompt <- planning$allocation_prompt
planning$run <- planning$allocation_run
planning$measurements_per_item <- planning$measurements_per_object
planning <- planning[order(planning$measurements_per_item, -planning$Phi), ]
rownames(planning) <- NULL
print(planning)
#>    design_id    kind outcome     Erho2   Erho2_se Erho2_lower Erho2_upper       Phi     Phi_se
#> 1          1 outcome quality 0.8789244 0.03560952   0.7902498   0.9332760 0.8311242 0.05452953
#> 2          4 outcome quality 0.8989630 0.03088554   0.8204331   0.9454336 0.8543855 0.05049529
#> 3          2 outcome quality 0.9241948 0.02382082   0.8622785   0.9595798 0.8905661 0.03849659
#> 4          7 outcome quality 0.9093288 0.02842758   0.8361284   0.9517192 0.8665114 0.04853116
#> 5         10 outcome quality 0.8985872 0.03138008   0.8185730   0.9456557 0.8508181 0.05218882
#> 6          3 outcome quality 0.9403393 0.01948089   0.8886455   0.9688759 0.9123156 0.03265003
#> 7          5 outcome quality 0.9390021 0.01957600   0.8873654   0.9678246 0.9095813 0.03309628
#> 8         13 outcome quality 0.9125791 0.02788504   0.8403014   0.9539379 0.8681573 0.04884324
#> 9          8 outcome quality 0.9465851 0.01741097   0.9002366   0.9720689 0.9193968 0.03047053
#> 10        11 outcome quality 0.9349508 0.02127879   0.8786388   0.9661408 0.9017489 0.03673160
#> 11        16 outcome quality 0.9197397 0.02612292   0.8513494   0.9582099 0.8770947 0.04728321
#> 12         6 outcome quality 0.9531530 0.01545102   0.9117091   0.9756624 0.9295996 0.02665837
#> 13         9 outcome quality 0.9596917 0.01340455   0.9235008   0.9791477 0.9384895 0.02371661
#> 14        12 outcome quality 0.9477350 0.01771580   0.8999582   0.9733703 0.9201084 0.03137259
#> 15        14 outcome quality 0.9463767 0.01777804   0.8988071   0.9722742 0.9173273 0.03180532
#> 16        17 outcome quality 0.9521950 0.01603064   0.9089930   0.9754426 0.9253201 0.02947603
#> 17        15 outcome quality 0.9582059 0.01419569   0.9196465   0.9786904 0.9349788 0.02569701
#> 18        18 outcome quality 0.9635285 0.01243621   0.9295932   0.9814341 0.9425957 0.02296281
#>    Phi_lower Phi_upper allocation_evaluator allocation_prompt allocation_run    scale
#> 1  0.6968108 0.9133367                    2                 2              1 observed
#> 2  0.7259000 0.9285695                    2                 3              1 observed
#> 3  0.7895704 0.9463806                    4                 2              1 observed
#> 4  0.7404140 0.9366003                    2                 4              1 observed
#> 5  0.7181184 0.9273660                    2                 2              2 observed
#> 6  0.8237973 0.9586001                    6                 2              1 observed
#> 7  0.8205098 0.9567797                    4                 3              1 observed
#> 8  0.7404664 0.9382622                    2                 3              2 observed
#> 9  0.8359359 0.9623144                    4                 4              1 observed
#> 10 0.8028545 0.9538842                    4                 2              2 observed
#> 11 0.7512927 0.9440057                    2                 4              2 observed
#> 12 0.8559650 0.9670397                    6                 3              1 observed
#> 13 0.8721195 0.9715377                    6                 4              1 observed
#> 14 0.8330411 0.9637470                    6                 2              2 observed
#> 15 0.8298542 0.9618948                    4                 3              2 observed
#> 16 0.8430236 0.9662015                    4                 4              2 observed
#> 17 0.8626345 0.9705244                    6                 3              2 observed
#> 18 0.8772614 0.9741760                    6                 4              2 observed
#>    interval_level uncertainty_available uncertainty_reason uncertainty_conditional
#> 1            0.95                  TRUE               <NA>                   FALSE
#> 2            0.95                  TRUE               <NA>                   FALSE
#> 3            0.95                  TRUE               <NA>                   FALSE
#> 4            0.95                  TRUE               <NA>                   FALSE
#> 5            0.95                  TRUE               <NA>                   FALSE
#> 6            0.95                  TRUE               <NA>                   FALSE
#> 7            0.95                  TRUE               <NA>                   FALSE
#> 8            0.95                  TRUE               <NA>                   FALSE
#> 9            0.95                  TRUE               <NA>                   FALSE
#> 10           0.95                  TRUE               <NA>                   FALSE
#> 11           0.95                  TRUE               <NA>                   FALSE
#> 12           0.95                  TRUE               <NA>                   FALSE
#> 13           0.95                  TRUE               <NA>                   FALSE
#> 14           0.95                  TRUE               <NA>                   FALSE
#> 15           0.95                  TRUE               <NA>                   FALSE
#> 16           0.95                  TRUE               <NA>                   FALSE
#> 17           0.95                  TRUE               <NA>                   FALSE
#> 18           0.95                  TRUE               <NA>                   FALSE
#>    uncertainty_conditioning Erho2_interval_available Phi_interval_available fixed_facets
#> 1                      <NA>                     TRUE                   TRUE             
#> 2                      <NA>                     TRUE                   TRUE             
#> 3                      <NA>                     TRUE                   TRUE             
#> 4                      <NA>                     TRUE                   TRUE             
#> 5                      <NA>                     TRUE                   TRUE             
#> 6                      <NA>                     TRUE                   TRUE             
#> 7                      <NA>                     TRUE                   TRUE             
#> 8                      <NA>                     TRUE                   TRUE             
#> 9                      <NA>                     TRUE                   TRUE             
#> 10                     <NA>                     TRUE                   TRUE             
#> 11                     <NA>                     TRUE                   TRUE             
#> 12                     <NA>                     TRUE                   TRUE             
#> 13                     <NA>                     TRUE                   TRUE             
#> 14                     <NA>                     TRUE                   TRUE             
#> 15                     <NA>                     TRUE                   TRUE             
#> 16                     <NA>                     TRUE                   TRUE             
#> 17                     <NA>                     TRUE                   TRUE             
#> 18                     <NA>                     TRUE                   TRUE             
#>    composite_weights measurements_per_object measurement_count_exact extrapolated batch_status
#> 1               <NA>                       4                    TRUE        FALSE not_declared
#> 2               <NA>                       6                    TRUE        FALSE not_declared
#> 3               <NA>                       8                    TRUE        FALSE not_declared
#> 4               <NA>                       8                    TRUE         TRUE not_declared
#> 5               <NA>                       8                    TRUE        FALSE not_declared
#> 6               <NA>                      12                    TRUE         TRUE not_declared
#> 7               <NA>                      12                    TRUE        FALSE not_declared
#> 8               <NA>                      12                    TRUE        FALSE not_declared
#> 9               <NA>                      16                    TRUE         TRUE not_declared
#> 10              <NA>                      16                    TRUE        FALSE not_declared
#> 11              <NA>                      16                    TRUE         TRUE not_declared
#> 12              <NA>                      18                    TRUE         TRUE not_declared
#> 13              <NA>                      24                    TRUE         TRUE not_declared
#> 14              <NA>                      24                    TRUE         TRUE not_declared
#> 15              <NA>                      24                    TRUE        FALSE not_declared
#> 16              <NA>                      32                    TRUE         TRUE not_declared
#> 17              <NA>                      36                    TRUE         TRUE not_declared
#> 18              <NA>                      48                    TRUE         TRUE not_declared
#>    evaluator prompt run measurements_per_item
#> 1          2      2   1                     4
#> 2          2      3   1                     6
#> 3          4      2   1                     8
#> 4          2      4   1                     8
#> 5          2      2   2                     8
#> 6          6      2   1                    12
#> 7          4      3   1                    12
#> 8          2      3   2                    12
#> 9          4      4   1                    16
#> 10         4      2   2                    16
#> 11         2      4   2                    16
#> 12         6      3   1                    18
#> 13         6      4   1                    24
#> 14         6      2   2                    24
#> 15         4      3   2                    24
#> 16         4      4   2                    32
#> 17         6      3   2                    36
#> 18         6      4   2                    48

Runs are counted within each evaluator/prompt pair, so the budget per item is evaluators x prompts x runs. A row using more levels than observed extrapolates the fitted variance model to the declared exchangeable population.

screen <- gt_dstudy_target(dstudy, target = 0.80, coefficient = "Phi",
  outcome = "quality")
screen[screen$meets_target %in% TRUE, ]
#>    design_id    kind outcome     Erho2   Erho2_se Erho2_lower Erho2_upper       Phi     Phi_se
#> 1          1 outcome quality 0.8789244 0.03560952   0.7902498   0.9332760 0.8311242 0.05452953
#> 2          2 outcome quality 0.9241948 0.02382082   0.8622785   0.9595798 0.8905661 0.03849659
#> 3          3 outcome quality 0.9403393 0.01948089   0.8886455   0.9688759 0.9123156 0.03265003
#> 4          4 outcome quality 0.8989630 0.03088554   0.8204331   0.9454336 0.8543855 0.05049529
#> 5          5 outcome quality 0.9390021 0.01957600   0.8873654   0.9678246 0.9095813 0.03309628
#> 6          6 outcome quality 0.9531530 0.01545102   0.9117091   0.9756624 0.9295996 0.02665837
#> 7          7 outcome quality 0.9093288 0.02842758   0.8361284   0.9517192 0.8665114 0.04853116
#> 8          8 outcome quality 0.9465851 0.01741097   0.9002366   0.9720689 0.9193968 0.03047053
#> 9          9 outcome quality 0.9596917 0.01340455   0.9235008   0.9791477 0.9384895 0.02371661
#> 10        10 outcome quality 0.8985872 0.03138008   0.8185730   0.9456557 0.8508181 0.05218882
#> 11        11 outcome quality 0.9349508 0.02127879   0.8786388   0.9661408 0.9017489 0.03673160
#> 12        12 outcome quality 0.9477350 0.01771580   0.8999582   0.9733703 0.9201084 0.03137259
#> 13        13 outcome quality 0.9125791 0.02788504   0.8403014   0.9539379 0.8681573 0.04884324
#> 14        14 outcome quality 0.9463767 0.01777804   0.8988071   0.9722742 0.9173273 0.03180532
#> 15        15 outcome quality 0.9582059 0.01419569   0.9196465   0.9786904 0.9349788 0.02569701
#> 16        16 outcome quality 0.9197397 0.02612292   0.8513494   0.9582099 0.8770947 0.04728321
#> 17        17 outcome quality 0.9521950 0.01603064   0.9089930   0.9754426 0.9253201 0.02947603
#> 18        18 outcome quality 0.9635285 0.01243621   0.9295932   0.9814341 0.9425957 0.02296281
#>    Phi_lower Phi_upper allocation_evaluator allocation_prompt allocation_run    scale
#> 1  0.6968108 0.9133367                    2                 2              1 observed
#> 2  0.7895704 0.9463806                    4                 2              1 observed
#> 3  0.8237973 0.9586001                    6                 2              1 observed
#> 4  0.7259000 0.9285695                    2                 3              1 observed
#> 5  0.8205098 0.9567797                    4                 3              1 observed
#> 6  0.8559650 0.9670397                    6                 3              1 observed
#> 7  0.7404140 0.9366003                    2                 4              1 observed
#> 8  0.8359359 0.9623144                    4                 4              1 observed
#> 9  0.8721195 0.9715377                    6                 4              1 observed
#> 10 0.7181184 0.9273660                    2                 2              2 observed
#> 11 0.8028545 0.9538842                    4                 2              2 observed
#> 12 0.8330411 0.9637470                    6                 2              2 observed
#> 13 0.7404664 0.9382622                    2                 3              2 observed
#> 14 0.8298542 0.9618948                    4                 3              2 observed
#> 15 0.8626345 0.9705244                    6                 3              2 observed
#> 16 0.7512927 0.9440057                    2                 4              2 observed
#> 17 0.8430236 0.9662015                    4                 4              2 observed
#> 18 0.8772614 0.9741760                    6                 4              2 observed
#>    interval_level uncertainty_available uncertainty_reason uncertainty_conditional
#> 1            0.95                  TRUE               <NA>                   FALSE
#> 2            0.95                  TRUE               <NA>                   FALSE
#> 3            0.95                  TRUE               <NA>                   FALSE
#> 4            0.95                  TRUE               <NA>                   FALSE
#> 5            0.95                  TRUE               <NA>                   FALSE
#> 6            0.95                  TRUE               <NA>                   FALSE
#> 7            0.95                  TRUE               <NA>                   FALSE
#> 8            0.95                  TRUE               <NA>                   FALSE
#> 9            0.95                  TRUE               <NA>                   FALSE
#> 10           0.95                  TRUE               <NA>                   FALSE
#> 11           0.95                  TRUE               <NA>                   FALSE
#> 12           0.95                  TRUE               <NA>                   FALSE
#> 13           0.95                  TRUE               <NA>                   FALSE
#> 14           0.95                  TRUE               <NA>                   FALSE
#> 15           0.95                  TRUE               <NA>                   FALSE
#> 16           0.95                  TRUE               <NA>                   FALSE
#> 17           0.95                  TRUE               <NA>                   FALSE
#> 18           0.95                  TRUE               <NA>                   FALSE
#>    uncertainty_conditioning Erho2_interval_available Phi_interval_available fixed_facets
#> 1                      <NA>                     TRUE                   TRUE             
#> 2                      <NA>                     TRUE                   TRUE             
#> 3                      <NA>                     TRUE                   TRUE             
#> 4                      <NA>                     TRUE                   TRUE             
#> 5                      <NA>                     TRUE                   TRUE             
#> 6                      <NA>                     TRUE                   TRUE             
#> 7                      <NA>                     TRUE                   TRUE             
#> 8                      <NA>                     TRUE                   TRUE             
#> 9                      <NA>                     TRUE                   TRUE             
#> 10                     <NA>                     TRUE                   TRUE             
#> 11                     <NA>                     TRUE                   TRUE             
#> 12                     <NA>                     TRUE                   TRUE             
#> 13                     <NA>                     TRUE                   TRUE             
#> 14                     <NA>                     TRUE                   TRUE             
#> 15                     <NA>                     TRUE                   TRUE             
#> 16                     <NA>                     TRUE                   TRUE             
#> 17                     <NA>                     TRUE                   TRUE             
#> 18                     <NA>                     TRUE                   TRUE             
#>    composite_weights measurements_per_object measurement_count_exact extrapolated batch_status
#> 1               <NA>                       4                    TRUE        FALSE not_declared
#> 2               <NA>                       8                    TRUE        FALSE not_declared
#> 3               <NA>                      12                    TRUE         TRUE not_declared
#> 4               <NA>                       6                    TRUE        FALSE not_declared
#> 5               <NA>                      12                    TRUE        FALSE not_declared
#> 6               <NA>                      18                    TRUE         TRUE not_declared
#> 7               <NA>                       8                    TRUE         TRUE not_declared
#> 8               <NA>                      16                    TRUE         TRUE not_declared
#> 9               <NA>                      24                    TRUE         TRUE not_declared
#> 10              <NA>                       8                    TRUE        FALSE not_declared
#> 11              <NA>                      16                    TRUE        FALSE not_declared
#> 12              <NA>                      24                    TRUE         TRUE not_declared
#> 13              <NA>                      12                    TRUE        FALSE not_declared
#> 14              <NA>                      24                    TRUE        FALSE not_declared
#> 15              <NA>                      36                    TRUE         TRUE not_declared
#> 16              <NA>                      16                    TRUE         TRUE not_declared
#> 17              <NA>                      32                    TRUE         TRUE not_declared
#> 18              <NA>                      48                    TRUE         TRUE not_declared
#>    coefficient target  estimate meets_target fewest_measurements
#> 1          Phi    0.8 0.8311242         TRUE                TRUE
#> 2          Phi    0.8 0.8905661         TRUE               FALSE
#> 3          Phi    0.8 0.9123156         TRUE               FALSE
#> 4          Phi    0.8 0.8543855         TRUE               FALSE
#> 5          Phi    0.8 0.9095813         TRUE               FALSE
#> 6          Phi    0.8 0.9295996         TRUE               FALSE
#> 7          Phi    0.8 0.8665114         TRUE               FALSE
#> 8          Phi    0.8 0.9193968         TRUE               FALSE
#> 9          Phi    0.8 0.9384895         TRUE               FALSE
#> 10         Phi    0.8 0.8508181         TRUE               FALSE
#> 11         Phi    0.8 0.9017489         TRUE               FALSE
#> 12         Phi    0.8 0.9201084         TRUE               FALSE
#> 13         Phi    0.8 0.8681573         TRUE               FALSE
#> 14         Phi    0.8 0.9173273         TRUE               FALSE
#> 15         Phi    0.8 0.9349788         TRUE               FALSE
#> 16         Phi    0.8 0.8770947         TRUE               FALSE
#> 17         Phi    0.8 0.9253201         TRUE               FALSE
#> 18         Phi    0.8 0.9425957         TRUE               FALSE
screen[screen$fewest_measurements %in% TRUE, ]
#>   design_id    kind outcome     Erho2   Erho2_se Erho2_lower Erho2_upper       Phi     Phi_se
#> 1         1 outcome quality 0.8789244 0.03560952   0.7902498    0.933276 0.8311242 0.05452953
#>   Phi_lower Phi_upper allocation_evaluator allocation_prompt allocation_run    scale interval_level
#> 1 0.6968108 0.9133367                    2                 2              1 observed           0.95
#>   uncertainty_available uncertainty_reason uncertainty_conditional uncertainty_conditioning
#> 1                  TRUE               <NA>                   FALSE                     <NA>
#>   Erho2_interval_available Phi_interval_available fixed_facets composite_weights
#> 1                     TRUE                   TRUE                           <NA>
#>   measurements_per_object measurement_count_exact extrapolated batch_status coefficient target
#> 1                       4                    TRUE        FALSE not_declared         Phi    0.8
#>    estimate meets_target fewest_measurements
#> 1 0.8311242         TRUE                TRUE

A threshold of 0.80 is illustrative; choose and justify your own criterion. Read the interval columns alongside the point projection: several allocations here have overlapping intervals, so the ordering of nearby rows is not established by these data.

The screening helper keeps every supplied candidate, including those that miss the target or have an unavailable estimate. It marks all ties for the fewest measurements among qualifying candidates. It performs no search outside this grid and does not guarantee that a future panel will meet the target.

plot(dstudy, coefficient = "Phi", outcome = "quality", target = 0.80,
  main = "Candidate measurement allocations")

Different allocations can use the same total number of measurements; they are separate points rather than one assumed smooth curve. The figure identifies allocations extending beyond the observed facet counts. Such projections keep the fitted source distributions fixed.

Both result tables contain ordinary scalar columns, suitable for CSV output:

write.csv(reliability_table, "reliability.csv", row.names = FALSE)
write.csv(as.data.frame(dstudy), "decision-study.csv", row.names = FALSE)
pdf("decision-study.pdf", width = 7, height = 5)
plot(dstudy, coefficient = "Phi", outcome = "quality", target = 0.80)
dev.off()

6. Save a portable analysis report

A report collects the design, fitted model context, diagnostic flags, coefficient and decision-study tables, and figures. It does not rerun optimization. Supplied coefficient results are checked against projections from the fit; supplied preflight information is checked for compatible aggregate specifications, which does not prove that it came from identical observations.

analysis_report <- gt_report(fit, preflight = preflight,
  reliability = reliability, dstudy = dstudy)
print(analysis_report)
#> G-theory analysis report | schema 1.0 
#> Backend: OpenMx exact balanced Gaussian factorial contrasts | Numerically accepted: TRUE 
#> Tables: 4 design variables | 1 reliability rows | 18 decision-study rows
#> Export with gt_export_report(report, path).
gt_export_report(analysis_report, "analysis-report.html")

The HTML file contains its figures and styling and can be opened offline. An existing destination is refused unless replacement is explicitly requested. The report omits observations, row examples and group-level identifiers, while keeping variable and source names and aggregate results. This is not an anonymization guarantee. Inspect study-specific names before sharing it. Retained fitting provenance and the report-generation environment are separate; missing fitting provenance is reported as unavailable.

7. Fixed facets

A facet is random when the study generalizes to a population of its levels, and fixed when the universe of generalization is exactly the levels used. LLM evaluation may target exactly a chosen temperature or prompt set, in which case those facets are fixed. Selecting levels deliberately does not by itself decide the universe of generalization.

Declare them with fixed. Following the mixed model of Brennan (2001), the object-by-fixed-facet variance is averaged over that facet’s levels and added to universe-score variance, and a source built only from fixed facets shifts every item equally and leaves the model.

fixed_prompt <- gt_reliability(fit, fixed = "prompt")
fixed_prompt
#> G-theory reliability | observed scale
#> Facet counts: evaluator=4, prompt=3, run=2 
#> Fixed facets: prompt | mixed (Brennan 2001): object-by-fixed-facet variance enters the universe score 
#>  outcome Erho2 Erho2_se Erho2_lower Erho2_upper    Phi  Phi_se Phi_lower Phi_upper
#>  quality 0.963  0.01261      0.9286      0.9811 0.9428 0.02464    0.8707    0.9758
#> Erho2: relative comparisons; Phi: absolute decisions.
#> 95% intervals: delta method on the logit scale from the fitted parameter covariance.
#> Full covariance matrices and fit diagnostics remain in the returned object.
fixed_prompt$source_roles
#>                                evaluator                                     item 
#>                         "absolute_error"                               "universe" 
#>                                   prompt                           item:evaluator 
#> "dropped_fixed_instrumentation_constant"            "relative_and_absolute_error" 
#>                              item:prompt                     evaluator:prompt:run 
#>   "universe_after_fixed_facet_averaging"                         "absolute_error" 
#>                                 Residual 
#>            "relative_and_absolute_error"

Treating the prompt set as fixed raises both coefficients, because item-by-prompt variation is now part of what the study is trying to measure rather than error. A fixed facet’s count cannot be changed, and a decision study may not project over it.

fixed acts at the reliability stage: it changes how the already-fitted components are aggregated into a coefficient. It does not change the G study, so it cannot repair a source whose variance was misspecified during fitting. Observed seed disagreement can change across temperatures even when latent variance is constant, because category probabilities can change. If diagnostics raise doubts about pooling temperatures, fitting within one temperature is a useful sensitivity analysis for a conditional estimand; declaring a facet fixed is a separate decision about generalization.

8. Match other outcomes to their observation model

Outcome Current scalar reliability support
Gaussian Observed continuous-score averages.
Binary or ordinal Explicit scale = "latent" after an accepted fit. This describes latent-response averages, not majority votes, label proportions, or observed ordinal-score averages.
Unordered categorical No implemented default scalar G/Phi. A scientific score or category-probability estimand is needed.

Declare category order and references explicitly. Do not recode binary, ordinal, or nominal labels as the continuous quality score used above. Discrete fitting uses a bounded dense Laplace engine with Gaussian latent random effects; gt_preflight() reports its current limits, and changing the link does not remove the random-effects distribution assumption. That engine reports point estimates only: it computes no standard errors. Analytic reliability and D studies require complete balanced coded panels even when discrete fitting admits missing whole cells.

At one observation per object-by-facets cell, a discrete fit cannot identify the full-cell source that the default design requests. Declare the same design without it:

gt_design("item", "rater", full_cell = FALSE)$terms_requested
#> [1] "item"  "rater"

A complete panel can still contain little observed outcome variation. This small, deliberately imbalanced example has one positive response. It is a data profile, not an estimation example or a recommended design.

information_data <- expand.grid(item = seq_len(6), rater = seq_len(3))
information_data$label <- as.integer(information_data$item == 1 &
  information_data$rater == 1)
information_design <- gt_design("item", "rater", random = ~ item + rater)
information <- gt_preflight(information_data, "label", information_design,
  gt_family("binary"), max_examples = 0)
information$outcome_profile$label$categories
#>   category count proportion observed
#> 1        0    17 0.94444444     TRUE
#> 2        1     1 0.05555556     TRUE
information$outcome_profile$label$by_variable
#>   variable        grouping_scope total_groups no_variation_groups single_row_groups
#> 1     item marginal_coded_levels            6                   5                 0
#> 2    rater marginal_coded_levels            3                   2                 0
#>   no_variation_proportion
#> 1               0.8333333
#> 2               0.6666667
plot(information, type = "outcomes", outcome = "label")

The heatmap describes the fraction of groups in which each category appears. A group with one observation necessarily has no observed variation. These summaries supply no minimum-count threshold, independence claim or guarantee of parameter recovery. Declared categories absent from the entire panel are shown with zero counts and block fitting; the preflight report does not merge them or silently change their order. No finite panel can reveal a category never observed and never declared.