This tutorial illustrates the study-planning workflows in source version 0.5.0. All observations and monetary amounts below are synthetic. No requests are executed and no provider prices are retrieved.
| Quantity | Purpose |
|---|---|
| Distinct pilot items | Information for estimating source variation and coefficients |
| Annotations per item | The averaging protocol whose reliability is projected |
| Items per request | Shared context and request overhead; not additional repeated measurements |
The first example uses independent-item annotations with item, rater and item-by-rater sources. Runs supply replication of those sources through distinct coded cells. This working model defines the example’s intended score universe.
library(Gtheory4LLM)
set.seed(2718)
panel <- expand.grid(item = 1:24, rater = 1:4, run = 1:3,
KEEP.OUT.ATTRS = FALSE)
item_rater <- as.integer(interaction(panel$item, panel$rater, drop = TRUE))
panel$quality <- rnorm(24)[panel$item] + rnorm(4, sd = 0.5)[panel$rater] +
rnorm(96, sd = 0.4)[item_rater] + rnorm(nrow(panel), sd = 0.8)
design <- gt_design("item", c("rater", "run"),
random = ~ item + rater + item:rater)
control <- gt_control(gaussian = list(extra_tries = 2L, retry_seed = 439L))
fit <- gt_fit(panel, "quality", design, control = control)
stopifnot(isTRUE(fit$numerically_accepted))
gt_reliability(fit)$per_trait
#> outcome Erho2 Erho2_se Erho2_lower Erho2_upper Phi Phi_se Phi_lower Phi_upper
#> 1 quality 0.8954299 0.03560501 0.8025272 0.9474857 0.864711 0.04998046 0.7345082 0.9365736Acceptance permits the supported projections. It does not establish a global optimum, adequate pilot precision, labeling accuracy or correct source selection for another study.
Here one annotation costs 0.01 illustrative dollars, with a setup cost of two dollars per rater. The cost callback could instead return separate request, token, human-review and expected-retry costs, provided all assumptions use the same currency and are recorded. Costs do not introduce a statistical benefit from review or change the population of raters represented by the fit.
plan <- gt_plan(fit,
candidates = list(rater = c(2, 4, 6), run = c(1, 3, 5)),
cost = function(d) list(
costs = data.frame(annotation = 0.01 * d$total_annotations,
setup = 2 * d$allocation_rater),
resources = data.frame(requests = d$total_annotations),
assumptions = "Illustrative prices; one item per request."),
currency = "USD", cost_date = "2026-10-09", study_items = 1000,
target = 0.80, coefficient = "Phi", baseline = 1L,
constraints = function(d) data.frame(budget = d$total_cost <= 180))
as.data.frame(plan)[, c("design_id", "total_cost", "screening_value",
"feasible", "minimum_cost", "pareto")]
#> design_id total_cost screening_value feasible minimum_cost pareto
#> 1 1 24 0.6486171 TRUE FALSE TRUE
#> 2 2 48 0.7868620 TRUE FALSE TRUE
#> 3 3 72 0.8470409 TRUE TRUE TRUE
#> 4 4 64 0.7616660 TRUE FALSE FALSE
#> 5 5 128 0.8647110 TRUE FALSE TRUE
#> 6 6 192 0.9055479 FALSE FALSE FALSE
#> 7 7 104 0.7891754 TRUE FALSE FALSE
#> 8 8 208 0.8821666 FALSE FALSE FALSE
#> 9 9 312 0.9182328 FALSE FALSE FALSE
head(plan$source_contributions)
#> design_id source role universe relative_error absolute_error
#> 1 1 item universe 0.900273 0.0000000 0.00000000
#> 2 1 rater absolute_error 0.000000 0.0000000 0.07143433
#> 3 1 item:rater relative_and_absolute_error 0.000000 0.1072665 0.10726651
#> 4 1 Residual relative_and_absolute_error 0.000000 0.3090147 0.30901465
#> 5 2 item universe 0.900273 0.0000000 0.00000000
#> 6 2 rater absolute_error 0.000000 0.0000000 0.03571716
#> relative_error_reduction absolute_error_reduction
#> 1 0 0.00000000
#> 2 0 0.00000000
#> 3 0 0.00000000
#> 4 0 0.00000000
#> 5 0 0.00000000
#> 6 0 0.03571716Minimum cost means minimum among the supplied candidates that meet
both the constraints and the selected reliability screen. All candidates
remain in the result, including cost ties and infeasible rows. Pareto
candidates have no feasible cheaper-or-equal alternative with
equal-or-better performance and at least one strict improvement.
screening = "lower" separately uses an available lower
confidence bound. It never replaces a missing bound with a point
estimate; these are pointwise intervals, not post-search or future-panel
guarantees.
study_items changes total cost, not the per-item
coefficient. Every candidate retains the same declared random/fixed
universe. In particular, the planner cannot make a protocol appear more
dependable by changing a random facet into a fixed one.
The next synthetic panel adds a shared shift for each call of six items. Every condition uses the same grouping. A recorded call identifier allows the fit to retain that layout even when rows are reordered.
batched <- panel
group <- (batched$item - 1L) %/% 6L + 1L
call_index <- as.integer(interaction(group, batched$rater, batched$run, drop = TRUE))
batched$call_id <- paste0("call_", call_index)
batched$quality <- batched$quality + rnorm(max(call_index), sd = 0.7)[call_index]
batch_design <- gt_design("item", c("rater", "run"),
random = ~ item + rater + item:rater, batch = gt_batch(6, id = "call_id"))
batch_fit <- gt_fit(batched, "quality", batch_design, control = control)
stopifnot(isTRUE(batch_fit$numerically_accepted))
within <- c("1" = 1, "2" = -1)
between <- c("1" = 1, "7" = -1)
corpus <- setNames(rep(1 / 24, 24), as.character(1:24))
batch_projection <- gt_batch_reliability(batch_fit,
contrasts = list(within = within, between = between), aggregate = list(corpus = corpus))
subset(as.data.frame(batch_projection), target_kind != "absolute_item")
#> target target_kind outcome outcome_kind universe_variance error_variance
#> 25 contrast:within contrast quality outcome 1.57975108 0.21306534
#> 26 contrast:between contrast quality outcome 1.57975108 0.30603717
#> 27 aggregate:corpus aggregate quality outcome 0.03291148 0.03200569
#> error_sd coefficient coefficient_name
#> 25 0.4615900 0.8811561 contrast_reliability
#> 26 0.5532063 0.8377139 contrast_reliability
#> 27 0.1789013 0.5069765 aggregate_dependability
subset(batch_projection$source_contributions,
source == "Call" & target_kind != "absolute_item")
#> target target_kind outcome outcome_kind source role kernel contraction
#> 124 contrast:within contrast quality outcome Call error same_batch 0.00
#> 129 contrast:between contrast quality outcome Call error same_batch 2.00
#> 134 aggregate:corpus aggregate quality outcome Call error same_batch 0.25
#> divisor variance
#> 124 12 0.00000000
#> 129 12 0.09297184
#> 134 12 0.01162148A shared call shift cancels from the same-batch difference, remains in the between-batch difference, and contributes to each absolute item score. A corpus mean has another error target: call shifts average over calls, while common rater shifts remain. Weighted aggregates can depend on how weights are spread across groups. None of these effects can be represented by one universal batch penalty.
batch_study <- gt_batch_dstudy(batch_fit,
expand.grid(rater = c(2, 4), run = c(1, 3)),
contrasts = list(within = within, between = between), aggregate = list(corpus = corpus))
subset(as.data.frame(batch_study), target_kind == "contrast")
#> design_id target target_kind outcome outcome_kind universe_variance
#> 1.24 1 contrast:within contrast quality outcome 1.579751
#> 1.25 1 contrast:between contrast quality outcome 1.579751
#> 2.24 2 contrast:within contrast quality outcome 1.579751
#> 2.25 2 contrast:between contrast quality outcome 1.579751
#> 3.24 3 contrast:within contrast quality outcome 1.579751
#> 3.25 3 contrast:between contrast quality outcome 1.579751
#> 4.24 4 contrast:within contrast quality outcome 1.579751
#> 4.25 4 contrast:between contrast quality outcome 1.579751
#> error_variance error_sd coefficient coefficient_name allocation_rater allocation_run
#> 1.24 0.8430954 0.9182022 0.6520228 contrast_reliability 2 1
#> 1.25 1.4009264 1.1836074 0.5299973 contrast_reliability 2 1
#> 2.24 0.4215477 0.6492670 0.7893629 contrast_reliability 4 1
#> 2.25 0.7004632 0.8369368 0.6928082 contrast_reliability 4 1
#> 3.24 0.4261307 0.6527868 0.7875594 contrast_reliability 2 3
#> 3.25 0.6120743 0.7823518 0.7207468 contrast_reliability 2 3
#> 4.24 0.2130653 0.4615900 0.8811561 contrast_reliability 4 3
#> 4.25 0.3060372 0.5532063 0.8377139 contrast_reliability 4 3
#> pilot_items batch_size annotations_per_item projected_calls extrapolated scale
#> 1.24 24 6 2 8 FALSE observed
#> 1.25 24 6 2 8 FALSE observed
#> 2.24 24 6 4 16 FALSE observed
#> 2.25 24 6 4 16 FALSE observed
#> 3.24 24 6 6 24 FALSE observed
#> 3.25 24 6 6 24 FALSE observed
#> 4.24 24 6 12 48 FALSE observed
#> 4.25 24 6 12 48 FALSE observed
#> uncertainty_available uncertainty_reason
#> 1.24 FALSE Point projections only; sampling intervals are not implemented.
#> 1.25 FALSE Point projections only; sampling intervals are not implemented.
#> 2.24 FALSE Point projections only; sampling intervals are not implemented.
#> 2.25 FALSE Point projections only; sampling intervals are not implemented.
#> 3.24 FALSE Point projections only; sampling intervals are not implemented.
#> 3.25 FALSE Point projections only; sampling intervals are not implemented.
#> 4.24 FALSE Point projections only; sampling intervals are not implemented.
#> 4.25 FALSE Point projections only; sampling intervals are not implemented.
#> interpretation
#> 1.24 Conditional on the fitted fixed item grouping, with all instrumentation facets random and exchangeable and equal averaging over the planned complete panel. The item main effect defines the universe score. Coefficients describe the explicitly weighted target, not a universal ranking coefficient. Aggregate error is annotation error conditional on these items, not corpus-sampling uncertainty. Counts do not change batch size, grouping, item count, or fitted source covariances.
#> 1.25 Conditional on the fitted fixed item grouping, with all instrumentation facets random and exchangeable and equal averaging over the planned complete panel. The item main effect defines the universe score. Coefficients describe the explicitly weighted target, not a universal ranking coefficient. Aggregate error is annotation error conditional on these items, not corpus-sampling uncertainty. Counts do not change batch size, grouping, item count, or fitted source covariances.
#> 2.24 Conditional on the fitted fixed item grouping, with all instrumentation facets random and exchangeable and equal averaging over the planned complete panel. The item main effect defines the universe score. Coefficients describe the explicitly weighted target, not a universal ranking coefficient. Aggregate error is annotation error conditional on these items, not corpus-sampling uncertainty. Counts do not change batch size, grouping, item count, or fitted source covariances.
#> 2.25 Conditional on the fitted fixed item grouping, with all instrumentation facets random and exchangeable and equal averaging over the planned complete panel. The item main effect defines the universe score. Coefficients describe the explicitly weighted target, not a universal ranking coefficient. Aggregate error is annotation error conditional on these items, not corpus-sampling uncertainty. Counts do not change batch size, grouping, item count, or fitted source covariances.
#> 3.24 Conditional on the fitted fixed item grouping, with all instrumentation facets random and exchangeable and equal averaging over the planned complete panel. The item main effect defines the universe score. Coefficients describe the explicitly weighted target, not a universal ranking coefficient. Aggregate error is annotation error conditional on these items, not corpus-sampling uncertainty. Counts do not change batch size, grouping, item count, or fitted source covariances.
#> 3.25 Conditional on the fitted fixed item grouping, with all instrumentation facets random and exchangeable and equal averaging over the planned complete panel. The item main effect defines the universe score. Coefficients describe the explicitly weighted target, not a universal ranking coefficient. Aggregate error is annotation error conditional on these items, not corpus-sampling uncertainty. Counts do not change batch size, grouping, item count, or fitted source covariances.
#> 4.24 Conditional on the fitted fixed item grouping, with all instrumentation facets random and exchangeable and equal averaging over the planned complete panel. The item main effect defines the universe score. Coefficients describe the explicitly weighted target, not a universal ranking coefficient. Aggregate error is annotation error conditional on these items, not corpus-sampling uncertainty. Counts do not change batch size, grouping, item count, or fitted source covariances.
#> 4.25 Conditional on the fitted fixed item grouping, with all instrumentation facets random and exchangeable and equal averaging over the planned complete panel. The item main effect defines the universe score. Coefficients describe the explicitly weighted target, not a universal ranking coefficient. Aggregate error is annotation error conditional on these items, not corpus-sampling uncertainty. Counts do not change batch size, grouping, item count, or fitted source covariances.The data frame includes allocation_rater,
allocation_run, batch size and uncertainty availability in
ordinary columns. These survive CSV export alongside the target and
coefficient labels. Save the projection object separately when the
complete item grouping and target weights are needed.
These are point projections conditional on the fitted grouping. Facet counts change under exchangeability assumptions; item count, grouping, batch size and fitted source covariances do not. All instrumentation facets remain random. Aggregate error is annotation error conditional on the selected items, not uncertainty from sampling a corpus. Contrast and aggregate ratios have explicit names; they are not universal item-ranking G coefficients. Sampling intervals are not provided.
Ordinary gt_reliability(), gt_dstudy() and
the monetary gt_plan() refuse modelled Call
fits. The separate batch functions make their target explicit. A fit
keeps the compact item-to-batch labels even with data retention
disabled; projection objects contain those labels and target weights.
Portable gt_report(fit) and HTML output exclude them.
More pilot information may improve coefficient estimation without changing the protocol eventually used for every item. The example varies pilot item and rater counts, but every refit estimates the same four-rater, three-run target.
precision <- gt_pilot_plan(fit,
grid = data.frame(n_items = c(16, 24), rater = c(3, 4)),
target_counts = c(rater = 4, run = 3),
nsim = 2, seed = 2026, coefficient = "Phi", width_target = 0.25)
precision$summary
#> design_id attempted accepted preparation_or_fit_errors simulation_errors refused
#> 1 1 2 2 0 0 0
#> 2 2 2 2 0 0 0
#> reliability_errors interval_available interval_unavailable conditional_intervals
#> 1 0 2 0 0
#> 2 0 2 0 0
#> acceptance_rate acceptance_mcse refusal_rate refusal_mcse error_rate error_mcse
#> 1 1 0 0 0 0 0
#> 2 1 0 0 0 0 0
#> interval_rate interval_mcse mean_width_given_interval mean_width_mcse_given_interval
#> 1 1 0 0.1801413 0.03352617
#> 2 1 0 0.2157597 0.03154154
#> width_target precision_successes precision_success_rate precision_success_mcse
#> 1 0.25 2 1 0
#> 2 0.25 2 1 0
#> precision_rate_given_interval precision_mcse_given_interval
#> 1 1 0
#> 2 1 0
precision$replicates[, c("design_id", "replicate_id", "status", "accepted",
"interval_available", "interval_conditional", "width",
"acceptance_failures", "selected_attempt")]
#> design_id replicate_id status accepted interval_available interval_conditional
#> 1 1 1 interval_available TRUE TRUE FALSE
#> 2 1 2 interval_available TRUE TRUE FALSE
#> 3 2 1 interval_available TRUE TRUE FALSE
#> 4 2 2 interval_available TRUE TRUE FALSE
#> width acceptance_failures selected_attempt
#> 1 0.2136675 <NA> 1
#> 2 0.1466152 <NA> 1
#> 3 0.2473012 <NA> 1
#> 4 0.1842181 <NA> 1Two simulations per candidate only exercise the software. They cannot support a pilot-size recommendation. Set a predeclared simulation budget and inspect Monte Carlo uncertainty before using the results for planning. Every attempted replicate remains in the ledger. Precision success requires an available interval meeting the width target, so errors, refusals and unavailable intervals are non-successes. Mean widths condition on an available interval and show that denominator. Conditional boundary intervals retain their qualification. Numerically refused rows retain the fit’s acceptance-failure codes and selected attempt identifier; errors before a fit is returned have unavailable diagnostics. The data and retry seeds allow an individual attempt to be replayed.
The simulator draws new items and new levels of every random facet from a fitted or explicitly supplied Gaussian parameter scenario. It does not calibrate coverage, integrate parameter uncertainty automatically, or provide bootstrap confidence intervals. The initial simulation scope excludes batch declarations and discrete outcomes. A model-based precision exercise does not establish study adequacy when its scenario or facet universe is inappropriate.