Plan annotation effort, costs, and pilot precision

This tutorial illustrates the study-planning workflows in source version 0.5.0. All observations and monetary amounts below are synthetic. No requests are executed and no provider prices are retrieved.

Quantity Purpose
Distinct pilot items Information for estimating source variation and coefficients
Annotations per item The averaging protocol whose reliability is projected
Items per request Shared context and request overhead; not additional repeated measurements

Fit a small Gaussian pilot

The first example uses independent-item annotations with item, rater and item-by-rater sources. Runs supply replication of those sources through distinct coded cells. This working model defines the example’s intended score universe.

library(Gtheory4LLM)
set.seed(2718)
panel <- expand.grid(item = 1:24, rater = 1:4, run = 1:3,
                     KEEP.OUT.ATTRS = FALSE)
item_rater <- as.integer(interaction(panel$item, panel$rater, drop = TRUE))
panel$quality <- rnorm(24)[panel$item] + rnorm(4, sd = 0.5)[panel$rater] +
  rnorm(96, sd = 0.4)[item_rater] + rnorm(nrow(panel), sd = 0.8)
design <- gt_design("item", c("rater", "run"),
                    random = ~ item + rater + item:rater)
control <- gt_control(gaussian = list(extra_tries = 2L, retry_seed = 439L))
fit <- gt_fit(panel, "quality", design, control = control)
stopifnot(isTRUE(fit$numerically_accepted))
gt_reliability(fit)$per_trait
#>   outcome     Erho2   Erho2_se Erho2_lower Erho2_upper      Phi     Phi_se Phi_lower Phi_upper
#> 1 quality 0.8954299 0.03560501   0.8025272   0.9474857 0.864711 0.04998046 0.7345082 0.9365736

Acceptance permits the supported projections. It does not establish a global optimum, adequate pilot precision, labeling accuracy or correct source selection for another study.

Compare costs within a declared allocation grid

Here one annotation costs 0.01 illustrative dollars, with a setup cost of two dollars per rater. The cost callback could instead return separate request, token, human-review and expected-retry costs, provided all assumptions use the same currency and are recorded. Costs do not introduce a statistical benefit from review or change the population of raters represented by the fit.

plan <- gt_plan(fit,
  candidates = list(rater = c(2, 4, 6), run = c(1, 3, 5)),
  cost = function(d) list(
    costs = data.frame(annotation = 0.01 * d$total_annotations,
                       setup = 2 * d$allocation_rater),
    resources = data.frame(requests = d$total_annotations),
    assumptions = "Illustrative prices; one item per request."),
  currency = "USD", cost_date = "2026-10-09", study_items = 1000,
  target = 0.80, coefficient = "Phi", baseline = 1L,
  constraints = function(d) data.frame(budget = d$total_cost <= 180))
as.data.frame(plan)[, c("design_id", "total_cost", "screening_value",
                       "feasible", "minimum_cost", "pareto")]
#>   design_id total_cost screening_value feasible minimum_cost pareto
#> 1         1         24       0.6486171     TRUE        FALSE   TRUE
#> 2         2         48       0.7868620     TRUE        FALSE   TRUE
#> 3         3         72       0.8470409     TRUE         TRUE   TRUE
#> 4         4         64       0.7616660     TRUE        FALSE  FALSE
#> 5         5        128       0.8647110     TRUE        FALSE   TRUE
#> 6         6        192       0.9055479    FALSE        FALSE  FALSE
#> 7         7        104       0.7891754     TRUE        FALSE  FALSE
#> 8         8        208       0.8821666    FALSE        FALSE  FALSE
#> 9         9        312       0.9182328    FALSE        FALSE  FALSE
head(plan$source_contributions)
#>   design_id     source                        role universe relative_error absolute_error
#> 1         1       item                    universe 0.900273      0.0000000     0.00000000
#> 2         1      rater              absolute_error 0.000000      0.0000000     0.07143433
#> 3         1 item:rater relative_and_absolute_error 0.000000      0.1072665     0.10726651
#> 4         1   Residual relative_and_absolute_error 0.000000      0.3090147     0.30901465
#> 5         2       item                    universe 0.900273      0.0000000     0.00000000
#> 6         2      rater              absolute_error 0.000000      0.0000000     0.03571716
#>   relative_error_reduction absolute_error_reduction
#> 1                        0               0.00000000
#> 2                        0               0.00000000
#> 3                        0               0.00000000
#> 4                        0               0.00000000
#> 5                        0               0.00000000
#> 6                        0               0.03571716

Minimum cost means minimum among the supplied candidates that meet both the constraints and the selected reliability screen. All candidates remain in the result, including cost ties and infeasible rows. Pareto candidates have no feasible cheaper-or-equal alternative with equal-or-better performance and at least one strict improvement. screening = "lower" separately uses an available lower confidence bound. It never replaces a missing bound with a point estimate; these are pointwise intervals, not post-search or future-panel guarantees.

study_items changes total cost, not the per-item coefficient. Every candidate retains the same declared random/fixed universe. In particular, the planner cannot make a protocol appear more dependable by changing a random facet into a fixed one.

Make the batch target explicit

The next synthetic panel adds a shared shift for each call of six items. Every condition uses the same grouping. A recorded call identifier allows the fit to retain that layout even when rows are reordered.

batched <- panel
group <- (batched$item - 1L) %/% 6L + 1L
call_index <- as.integer(interaction(group, batched$rater, batched$run, drop = TRUE))
batched$call_id <- paste0("call_", call_index)
batched$quality <- batched$quality + rnorm(max(call_index), sd = 0.7)[call_index]
batch_design <- gt_design("item", c("rater", "run"),
  random = ~ item + rater + item:rater, batch = gt_batch(6, id = "call_id"))
batch_fit <- gt_fit(batched, "quality", batch_design, control = control)
stopifnot(isTRUE(batch_fit$numerically_accepted))
within <- c("1" = 1, "2" = -1)
between <- c("1" = 1, "7" = -1)
corpus <- setNames(rep(1 / 24, 24), as.character(1:24))
batch_projection <- gt_batch_reliability(batch_fit,
  contrasts = list(within = within, between = between), aggregate = list(corpus = corpus))
subset(as.data.frame(batch_projection), target_kind != "absolute_item")
#>              target target_kind outcome outcome_kind universe_variance error_variance
#> 25  contrast:within    contrast quality      outcome        1.57975108     0.21306534
#> 26 contrast:between    contrast quality      outcome        1.57975108     0.30603717
#> 27 aggregate:corpus   aggregate quality      outcome        0.03291148     0.03200569
#>     error_sd coefficient        coefficient_name
#> 25 0.4615900   0.8811561    contrast_reliability
#> 26 0.5532063   0.8377139    contrast_reliability
#> 27 0.1789013   0.5069765 aggregate_dependability
subset(batch_projection$source_contributions,
       source == "Call" & target_kind != "absolute_item")
#>               target target_kind outcome outcome_kind source  role     kernel contraction
#> 124  contrast:within    contrast quality      outcome   Call error same_batch        0.00
#> 129 contrast:between    contrast quality      outcome   Call error same_batch        2.00
#> 134 aggregate:corpus   aggregate quality      outcome   Call error same_batch        0.25
#>     divisor   variance
#> 124      12 0.00000000
#> 129      12 0.09297184
#> 134      12 0.01162148

A shared call shift cancels from the same-batch difference, remains in the between-batch difference, and contributes to each absolute item score. A corpus mean has another error target: call shifts average over calls, while common rater shifts remain. Weighted aggregates can depend on how weights are spread across groups. None of these effects can be represented by one universal batch penalty.

batch_study <- gt_batch_dstudy(batch_fit,
  expand.grid(rater = c(2, 4), run = c(1, 3)),
  contrasts = list(within = within, between = between), aggregate = list(corpus = corpus))
subset(as.data.frame(batch_study), target_kind == "contrast")
#>      design_id           target target_kind outcome outcome_kind universe_variance
#> 1.24         1  contrast:within    contrast quality      outcome          1.579751
#> 1.25         1 contrast:between    contrast quality      outcome          1.579751
#> 2.24         2  contrast:within    contrast quality      outcome          1.579751
#> 2.25         2 contrast:between    contrast quality      outcome          1.579751
#> 3.24         3  contrast:within    contrast quality      outcome          1.579751
#> 3.25         3 contrast:between    contrast quality      outcome          1.579751
#> 4.24         4  contrast:within    contrast quality      outcome          1.579751
#> 4.25         4 contrast:between    contrast quality      outcome          1.579751
#>      error_variance  error_sd coefficient     coefficient_name allocation_rater allocation_run
#> 1.24      0.8430954 0.9182022   0.6520228 contrast_reliability                2              1
#> 1.25      1.4009264 1.1836074   0.5299973 contrast_reliability                2              1
#> 2.24      0.4215477 0.6492670   0.7893629 contrast_reliability                4              1
#> 2.25      0.7004632 0.8369368   0.6928082 contrast_reliability                4              1
#> 3.24      0.4261307 0.6527868   0.7875594 contrast_reliability                2              3
#> 3.25      0.6120743 0.7823518   0.7207468 contrast_reliability                2              3
#> 4.24      0.2130653 0.4615900   0.8811561 contrast_reliability                4              3
#> 4.25      0.3060372 0.5532063   0.8377139 contrast_reliability                4              3
#>      pilot_items batch_size annotations_per_item projected_calls extrapolated    scale
#> 1.24          24          6                    2               8        FALSE observed
#> 1.25          24          6                    2               8        FALSE observed
#> 2.24          24          6                    4              16        FALSE observed
#> 2.25          24          6                    4              16        FALSE observed
#> 3.24          24          6                    6              24        FALSE observed
#> 3.25          24          6                    6              24        FALSE observed
#> 4.24          24          6                   12              48        FALSE observed
#> 4.25          24          6                   12              48        FALSE observed
#>      uncertainty_available                                              uncertainty_reason
#> 1.24                 FALSE Point projections only; sampling intervals are not implemented.
#> 1.25                 FALSE Point projections only; sampling intervals are not implemented.
#> 2.24                 FALSE Point projections only; sampling intervals are not implemented.
#> 2.25                 FALSE Point projections only; sampling intervals are not implemented.
#> 3.24                 FALSE Point projections only; sampling intervals are not implemented.
#> 3.25                 FALSE Point projections only; sampling intervals are not implemented.
#> 4.24                 FALSE Point projections only; sampling intervals are not implemented.
#> 4.25                 FALSE Point projections only; sampling intervals are not implemented.
#>                                                                                                                                                                                                                                                                                                                                                                                                                                                                                     interpretation
#> 1.24 Conditional on the fitted fixed item grouping, with all instrumentation facets random and exchangeable and equal averaging over the planned complete panel. The item main effect defines the universe score. Coefficients describe the explicitly weighted target, not a universal ranking coefficient. Aggregate error is annotation error conditional on these items, not corpus-sampling uncertainty. Counts do not change batch size, grouping, item count, or fitted source covariances.
#> 1.25 Conditional on the fitted fixed item grouping, with all instrumentation facets random and exchangeable and equal averaging over the planned complete panel. The item main effect defines the universe score. Coefficients describe the explicitly weighted target, not a universal ranking coefficient. Aggregate error is annotation error conditional on these items, not corpus-sampling uncertainty. Counts do not change batch size, grouping, item count, or fitted source covariances.
#> 2.24 Conditional on the fitted fixed item grouping, with all instrumentation facets random and exchangeable and equal averaging over the planned complete panel. The item main effect defines the universe score. Coefficients describe the explicitly weighted target, not a universal ranking coefficient. Aggregate error is annotation error conditional on these items, not corpus-sampling uncertainty. Counts do not change batch size, grouping, item count, or fitted source covariances.
#> 2.25 Conditional on the fitted fixed item grouping, with all instrumentation facets random and exchangeable and equal averaging over the planned complete panel. The item main effect defines the universe score. Coefficients describe the explicitly weighted target, not a universal ranking coefficient. Aggregate error is annotation error conditional on these items, not corpus-sampling uncertainty. Counts do not change batch size, grouping, item count, or fitted source covariances.
#> 3.24 Conditional on the fitted fixed item grouping, with all instrumentation facets random and exchangeable and equal averaging over the planned complete panel. The item main effect defines the universe score. Coefficients describe the explicitly weighted target, not a universal ranking coefficient. Aggregate error is annotation error conditional on these items, not corpus-sampling uncertainty. Counts do not change batch size, grouping, item count, or fitted source covariances.
#> 3.25 Conditional on the fitted fixed item grouping, with all instrumentation facets random and exchangeable and equal averaging over the planned complete panel. The item main effect defines the universe score. Coefficients describe the explicitly weighted target, not a universal ranking coefficient. Aggregate error is annotation error conditional on these items, not corpus-sampling uncertainty. Counts do not change batch size, grouping, item count, or fitted source covariances.
#> 4.24 Conditional on the fitted fixed item grouping, with all instrumentation facets random and exchangeable and equal averaging over the planned complete panel. The item main effect defines the universe score. Coefficients describe the explicitly weighted target, not a universal ranking coefficient. Aggregate error is annotation error conditional on these items, not corpus-sampling uncertainty. Counts do not change batch size, grouping, item count, or fitted source covariances.
#> 4.25 Conditional on the fitted fixed item grouping, with all instrumentation facets random and exchangeable and equal averaging over the planned complete panel. The item main effect defines the universe score. Coefficients describe the explicitly weighted target, not a universal ranking coefficient. Aggregate error is annotation error conditional on these items, not corpus-sampling uncertainty. Counts do not change batch size, grouping, item count, or fitted source covariances.

The data frame includes allocation_rater, allocation_run, batch size and uncertainty availability in ordinary columns. These survive CSV export alongside the target and coefficient labels. Save the projection object separately when the complete item grouping and target weights are needed.

These are point projections conditional on the fitted grouping. Facet counts change under exchangeability assumptions; item count, grouping, batch size and fitted source covariances do not. All instrumentation facets remain random. Aggregate error is annotation error conditional on the selected items, not uncertainty from sampling a corpus. Contrast and aggregate ratios have explicit names; they are not universal item-ranking G coefficients. Sampling intervals are not provided.

Ordinary gt_reliability(), gt_dstudy() and the monetary gt_plan() refuse modelled Call fits. The separate batch functions make their target explicit. A fit keeps the compact item-to-batch labels even with data retention disabled; projection objects contain those labels and target weights. Portable gt_report(fit) and HTML output exclude them.

Examine pilot precision for one final protocol

More pilot information may improve coefficient estimation without changing the protocol eventually used for every item. The example varies pilot item and rater counts, but every refit estimates the same four-rater, three-run target.

precision <- gt_pilot_plan(fit,
  grid = data.frame(n_items = c(16, 24), rater = c(3, 4)),
  target_counts = c(rater = 4, run = 3),
  nsim = 2, seed = 2026, coefficient = "Phi", width_target = 0.25)
precision$summary
#>   design_id attempted accepted preparation_or_fit_errors simulation_errors refused
#> 1         1         2        2                         0                 0       0
#> 2         2         2        2                         0                 0       0
#>   reliability_errors interval_available interval_unavailable conditional_intervals
#> 1                  0                  2                    0                     0
#> 2                  0                  2                    0                     0
#>   acceptance_rate acceptance_mcse refusal_rate refusal_mcse error_rate error_mcse
#> 1               1               0            0            0          0          0
#> 2               1               0            0            0          0          0
#>   interval_rate interval_mcse mean_width_given_interval mean_width_mcse_given_interval
#> 1             1             0                 0.1801413                     0.03352617
#> 2             1             0                 0.2157597                     0.03154154
#>   width_target precision_successes precision_success_rate precision_success_mcse
#> 1         0.25                   2                      1                      0
#> 2         0.25                   2                      1                      0
#>   precision_rate_given_interval precision_mcse_given_interval
#> 1                             1                             0
#> 2                             1                             0
precision$replicates[, c("design_id", "replicate_id", "status", "accepted",
                         "interval_available", "interval_conditional", "width",
                         "acceptance_failures", "selected_attempt")]
#>   design_id replicate_id             status accepted interval_available interval_conditional
#> 1         1            1 interval_available     TRUE               TRUE                FALSE
#> 2         1            2 interval_available     TRUE               TRUE                FALSE
#> 3         2            1 interval_available     TRUE               TRUE                FALSE
#> 4         2            2 interval_available     TRUE               TRUE                FALSE
#>       width acceptance_failures selected_attempt
#> 1 0.2136675                <NA>                1
#> 2 0.1466152                <NA>                1
#> 3 0.2473012                <NA>                1
#> 4 0.1842181                <NA>                1

Two simulations per candidate only exercise the software. They cannot support a pilot-size recommendation. Set a predeclared simulation budget and inspect Monte Carlo uncertainty before using the results for planning. Every attempted replicate remains in the ledger. Precision success requires an available interval meeting the width target, so errors, refusals and unavailable intervals are non-successes. Mean widths condition on an available interval and show that denominator. Conditional boundary intervals retain their qualification. Numerically refused rows retain the fit’s acceptance-failure codes and selected attempt identifier; errors before a fit is returned have unavailable diagnostics. The data and retry seeds allow an individual attempt to be replayed.

The simulator draws new items and new levels of every random facet from a fitted or explicitly supplied Gaussian parameter scenario. It does not calibrate coverage, integrate parameter uncertainty automatically, or provide bootstrap confidence intervals. The initial simulation scope excludes batch declarations and discrete outcomes. A model-based precision exercise does not establish study adequacy when its scenario or facet universe is inappropriate.