Skip to contents

Linking two language groups requires an assumption about which items are free of DIF, and when DIF is pervasive on a short test the data may not settle it. This function makes that dependence visible:

Linking assumptions

The ability difference under the weighted mode (DIF-free items form the densest cluster), iterative purification, and all items as anchors (DIF cancels out), with the DIF model refitted under each and, if `cut` is given, the pass-rate impact under each.

Stability of the mode

A parametric bootstrap redraws each item's between-language difference from its sampling distribution and re-estimates the mode. When DIF-free items form a clear cluster the bootstrap modes stay close to the estimate; when they do not, the mode jumps between clusters, which is exactly the failure a single estimate hides.

Linking-sensitive items

Items whose DIF flag changes across the linking assumptions. These are the items to prioritize for review by content and translation experts.

The overall verdict is `"robust"` when the linkings agree within `tolerance` logits and the bootstrap keeps at least `stable_share` of the modes within `tolerance` of the estimate; otherwise `"sensitive"`.

Usage

td_sensitivity(
  dif,
  cut = NULL,
  tolerance = 0.15,
  stable_share = 0.8,
  B = 200,
  n_draws = 100,
  seed = NULL
)

Arguments

dif

A `td_dif` object (from [td_dif()]).

cut

Optional raw-score passing standard for pass-rate impact.

tolerance

Difference in the linking shift (logits) considered substantively negligible.

stable_share

Minimum share of bootstrap modes within `tolerance` of the estimate for the mode to count as stable.

B

Bootstrap replicates.

n_draws

Posterior draws for each pass-rate impact.

seed

Optional seed.

Value

A `td_sensitivity` object: `$linkings` (one row per assumption: `assumption`, `shift`, `focal_mean`, `se`, `n_flagged`, and pass-rate impact columns when `cut` is given), `$bootstrap` (`sd`, `lower`, `upper` (90 linking and a `linking_sensitive` indicator), `$range` (largest difference between linkings) and `$verdict`.

Details

**What the verdict means.** It is a statement about *dependence on an untestable assumption*, not an estimate of which linking is correct. In the package's known-truth simulations, data sets judged "sensitive" did not have larger linking errors for the mode than those judged "robust". Disagreement between linkings arose mostly because mean and purified linking are biased under heavy directional DIF. Bootstrap instability was only weakly associated with error. Use the verdict to decide when to report results under several assumptions and to send the linking-sensitive items to content and translation experts; only that review can settle which items should anchor the scale.

Examples

sim <- td_simulate(n_ref = 400, n_focal = 120, n_items = 20, seed = 5)
dif <- td_dif(td_calibrate(sim$responses, sim$group))
sens <- td_sensitivity(dif, cut = 12, B = 50, n_draws = 20, seed = 1)
sens
#> <td_sensitivity> verdict: SENSITIVE (tolerance 0.15 logits)
#> Linkings differ by up to 0.011 logits; 76% of 50 bootstrap modes fall within 0.15 of the estimate (90% interval 0.451 to 0.815)
#> 
#>  assumption shift focal_mean    se n_flagged pass_rate_change change_lower
#>        mode 0.616     -0.616 0.246         3         0.000737      -0.1061
#>    purified 0.621     -0.621 0.125         3         0.002383      -0.0181
#>   all_items 0.610     -0.610 0.123         3        -0.000853      -0.0497
#>  change_upper
#>        0.0636
#>        0.0442
#>        0.0283
#> 
#> 2 linking-sensitive item(s):
#>  item      d flag_mode flag_purified flag_all_items linking_sensitive
#>   Q14 0.0474     FALSE          TRUE          FALSE              TRUE
#>   Q16 1.2312      TRUE         FALSE           TRUE              TRUE