Not “did it drift?” but “when, how, and what should we do?”
Most drift checks compare two calibrations at equating time. Continuous testing programs and pre-equated banks need ongoing surveillance. driftwatch borrows sequential change detection from statistical process control.
library(driftwatch)
sim <- dw_simulate(seed = 1) # or your responses + bank
est <- dw_estimate(sim$responses, sim$bank) # item x window estimates
tu <- dw_tune(est, target = 0.01) # threshold for 1% false alarms/item
mon <- dw_monitor(est, h = tu$h) # CUSUM + change points + type
mon
dw_actions(mon, anchors = my_anchor_ids) # recommendations + audit log
dw_impact(form_ids, sim$bank, current_b, flagged) # score / pass-rate impactInstallation
From CRAN:
install.packages("driftwatch")Development version from GitHub:
install.packages("pak")
pak::pak("edidatasolutions/driftwatch")How it works
-
Window estimates. Rasch difficulty per item per window, with examinee ability known from operational scoring. Standardized deviation from the bank:
z = (b_t - b_bank) / sqrt(SE_t^2 + SE_bank^2). -
Two-sided CUSUM on
z(reference valuek, thresholdh). -
Threshold tuning on your own design.
dw_tune()regenerates responses under no drift for the program’s actual items, windows, sample sizes and examinees, then reruns estimation and CUSUM. This matters: each item’s bank error is shared by every window and accumulates in the CUSUM. The design-based threshold (h ≈ 9.5) is far above the iid-normal one (≈ 6.9), and only the former meets the false-alarm target. -
Gradual vs abrupt. Step and ramp models are fitted with a profiled change point, and the log likelihood ratio decides between them. With weak evidence the type is
undeterminedrather than guessed. - Impact and action. True-score scoring impact of keeping, removing or recalibrating flagged items. Actions are retire, recalibrate or remove from the anchor set, each with a written rationale, threshold settings, analyst and timestamp.
Validation (known truth, 100 replications)
300 items, 40 windows, ~80 responses per item per window. 10% of items drift gradually (0.02–0.06 logits/window) and 5% jump (0.4–1.0 logits).
| continuous (driftwatch) | two-point (first vs last 5 windows) | |
|---|---|---|
| false alarms, stable items (target 1%) | 0.86% (iid-normal threshold: 4.0%) | 0.27% |
| abrupt detected | 99%, median 4.1 windows after onset | 72%, at the end |
| gradual detected | 83%, median 12.5 windows after onset | 73%, at the end |
| abrupt onset error (windows) | 1.1 | — |
Drift-type classification:
| data used | abrupt: typed / accuracy | gradual: typed / accuracy |
|---|---|---|
| alarm + 3 windows | 78% / 94% | 51% / 65% |
| all 40 windows | 94% / 99% | 78% / 89% |
Gradual drift is hard to type soon after an alarm, because a short ramp looks like a step. Reclassify as windows accumulate.
Score impact (60-item form with 12 drifted items, cut at theta = 0.5): keeping banked values mis-states the pass rate by -0.8 points; removing or recalibrating the flagged items brings the error to about 0 (within ±0.01 points).
Status and assumptions
Done: dw_simulate, dw_estimate, dw_tune, dw_monitor, dw_impact, dw_actions, dw_twopoint. Rasch difficulty only; examinee ability treated as known. Next: 2PL discrimination drift, ability uncertainty, anchor-set re-linking after removals, and a per-item run-length (ARL) view.
Getting help and contributing
Questions and bug reports: https://github.com/edidatasolutions/driftwatch/issues. See CONTRIBUTING.md for how to report problems, get help, or contribute code.