Does a score mean the same thing in both languages?
A comparability workflow for translated and adapted exams, built for the conditions where standard DIF tools break down: a small translated-language group (50–200), a group that differs in ability, and DIF that runs mostly in one direction, so anchors themselves are biased.
library(transDIF)
sim <- td_simulate(n_focal = 150, seed = 1) # or your 0/1 matrix + group
cal <- td_calibrate(sim$responses, sim$group) # Rasch MML per language group
dif <- td_dif(cal) # robust linking + small-sample DIF
imp <- td_impact(dif, cut = 36) # does DIF change pass rates?
fea <- td_features(dif, sim$features) # which features predict DIF?
cat(td_report(dif, imp, fea, c("English", "French")))Installation
From CRAN (once released):
install.packages("transDIF")Development version from GitHub:
install.packages("pak")
pak::pak("edidatasolutions/transDIF")Method
- Calibration. Rasch MML in each group, with SEs from the exact marginal Hessian. DIF uses relative SEs: the scale-location uncertainty shared by all items belongs to the linking shift, and is added to its SE there.
-
Robust linking. Item differences are modeled as
d_i ~ N(c + delta_i, se_i^2). The shiftc(the ability difference) is the precision-weighted mode of thed_i: the center of the densest cluster of items. It assumes that DIF-free items form the largest cluster, not that DIF cancels out. Its SE is a sandwich estimate plus the scales’ location uncertainty. -
Ambiguity warning. If a second cluster of items is nearly as dense, the data cannot say which cluster is DIF-free.
td_difthen warns, and the report shows both candidate shifts. -
Small-sample DIF. Given
c, a spike-and-slab mixture (a DIF-free majority with negligibletau0, plus a directional DIF slab) gives each item a posterior DIF probability, a shrunken DIF estimate and a local-FDR flag. - Aggregate impact. The group’s pass-rate change and the raw-score shift at the cut, from exact score distributions. Intervals propagate both linking and DIF uncertainty.
- Feature explanation. Random-effects meta-regression of DIF on item features, giving concrete guidance for translators.
- Baselines for comparison: mean linking, iterative purification, Mantel–Haenszel with purification and BH.
Validation (known truth, 100 replications per size)
60 items; reference n = 2,000; translated group 0.5 logits lower; ~35% of items carry adaptation features that make them mostly harder in translation.
DIF detection (items with DIF > 0.2 logits, target FDR 10%):
| focal n | EB power / FDP | purified z power / FDP | Mantel–Haenszel power / FDP |
|---|---|---|---|
| 50 | 7.7% / 6.3% | 6.7% / 8.3% | 3.9% / 4.8% |
| 100 | 27.8% / 6.0% | 24.7% / 8.3% | 17.4% / 5.1% |
| 200 | 58.8% / 8.4% | 55.1% / 10.1% | 45.7% / 8.6% |
EB has the most power at every size and keeps its false discovery proportion below the 10% target. With 50 translated-language candidates, no method finds more than a few percent of DIF items, so treat flags there as candidates for expert review.
DIF effect estimation (RMSE, logits), EB shrinkage vs raw differences on DIF-free items: 0.105 vs 0.363 (n = 50), 0.092 vs 0.260 (n = 100), 0.083 vs 0.188 (n = 200). Genuinely large DIF is shrunk somewhat toward the slab mean.
Linking (true shift 0.5):
| focal n | bias: mode / mean / purified | RMSE: mode / mean / purified |
|---|---|---|
| 50 | 0.001 / 0.044 / 0.023 | 0.151 / 0.151 / 0.154 |
| 100 | 0.013 / 0.071 / 0.030 | 0.112 / 0.132 / 0.115 |
| 200 | 0.008 / 0.087 / 0.018 | 0.084 / 0.118 / 0.088 |
Mean linking is biased by directional DIF, and more data does not fix that. The weighted mode is essentially unbiased and has the lowest RMSE at every size. In a broader benchmark (180 data sets: balanced, directional and heavy DIF at n = 50/100/200; dev/bench_link.R) it had the lowest overall RMSE of nine estimators (0.114, vs 0.119 for purification, 0.137 for mean linking and 0.164 for the joint mixture used in earlier development). Its 90% intervals cover the true shift 89–90% of the time.
Impact: 90% intervals for the pass-rate change cover the truth in 94% of replications at n = 150 (true mean change -3.0 points, estimated -2.2).
Features: idiom, cultural referent, units and vocabulary effects are all recovered (true 0.60 / 0.50 / -0.40 / 0.35; mean estimates 0.60 / 0.48 / -0.38 / 0.35), with 95% CI coverage of 92–97%.
When the linking is not identified. Linking assumes DIF-free items form the largest cluster. On short tests with pervasive DIF this can fail: in real cross-national data (14 FIMS mathematics items, Australia vs Japan), the linkings disagreed by 0.3 logits and the estimated impact of DIF on the pass rate ranged from -8.0 points to -0.6 (interval including zero). td_sensitivity() reports this dependence; see the vignette.
Status
Done: td_simulate, td_calibrate, td_dif, td_mh, td_impact, td_features, td_report, td_sensitivity (development version). Next: uniform vs non-uniform DIF (2PL), polytomous items, and validation on real French–English data.
Getting help and contributing
Questions and bug reports: https://github.com/edidatasolutions/transDIF/issues. See CONTRIBUTING.md for how to report problems, get help, or contribute code.