Skip to contents

Does a score mean the same thing in both languages?

A comparability workflow for translated and adapted exams, built for the conditions where standard DIF tools break down: a small translated-language group (50–200), a group that differs in ability, and DIF that runs mostly in one direction, so anchors themselves are biased.

library(transDIF)

sim <- td_simulate(n_focal = 150, seed = 1)          # or your 0/1 matrix + group
cal <- td_calibrate(sim$responses, sim$group)        # Rasch MML per language group
dif <- td_dif(cal)                                   # robust linking + small-sample DIF
imp <- td_impact(dif, cut = 36)                      # does DIF change pass rates?
fea <- td_features(dif, sim$features)                # which features predict DIF?
cat(td_report(dif, imp, fea, c("English", "French")))

Installation

From CRAN (once released):

install.packages("transDIF")

Development version from GitHub:

install.packages("pak")
pak::pak("edidatasolutions/transDIF")

Method

  • Calibration. Rasch MML in each group, with SEs from the exact marginal Hessian. DIF uses relative SEs: the scale-location uncertainty shared by all items belongs to the linking shift, and is added to its SE there.
  • Robust linking. Item differences are modeled as d_i ~ N(c + delta_i, se_i^2). The shift c (the ability difference) is the precision-weighted mode of the d_i: the center of the densest cluster of items. It assumes that DIF-free items form the largest cluster, not that DIF cancels out. Its SE is a sandwich estimate plus the scales’ location uncertainty.
  • Ambiguity warning. If a second cluster of items is nearly as dense, the data cannot say which cluster is DIF-free. td_dif then warns, and the report shows both candidate shifts.
  • Small-sample DIF. Given c, a spike-and-slab mixture (a DIF-free majority with negligible tau0, plus a directional DIF slab) gives each item a posterior DIF probability, a shrunken DIF estimate and a local-FDR flag.
  • Aggregate impact. The group’s pass-rate change and the raw-score shift at the cut, from exact score distributions. Intervals propagate both linking and DIF uncertainty.
  • Feature explanation. Random-effects meta-regression of DIF on item features, giving concrete guidance for translators.
  • Baselines for comparison: mean linking, iterative purification, Mantel–Haenszel with purification and BH.

Validation (known truth, 100 replications per size)

60 items; reference n = 2,000; translated group 0.5 logits lower; ~35% of items carry adaptation features that make them mostly harder in translation.

DIF detection (items with DIF > 0.2 logits, target FDR 10%):

focal n EB power / FDP purified z power / FDP Mantel–Haenszel power / FDP
50 7.7% / 6.3% 6.7% / 8.3% 3.9% / 4.8%
100 27.8% / 6.0% 24.7% / 8.3% 17.4% / 5.1%
200 58.8% / 8.4% 55.1% / 10.1% 45.7% / 8.6%

EB has the most power at every size and keeps its false discovery proportion below the 10% target. With 50 translated-language candidates, no method finds more than a few percent of DIF items, so treat flags there as candidates for expert review.

DIF effect estimation (RMSE, logits), EB shrinkage vs raw differences on DIF-free items: 0.105 vs 0.363 (n = 50), 0.092 vs 0.260 (n = 100), 0.083 vs 0.188 (n = 200). Genuinely large DIF is shrunk somewhat toward the slab mean.

Linking (true shift 0.5):

focal n bias: mode / mean / purified RMSE: mode / mean / purified
50 0.001 / 0.044 / 0.023 0.151 / 0.151 / 0.154
100 0.013 / 0.071 / 0.030 0.112 / 0.132 / 0.115
200 0.008 / 0.087 / 0.018 0.084 / 0.118 / 0.088

Mean linking is biased by directional DIF, and more data does not fix that. The weighted mode is essentially unbiased and has the lowest RMSE at every size. In a broader benchmark (180 data sets: balanced, directional and heavy DIF at n = 50/100/200; dev/bench_link.R) it had the lowest overall RMSE of nine estimators (0.114, vs 0.119 for purification, 0.137 for mean linking and 0.164 for the joint mixture used in earlier development). Its 90% intervals cover the true shift 89–90% of the time.

Impact: 90% intervals for the pass-rate change cover the truth in 94% of replications at n = 150 (true mean change -3.0 points, estimated -2.2).

Features: idiom, cultural referent, units and vocabulary effects are all recovered (true 0.60 / 0.50 / -0.40 / 0.35; mean estimates 0.60 / 0.48 / -0.38 / 0.35), with 95% CI coverage of 92–97%.

When the linking is not identified. Linking assumes DIF-free items form the largest cluster. On short tests with pervasive DIF this can fail: in real cross-national data (14 FIMS mathematics items, Australia vs Japan), the linkings disagreed by 0.3 logits and the estimated impact of DIF on the pass rate ranged from -8.0 points to -0.6 (interval including zero). td_sensitivity() reports this dependence; see the vignette.

Status

Done: td_simulate, td_calibrate, td_dif, td_mh, td_impact, td_features, td_report, td_sensitivity (development version). Next: uniform vs non-uniform DIF (2PL), polytomous items, and validation on real French–English data.

Getting help and contributing

Questions and bug reports: https://github.com/edidatasolutions/transDIF/issues. See CONTRIBUTING.md for how to report problems, get help, or contribute code.