When an exam is offered in a second language, the question is whether a score means the same thing in both. Standard DIF tools struggle when the translated group is small, differs in ability, and many items shift in the same direction. transDIF is built for that situation.
Data
A source-language group of 2,000 and a translated-language group of 150 who are 0.5 logits lower on average. Items with idioms, cultural referents, measurement units or heavy vocabulary tend to shift in translation.
library(transDIF)
sim <- td_simulate(n_ref = 2000, n_focal = 150, n_items = 40, seed = 5)
colSums(sim$features[-1])
#> idiom cultural units vocabulary
#> 4 8 4 2Calibrate each group, then link and detect DIF together
cal <- td_calibrate(sim$responses, sim$group)
dif <- td_dif(cal)
dif
#> <td_dif> 40 items | linking shift c = 0.479 (SE 0.116, mode); mean-linking c = 0.623
#> estimated DIF-free share 0.70; DIF component mean +0.51, SD 0.05
#> 9 items flagged at Bayesian FDR 0.1 | 24 clean anchors
#>
#> item d p_dif dif_mean dif_sd
#> Q25 1.221 0.995 0.522 0.0591
#> Q13 1.192 0.996 0.522 0.0578
#> Q09 1.167 0.993 0.519 0.0623
#> Q30 1.039 0.962 0.495 0.1036
#> Q15 1.126 0.942 0.488 0.1239
#> Q31 1.134 0.916 0.475 0.1443
#> Q03 0.979 0.921 0.472 0.1371
#> Q36 1.044 0.803 0.415 0.2016
#> Q29 0.875 0.727 0.372 0.2182The linking shift (the ability difference) is the center of the densest cluster of items, not the mean of all items, so it does not assume that DIF cancels out. Compare with linking on all items:
c(true = -sim$truth$focal_mean, robust = dif$link[["c"]], all_items = dif$c_mean)
#> true robust all_items
#> 0.5000000 0.4791316 0.6228016Does DIF change pass rates?
td_impact(dif, cut = 24, n_draws = 100, seed = 1)
#> quantity estimate lower upper
#> 1 pass_rate_fair 0.19513094 0.14950118 0.24928711
#> 2 pass_rate_translated 0.15933817 0.12855365 0.18308398
#> 3 pass_rate_change -0.03579277 -0.06834324 -0.01838312
#> 4 score_shift_at_cut -1.16064342 -1.99479015 -0.64982880How much do the conclusions depend on the linking?
Linking assumes that DIF-free items form the largest cluster. When
many items show DIF, the data may not settle which items are DIF-free.
td_sensitivity() shows the results under each linking
assumption, how stable the mode is, and which items’ DIF verdicts depend
on the choice:
td_sensitivity(dif, cut = 24, B = 100, n_draws = 50, seed = 1)
#> <td_sensitivity> verdict: SENSITIVE (tolerance 0.15 logits)
#> Linkings differ by up to 0.144 logits; 77% of 100 bootstrap modes fall within 0.15 of the estimate (90% interval 0.369 to 0.753)
#>
#> assumption shift focal_mean se n_flagged pass_rate_change change_lower
#> mode 0.479 -0.479 0.1163 9 -0.03579 -0.0646
#> purified 0.519 -0.519 0.0964 8 -0.02903 -0.0491
#> all_items 0.623 -0.623 0.0949 3 -0.00459 -0.0197
#> change_upper
#> -0.0188
#> -0.0152
#> 0.0155
#>
#> 6 linking-sensitive item(s):
#> item d flag_mode flag_purified flag_all_items linking_sensitive
#> Q03 0.979 TRUE TRUE FALSE TRUE
#> Q15 1.126 TRUE TRUE FALSE TRUE
#> Q29 0.875 TRUE FALSE FALSE TRUE
#> Q30 1.039 TRUE TRUE FALSE TRUE
#> Q31 1.134 TRUE TRUE FALSE TRUE
#> Q36 1.044 TRUE TRUE FALSE TRUEA “sensitive” verdict does not say which linking is right. It says that a conclusion depends on an assumption the data cannot check, so report the alternatives and send the listed items to expert review.
What should translators look at?
td_features(dif, sim$features)
#> term estimate se z p_value
#> 1 (Intercept) 0.02073289 0.03984978 0.5202762 6.028711e-01
#> 2 idiom 0.48388035 0.10197468 4.7451028 2.084004e-06
#> 3 cultural 0.51599144 0.08421317 6.1272061 8.943560e-10
#> 4 units -0.28103451 0.11066300 -2.5395526 1.109944e-02
#> 5 vocabulary 0.01735354 0.16537924 0.1049318 9.164299e-01A report to start from
cat(td_report(dif, languages = c("English", "French")))
#> # Score comparability: English vs French forms
#>
#> ## Samples
#> - English: 2000 examinees; French: 150 examinees.
#>
#> ## Linking and anchors
#> - The French group's mean ability is estimated at -0.48 logits relative to the English group (SE 0.12).
#> - This estimate relies on DIF-free items forming the largest cluster of items (estimated DIF-free share: 70.3%), not on DIF cancelling out. Linking with all items as anchors would have given -0.62.
#> - 24 items qualify as clean anchors (posterior DIF probability below 0.2).
#>
#> ## Item-level DIF
#> - 9 of 40 items are flagged (Bayesian false discovery rate 10.0%).
#>
#> | item | DIF (logits) | posterior SD | P(DIF) | direction |
#> |---|---|---|---|---|
#> | Q25 | +0.52 | 0.06 | 1.00 | harder in French |
#> | Q13 | +0.52 | 0.06 | 1.00 | harder in French |
#> | Q09 | +0.52 | 0.06 | 0.99 | harder in French |
#> | Q30 | +0.50 | 0.10 | 0.96 | harder in French |
#> | Q15 | +0.49 | 0.12 | 0.94 | harder in French |
#> | Q31 | +0.47 | 0.14 | 0.92 | harder in French |
#> | Q03 | +0.47 | 0.14 | 0.92 | harder in French |
#> | Q36 | +0.42 | 0.20 | 0.80 | harder in French |
#> | Q29 | +0.37 | 0.22 | 0.73 | harder in French |
#>
#> ## Notes for review
#> - Flagged items call for expert review of the translation before any statistical adjustment.
#> - Conclusions assume the Rasch model holds in both groups and that DIF-free items form the largest cluster of items; td_sensitivity() shows how much they depend on that assumption.
#> - Frame decisions using the fairness chapter of the Standards (AERA, APA, NCME, 2014) and the ITC Guidelines for Translating and Adapting Tests (2nd ed., 2017).With very small translated-language groups (around 50 candidates), every method has little power to detect DIF, so treat flags as candidates for expert review rather than conclusions.