Counterfactual pass probabilities: would this candidate have passed with different raters?
Source:R/counterfactual.R
df_counterfactual.RdFor each candidate, computes the probability of passing a re-rating under (a) the observed panel, (b) a panel of average-severity raters, and (c) a panel drawn at random from the rater pool, all under the same decision rule. Probabilities are exact (recursive convolution of the model's category probabilities); the only approximation is sampling panels when the pool is too large to enumerate.
Arguments
- object
A `df_fit` (estimated parameters) or `df_sim` (true parameters).
- cut
A `df_cut`.
- theta
How candidate ability enters: `"posterior"` integrates over the grid posterior given the candidate's observed ratings (the default for a `df_fit`, so probabilities include measurement error); `"point"` plugs in `object$par$theta` (the default for a `df_sim`, giving the known truth).
- max_panels
Enumerate all rater panels when there are at most this many; otherwise sample this many panels.
- flag_delta
Minimum rater advantage (see below) for a flag.
- grid
Theta grid for the posterior.
- prior_mean, prior_sd
Normal prior for the posterior; default to the fitted population prior when the fit supplies one (`par$theta_prior`), otherwise the mean and SD of the person estimates.
- seed
Optional seed for panel sampling.
Value
A `df_counterfactual` data frame, one row per candidate: `person`, `panel`, `total`, `raw_cut`, `pass_observed` (actual decision), `p_observed`, `p_average`, `p_random`, `p_min`, `p_max` (worst and best panel in the pool), `delta` (= p_observed - p_random), `advantage`, `direction` and `rater_dependent`.
`advantage` is how much the assigned panel pushed the candidate toward the outcome they actually received: `delta` for a pass, `-delta` for a fail. Equivalently, it is the increase in the probability that a re-rating would reverse the decision when the observed panel is swapped for a random one; decision reversals from measurement error alone cancel out. `rater_dependent` is `advantage >= flag_delta`, and `direction` labels flagged cases `"lenient_panel_pass"` (board's false-pass exposure) or `"harsh_panel_fail"` (the appeal case).
Examples
sim <- df_simulate(n_persons = 200, n_items = 3, n_raters = 6, seed = 1)
fit <- df_fit(sim$data, engine = "jmle")
cf <- df_counterfactual(fit, df_cut(12, "raw_total"))
cf # most rater-dependent candidates first
#> <df_counterfactual> rule = raw_total | cut = 12 | theta = posterior | 15 panels (all)
#> 36 of 200 candidates flagged as rater-dependent
#>
#> person panel total raw_cut pass_observed p_observed p_average p_random
#> P0124 R01|R02 13 12 TRUE 0.7451 0.218 0.281
#> P0116 R01|R02 15 12 TRUE 0.9042 0.452 0.464
#> P0156 R02|R06 13 12 TRUE 0.7389 0.292 0.341
#> P0069 R02|R06 13 12 TRUE 0.7389 0.292 0.341
#> P0119 R03|R05 9 12 FALSE 0.1838 0.600 0.574
#> P0173 R01|R02 16 12 TRUE 0.9489 0.586 0.565
#> P0075 R02|R06 12 12 TRUE 0.6241 0.191 0.257
#> P0057 R02|R05 13 12 TRUE 0.7339 0.363 0.396
#> P0170 R03|R04 7 12 FALSE 0.0583 0.355 0.388
#> P0005 R03|R06 10 12 FALSE 0.2889 0.650 0.613
#> p_min p_max delta advantage rater_dependent direction
#> 0.0237 0.745 0.464 0.464 TRUE lenient_panel_pass
#> 0.0931 0.904 0.441 0.441 TRUE lenient_panel_pass
#> 0.0400 0.812 0.398 0.398 TRUE lenient_panel_pass
#> 0.0400 0.812 0.398 0.398 TRUE lenient_panel_pass
#> 0.1684 0.952 -0.391 0.391 TRUE harsh_panel_fail
#> 0.1635 0.949 0.384 0.384 TRUE lenient_panel_pass
#> 0.0188 0.712 0.367 0.367 TRUE lenient_panel_pass
#> 0.0604 0.858 0.338 0.338 TRUE lenient_panel_pass
#> 0.0583 0.846 -0.329 0.329 TRUE harsh_panel_fail
#> 0.2026 0.964 -0.324 0.324 TRUE harsh_panel_fail
summary(cf)
#> decision_rule cut n pass_rate_observed expected_pass_rate_random_panel
#> 1 raw_total 12 200 0.57 0.5535833
#> mean_abs_delta n_rater_dependent n_lenient_panel_pass n_harsh_panel_fail
#> 1 0.1179387 36 22 14
#> n_panel_sensitive
#> 1 107
# Known truth: the same analysis with the true parameters
summary(df_counterfactual(sim, df_cut(12, "raw_total")))
#> decision_rule cut n pass_rate_observed expected_pass_rate_random_panel
#> 1 raw_total 12 200 0.57 0.53669
#> mean_abs_delta n_rater_dependent n_lenient_panel_pass n_harsh_panel_fail
#> 1 0.1294924 40 21 19
#> n_panel_sensitive
#> 1 118