Would this candidate have passed with a different set of raters?
Decision accuracy and consistency for rater-mediated exams (oral exams, OSCEs, essay and constructed-response scoring), built on the many-facet Rasch rating scale model.
library(decisionfacets)
sim <- df_simulate(n_persons = 1000, n_items = 4, n_raters = 12,
raters_per_person = 2, severity_sd = 0.6, seed = 1)
fit <- df_fit(sim$data) # TAM if installed, else built-in JMLE
cut <- df_cut(16, decision_rule = "raw_total")
cf <- df_counterfactual(fit, cut) # per-candidate pass probabilities
cf # most rater-dependent candidates first
df_attribute(cf) # measurement error vs rater assignmentInstallation
From CRAN:
install.packages("decisionfacets")Development version from GitHub:
install.packages("pak")
pak::pak("edidatasolutions/decisionfacets")The decision rule is the point
How much rater severity harms a candidate depends on what the board decides on:
decision_rule |
severity reaches the decision? |
|---|---|
raw_total |
yes, directly |
fair_average, measure
|
no; only through measurement precision |
Internally every rule reduces to a raw-total cut that depends on the panel (harsher panels get lower cuts under measure-based rules), so pass probabilities are computed exactly by convolving category probabilities.
What each candidate gets
df_counterfactual() returns pass probabilities under the observed panel, an average-severity panel, and a random panel from the pool (enumerated when feasible), plus the best and worst panel. advantage is how far the assigned panel pushed the candidate toward the outcome they received; candidates with advantage >= 0.2 are flagged lenient_panel_pass or harsh_panel_fail.
df_attribute() splits expected misclassifications into measurement error (present even with average raters), rater assignment (observed minus average), and the expected cost of the random assignment design.
Validation (known truth, 100 replications)
1,000 candidates, 4 items, 12 raters (severity SD 0.6), 2 raters per candidate, cut = mean rating of 2. Means over 100 replications:
- Raw totals: 176 truly rater-dependent candidates per 1,000. TAM flags 169, recovering 71% of true flags (75% of TAM’s flags are real). An oracle with the true parameters also recovers 71%, so the remaining gap is measurement error from 8 ratings per candidate, not estimation error.
- Expected misclassifications per 1,000: 160 under raw totals (116 measurement, 44 rater assignment) versus 117 under the measure rule (assignment about 1). TAM estimates: 160 / 116 / 44. Realized errors in the simulation: 160.
- Under measure-based and fair-average rules, no candidate is flagged as rater-dependent.
Status
| layer | functions | status |
|---|---|---|
| data / model |
df_data, df_fit (TAM, JMLE), df_rater_effects
|
done |
| decision |
df_cut, df_classify
|
done |
| flagship |
df_counterfactual, df_attribute (severity) |
done |
| forward-looking |
df_simulate done; df_design, df_evaluate_design
|
stubs |
| monitoring / output |
df_drift, df_report
|
stubs |
Current design limits: every panel rater scores every item for a candidate, panel sizes are equal, no missing ratings, severity only (no halo/drift terms). With few ratings per person, JMLE stretches the logit scale (~15% in validation); prefer TAM.
Getting help and contributing
Questions and bug reports: https://github.com/edidatasolutions/decisionfacets/issues. See CONTRIBUTING.md for how to report problems, get help, or contribute code.