The exact rules the analysis applies — thresholds, severities and what to do about each. Every example below is executed against the engine by the test suite, so this reference cannot drift from what the tool actually does.
F1 — Too easy
WarningFacility ≥ 0.80 — candidates averaged 80%+ of the item's maximum mark.
Example: On "Introduces self and confirms identity", 15 of 16 candidates score the full 2 points: facility 0.97.
What to do: 80%+ of the maximum mark was achieved on average. Item may not discriminate; consider reviewing.
Pell et al. 2010, AMEE Guide 49
F2 — Too hard
WarningFacility ≤ 0.30 — the average score is below 30% of the item's maximum.
Example: Only 4 of 16 candidates score anything on a 2-point item: facility 0.13. Is the task unclear, or were examiners marking to different standards?
What to do: Average score is below 30% of the maximum. Verify the item wording, scoring scale and examiner calibration.
Pell et al. 2010, AMEE Guide 49
F3 — Negative discrimination
CriticalDiscrimination below 0 — the bottom third of candidates outscored the top third on this item.
Example: The five strongest candidates average 0 on the item while the five weakest average 2: discrimination −1.0. Classic signs: a reversed scale, a miskeyed checklist row, or examiners interpreting the item differently.
What to do: Lower-performing candidates scored higher than top performers on this item. Possible scoring or wording problem.
Pell et al. 2010, AMEE Guide 49
F4 — Poor discrimination
WarningDiscrimination in [0, 0.20) — exactly 0.20 is not flagged. The item barely separates strong from weak candidates.
Example: Top-third and bottom-third candidates average almost the same score: discrimination 0.10.
What to do: Item does not separate strong from weak candidates effectively. Consider revision.
Pell et al. 2010, AMEE Guide 49
F5 — Low item-total correlation
WarningItem-rest correlation below 0.20 — exactly 0.20 is not flagged. The item's scores barely track the rest of the station.
Example: Scores on the item are unrelated to how candidates did on everything else (r = 0.05) — it may be measuring a different construct, or examiner noise.
What to do: Item correlates weakly with the rest of the station score; it may be measuring something different.
Pell et al. 2010, AMEE Guide 49
B1 — Easy band
InfoFacility ≥ 0.80 (0.80 itself is Easy).
Example: Facility 0.80 lands in Easy; 0.79 is Moderate.
What to do: A few easy items are fine — they reassure candidates and catch absolute non-performance — but an easy-heavy station wastes testing time.
Khan et al. 2013, AMEE Guide 81 Part II
B2 — Moderate band
InfoFacility between 0.30 and 0.80, both exclusive.
Example: Facility 0.55 is Moderate.
What to do: The productive zone — moderate items carry most of the measurement information.
Khan et al. 2013, AMEE Guide 81 Part II
B3 — Hard band
InfoFacility ≤ 0.30 (0.30 itself is Hard).
Example: Facility 0.30 lands in Hard; 0.31 is Moderate.
What to do: Check hard items before blaming candidates: unclear task wording and uncalibrated examiners produce the same numbers as a genuinely difficult skill.
Khan et al. 2013, AMEE Guide 81 Part II
N1 — Discrimination needs ~10+ candidates
InfoDiscrimination compares the top and bottom 33% and needs at least 3 candidates in each group: with fewer than 10 candidates, group size floor(N × 0.33) drops below 3 and the metric shows “—” instead of a misleading number.
Example: With 8 candidates, group size is floor(8 × 0.33) = 2, so discrimination is not computed.
What to do: Don't read discrimination on tiny cohorts — collect more administrations before acting on it.
Daniels & Pugh 2018, Twelve tips for developing an OSCE
N2 — Item-rest needs variation
InfoThe item-rest correlation needs at least 3 scored candidates AND variation on both sides — if every candidate got the same item score (or the same rest score), it shows “—”.
Example: Every candidate scored 2 on the item: no variance, so no correlation exists.
What to do: A dash here on an item everyone scored identically is expected, not a data problem.
N3 — Reliability needs complete cases
InfoCronbach's α uses only candidates scored on every item, and needs at least 2 such candidates and 2 items — otherwise the report says “Insufficient data”.
Example: One complete-case candidate is not enough to estimate internal consistency.
What to do: Blank cells (not-assessed) remove that candidate from the α calculation only; their scored items still count everywhere else.
Pell et al. 2010, AMEE Guide 49