MCQ Item-Flaw Checker

Free · no signup · your questions never upload

MCQ Item-Flaw Checker

Paste a multiple-choice question and get an NBME-style structural flaw report — every flag cites the item-writing guideline behind it, with a suggested fix. Check up to 50 questions at once and see how you compare to the published flaw benchmark. Everything runs in your browser.

Check your question

Mark the intended answer to unlock the longest-option check.

Questions are analyzed locally — open your browser's network tab to verify nothing is uploaded.

A worked example

This is the tool's built-in sample — a plausible-looking question carrying three classic flaws. The report below is computed live by the same engine you use above:

A patient newly diagnosed with type 2 diabetes asks about lifestyle measures. Which of the following is the best initial recommendation?

  1. A) Exercise only when symptoms appear
  2. B) A structured program of regular aerobic exercise combined with dietary changes and ongoing glucose monitoringcorrect answer
  3. C) Bariatric surgery
  4. D) All of the above
DScreening gradeStructural screening only — content validity needs human review.

3 structural flaws found:

  • Absolute termsWarning · option A

    Option A uses an absolute term — testwise students avoid options with "only".

    Fix: Reword without the specific determiner, or make the option defensibly absolute. Haladyna, Downing & Rodriguez 2002, Applied Measurement in Education

  • All of the aboveWarning · option D

    “All of the above” is answerable from partial knowledge — recognizing two correct options gives it away.

    Fix: Replace with a substantive option. NBME Item-Writing Guide

  • Longest option is the keyCritical · option B

    The correct answer is 4.27× the median option length — “pick the longest answer” is the best-known testwiseness cue.

    Fix: Trim the key or lengthen the distractors until they're parallel. NBME Item-Writing Guide

Load this sample in the checker above to experiment with fixes.

What each check looks for

The complete rule set — what triggers each check, why the flaw matters, and the guideline it comes from. Every example below is run against the real engine in our test suite, so this documentation can't drift from what the tool actually does.

Absolute terms

Warning · checks the options

An option contains a specific determiner — always, never, all, none, only, must.

Few things in any field are absolute, and testwise students know it: options with absolutes are disproportionately wrong, so they can be eliminated without knowledge.

Example that trips it:

Which statement about beta blockers is correct?

  1. A) They always lower heart rate
  2. B) They can mask hypoglycemia
  3. C) They are contraindicated in hypertension

Trigger: always

Fix: Reword without the determiner, or make the option defensibly absolute. Haladyna, Downing & Rodriguez 2002, Applied Measurement in Education

Vague frequency terms

Warning · checks stem and options

The stem or an option hedges with often, sometimes, usually, rarely, frequently, commonly or “may occur”.

Imprecise frequency words mean different things to different examinees — the item becomes an argument about vocabulary, not knowledge.

Example that trips it:

Which finding commonly occurs in early sepsis?

  1. A) Tachycardia
  2. B) Bradycardia
  3. C) Hypothermia

Trigger: commonly

Fix: Replace with a specific frequency, proportion or condition. NBME Item-Writing Guide

Negative stem

Warning · checks the stem

The stem asks for the wrong answer — not, except, false, least. A second, minor flag fires when the negative word isn't capitalized.

Negative stems measure careful reading as much as knowledge; students who know the content still miss them under time pressure.

Example that trips it:

Which of the following is not a beta blocker?

  1. A) Atenolol
  2. B) Amlodipine
  3. C) Metoprolol

Trigger: not

Fix: Rephrase positively, or capitalize the negative (NOT) if it must stay. Tarrant et al. 2006, Nurse Education Today (item-flaw frequency study)

All / none of the above

Warning · checks the options

An option is literally “All of the above” or “None of the above”.

“All of the above” is answerable from partial knowledge — recognizing two correct options gives it away. “None of the above” tests recognizing wrongness rather than knowing the answer.

Example that trips it:

Which are established cardiovascular risk factors?

  1. A) Smoking
  2. B) Hypertension
  3. C) All of the above

Trigger: All of the above

Fix: Replace with a substantive option. NBME Item-Writing Guide

Longest option is the key

Critical · checks the options

The marked correct answer is more than 1.5× the median option length. Requires the key to be marked.

“Pick the longest answer” is the best-known testwiseness strategy — correct answers grow long because item writers qualify them until they're defensible.

Example that trips it:

What is the best initial management?

  1. A) Watchful waiting
  2. B) A structured multidisciplinary plan with regular specialist follow-up and staged escalation of therapy
  3. C) Immediate surgery
  4. D) Discharge

Trigger: the long correct answer

Fix: Trim the key or lengthen the distractors until they're parallel. NBME Item-Writing Guide

Option length imbalance

Info · checks the options

Any option (key or not) is more than 1.5× the median length — flagged even when no key is marked.

Options that stand out visually attract attention regardless of content; parallel construction keeps the choice about knowledge.

Example that trips it:

Which is a first-line treatment?

  1. A) Metformin
  2. B) A carefully titrated combination of several second-line agents chosen case by case
  3. C) Insulin
  4. D) Diet

Trigger: the long distractor

Fix: Balance option lengths so none stands out. Haladyna, Downing & Rodriguez 2002, Applied Measurement in Education

Grammatical cue (heuristic)

Warning · checks stem and options

The stem's grammar rules out some options before any knowledge is applied — e.g. a stem ending in “an” when only some options start with a vowel, or a plural verb with singular options.

Grammatical disagreement eliminates distractors for free. This check is a heuristic: it flags likely cases and says so.

Example that trips it:

Scurvy is caused by a deficiency of an

  1. A) ascorbic acid
  2. B) thiamine
  3. C) riboflavin

Trigger: an

Fix: End the stem so every option fits grammatically (e.g. “…a deficiency of which vitamin?”). NBME Item-Writing Guide

Clang association

Warning · checks stem and options

Exactly one option repeats a distinctive word (6+ letters) from the stem that no other option contains.

Word echoes between stem and one option cue the answer — students pick the option that “sounds like” the question.

Example that trips it:

Which hormone regulates calcium homeostasis?

  1. A) Parathyroid hormone, which controls calcium directly
  2. B) Insulin
  3. C) Glucagon

Trigger: calcium

Fix: Remove the repeated word from the option, or repeat it across several options. NBME Item-Writing Guide

Non-homogeneous options

Info · checks the options

Options mix numeric and verbal answers, numeric options aren't sorted, or numeric ranges overlap.

Heterogeneous options leak information (the odd one out draws attention); overlapping ranges can make two options simultaneously true.

Example that trips it:

What is the usual adult starting dose?

  1. A) 500 mg
  2. B) 250 mg
  3. C) 1000 mg

Trigger: unsorted doses

Fix: Keep options the same kind, sort numerics, make ranges mutually exclusive. Haladyna, Downing & Rodriguez 2002, Applied Measurement in Education

Option-count sanity

Info · checks the options

Fewer than 3 options (Info), or two options with identical text (Warning).

Research shows 3 well-built options suffice — but 2 makes guessing a coin flip, and duplicates are outright defects.

Example that trips it:

Is aspirin an antiplatelet agent?

  1. A) Yes
  2. B) No

Trigger: 2 options

Fix: Add at least one plausible distractor; replace duplicates. Haladyna, Downing & Rodriguez 2002, Applied Measurement in Education

The rules, and where they come from

Every check implements a documented item-writing guideline, cited next to its flag. The screening grade weights flags by severity — the exact weights, so you can disagree with them:

Critical = 3 points · Warning = 1 · Info = 0.25. Grade A = 0 points, B ≤ 1, C ≤ 3, D above that. The grade screens for structure; it is not a psychometric judgment.

This is structural screening against published item-writing guidelines. It cannot judge whether the content is correct, relevant or fair — that still needs a human (or your exam committee).

Prefer a printable version? The item-writing checklist template is this checker's static sibling — the same ten checks with the same citations, as a ready-to-print XLSX.

Get the checklist template
Share this tool

Frequently asked questions

Are my questions uploaded anywhere?

No. The rules run entirely in your browser. The only network calls this page makes are anonymous aggregate counters (which tool, which action) and — only if you use it — the export form, which sends your email address and nothing else. Verify it in your browser's network tab.

What does this check — and what doesn't it?

It checks the structural item-writing flaws that are detectable deterministically: testwiseness cues (absolute terms, the longest-option giveaway, grammatical cues, word echoes), ambiguity (vague qualifiers, negative stems), and format problems (all/none-of-the-above, unbalanced or non-homogeneous options, duplicates). It cannot tell you whether the content is medically or factually right, whether the distractors are plausible, or whether the question tests what you intend — that needs expert review.

How does bulk mode work?

Paste up to 50 questions, one per block separated by a blank line. Option lines start with a letter — “A)” or “A.” — and a leading * marks the correct answer (marking the key unlocks the strongest check, the longest-option cue). You get a per-question grade table and your overall flaw rate against the published 46.2% benchmark.

What does the 46.2% benchmark mean?

A peer-reviewed audit of multiple-choice questions used in high-stakes nursing assessments (Tarrant et al. 2006) found 46.2% contained at least one item-writing flaw. One honest caveat: that audit used human reviewers checking a broader flaw list than the ten structural rules this tool can run, so your rate here will naturally come out lower than their methodology would find. Treat a rate below the benchmark as encouraging, not conclusive — and use the flag list to see exactly what to fix.

Is this really free?

Yes. Checking and the copyable report need nothing at all; the print-ready PDF export asks only for an email address. The tool exists so you can see the kind of quality checks StudyDrome applies to real exams — before and after delivery.