Citation timeline · 1990–2026 · 35 entries · 8 recurring authors

Is cloze deletion the right default?

Thirty-six years of doctrine, decks, and lab results on how a flashcard should ask its question. The premise under examination: fill-in-the-blank is the right default format for medical and MCAT flashcards. Every entry below compares one way of asking against another, or measures what the decks people actually use contain. Three separable claims. The evidence treats them very differently.

Claim A · supported, with a boundary

Cloze cards work. Testing beats restudy, and Anki use tracks Step 1. But the measured benefit is verbatim: the exact blank you practised, and no further.

Claim B · needs revision

Cloze is the right default. Head-to-head, the winners are formats that add rivals, production with feedback, or an explanation on the back, not a re-shown sentence.

Claim C · untested

The direct comparison. Zero trials have randomized cloze against Q&A cards on identical content. The default was settled by inheritance, not by experiment.

Verdict on each claim

Where the timeline lands, claim by claim.

Supported

Claim A: cloze cards work

Supported, narrowly. Repeated testing with electronic flashcards beats restudying the same material (Schmidmaier et al., 2011), and a decade of observational studies links Anki volume to Step 1 scores, with a dose-response on cards reviewed (Chernov & Alben, 2026).

But the boundary is sharp. Fill-in-the-blank retrieval helps only the items repeated verbatim (Hinze & Wiley, 2011), and the specificity is item-level (Rickard & Pan, 2019). So the claim survives for verbatim recall, not for the reworded, transfer-heavy questions Step and the MCAT actually ask.

Revise

Claim B: cloze is the right default for medical and MCAT decks

This is the weak link. Three problems:

  1. Kang, McDermott & Roediger (2007) found production beats selection only when feedback closes the loop, and a cloze back that re-shows the same sentence with the blank filled adds no information.
  2. Little et al. (2012) showed multiple choice with competitive distractors also teaches the material attached to the wrong answers. Cued recall, cloze's category, does not.
  3. Pan & Rickard (2018) found the testing effect roughly halves without response congruency, and in the field, the Anki benefit stops at the recall exam: present for Step 1, absent for Step 2 CK (Wothe et al., 2023) and the MCAT (Rowe et al., 2025).

Revised: cloze is a speed default, not an outcome default. The formats with head-to-head wins add rivals to reject, production with real feedback, or an explanation on the back.

Untested

Claim C: the direct comparison

No study has randomized cloze-format cards against Q&A-format cards, on identical medical content, inside a spaced-repetition system, with exam outcomes. The closest attempt, van Wijk et al. (2024), omitted feedback entirely and found null, exactly the condition Kang predicted would erase the production advantage.

The field's own taxonomy paper (Balczewski et al., 2025) concedes in print that existing studies exclude card design, and notes that because Anki logs everything, the comparison could be run quickly.

What the evidence actually specifies

Not "cloze everything." Five properties, each traceable to a result.

  1. Feedback must add information

    Production only beats selection when feedback closes the loop, and a back that re-shows the sentence is not feedback. Kang, McDermott & Roediger (2007).

  2. If you use recognition, write competitive distractors

    Plausible rivals are the active ingredient, not the multiple-choice label: four throwaway options do nothing. Little & Bjork (2015).

  3. An explanation on the back, for transfer

    Explanation feedback buys nothing on the same question and everything on a different one. Butler, Godbole & Marsh (2013).

  4. Match practice format to the exam's format

    Response congruency is the dominant moderator of transfer: d = 0.28 without it, 0.58 with it. Pan & Rickard (2018).

  5. Write cards above the retrieval floor

    Rewriting cards to target comprehension and application changed exam scores in a randomized classroom trial (Senzaki et al., 2017), while 80% of the popular decks' cards elicit only retrieval (Rajpurkar et al., 2026).