Where the timeline lands, claim by claim.
Supported
Claim A: cloze cards work
Supported, narrowly. Repeated testing with electronic flashcards beats restudying the same material (Schmidmaier et al., 2011), and a decade of observational studies links Anki volume to Step 1 scores, with a
dose-response on cards reviewed
(Chernov & Alben, 2026).
But the boundary is sharp. Fill-in-the-blank retrieval helps only the items repeated
verbatim (Hinze & Wiley, 2011), and the specificity is item-level
(Rickard & Pan, 2019).
So the claim survives for verbatim recall, not for the reworded, transfer-heavy
questions Step and the MCAT actually ask.
Revise
Claim B: cloze is the right default for medical and MCAT decks
This is the weak link. Three problems:
- Kang, McDermott & Roediger (2007)
found production beats selection only when feedback closes the loop, and a cloze
back that re-shows the same sentence with the blank filled adds no information.
- Little et al. (2012)
showed multiple choice with competitive distractors also teaches the material attached to
the wrong answers. Cued recall, cloze's category, does not.
- Pan & Rickard (2018) found
the testing effect roughly halves without response congruency, and in the field,
the Anki benefit stops at the recall exam: present for Step 1, absent for Step 2 CK
(Wothe et al., 2023) and the MCAT
(Rowe et al., 2025).
Revised: cloze is a speed default, not an outcome default. The
formats with head-to-head wins add rivals to reject, production with real feedback, or an
explanation on the back.
Untested
Claim C: the direct comparison
No study has randomized cloze-format cards against Q&A-format cards, on identical
medical content, inside a spaced-repetition system, with exam outcomes. The closest attempt, van Wijk et al. (2024), omitted
feedback entirely and found null, exactly the condition Kang predicted would erase the
production advantage.
The field's own taxonomy paper
(Balczewski et al., 2025)
concedes in print that existing studies exclude card design, and notes that because Anki
logs everything, the comparison could be run quickly.