Writing quiz questions that measure understanding, not recall
The easiest question to write is the worst one to publish. A practical guide to turning recall questions into decision questions: choosing the type, writing the wrong options, doing the guessing arithmetic, and setting a pass mark that means something.
· فريق دورة
In short
A recall question measures who still had the slide open; an understanding question puts the learner in a situation and asks for a decision. The practical test: if someone who never took your course could answer it after a thirty-second search, the question is measuring search skill. Write for a quiz that is open book by default, pick the question type by what you actually want to measure, and work out the guessing floor before you set a pass mark.
A trainer sent me the quiz from her digital advertising course. Twelve questions, nine of them opening with "which of the following best defines". Ninety-eight percent passed. Two weeks later, in the live session, she asked the group to pick an ad channel for a specific budget, and the room went silent.
The grading was not wrong. The questions were. They measured who still had the last slide open, not who understood it. That is the default outcome when a quiz gets written in the hour before publishing.
This article is about the question itself: how to tell whether yours measures understanding, how to turn recall into a decision, and what a quiz cannot measure however well you write it.
Why the recall question wins, and why it no longer holds
The recall question does not win because anyone believes in it. It wins because it is cheap. Ten minutes with your own slides produces ten definition questions. One question that measures judgement asks you to remember where people actually go wrong, build a situation around it, and write three wrong answers that each look reasonable. The cost gap is the whole explanation.
Then comes the false reward. A high pass rate looks like proof of good teaching when it is proof of an easy question. A quiz everyone passes tells you nothing you can use, because it never separates the learners who understood from the ones who did not, and that is the definition of a broken instrument.
So run a fast test on any question you have written: if someone who never took your course could answer it after a thirty-second search, the question measures search skill, not your course.
That test bites harder now than it used to, because the quiz is open book whether you like it or not. Your learner sits it in a browser. In the next tab is your material, a search engine, and an AI assistant that answers in two seconds. No platform closes those tabs, and a quiz designed as if they were closed is built on ground that does not exist.
That is not a problem. It is a constraint that improves your questions. A question that loses its value the moment a learner opens a new tab never had value. A question whose answer is a decision survives, because a search hands your learner information and never hands them judgement. The difference is literal:
- A question that collapses: "What is the default attribution window?" One line, found in seconds.
- A question that holds: "A client sells a monthly subscription and most purchases happen eleven days after the first visit. Which setting do you check first, and why?" The answer is in your material too, but reaching it means understanding how a long buying cycle relates to that window.
Stop asking for the fact. Ask about the situation where the fact is needed.
Every question type measures something and misses something
Quizzes and the question bank give you five types, and picking the type is half the work. The wrong type makes a good question look ambiguous. The right one exposes understanding without any extra words.
| Type | Good at measuring | Fails when | Practical note |
|---|---|---|---|
| Multiple choice, one answer | Telling close alternatives apart | The wrong options are invented | Its quality is its wrong options |
| Multi-select | Completeness: three conditions, not two | Learners miss that several apply | Harder than it looks, so use it deliberately |
| True or false | Breaking one common misconception | You need a precise distinction | The guessing floor is 50% |
| Matching two columns | Concept to example, tool to job | The relation is causal or sequential | Best for terms learners confuse |
| Ordering | A procedure whose order has a reason | The order is convention, not logic | Ask whether swapping two steps really hurts |
The practical rule: start from what you want to happen inside your learner's head, then pick the type that forces it. If you want them to separate two lookalike options, single choice is enough. If you want them to know a third condition exists, only multi-select will reveal whether they do. The worst way to choose is to pick whichever type is fastest for you to write.
The wrong options are the real question
A good question gets written twice: once for the correct answer, once for the errors. A good wrong option has a single definition, a mistake a real learner actually makes. You already own the sources for those: the errors that repeat in your graded assignments, the questions that come up in every session, and the messages that land after each cohort.
These habits quietly ruin a question:
- The correct option is longer and more hedged than its siblings. Experienced test-takers pick it without reading the question.
- "All of the above" and "none of the above". Checking two options gets a learner to the answer without understanding the third.
- The joke option. It is eliminated instantly, so four options are really three and the guessing floor is 33%.
- The correct answer sitting in the same position every time. Detected by question three.
- Double negatives: "which of the following is not incorrect". That measures parsing, not your material.
The questions that survive that review are worth keeping, because a well-observed wrong option is the most expensive thing in the whole quiz. The bank stores them, so you reuse them in another quiz or another course without retyping a word. Over time you hold a stock of tested questions that assembles a full quiz in minutes and lets you vary it from cohort to cohort instead of shipping the same paper every round.
The pass mark and the gate are teaching decisions
Before you pick any number, do the guessing arithmetic instead of ignoring it. A four-option question has a 25% guessing floor, and true or false has 50%. Ten true or false questions with a 60% pass mark are cleared by random answering in 38% of attempts. If that quiz is what unlocks the next lesson, you are unlocking it with a coin.
The settings are yours: how many questions appear, what counts as a pass, what your learner sees after submitting, and whether the next lesson stays locked until they clear it. The trouble is that these numbers usually get set by instinct, when in fact they decide who finishes your course and who stops.
On the pass mark, derive it rather than rounding it. Ask of each question whether getting it wrong is acceptable in someone about to move on. The number of questions where the answer is no is your pass mark. Count in questions, not percentages: in a five-question quiz, 80% allows exactly one mistake, which looks moderate on paper and is strict in practice.
On the gate, turn it on where the next lesson genuinely depends on this one. Where the order is a suggestion rather than a prerequisite, leave the quiz as a measurement rather than a barrier. A lock in the wrong place teaches your learner to guess, because now they want passage rather than understanding.
On what they see after submitting, the trade is simple: the more you reveal, the more they learn from the attempt, and the faster your questions leave your hands. Reveal freely in a short formative quiz inside a lesson. For a final quiz in a course that takes cohort after cohort, varying the questions matters more than showing the answers straight away.
AI writes the draft, the judgement stays yours
AI generation here is bounded to your own sources rather than general knowledge, and that constraint is exactly what makes it useful: the questions come out of the material you uploaded and stay inside what you actually taught. It is excellent at the tedious part, handing you a twenty-question draft on the chapter you kept postponing, so your bank grows faster than you could type it.
But know the limit. The model knows your material and does not know your learners. The options it writes are plausible rather than observed, and that gap is most of what makes a question strong. The worst case is a question with two defensible answers: an option the model marked wrong that is right in practice, which comes back as a complaint from a learner who has a point.
So you are the final editor. Read every generated question while picturing a learner disputing it in writing, then fix the wording, fix the options, or delete it.
And your final review is not a read, it is a sitting. Simulation mode walks you through the quiz exactly as your learner walks it: the order of the questions, answering them, the result screen. That single pass reveals what reading cannot, because you read a question knowing the answer and you sit it not knowing.
Note four things on the way: a question you hesitated on, two options that say the same thing in different words, a question unanswerable without the exact phrasing from the slide, and a pass mark you did not reach yourself.
Recurring mistakes, and limits no quiz crosses
- One quiz at the end. Two questions after each lesson tell you exactly where understanding broke. Thirty questions at the end tell you only that something broke.
- Measuring phrasing instead of the idea. If swapping a word for its synonym changes the answer, you are testing memorised wording.
- The same paper for every cohort. After two rounds the questions circulate in learner group chats and your results stop meaning anything.
- Gating every single lesson. Three locks in a row teach a learner to guess until they get through.
- Publishing AI output as it came. The draft is not the quiz, and reviewing it is not an optional step.
Avoid all five and there are still limits no quiz crosses, however well you write it:
- No platform judges the quality of your question. Grading is automatic. Judging the question stays with you alone.
- A quiz does not measure production. Someone who picks the right definition may still build the wrong campaign. What is proven by doing belongs in an assignment, not a question.
- Nothing prevents searching during an attempt. Write questions that survive it instead of waiting for a lock that does not exist.
- Today's score is not learning a month from now. A question straight after the lesson measures short-term memory, so ask it again later to see what stayed.
- Results tell you who failed, not why. The why comes from the pattern across questions and from asking them in the session.
Do not rewrite everything today. Take the last course you published and ask which three decisions a graduate has to get right. Write one question per decision that puts the learner in a situation where they have to make it. Three of those tell you more than thirty definition questions, and after five courses they are also the first fifteen questions in your bank.
Then read the results as curriculum, not as grades. Each result lands on the learner's record next to their progress: who passed, who needs another attempt, who never started. The question most of them got wrong is not a problem with your learners, it is a lesson to re-explain in your next live session. The ones who need to apply an idea rather than answer about it get an assignment. And for those who genuinely cleared it, the certificate deserves to be tied to passing the quiz rather than to finishing the video, which is the whole reason it means anything.
Quizzes and the question bank are on the Pro plan, and the 14-day free trial needs no card and opens the Pro features so you can try them before subscribing. Try one well-written question on your next cohort and you will see the difference in the first set of results.
Frequently asked questions
How do I know if my question measures understanding or recall?
Run one test: could someone who never took your course answer it after a thirty-second search? If yes, the question measures search skill. A question that measures understanding puts the learner in a situation and asks for a decision, so having the fact is not enough to answer it.
How many questions should a quiz have?
The count is not the issue, the distribution is. Two or three questions after each lesson tell you exactly where understanding broke, while one large quiz at the end tells you only that something broke. How many questions appear is a setting you control on the quiz.
Should I gate every lesson behind a pass?
No. Turn gating on where the next lesson genuinely depends on this one, and it stays locked until the learner reaches the pass mark you set. Where the order is a suggestion rather than a prerequisite, leave the quiz as a measurement, because a lock in the wrong place teaches fast guessing.
Can I trust AI-generated questions?
Trust them as a draft. Generation is bounded to your own uploaded sources rather than general knowledge, so the questions stay inside what you taught. But the model knows your material and not your learners' mistakes, and it can produce a question with two defensible answers. Review every one before publishing.
How do I set the right pass mark?
Derive it from the questions instead of rounding to a familiar number. Ask of each question whether getting it wrong is acceptable in someone moving on. The count where the answer is no is your pass mark. Count in questions, not percentages: in a five-question quiz, 80% allows exactly one mistake.