There is a particular silence that follows a badly built quiz question. Not the good silence, the one where six people hunch over a beer mat and argue productively for thirty seconds. The other kind: the flat, deflated pause in which everyone has privately concluded the question was never really answerable, and the only suspense left is how loudly they are allowed to grumble when the answer is read out.
Avoiding that pause turns out to be closer to engineering than to taste. People who set questions for a living, whether for an exam board, a licensing authority or a Tuesday night in the back room of a pub, judge their work against standards that have little to do with whether they personally find the subject fascinating. A question can be assessed on its own account, separately from the people answering it, and quite often it fails.
Hard is a compliment. Sneaky is not.
The first line any setter draws is between difficulty and trickery, and the two are frequently confused by people who have never had to write forty questions by Thursday. Difficulty is the amount of knowledge or reasoning standing between a reader and the answer. Trickery is a penalty for reading in the ordinary way: a negative buried in the middle of a sentence, a plural that quietly changes the target, a clause that reverses everything if you happen to notice a comma.
The diagnostic is simple and slightly brutal. When the answer is announced, does the room say "of course" or does it say "hang on"? A genuinely hard question produces the first noise even from the tables that missed it. A tricky one produces the second even from the tables that got it, because they know perfectly well they got it by accident.
Assessment researchers have spent decades codifying that instinct. In a widely cited review, Thomas Haladyna, Steven Downing and Michael Rodriguez worked through 27 textbooks on educational testing and a further 27 published studies to assemble a taxonomy of 31 guidelines for writing multiple-choice items. A striking number of them counsel restraint rather than cleverness. Phrase the stem as a complete question rather than a sentence with a hole in it. Move any wording repeated across the options up into the stem. Avoid negatives, and where you cannot avoid them, make them impossible to miss. Never write an answer that is correct only on a technicality, however satisfying the technicality feels at midnight.
The question gets marked too
What lifts this above opinion is that questions carry numbers of their own, and two of them do most of the work.
The first is the difficulty index, generally written as a p-value: the proportion of people who answered correctly. A question that 95 per cent of the room gets and one that 4 per cent gets are opposite-looking failures, but both are failures, because neither told you anything about the people in front of you. Setters sorting a mixed field want most items landing somewhere in the middle, with a deliberate scattering at either end to give the strongest and the shakiest something to do.
The second number is discrimination, and it is the one that separates a professional from an enthusiast. Take the highest-scoring group overall and the lowest-scoring group, then subtract the proportion of the weak group who answered the item correctly from the proportion of the strong group who did. If the people who know the most are getting it and the people who know the least are not, the figure is positive and healthy. Testing handbooks commonly treat anything around 0.3 or above as sound and anything under roughly 0.2 as due for a rewrite.
A negative discrimination value is a klaxon. It means the best-informed answerers were systematically pulled away from the answer you marked correct, which nearly always means the wording is ambiguous, a second option is defensible, or the setter has made an error that only the well-read noticed. Anyone who has watched a quiz collapse into a twenty-minute row about whether a country counts as landlocked has seen it happen.
The wrong answers do most of the work
Ask a setter which part of a multiple-choice item takes longest and they will not say the question. They will say the distractors. The wrong options are where the difficulty actually lives, and a weak one throws the whole item away.
Analysts talk about non-functioning distractors: options chosen by so few people, conventionally under about five per cent, that they are effectively invisible. A four-option question with two dead distractors is a two-option question wearing a disguise, and it hands roughly half the marks to anyone willing to guess. This is why Michael Rodriguez, reviewing eighty years of research in a well-known meta-analysis, concluded that three options are usually optimal. Writing a genuinely plausible fourth wrong answer costs real effort and, more often than not, produces a dud. Better to write three that all bite.
Good distractors are harvested rather than invented. They come from the mistakes people actually make: the adjacent fact, the thing everybody half-remembers from school, the answer that would be perfectly correct to a slightly different question. Pub setters read the marked sheets afterwards and steal the best wrong answers for next time. Exam boards pre-test items on live cohorts and watch where the errors cluster. A distractor should be a place a reasonable person could honestly end up.
The tells that give it away
The counterweight to all this is testwiseness: the habits by which a canny answerer extracts marks from the shape of a question rather than its content. Setters spend real energy closing these leaks.
The longest, most carefully hedged option is very often correct, because the setter took pains to make it precisely true and rattled off the wrong ones. Grammatical agreement betrays answers constantly: an "an" before the options quietly eliminates every choice beginning with a consonant. Absolute words such as always or never tend to mark a wrong answer, since reality is rarely that tidy, while cautious words such as usually tend to mark a right one. Two old standbys attract particular scorn. "All of the above" hands the item to anyone who spots two correct choices and stops reading, and "none of the above" establishes only that the other options are wrong, which is not the same as establishing that the answerer knows anything.
Obscure is cheap. Interesting is expensive.
Here is the distinction that costs setters the most sleep. Obscurity is free. Anyone with a reference shelf and ten idle minutes can produce a question nobody can answer, and the resulting difficulty is entirely fake, because it measures access to the reference shelf rather than anything about the answerer.
Interest is expensive because it requires the answer to be reachable from something the reader already owns. The finest questions are about a relationship rather than a fact: why a border does something peculiar, why two unrelated-looking words share a root, why a familiar object has a feature nobody can explain. The answerer does not retrieve the answer so much as assemble it, and the pleasure of assembly is the whole product.
The shape of an evening
Individual items are only half the craft; the other half is sequence. A round that opens with its hardest question is a door slammed in the room's face. Experienced setters open with something almost everybody gets, not out of charity but because it is a handshake. It puts pens on paper, settles the table into a rhythm, and signals that the quiz is broadly on their side.
From there the curve climbs, with the sharpest item saved for last so the round ends on the sound setters actually work for: a collective groan of near-misses mixed with one delighted shriek from the far corner. Across the night the aim is a spread. If every team finishes a round on seven or eight out of ten, that round was decoration. Picture and music rounds redistribute expertise, handing the hero role to somebody who has not contributed since the first question.
Underneath the statistics, that is what the discipline is for. A well-made question is a small act of hospitality: hard enough to be worth answering, fair enough that nobody feels swindled, and interesting enough that even the table in last place leans in to hear how it turned out.







