A student who can execute every inference procedure in the course can still lose the question, because choosing the procedure happens before any of that skill is used. The framework treats selection as its own thing, attaching the skill identify appropriate statistical inference methods to eleven separate topics, and the reason is that the choice is settled by the study design rather than by the numbers.

Twelve procedures survive in the revised course: six families, each available as an interval and as a test. That sounds like a lot to sort until the sorting is written down, at which point it is three questions.

Three questions, in order

What kind of response was recorded? If each individual contributes a category — yes or no, brand A or B, resolved or not — the parameter is a proportion. If each individual contributes a number, it is a mean. Nothing else about the problem matters until this is answered, because it decides which half of the course you are in.

How many groups, and where did they come from? One group gives a one-sample procedure. Two independently collected groups give a two-sample one. A single group classified two ways gives a two-way table and a chi-square procedure.

If there are two sets of numbers, are they paired? Two measurements on the same individuals, or on individuals deliberately matched, is not two samples. It is one sample of differences.

Only after those three is there a fourth: is the question asking what the parameter is, or whether a claim about it survives? The first is an interval and the second is a test.

The drill

Classify the design. No arithmetic, and no numbers are supplied for any.

Twelve scenarios, cycling. Both rows have to be answered before the verdict appears, because the two decisions are independent: the family comes from how the data were collected and the purpose comes from what the question asks for. The feedback never names the arithmetic. It names the feature of the design that settled the choice, which is the only part that transfers to a scenario you have not seen. Four of the twelve are built as confusable pairs and are worth returning to.

The pair that looks like two samples

Scenarios three and eight are paired designs, and both invite the two-sample answer because two sets of numbers arrive. The framework’s instruction is unusually direct about what to do instead: for a matched pairs design with two dependent samples, the appropriate analysis calculates differences between pairs of values to produce one sample of differences, and the procedure is a one-sample \(t\)-interval for a population mean difference.

So a paired design does not get its own family of formulas. It gets converted into a one-sample problem before any formula is used, and the parameter changes with it — from \(\mu_1 - \mu_2\), the difference of two population means, to \(\mu_d\), the mean of a population of differences. Those are different quantities, and defining the parameter correctly is where the distinction is graded.

The tell is never the arithmetic. It is whether the two numbers in a pair came from the same individual, or from two individuals matched deliberately. Forty runners timed twice is forty differences. Forty city employees and forty suburban ones is two samples, and no amount of equal sample size makes them pairs.

The pair that looks like one table

Scenarios five and six both end in a two-way table, both use \(\chi^2 = \textstyle\sum (O-E)^2/E\), and both have \((r-1)(c-1)\) degrees of freedom. The tables can be identical. Only the design separates them: one sample classified two ways is a test for independence, and several separately collected samples compared on one variable is a test for homogeneity.

The framework keeps the distinction alive even in the conditions, where the randomisation requirement is worded one way for independence and another for homogeneity. A procedure whose arithmetic is identical and whose conditions differ is a procedure that is really two.

Estimate, or judge a claim

The second row of buttons is a separate decision and it fails separately. A question that supplies a specific value to argue about — an advertised 600 calories, a claimed 60% — is offering a null hypothesis and wants a test. A question that asks how large something is, or by how much two things differ, wants an interval.

The wording is reliable enough to use as a rule. Is there convincing evidence that, do the data support, and test the claim are tests. Estimate, find a range of plausible values, and by how much are intervals. Scenario seven mentions 500 millilitres and is still an interval, because the number arrives as context rather than as a claim to be judged.

A drill in the same spirit as the tool, but harder: take a released free-response question, cover everything after the first sentence of the stem, and name the procedure from the design alone. Then uncover the rest and check. The scenarios that resist are the ones where the design is described last, which is a writing choice rather than a statistical one, and noticing it is worth more under time pressure than any formula on the reference sheet.