Resources
The two ways a test can be wrong
The probability of a Type I error is a number chosen in advance. The probability of a Type II error is inherited from the situation, and in a real study it is never observed at all. That asymmetry is the whole of this topic, and it accounts for a lopsidedness students notice without being able to name: \(\alpha\) is announced before the data are collected, and its counterpart usually is not announced anywhere.
Two errors, named for what they get wrong
The framework defines both by the verdict rather than by the arithmetic. A Type I error occurs when there is convincing statistical evidence that the alternative hypothesis is true, on the strength of a small p-value, but it is not. A Type II error occurs when there is not convincing evidence that the alternative is true, on the strength of a large p-value, but it is.
Each error pairs a verdict with a truth, and the two are the only ways that pair can disagree. A test that rejects a true null has raised a false alarm; a test that fails to reject a false null has missed something real.
A third definition completes the set: the power of a test is the probability that it correctly rejects a false null hypothesis. Power and the Type II error are complements, and the framework states the relationship as a formula rather than as a picture:
\[P(\text{Type II error}) = 1 - \text{power}.\]One cut, two curves
A production line is treated as in control when at most 10% of parts need rework. An inspector draws a random sample and tests \(H_0: p = 0.10\) against \(H_a: p > 0.10\), rejecting when the sample proportion lands far enough above 0.10.
Both panels below show the same cut, drawn in the same place. What changes between them is which world the sample was drawn from. Move the true rate and watch only the lower panel respond.
Dark shading is where the test rejects, and it is the same region in both panels, because the cut does not know which world it is in. Top: the line really is in control, so every dark outcome is a false alarm and the dark area is exactly the significance level. Bottom: the line really is running at the chosen rate, so the dark area is now the test working and the pale area is the failure to notice. Heights are scaled to fit; only the shaded areas are probabilities. Raise n and both curves narrow while the top area holds, which is the one thing on screen that never moves on its own.
What each control moves, and what it leaves alone
The framework lists four things that raise power, and attaches the same clause to all four: provided the others do not change. The three controls above are those four factors, because sample size and standard error are one lever rather than two.
Raise \(n\) from 200 to 400 and power climbs from 0.7252 to 0.9220. Both curves narrow, the cut slides left, and the top panel’s dark area stays at 0.05 throughout. That is the first thing worth watching. The significance level is a constraint the cut is built to satisfy, so it cannot drift; everything else on screen is free to move.
Drag the true rate from 0.15 out to 0.20 and power runs to 0.9893, with nothing about the test changed. A test is not powerful or weak on its own, only powerful against a particular alternative, and a discrepancy far from the null is easier to see than one nearby. Drag the other way, to 0.12, and power falls to 0.2585 — the same test, now missing a real problem three times in four.
Loosening \(\alpha\) to 0.10 raises power to 0.8169, and tightening it to 0.01 drops power to 0.5103. This is the only one of the four that is a trade rather than a gain. The others buy power with sample size or receive it as a gift from the truth; this one buys it by accepting more false alarms.
Power at the null
Slide the true rate all the way down to 0.10, so the two panels show the same curve. The readout stops calling the lower area power in any useful sense and reports a number equal to \(\alpha\) exactly — 0.0500 at the default setting, 0.0100 if the level is tightened.
The coincidence is not a coincidence. Power is the probability of rejecting, evaluated at whatever the truth is; at the null value that is the probability of rejecting a true null, which is the definition of \(\alpha\). So \(\alpha\) is the beginning of the power curve rather than a separate quantity, and the two errors are measured on one continuum instead of belonging to different worlds.
What the exam asks of this
Less arithmetic than the picture suggests. The relationship \(P(\text{Type II error}) = 1 - \text{power}\) is the calculation, so a question supplying a power of 0.80 is asking for 0.20 and nothing more. Nothing on the exam requires shading a normal curve to produce \(\beta\) from scratch.
What is assessed is identification, the four factors, and consequences. Consequences are where the two errors stop being symmetric: halting a line that is running properly costs production, while letting a faulty line run costs the parts that reach customers, and the framework asks which is worse to be settled before the study rather than after. That judgment sets \(\alpha\), because \(\alpha\) is the probability of a Type I error, and it sets \(n\), because sample size drives the probability of the other. Neither number is chosen for a statistical reason, and both are fixed before any data exist — the same discipline that makes a p-value interpretable at all once it arrives.
A prediction to test against the interactive: set the true rate to 0.15 and the level to 0.05, then find the smallest sample size that lifts power to the 0.80 benchmark. Guess first, nearer 250 or nearer 500. It is 253. Now ask what that number becomes if the inspector cares about a rate of 0.12 instead, and notice the question cannot be answered without naming an alternative. Power is never a property of a test alone, which is why a study reporting it always reports the discrepancy it was built to catch.