QED
Statistics · step 9 of 13

Hypothesis testing & p-values

A hypothesis test asks whether the data are surprising under a null hypothesis H₀. The p-value is the probability of a test statistic at least as extreme as the observed one ASSUMING H₀ is true — it is not the probability that H₀ is true, and that misreading is the single most common statistical error in published work.

Unlimited questions · marked criterion by criterion · no card needed

Method: how to approach it

The order below is what examiners expect to see, and each step carries its own marks.

  1. State H₀ and H₁ symbolicallyH₀ always contains the equality. Decide one- or two-tailed BEFORE seeing the data.
  2. Compute the test statisticFor a mean, t = (x̄ − μ₀)/(s/√n) with df = n − 1.
  3. Find the p-valueThe tail area beyond the statistic; double it for a two-tailed test.
  4. Conclude in contextCompare with α, then write "there is (in)sufficient evidence at the 5% level to conclude …" — never just "reject H₀".

Worked example

A machine should fill 500 ml. A sample of 16 bottles gives x̄ = 495 and s = 8. Test at the 5% level whether the mean differs from 500.

  1. H₀: μ = 500 versus H₁: μ ≠ 500, a two-tailed test.
  2. t = (495 − 500)/(8/√16) = −5/2 = −2.5, with df = 15.
  3. The critical value t₀.₀₂₅,₁₅ ≈ 2.131, and |−2.5| > 2.131.
  4. The two-tailed p-value is approximately 0.024 < 0.05.

Answer. Reject H₀: there is sufficient evidence at the 5% level that the mean fill differs from 500 ml.

Where marks get dropped

These are the specific errors that cost credit on hypothesis testing & p-values questions — QED's rubric penalises each of them separately.

Practise this until it is automatic

Unlimited fresh questions

QED generates new hypothesis testing & p-values problems on demand at warm-up, exam and challenge level, so you can drill this one skill until it stops costing you marks.

Marked like an examiner

Every answer is scored against a point-by-point rubric with partial credit, so you see exactly which step of the method broke down — not just a tick or a cross.

Answer in real notation

A one-tap symbol palette, a visual equation editor and a truth-table builder — or photograph your handwritten working and QED converts it to LaTeX.

Saved to your library

Every question you generate is kept and re-takeable as a timed exam, and your Statistics mastery is tracked so you know when this is exam-ready.

Hypothesis testing & p-values — frequently asked questions

What exactly is a p-value?

The probability, assuming H₀ is true, of obtaining a test statistic at least as extreme as the one observed. Small means the data are surprising under H₀.

Why is 0.05 the threshold?

Pure convention, from Fisher. It carries no special mathematical status, and many fields now prefer reporting effect sizes and intervals instead.

What is p-hacking?

Testing many hypotheses or stopping when a result becomes significant. With 20 independent tests at α = 0.05, one false positive is expected by chance alone.

The rest of Statistics

Describing data, distributions, estimation and hypothesis tests. Each subtopic below has its own method, worked example and mark-losing traps.

Ready to make hypothesis testing & p-values exam-proof?

Generate your first questions free — no card, no setup, no personal data stored. Practise until the method is second nature.

Start practising free →