These worksheets are free forever. Want lessons that adapt to your child as they learn, plus progress tracking? Try Ignition Learning free.

Sign up free

Ignition Learning — Activity Sheet

Hypothesis testing & statistical significance

Mathematics · Year 12

Name: ______________________Date: ____________

Hypothesis testing provides a formal statistical framework for deciding whether observed data provides enough evidence to reject a null hypothesis (typically, a claim of 'no effect' or 'no difference') in favour of an alternative hypothesis. A p-value represents the probability of observing data at least as extreme as what was actually found, assuming the null hypothesis is true — a small p-value (conventionally below 0.05) suggests the observed result would be unlikely if there really were no effect, providing evidence against the null hypothesis. Statistical significance is not the same as practical significance: a result can be statistically significant (unlikely to be due to chance) while still being too small in real-world magnitude to actually matter practically.

Example

A trial testing whether a new teaching method improves test scores might find a p-value of 0.03, meaning there's only a 3% chance of seeing a difference this large (or larger) between groups if the teaching method genuinely had no effect — small enough to reject the null hypothesis at the conventional 0.05 threshold, but the actual size of the improvement (say, half a percentage point) would still need to be considered separately to judge whether it's practically meaningful.

Key terms

Null hypothesis:
A claim of 'no effect' or 'no difference' that a statistical test evaluates evidence against.
P-value:
The probability of observing data at least as extreme as what was found, assuming the null hypothesis is true.

Questions

  1. 1. A null hypothesis typically claims:

    • No effect or no difference
    • A guaranteed, definite effect
    • Nothing related to the data being tested
    • That the alternative hypothesis is automatically true
  2. 2. A p-value represents:

    • The probability of observing data at least as extreme, assuming the null hypothesis is true
    • A guaranteed measure of practical importance
    • The exact size of an observed effect
    • Something unrelated to probability
  3. 3. A conventional threshold for statistical significance is a p-value below:

    • 0.05
    • 5
    • 50
    • 500
  4. 4. A small p-value suggests:

    • The observed result would be unlikely if the null hypothesis were true
    • The null hypothesis is definitely, absolutely true
    • Nothing meaningful about the data
    • The result has no connection to probability at all
  5. 5. Statistical significance is:

    • Not the same as practical significance
    • Always exactly identical to practical significance
    • A concept unrelated to practical importance
    • Only relevant to very large real-world effects
  6. 6. Hypothesis testing provides a framework for deciding whether:

    • Data provides enough evidence to reject the null hypothesis
    • Data should always be immediately discarded
    • The alternative hypothesis is automatically correct with no test needed
    • No decision about the data is ever required
  7. 7. A statistically significant result can still be:

    • Too small in real-world magnitude to practically matter
    • Always guaranteed to be practically important
    • Completely unrelated to the size of the observed effect
    • Impossible to occur in real research
  8. 8. Why might a p-value of 0.03 lead a researcher to reject the null hypothesis at the conventional 0.05 threshold?

    • A p-value below 0.05 suggests the observed result would be quite unlikely if there really were no effect, providing evidence against the null hypothesis
    • A p-value of 0.03 always proves with complete certainty that the null hypothesis is false
    • The specific threshold used for statistical significance has no bearing on whether a null hypothesis is rejected
    • A p-value of 0.03 indicates the observed result is highly likely to occur even if the null hypothesis is true
  9. 9. Why might a very large sample size sometimes produce a statistically significant result even for a genuinely tiny, practically unimportant effect?

    • Larger sample sizes make it easier to detect even very small effects as statistically significant, since more data reduces the role of random chance in the result
    • Sample size has no bearing on whether a genuinely small effect can be detected as statistically significant
    • Statistically significant results are always automatically large and practically important, regardless of sample size
    • A very large sample size always makes it harder, not easier, to detect a statistically significant effect
  10. 10. Why is the distinction between statistical and practical significance important when interpreting research findings reported in the media?

    • A study result being labelled "statistically significant" doesn't automatically mean the effect is large enough to be practically meaningful or worth acting on
    • Statistical significance and practical significance always mean exactly the same thing with no meaningful distinction
    • Media reports on research findings never actually need to make this distinction clear to readers
    • A result that is not statistically significant can never be practically significant either, and vice versa
  11. 11. Why might a researcher choose a one-tailed hypothesis test (checking for an effect in only one specific direction) rather than a two-tailed test (checking for an effect in either direction) in some situations?

    • If there is a strong, specific prior reason to expect an effect only in one particular direction, a one-tailed test can be more appropriately targeted, though it requires that assumption to be genuinely justified
    • One-tailed and two-tailed tests always produce exactly identical results regardless of which direction of effect is actually being investigated
    • A one-tailed test is always the more appropriate choice for absolutely every hypothesis testing situation, regardless of context
    • The choice between a one-tailed and two-tailed test has no genuine bearing on how a hypothesis test should be conducted or interpreted
  12. 12. Why might comparing the p-values of two separate studies testing similar hypotheses NOT be a reliable way to directly judge which study found the "bigger" or "more important" effect?

    • A p-value reflects the strength of evidence against the null hypothesis given the specific sample, but doesn't directly measure effect size, so two very different p-values could still correspond to similarly sized effects (or vice versa)
    • A p-value always directly and reliably indicates the actual size of the effect being studied, making this kind of comparison perfectly valid
    • Two studies with different p-values always have effects that are correspondingly different in their actual real-world size and importance
    • P-values from different studies can always be meaningfully and directly compared to determine which effect is more practically important
  13. 13. Why might researchers set the significance threshold at 0.05 (rather than, say, 0.5) as the conventional standard for rejecting a null hypothesis?

    • A stricter threshold like 0.05 reduces the chance of concluding there is a genuine effect when the result was actually just due to random chance
    • A threshold of 0.5 would provide exactly the same level of confidence in rejecting a null hypothesis as a threshold of 0.05
    • The specific threshold used for statistical significance has no bearing on how confident researchers can be in their conclusions
    • Using a higher, less strict threshold like 0.5 would always be the more scientifically rigorous standard to use
  14. 14. Why might failing to reject the null hypothesis not be the same as proving the null hypothesis is definitely true?

    • Insufficient evidence to reject the null hypothesis could reflect a genuinely small sample, high variability, or a real effect too small to detect, not necessarily proof that no effect exists at all
    • Failing to reject the null hypothesis always definitively proves that the null hypothesis is completely true
    • There is no meaningful distinction between "failing to reject" a hypothesis and "proving" that hypothesis is true
    • A hypothesis test always either definitively proves or definitively disproves the null hypothesis with complete certainty
  15. 15. Why might "p-hacking" (repeatedly testing data in different ways until a p-value below 0.05 is found) be considered a serious problem in research integrity?

    • Testing data in many different ways increases the chance of finding a "significant" result purely by chance, even when no genuine underlying effect actually exists
    • Repeatedly testing data in different ways always increases the genuine reliability and validity of a research finding
    • P-hacking has no real connection to whether a statistically significant result reflects a genuine underlying effect
    • Finding a p-value below 0.05 through repeated testing always indicates a real, and never a coincidental, effect
  16. 16. Why might a single study reporting statistical significance be considered less reliable evidence than a consistent pattern of similar results replicated across multiple independent studies?

    • Replication across independent studies reduces the chance that a single result was a false positive due to random chance, providing stronger cumulative evidence for a genuine effect
    • A single study reporting statistical significance is always exactly as reliable as a consistent pattern replicated across many independent studies
    • Replication of a result across multiple studies never actually adds any additional confidence beyond a single significant study
    • The number of independent studies supporting a finding has no bearing on how reliable that finding should be considered
  17. 17. Why might researchers report a confidence interval alongside a p-value, rather than relying on the p-value alone?

    • A confidence interval provides additional information about the plausible size and precision of an effect, which a simple significant/not-significant p-value result doesn't communicate on its own
    • A p-value alone always provides exactly as much useful information as reporting a p-value alongside a confidence interval
    • Confidence intervals and p-values always convey identical information with no meaningful additional insight from including both
    • The size and precision of an observed effect are never relevant considerations once a p-value has been calculated
  18. 18. Why might understanding hypothesis testing and statistical significance be considered an important skill for critically evaluating claims in areas like medicine, psychology or public policy?

    • Many claims in these fields are based on statistical research, so understanding how to interpret significance (and its limitations) supports more genuinely informed evaluation of the evidence behind such claims
    • Statistical significance and hypothesis testing have no real relevance to evaluating claims made in medicine, psychology or public policy
    • Claims made in these fields are never actually based on any kind of statistical research or hypothesis testing
    • Understanding hypothesis testing provides no additional benefit for critically evaluating research-based claims in any field
  19. 19. Why might a "false positive" (incorrectly rejecting a true null hypothesis) and a "false negative" (incorrectly failing to reject a false null hypothesis) both represent genuine risks that hypothesis testing cannot completely eliminate?

    • Any statistical test based on a sample carries some inherent chance of error in either direction, so researchers must weigh and accept these risks rather than assuming a test result is always perfectly correct
    • Hypothesis testing always produces a completely error-free result with absolutely no risk of either type of mistake occurring
    • Only false positives are ever a genuine concern in hypothesis testing, while false negatives never actually occur
    • The risk of drawing an incorrect conclusion from a hypothesis test can always be completely eliminated with sufficiently careful analysis
  20. 20. Why might a medical trial specifically pre-register its hypothesis and analysis plan before collecting any data, rather than deciding what to test after seeing the results?

    • Pre-registration helps prevent researchers from selectively choosing which analyses to report based on what happens to reach statistical significance, improving the genuine reliability of the reported findings
    • Pre-registering a hypothesis and analysis plan before data collection provides no genuine benefit to the reliability of a medical trial's findings
    • Deciding on an analysis approach after already seeing the data always produces exactly as reliable a result as pre-registering it beforehand
    • This kind of pre-registration practice is never actually used in real medical or scientific research
  21. 21. Understanding hypothesis testing and statistical significance mainly helps you to:

    • Evaluate whether observed data provides meaningful evidence against a null hypothesis, while distinguishing statistical from practical significance
    • Assume a statistically significant result always indicates a large, practically important effect
    • Ignore the role sample size plays in determining whether a small effect reaches statistical significance
    • Treat failing to reject a null hypothesis as definitive proof that no effect exists

Answer key (parent copy)

  1. 1. No effect or no difference
  2. 2. The probability of observing data at least as extreme, assuming the null hypothesis is true
  3. 3. 0.05
  4. 4. The observed result would be unlikely if the null hypothesis were true
  5. 5. Not the same as practical significance
  6. 6. Data provides enough evidence to reject the null hypothesis
  7. 7. Too small in real-world magnitude to practically matter
  8. 8. A p-value below 0.05 suggests the observed result would be quite unlikely if there really were no effect, providing evidence against the null hypothesis
  9. 9. Larger sample sizes make it easier to detect even very small effects as statistically significant, since more data reduces the role of random chance in the result
  10. 10. A study result being labelled "statistically significant" doesn't automatically mean the effect is large enough to be practically meaningful or worth acting on
  11. 11. If there is a strong, specific prior reason to expect an effect only in one particular direction, a one-tailed test can be more appropriately targeted, though it requires that assumption to be genuinely justified
  12. 12. A p-value reflects the strength of evidence against the null hypothesis given the specific sample, but doesn't directly measure effect size, so two very different p-values could still correspond to similarly sized effects (or vice versa)
  13. 13. A stricter threshold like 0.05 reduces the chance of concluding there is a genuine effect when the result was actually just due to random chance
  14. 14. Insufficient evidence to reject the null hypothesis could reflect a genuinely small sample, high variability, or a real effect too small to detect, not necessarily proof that no effect exists at all
  15. 15. Testing data in many different ways increases the chance of finding a "significant" result purely by chance, even when no genuine underlying effect actually exists
  16. 16. Replication across independent studies reduces the chance that a single result was a false positive due to random chance, providing stronger cumulative evidence for a genuine effect
  17. 17. A confidence interval provides additional information about the plausible size and precision of an effect, which a simple significant/not-significant p-value result doesn't communicate on its own
  18. 18. Many claims in these fields are based on statistical research, so understanding how to interpret significance (and its limitations) supports more genuinely informed evaluation of the evidence behind such claims
  19. 19. Any statistical test based on a sample carries some inherent chance of error in either direction, so researchers must weigh and accept these risks rather than assuming a test result is always perfectly correct
  20. 20. Pre-registration helps prevent researchers from selectively choosing which analyses to report based on what happens to reach statistical significance, improving the genuine reliability of the reported findings
  21. 21. Evaluate whether observed data provides meaningful evidence against a null hypothesis, while distinguishing statistical from practical significance