Two variables being associated (correlated) — tending to change together — is not the same as one causing the other. Genuine cause-and-effect requires stronger evidence than just observing a pattern; a third, hidden variable might be driving both, or the relationship could simply be coincidental. Two-way tables organise data by two categorical variables at once (like survey responses split by two different groups), making it possible to investigate whether a relationship actually exists between them, and how strong it appears to be — an important step before jumping to any conclusion about cause and effect.
Example
Ice cream sales and drowning incidents both rise in summer, showing a clear association — but ice cream doesn't cause drowning; a third variable, hot weather, drives both more people buying ice cream and more people swimming (and therefore more drowning risk), a classic example of confusing correlation with causation.
Key terms
Association:
A pattern where two variables tend to change together, without necessarily one causing the other.
Confounding variable:
A hidden third variable that influences both variables being studied, creating a misleading appearance of cause and effect.
Questions
1. Association between two variables means:
They tend to change together
One always causes the other
They have no relationship whatsoever
They are always measured identically
2. Causation means:
One variable directly causes a change in another
Two variables happen to change together with no real link
A type of chart only
Something that can never be tested
3. A confounding variable is:
A hidden third variable influencing both variables being studied
The same as the two variables already being studied
Something that never affects any study
Always the main cause being tested for
4. Ice cream sales and drowning incidents rising together in summer is an example of:
Association without direct causation
Direct proof that ice cream causes drowning
A relationship with no explanation possible
A two-way table
5. A two-way table organises data by:
Two categorical variables at once
Only one single variable
No variables at all
Exactly three variables always
6. Observing a pattern between two variables is:
Not, by itself, proof of causation
Always sufficient proof of causation
Completely unrelated to causation
The only type of evidence ever needed
7. Two-way tables can help investigate:
Whether a relationship exists between two categorical variables
Nothing meaningful about data
Only numerical, non-categorical data
A single variable in isolation
8. Why is a confounding variable (like hot weather in the ice cream/drowning example) important to consider before concluding causation?
It can independently explain why both variables rise together, without either one directly causing the other
Confounding variables never actually explain any observed association between two variables
A confounding variable always confirms, rather than complicates, a claim of causation
Considering additional variables never changes how an association should be interpreted
9. Why might a two-way table be useful for investigating whether smoking is associated with a health outcome across two groups (smokers vs non-smokers)?
It lets you directly compare outcome rates between the two groups, showing whether a pattern of association actually exists in the data
Two-way tables can only ever be used with data involving exactly one single group
This kind of comparison provides no useful information about a potential health association
Two-way tables are only appropriate for numerical data, never categorical groupings like smoker/non-smoker
10. Why might researchers use controlled experiments (rather than just observing existing data) when they specifically want to establish causation, not just association?
Controlled experiments can isolate and directly test the effect of one variable while accounting for or eliminating other potential influences
Observational data always provides exactly the same strength of evidence for causation as a controlled experiment
Controlled experiments provide no additional evidence for causation compared to simply observing existing data
Isolating variables through controlled conditions never actually helps establish causation
11. Why might a strong, consistent association observed across many different studies still fall short of proving causation?
Even a strong, repeated association could still be explained by an unaccounted-for confounding variable rather than a direct causal link
A strong, repeated association across many studies always automatically proves causation
The number of studies showing an association has no bearing on whether it constitutes proof of causation
Confounding variables can only ever affect a single study, never a pattern seen across many studies
12. Why might reversed causation (where the assumed "effect" is actually part of the cause) be a further complication when interpreting an observed association?
It's possible the direction of the relationship is the opposite of what seems intuitive, so the assumed cause might actually be a consequence
Reversed causation is a purely theoretical idea that never actually occurs in real research
The direction of a causal relationship is always obvious and never needs to be questioned
Reversed causation has no real connection to how an observed association should be interpreted
13. Why might a two-way table showing survey responses split by age group be useful for checking whether an apparent overall association actually holds true within each individual subgroup?
Breaking data down by subgroup can reveal whether a pattern seen in the combined data is consistent, or whether it disappears or reverses within specific groups
Splitting data by subgroup never reveals any additional information beyond what the combined, overall data already shows
An association seen in combined data always holds identically true within every individual subgroup
Two-way tables can only ever display combined, non-subdivided data with no possibility of further breakdown
14. Why might news headlines that state "X causes Y" based only on an observed association be considered scientifically premature?
Establishing genuine causation typically requires more rigorous evidence than a single observed pattern of association
Any observed association between two variables is always sufficient to confidently claim causation
News headlines about research findings are always completely accurate representations of the underlying evidence
The distinction between association and causation has no bearing on how research findings should be reported
15. Why might researchers specifically look for a "dose-response relationship" (where more of variable A consistently corresponds to more of variable B) as one piece of evidence supporting causation, even though it still isn't conclusive proof on its own?
A graded, consistent relationship adds supporting evidence beyond simple association, though it still doesn't rule out all other explanations like confounding factors
A dose-response relationship on its own is always considered complete, conclusive proof of causation
This type of relationship provides no additional evidence toward establishing causation whatsoever
Dose-response patterns are irrelevant to how scientists evaluate potential causal relationships
16. Why might identifying a plausible biological or mechanical explanation for HOW one variable could cause another strengthen a causal claim beyond just statistical association?
A credible mechanism helps explain why the pattern might be genuinely causal, rather than coincidental or explained by an unrelated confounding factor
A plausible mechanism for causation is always irrelevant once a statistical association has been observed
Statistical association alone is always sufficient to fully explain how or why a causal relationship might exist
Understanding a potential mechanism never actually adds any strength to a causal claim
17. Why might large-scale public health research increasingly combine multiple types of evidence (observational studies, controlled trials, and biological mechanisms) rather than relying on just one type?
Combining different, independent types of evidence increases confidence in a causal conclusion beyond what any single study or method alone could provide
Combining multiple types of evidence always weakens the overall strength of a research conclusion
A single type of study or evidence is always considered fully sufficient for establishing genuine causation
Different types of evidence in research always simply repeat and add nothing beyond what a single study shows
18. Why is distinguishing association from causation considered an essential skill for critically evaluating health, environmental or social claims reported in the media?
Claims are frequently based only on association, so recognising this distinction helps avoid drawing overconfident conclusions from limited evidence
The distinction between association and causation has no real relevance to how such claims should be critically evaluated
Media reports on health, environmental or social topics never actually confuse association with causation
This distinction only matters for professional scientists, never for evaluating claims as a general reader
19. Why might a randomised controlled trial (where participants are randomly assigned to different groups) provide stronger evidence for causation than an observational study of already-existing groups?
Random assignment helps ensure that, on average, confounding factors are evenly spread across groups, isolating the effect of the variable actually being tested
Random assignment has no real advantage over simply observing pre-existing groups when trying to establish causation
Observational studies of pre-existing groups always provide exactly as strong evidence for causation as a randomised trial
Confounding factors are automatically eliminated in any study, regardless of whether participants are randomly assigned
20. Why might it be misleading to claim "no evidence of an association" is the same as "evidence of no association" when interpreting a study's findings?
A study might simply have failed to detect a real association (due to small sample size or other limitations), which is different from proving no relationship genuinely exists
These two statements always mean exactly the same thing in every research context
Failing to detect an association in a study always proves conclusively that no relationship exists
The distinction between these two statements has no real bearing on how research findings should be interpreted
21. Understanding association versus causation and two-way tables mainly helps you to:
Critically evaluate claimed relationships between variables, distinguishing genuine causation from mere association
Assume any observed pattern between two variables automatically proves one causes the other
Ignore the possibility of confounding variables when interpreting a relationship
Treat two-way tables as unable to reveal any relationship between categorical variables
Answer key (parent copy)
1. They tend to change together
2. One variable directly causes a change in another
3. A hidden third variable influencing both variables being studied
4. Association without direct causation
5. Two categorical variables at once
6. Not, by itself, proof of causation
7. Whether a relationship exists between two categorical variables
8. It can independently explain why both variables rise together, without either one directly causing the other
9. It lets you directly compare outcome rates between the two groups, showing whether a pattern of association actually exists in the data
10. Controlled experiments can isolate and directly test the effect of one variable while accounting for or eliminating other potential influences
11. Even a strong, repeated association could still be explained by an unaccounted-for confounding variable rather than a direct causal link
12. It's possible the direction of the relationship is the opposite of what seems intuitive, so the assumed cause might actually be a consequence
13. Breaking data down by subgroup can reveal whether a pattern seen in the combined data is consistent, or whether it disappears or reverses within specific groups
14. Establishing genuine causation typically requires more rigorous evidence than a single observed pattern of association
15. A graded, consistent relationship adds supporting evidence beyond simple association, though it still doesn't rule out all other explanations like confounding factors
16. A credible mechanism helps explain why the pattern might be genuinely causal, rather than coincidental or explained by an unrelated confounding factor
17. Combining different, independent types of evidence increases confidence in a causal conclusion beyond what any single study or method alone could provide
18. Claims are frequently based only on association, so recognising this distinction helps avoid drawing overconfident conclusions from limited evidence
19. Random assignment helps ensure that, on average, confounding factors are evenly spread across groups, isolating the effect of the variable actually being tested
20. A study might simply have failed to detect a real association (due to small sample size or other limitations), which is different from proving no relationship genuinely exists
21. Critically evaluate claimed relationships between variables, distinguishing genuine causation from mere association