A box plot (or box-and-whisker plot) summarises a dataset's spread using five key values: the minimum, lower quartile (25th percentile), median, upper quartile (75th percentile), and maximum — making it efficient for comparing the spread and centre of multiple datasets side by side. Not every chart is a fair, accurate representation of data, though: misleading representations in media often use tricks like a broken or non-zero y-axis (exaggerating small differences), inconsistent scales, or cherry-picked data ranges that don't actually relate to the claim being made. Learning to spot these techniques is an essential critical numeracy skill for reading news, advertising, and social media.
Example
A bar graph with a y-axis starting at 90 (instead of 0) can make a difference between 91% and 94% look dramatic and enormous, when the actual real-world difference is only 3 percentage points — a classic misleading technique that relies on readers not checking the axis carefully.
Key terms
Box plot:
A chart summarising a dataset's spread using minimum, quartiles, median and maximum.
Broken axis:
A y-axis that doesn't start at zero, which can visually exaggerate small differences.
Questions
1. A box plot summarises a dataset using:
Minimum, quartiles, median and maximum
Only the mean of the dataset
A single data point only
No numerical values at all
2. The median in a box plot represents:
The middle value of the dataset
The smallest value only
The largest value only
The average of just the first two values
3. Box plots are useful for:
Comparing the spread of multiple datasets side by side
Displaying only a single number with no context
Hiding all information about a dataset
Only working with datasets of exactly one value
4. A broken or non-zero y-axis on a graph can:
Exaggerate small differences visually
Always represent data perfectly fairly
Have no effect on how a graph looks
Only ever make differences look smaller
5. Cherry-picked data ranges in a misleading chart involve:
Selecting a range that supports a claim while ignoring the fuller picture
Using the complete, unfiltered dataset always
Randomly selecting data with no intention
Always representing data in the fairest way possible
6. The lower quartile in a box plot represents:
The 25th percentile of the data
The very highest value in the dataset
The mean of the dataset
The 90th percentile of the data
7. Spotting misleading chart techniques is an important skill for:
Critically reading news, advertising and social media
Something with no real-world use
Only professional statisticians, never regular readers
Ignoring all graphs and charts entirely
8. Why might a y-axis that doesn't start at zero distort a viewer's perception of the actual size of a difference between two values?
It can make a visually large gap on the chart represent only a small real numerical difference, exaggerating the apparent change
A non-zero y-axis always represents differences with perfect visual accuracy
The starting point of a y-axis has no bearing on how a viewer perceives the data
This technique always makes differences appear smaller than they actually are
9. Why is comparing the interquartile range (the box itself) between two box plots useful for understanding data spread?
It shows how concentrated or spread out the middle 50% of each dataset is, allowing a direct visual comparison
The interquartile range provides no useful information about a dataset's spread
Box plots cannot be used to compare more than one dataset at a time
The middle 50% of a dataset is never relevant to understanding its overall spread
10. Why might using inconsistent scales between two side-by-side charts intended for direct comparison be considered misleading?
Different scales can make two genuinely different-sized quantities appear visually similar, or vice versa, distorting a fair comparison
Inconsistent scales always make a comparison between two charts clearer and more accurate
Scale consistency between charts has no bearing on how fairly a comparison can be made
Readers always automatically notice and mentally correct for inconsistent scales
11. Why might selecting a narrow, cherry-picked time range in a chart (rather than the full available data) mislead a viewer about an overall trend?
A short, selectively chosen window might show an unrepresentative pattern that doesn't reflect the fuller, longer-term trend
A narrow time range always accurately represents the fuller, longer-term trend
The time range chosen for a chart has no bearing on how representative it is of an overall trend
Cherry-picked data ranges are always immediately obvious to any viewer
12. Why might a box plot showing overlapping ranges between two datasets suggest less of a meaningful difference than a bar chart comparing only their averages?
A box plot reveals the full spread of each dataset, showing whether individual values genuinely differ, not just the summary averages
A box plot always hides any meaningful difference between two datasets, unlike a bar chart
Comparing spread and comparing averages always lead to exactly the same conclusion about two datasets
Overlapping ranges in a box plot have no bearing on how meaningfully two datasets actually differ
13. Why might a chart that uses inconsistent bar widths (not just heights) to represent different-sized categories be considered misleading?
Varying width alongside height can distort the visual sense of area and proportion beyond what the actual data values represent
Bar width has no visual effect on how a viewer perceives the relative size of different categories
Inconsistent bar widths always make a chart's data easier and more accurate to interpret
Only a bar's height, never its width, can ever influence how data proportions are perceived
14. Why might reading only a chart's headline claim, without checking its axes and data source, lead to being misled by a deceptive visual?
Deceptive charts often rely on visual impression alone, so verifying the actual scale and source is necessary to judge a claim accurately
A chart's headline claim always accurately reflects exactly what its axes and underlying data show
Checking a chart's axes and source provides no additional useful information beyond the headline
Deceptive charts are always immediately obvious without needing to check any further detail
15. Why might a box plot be a more honest way to compare two groups' data than a single bar showing only each group's average?
A box plot reveals whether high variability or outliers exist within each group, information a single average value would hide
A box plot always hides more information about a dataset than a single average value does
Comparing spread rather than just averages never actually reveals anything different or additional
A single average value always tells the complete, honest story of a dataset on its own
16. Why might a chart designer combine multiple misleading techniques (like a broken axis AND a cherry-picked range) rather than using just one?
Combining techniques can compound the distorting effect, making a claim appear even more dramatic or convincing than either technique alone
Using multiple misleading techniques together always cancels out their individual distorting effects
Combining techniques has no additional effect beyond using just a single misleading technique
Chart designers are never able to combine more than one distorting technique in the same chart
17. Why might outliers shown in a box plot (points falling well outside the whiskers) be particularly important to investigate rather than simply ignore?
Outliers can indicate genuinely unusual cases, data entry errors, or important edge cases that a summary of typical values alone would miss
Outliers in a box plot are always simply errors that should be automatically deleted from the dataset
Investigating outliers never actually adds any useful insight beyond looking at the typical range of data
Box plots never actually reveal or highlight any outliers within a dataset
18. Why might media literacy education increasingly emphasise "checking the axis" as a specific, practical skill, rather than just general scepticism about statistics?
A specific, actionable check (like examining axis scale and starting point) gives people a concrete way to catch a very common and effective misleading technique
Checking a chart's axis provides no practical benefit compared to general scepticism alone
This specific skill has no real connection to how misleading charts are commonly constructed
General scepticism about statistics is always sufficient without any specific, practical checking skills
19. Why might a chart designer choose 3D visual effects on what is really 2D data (like a 3D-styled pie chart) in a way that could distort a viewer's perception of relative sizes?
3D perspective effects can distort the apparent size of slices or bars depending on their position, misrepresenting the true proportions of the underlying data
3D visual effects always represent the true proportions of data more accurately than a simple 2D chart
Visual styling choices like 3D effects never actually have any impact on how data proportions are perceived
A chart's dimensional styling has no bearing on how accurately a viewer can judge relative data sizes
20. Why might a chart that omits error bars or any indication of data variability give a false impression of certainty about a reported figure?
Without showing the spread or uncertainty in the underlying data, a single reported value can appear more precise and definitive than the data actually supports
Omitting error bars or variability indicators never actually changes how certain a reported figure appears
A chart showing only a single value always accurately conveys the true level of certainty in the underlying data
Data variability has no bearing on how confidently a single reported figure should be interpreted
21. Understanding box plots and misleading data representations mainly helps you to:
Critically evaluate how data is visually presented, comparing spread fairly and spotting distorting techniques
Assume every chart or graph presents data with complete honesty and fairness
Ignore quartiles and spread when interpreting a dataset's distribution
Treat a chart's headline claim as always accurately reflecting its underlying data
Answer key (parent copy)
1. Minimum, quartiles, median and maximum
2. The middle value of the dataset
3. Comparing the spread of multiple datasets side by side
4. Exaggerate small differences visually
5. Selecting a range that supports a claim while ignoring the fuller picture
6. The 25th percentile of the data
7. Critically reading news, advertising and social media
8. It can make a visually large gap on the chart represent only a small real numerical difference, exaggerating the apparent change
9. It shows how concentrated or spread out the middle 50% of each dataset is, allowing a direct visual comparison
10. Different scales can make two genuinely different-sized quantities appear visually similar, or vice versa, distorting a fair comparison
11. A short, selectively chosen window might show an unrepresentative pattern that doesn't reflect the fuller, longer-term trend
12. A box plot reveals the full spread of each dataset, showing whether individual values genuinely differ, not just the summary averages
13. Varying width alongside height can distort the visual sense of area and proportion beyond what the actual data values represent
14. Deceptive charts often rely on visual impression alone, so verifying the actual scale and source is necessary to judge a claim accurately
15. A box plot reveals whether high variability or outliers exist within each group, information a single average value would hide
16. Combining techniques can compound the distorting effect, making a claim appear even more dramatic or convincing than either technique alone
17. Outliers can indicate genuinely unusual cases, data entry errors, or important edge cases that a summary of typical values alone would miss
18. A specific, actionable check (like examining axis scale and starting point) gives people a concrete way to catch a very common and effective misleading technique
19. 3D perspective effects can distort the apparent size of slices or bars depending on their position, misrepresenting the true proportions of the underlying data
20. Without showing the spread or uncertainty in the underlying data, a single reported value can appear more precise and definitive than the data actually supports
21. Critically evaluate how data is visually presented, comparing spread fairly and spotting distorting techniques