Producing trustworthy scientific results depends on careful attention to potential sources of error — choosing appropriate equipment, controlling variables, and accounting for possible confounding variables that could distort results. Instruments need to be correctly calibrated (checked against a known standard) before use, since even accurate-looking readings can be systematically wrong if an instrument isn't properly calibrated. Once data is collected, describing its properties (mean, median, range, and any large gaps or outliers) helps summarise it meaningfully, while evaluating the strength of a conclusion means honestly considering how well the data actually supports it — and distinguishing between random error (unpredictable variation) and systematic error (a consistent bias affecting every measurement).
Example
A thermometer that reads consistently 2 degrees too high, if used uncalibrated across an entire experiment, would introduce a systematic error — every single reading would be wrong by roughly the same amount in the same direction, unlike random error, which would cause readings to scatter unpredictably both above and below the true value.
Key terms
Calibration:
Checking and adjusting an instrument against a known standard before use.
Systematic error:
A consistent, repeated bias affecting every measurement in the same direction.
Random error:
Unpredictable variation in measurements, scattering both above and below the true value.
Questions
1. Calibrating an instrument means:
Checking and adjusting it against a known standard
Using it without any checks at all
Making it deliberately less accurate
A step that is never actually necessary
2. A systematic error is:
A consistent bias affecting every measurement in the same direction
Completely random and unpredictable
Something that never occurs in real experiments
Always caused by the weather
3. A random error causes measurements to:
Scatter unpredictably above and below the true value
Always be wrong in exactly the same direction
Always be perfectly accurate
Have no effect on results whatsoever
4. A confounding variable is:
A factor that could distort results if not accounted for
Something with no effect on any experiment
The main variable being tested
A type of measuring instrument
5. Describing a dataset's properties can include its:
Mean, median and range
Only its colour
Only its file size
Nothing measurable at all
6. Evaluating the strength of a conclusion means:
Honestly considering how well the data actually supports it
Always assuming a conclusion is completely correct
Ignoring the data entirely
Never questioning any research finding
7. An uncalibrated instrument can produce readings that are:
Systematically wrong, even if they look precise
Always perfectly accurate regardless of calibration
Completely unrelated to the instrument's condition
Impossible to ever be wrong
8. Why might a thermometer reading consistently 2 degrees too high introduce a systematic, rather than random, error into an experiment?
The error is consistent and in the same direction across every reading, rather than scattering unpredictably
A consistent, repeated bias is the defining characteristic of a random error, not a systematic one
Systematic and random errors always produce identical patterns in a dataset
An uncalibrated instrument can never actually introduce any kind of error into results
9. Why is it important to control variables (keeping everything except the one being tested constant) in a scientific investigation?
It helps ensure that any observed effect can be attributed to the variable being tested, rather than an unaccounted-for confounding factor
Controlling variables has no real effect on the reliability of an investigation's results
A confounding variable can never actually distort the outcome of an experiment
Any change in results can always be safely attributed to the tested variable, regardless of other conditions
10. Why might a large gap visible in a dataset (a noticeable jump between clusters of values) be worth investigating rather than ignoring?
It could indicate a genuinely meaningful pattern, an outlier, or even an error in the data collection process worth understanding
Gaps in a dataset never actually carry any useful or meaningful information
All datasets are always perfectly smooth and continuous with no notable gaps
Investigating unusual patterns in a dataset never actually reveals anything useful
11. Why might increasing the sample size (number of measurements) of an experiment help reduce the impact of random error on the overall result?
With more measurements, random variations tend to average out, bringing the overall result closer to the true value
Sample size has no bearing whatsoever on the impact of random error in an experiment
A larger sample size always makes systematic error worse, never random error better
Random error always increases in direct proportion to how many measurements are taken
12. Why is recalibrating an instrument between uses (not just once at the very start) sometimes necessary for maintaining reliable results?
Instruments can drift out of calibration over time or with repeated use, so periodic recalibration helps maintain ongoing accuracy
An instrument, once calibrated, always remains perfectly calibrated forever with no possibility of drifting
Recalibrating an instrument partway through a series of measurements never actually improves reliability
Calibration is only ever relevant the very first time an instrument is used, with no ongoing need
13. Why might using spreadsheet software to analyse data (rather than manual calculation) reduce the risk of certain types of error in a large investigation?
Automated calculation reduces the risk of manual arithmetic mistakes, especially when processing large amounts of data
Spreadsheet software always introduces more errors into an analysis than manual calculation would
The method used to analyse data has no bearing on the likelihood of a calculation error occurring
Manual calculation is always more reliable and accurate than using spreadsheet software for data analysis
14. Why might reporting both the mean AND the range of a dataset give a more complete picture than reporting the mean alone?
The range reveals how spread out or variable the data actually is, information the mean alone (a single central value) doesn't capture
The mean alone always provides a fully complete picture of a dataset with no need for any additional measure
Range and mean always provide exactly identical information about a dataset
Reporting the range alongside the mean never actually adds any useful additional context
15. Why might a conclusion drawn from a very small dataset (with only a few data points) be considered weaker than one drawn from a large, robust dataset, even if both show a similar apparent pattern?
A small dataset is more vulnerable to random variation and outliers skewing the apparent pattern, making the underlying conclusion less reliable
The size of a dataset has no bearing whatsoever on how reliable a conclusion drawn from it might be
Small and large datasets always provide exactly equally strong support for any given conclusion
A pattern seen in a small dataset is always more reliable and trustworthy than one seen in a large dataset
16. Why might systematic error be considered more concerning in some ways than random error, even though both reduce accuracy?
Systematic error consistently biases results in one direction, potentially leading to a confidently wrong conclusion, whereas random error is often more visible as scatter and can average out with more data
Systematic error and random error always have exactly identical effects on the reliability of results
Random error is always more concerning than systematic error in every possible scientific context
Systematic error is always immediately obvious and easy to detect, unlike random error
17. Why might a researcher deliberately test whether an observed pattern holds true across multiple independent experiments or datasets before considering a conclusion well-supported?
Replicating a finding across independent studies helps rule out the possibility that the original result was due to chance, error or a specific unaccounted-for confounding factor
A single experiment always provides exactly as strong evidence for a conclusion as multiple independent replications
Replicating a study across multiple independent experiments never actually strengthens confidence in a finding
A pattern seen in just one single experiment should always be considered fully proven and well-supported
18. Why might genuinely evaluating the "strength" of a conclusion involve considering not just whether the data supports it, but also what alternative explanations might exist?
A thorough evaluation considers whether other factors (like confounding variables or error) could equally explain the observed pattern, not just whether the preferred explanation fits
Considering alternative explanations for a pattern never actually strengthens or weakens confidence in a conclusion
A conclusion is always considered equally strong regardless of whether any alternative explanations exist
If data appears to support a conclusion, no other explanation for the pattern needs to ever be considered
19. Why might a well-designed experiment include a control group (a group not exposed to the tested variable) specifically to help distinguish a genuine effect from natural variation or error?
Comparing the tested group against a control group that experienced everything else identically helps isolate whether an observed difference is actually due to the tested variable
A control group provides no useful basis for comparison when evaluating whether an observed effect is genuine
Natural variation and error affect only the tested group, never a control group, making comparison unnecessary
Experiments without any control group always produce results that are exactly as reliable as those with one
20. Why might publishing full raw data alongside a scientific paper's conclusions (rather than just the summary and conclusion) support more rigorous evaluation of that research by others?
Access to the underlying raw data allows other scientists to independently check calculations, look for errors, and assess whether the conclusion is genuinely well-supported
Raw data provides no additional value to other scientists beyond what a paper's summarised conclusion already offers
Independently checking a paper's underlying data never actually reveals anything useful about the strength of its conclusions
Scientific papers should generally avoid publishing raw data, since summaries alone are always sufficient for evaluation
21. Understanding how to evaluate scientific evidence and data mainly helps you to:
Critically assess data quality, error sources and the genuine strength of scientific conclusions
Assume any data collected during an investigation is automatically free from error
Ignore the distinction between systematic and random sources of error
Treat every conclusion drawn from data as equally well-supported regardless of dataset size or quality
Answer key (parent copy)
1. Checking and adjusting it against a known standard
2. A consistent bias affecting every measurement in the same direction
3. Scatter unpredictably above and below the true value
4. A factor that could distort results if not accounted for
5. Mean, median and range
6. Honestly considering how well the data actually supports it
7. Systematically wrong, even if they look precise
8. The error is consistent and in the same direction across every reading, rather than scattering unpredictably
9. It helps ensure that any observed effect can be attributed to the variable being tested, rather than an unaccounted-for confounding factor
10. It could indicate a genuinely meaningful pattern, an outlier, or even an error in the data collection process worth understanding
11. With more measurements, random variations tend to average out, bringing the overall result closer to the true value
12. Instruments can drift out of calibration over time or with repeated use, so periodic recalibration helps maintain ongoing accuracy
13. Automated calculation reduces the risk of manual arithmetic mistakes, especially when processing large amounts of data
14. The range reveals how spread out or variable the data actually is, information the mean alone (a single central value) doesn't capture
15. A small dataset is more vulnerable to random variation and outliers skewing the apparent pattern, making the underlying conclusion less reliable
16. Systematic error consistently biases results in one direction, potentially leading to a confidently wrong conclusion, whereas random error is often more visible as scatter and can average out with more data
17. Replicating a finding across independent studies helps rule out the possibility that the original result was due to chance, error or a specific unaccounted-for confounding factor
18. A thorough evaluation considers whether other factors (like confounding variables or error) could equally explain the observed pattern, not just whether the preferred explanation fits
19. Comparing the tested group against a control group that experienced everything else identically helps isolate whether an observed difference is actually due to the tested variable
20. Access to the underlying raw data allows other scientists to independently check calculations, look for errors, and assess whether the conclusion is genuinely well-supported
21. Critically assess data quality, error sources and the genuine strength of scientific conclusions