Math101learn.math101.caBias in Data
Statistical bias is a systematic tendency for a collection or analysis method to miss the truth in a particular direction.
Random error creates scatter; bias systematically shifts what the study tends to observe or conclude.
Bias versus variability
Sampling variability changes results unpredictably from sample to sample. Bias pushes results systematically because of design, measurement, processing, or analysis choices.
A very large sample reduces random error but can estimate the wrong quantity extremely precisely when bias remains.
Selection bias
Selection bias occurs when inclusion relates to the variable of interest. Convenience and voluntary-response samples often overrepresent people who are available or strongly motivated.
To reduce it, use a probability sampling method, a suitable frame, and transparent eligibility rules.
Undercoverage
Undercoverage occurs when parts of the target population are missing or rarely reached by the sampling frame. An online-only survey may miss people with limited internet access; a daytime phone survey may miss some workers.
Name who could be absent and how their outcomes might differ.
Nonresponse bias
Nonresponse becomes bias when selected nonrespondents differ meaningfully from respondents. A high response rate helps but does not guarantee absence of bias, and a low rate does not quantify its direction by itself.
Use follow-ups, multiple contact modes, accessible design, and adjustment methods when justified.
Response and wording bias
People may alter answers due to social desirability, memory limits, interviewer presence, privacy concerns, or leading wording.
Questions such as “Do you agree with the responsible policy…” frame a preferred response. Use neutral language, balanced options, clear time frames, and one idea per question.
Worked example: online poll
A larger click count would not solve these design problems.
Measurement bias
An instrument or procedure can systematically over- or under-measure. Examples include a miscalibrated scale, inconsistent coding, different testing conditions, or a proxy that does not validly represent the intended concept.
Calibration, standardized protocols, blinding, and validated measures help reduce measurement bias.
Confounding
A confounder is related to both the explanatory variable and outcome and can create a misleading association. Observational studies are especially vulnerable.
Design controls, random assignment, matching, stratification, or statistical adjustment may help, but unmeasured confounding can remain.
Survivorship and reporting bias
Survivorship bias focuses on visible successes while missing failures. Publication or reporting bias makes some results more likely to appear than others, often emphasizing statistically significant or dramatic findings.
Pre-registration, complete outcome reporting, and searching for missing cases improve credibility.
Processing and analysis choices
Bias can enter through selective data cleaning, changing outcomes after seeing results, excluding inconvenient observations, choosing favourable graph scales, or testing many hypotheses but reporting only a few.
Document decisions, preserve raw data, and distinguish planned from exploratory analysis.
Detecting and communicating bias
Ask who is missing, who responded, how variables were measured, what incentives existed, what was adjusted, and which results were not shown. Compare sample demographics with the target population and inspect missing-data patterns.
Limitations should specify likely direction and consequence when possible, not merely say “bias may exist.”
Common mistakes
Using “bias” to mean personal disagreement. Statistical bias is systematic error in a process or estimator.
Claiming a large sample is representative. Selection mechanism matters.
Assuming neutral-looking wording is neutral to all respondents. Pilot and test questions.
Removing outliers without documented reasons. This can introduce analysis bias.
Listing limitations without changing the strength of the conclusion. Claims should match evidence quality.
Quick self-check
- Who is in the target population, frame, invited sample, and responding sample?
- Could selection, undercoverage, or nonresponse shift results?
- Are wording and measurements neutral, reliable, and valid?
- Could confounding explain the association?
- Were cleaning, exclusions, outcomes, and analyses pre-specified and transparent?
- Does the conclusion explicitly respect the likely biases and study design?
Related topics
Try it yourself
Hints are part of learning. Open one whenever it makes the next step feel possible.
A news website lets visitors voluntarily click 'excellent' or 'terrible' about a policy and generalizes to the public. What is the strongest concern?
- Visitors self-select and site users may differ from the public.
- Loaded, incomplete response options add measurement concerns.
- More biased responses do not remove systematic error.
End of lesson
Nice work making it this far.
Understanding grows through return visits. Save this lesson, try the practice, or continue when you are ready.
