Math101learn.math101.caBias in Sampling
A rigorous guide to coverage, selection, nonresponse, response, and convenience bias in samples.
Precise definition
Sampling bias is systematic error caused by a sampling or response process that tends to make the sample statistic differ from the population parameter in a particular direction. Unlike random sampling variability, increasing the size of a biased sample does not reliably remove the error.
Notation and mathematical language
Coverage bias arises when the sampling frame omits groups; voluntary-response and convenience samples select participants nonrandomly; nonresponse bias occurs when responders differ relevantly from nonresponders; response bias arises from wording, interviewer effects, or inaccurate answers. Selection probability is each unit's chance of inclusion.
Conceptual picture
A sample can be very large and very precise about the wrong subset. Random selection aims to balance unmeasured characteristics in expectation, while random assignment addresses causal comparison after sampling; the two forms of randomization solve different problems.
Conditions and key results
Bias is about the process, not merely whether one sample estimate misses the truth. A probability sample can by chance be unrepresentative without being systematically biased. Weighting can correct known unequal selection or response patterns only when the adjustment variables and model are adequate.
A reliable strategy
- Define the target population, sampling frame, unit, and parameter.
- Trace how units enter, decline, or are excluded at every stage.
- Identify mechanisms linking inclusion or response to the measured outcome.
- Redesign with probability sampling, follow-up, neutral measurement, or justified weighting and state remaining limitations.
Fully worked example
Interpretation and application
Sampling bias affects polls, health studies, customer analytics, and school surveys. A result should be generalized only to a population supported by the sampling design. Association within a biased sample can also differ from the target population and does not establish causation.
Common mistakes
Verification and reasonableness
- Compare frame demographics and response rates with credible population benchmarks.
- Perform follow-up on nonresponders or alternative recruitment channels.
- Repeat estimates under plausible weighting and missingness assumptions to assess sensitivity.
Practice
- Does a larger convenience sample eliminate selection bias?
- What bias can an outdated phone list create?
- What does random sampling primarily support?
Answers and brief solutions
- No.
- Coverage bias.
- Generalization from the sample to the target population under the design.
Further deduction
The margin of error commonly reported for a simple random sample measures sampling variability under the assumed design. It excludes systematic errors from undercoverage, nonresponse, question wording, fabrication, or measurement. Reporting a narrow margin without discussing these sources can convey false confidence even when its formula is computed correctly.
Post-stratification weights can align a sample with known population totals, but extreme weights increase variance and make estimates depend heavily on a few respondents. Trimming weights trades some bias for stability. A transparent report describes the weighting variables, range of weights, and sensitivity of conclusions rather than saying the sample was simply 'made representative.'
Nonresponse adjustments rely on information observed for both respondents and nonrespondents. If response depends on an unmeasured outcome even after adjustment variables, bias can remain. Sensitivity analysis should vary plausible outcomes for missing units and show how conclusions change.
Related topics
Try it yourself
Hints are part of learning. Open one whenever it makes the next step feel possible.
If a survey samples only club members, can doubling that sample remove the coverage bias? Enter 1 for yes, 0 for no.
- The sampling frame still excludes nonmembers.
- Doubling size reduces some random variability but leaves systematic coverage bias, so the answer is 0.
End of lesson
Nice work making it this far.
Understanding grows through return visits. Save this lesson, try the practice, or continue when you are ready.
