Math101learn.math101.caData Collection
Good data collection begins with a precise question, defined population and variables, ethical design, and a plan to reduce error.
A sophisticated analysis cannot rescue data gathered with a vague question, biased process, or unreliable measurement.
Begin with a research question
A strong question identifies the population, variables, and purpose. “Do students sleep enough?” is vague; “What proportion of Grade 12 students at this school report at least eight hours of sleep on weeknights?” is measurable.
The wording of the question determines what data are relevant and what conclusions are possible.
Population, sample, and unit
The population is the full group of interest. A sample is the subset actually observed. The observational unit is the individual person, object, event, or period on which variables are recorded.
Define inclusion and exclusion rules before collecting data so the target does not shift after results are seen.
Variables and operational definitions
Categorical variables place units into groups; quantitative variables record numbers with meaningful arithmetic. Quantitative variables may be discrete counts or continuous measurements.
An operational definition states exactly how a variable is measured. “Academic success” might mean course average, credit completion, or another defined outcome; these are related but not interchangeable.
Observational studies and experiments
An observational study records existing conditions without assigning treatments. It can reveal associations but is vulnerable to confounding.
An experiment deliberately assigns treatments. Random assignment helps balance lurking variables and supports causal inference when the design and implementation are sound.
Worked design example
Random assignment and random sampling solve different problems.
Surveys
Survey questions should be neutral, specific, understandable, and limited to one idea at a time. Response options should be exhaustive and non-overlapping when possible.
Pilot the survey with a small group to identify confusing terms, missing options, and technical barriers before launch.
Measurement quality
Reliability concerns consistency; validity concerns whether the method measures the intended construct. A scale can produce consistent readings that are all miscalibrated, making it reliable but invalid.
Standardize instruments, training, timing, and instructions. Record missing data and deviations from protocol.
Primary and secondary data
Primary data are collected for the current question. Secondary data were gathered for another purpose. Secondary sources can be efficient, but check definitions, coverage, dates, missingness, ownership, and collection methods.
Document provenance so another person can understand where each value came from.
Ethics and privacy
Minimize harm, collect only necessary information, explain consent and withdrawal, and protect confidentiality. Sensitive personal or student information may require institutional approval and specific legal safeguards.
Anonymization reduces risk but is not always irreversible when datasets can be linked. Access and retention plans matter.
Data management plan
Before collection, define variable names, units, valid ranges, coding for missing values, file versions, and quality checks. Avoid mixing “0,” blank, and “not applicable” without explanation.
A clean data dictionary prevents silent inconsistencies during analysis.
Common mistakes
Collecting first and defining the question later. This encourages selective analysis.
Confusing random assignment with random sampling. One supports causation; the other generalization.
Using an undefined construct. Operationalize what will actually be measured.
Treating missing responses as zero. Missingness needs its own code and analysis.
Ignoring consent and privacy because the project is small. Ethical obligations still apply.
Quick self-check
- Is the research question specific and answerable?
- Are population, sample, unit, and variables defined?
- Is the design observational or experimental, and what claims can it support?
- Are measurements reliable, valid, and standardized?
- Are sampling, missingness, and data-management plans documented?
- Are consent, privacy, security, and potential harm addressed?
Related topics
Try it yourself
Hints are part of learning. Open one whenever it makes the next step feel possible.
Which design feature most directly supports a causal comparison between two study routines?
- Random assignment makes treatment groups comparable on average.
- Random sampling instead mainly supports generalization to a population.
End of lesson
Nice work making it this far.
Understanding grows through return visits. Save this lesson, try the practice, or continue when you are ready.
