Math101learn.math101.caOne-Variable Data
One-variable data analysis describes the distribution of one measured characteristic through shape, centre, spread, and unusual values.
A distribution is more than an average: describe its shape, centre, spread, and unusual features together.
Variable type
Categorical data place observations into groups. Quantitative data record numerical amounts. Quantitative variables can be discrete counts or continuous measurements.
The variable type determines meaningful graphs and statistics. Averaging category codes is not meaningful.
Frequency tables
A frequency table lists values or intervals with counts. Relative frequency divides each count by the total and can be expressed as a proportion or percent.
Class intervals should be non-overlapping and cover the relevant data range. Changing bin width can change the appearance of a histogram.
Displays
- Bar charts: categorical counts or percentages.
- Dot plots: small quantitative datasets and exact values.
- Histograms: quantitative distributions grouped into intervals.
- Box plots: median, quartiles, spread, and potential outliers.
- Stem-and-leaf plots: shape while retaining individual values.
Choose a display that answers the question without distorting scale.
Measures of centre
The mean is the arithmetic balance point, the median the middle ordered value, and the mode the most frequent value.
Mean uses every numerical value and is sensitive to outliers. Median is resistant and often better for skewed distributions. Mode is useful for common categories or values.
Measures of spread
Range is maximum minus minimum. Interquartile range is
Standard deviation measures spread around the mean. Pair median with IQR for skewed/outlier-heavy data and mean with standard deviation for roughly symmetric data without strong outliers.
Worked example
Always order data before finding median and quartiles.
Shape
Describe a distribution as symmetric or skewed, unimodal or multimodal, and note gaps or clusters. In right-skewed data, the long tail points right and often pulls mean above median.
Shape affects which summaries are most informative.
Potential outliers
The $1.5IQR$ rule marks values below
or above
as potential outliers. Investigate them; do not automatically delete them. They may be errors or important genuine cases.
Comparing groups
Use the same scale and compatible summaries. Compare centre, spread, shape, overlap, and unusual values—not only which mean is larger.
Sample size and data-collection method affect how confidently differences can be generalized.
Technology and checking
Technology can calculate statistics and draw plots, but data entry, settings, population/sample choice, and graph scales still need checking. Compare output with rough mental estimates.
Report units and avoid unsupported decimal precision.
Common mistakes
Calculating median before ordering data. Order first.
Using a histogram for categorical data. Use bars with separated categories.
Reporting mean alone for a skewed distribution. Include shape and resistant spread.
Calling every distant value an error. Investigate potential outliers.
Generalizing from a biased sample because the graph looks clear. Design matters.
Quick self-check
- Is the variable categorical or quantitative, discrete or continuous?
- Does the display match the variable type?
- What are shape, centre, spread, and unusual features?
- Are mean/SD or median/IQR the better pair?
- Were quartiles and possible outliers found consistently?
- Do comparison and generalization claims respect sample size and collection method?
Related topics
Try it yourself
Hints are part of learning. Open one whenever it makes the next step feel possible.
Find the mean of 4, 5, 6, 6, and 9.
- 4 + 5 + 6 + 6 + 9 = 30.
- 30/5 = 6.
End of lesson
Nice work making it this far.
Understanding grows through return visits. Save this lesson, try the practice, or continue when you are ready.
