Math101learn.math101.caHistograms
A rigorous guide to histogram bins, density, distribution shape, and defensible comparison.
Precise definition
A histogram displays a quantitative variable by partitioning its number line into adjacent intervals and using bar area to represent frequency or relative frequency. With equal bin widths, heights may be counts; with unequal widths, height must be frequency density so area remains proportional to count.
Notation and mathematical language
For class $i$, density is $f_i/w_i$, where $f_i$ is frequency and $w_i$ width. Boundaries such as $[a,b)$ must be consistent. Unlike a bar graph, the horizontal position and touching bars express numerical adjacency and continuity of intervals.
Conceptual picture
A histogram approximates distribution shape: modes, skew, centre, spread, gaps, and tails. The bins aggregate observations, so the picture depends on origin and width. A striking pattern that disappears under reasonable alternative bins may be a display artifact.
Conditions and key results
Histogram shape is descriptive and does not prove a named population distribution. Small samples and broad bins can hide structure. Comparing groups requires compatible scales, bin boundaries, and usually relative frequency or density when sample sizes differ.
A reliable strategy
- Confirm the variable is quantitative and inspect its range, sample size, and measurement resolution.
- Choose documented, nonoverlapping bins that cover all observations.
- Compute counts and, for unequal widths, densities; label whether area represents count or proportion.
- Examine several sensible bin choices and describe shape without causal or inferential overreach.
Fully worked example
Interpretation and application
Histograms summarize response times, measurements, incomes, and scores. Observed skew can guide robust summaries or transformations, but explanations for the skew require domain knowledge and appropriate study design.
Common mistakes
Verification and reasonableness
- Sum bin counts and compare with the sample size.
- Multiply each density height by width and recover its count or proportion.
- Compare with a dot plot, ECDF, or alternative bin widths to assess hidden structure.
Practice
- A bin of width 4 contains 12 observations. Find count density.
- What should histogram area encode?
- Why do bins usually touch?
Answers and brief solutions
- $3$.
- Frequency or relative frequency.
- They represent adjacent numerical intervals.
Further deduction
Relative-frequency density divides by both width and total sample size, so the entire histogram area equals 1. This makes histograms from different sample sizes comparable and parallels a probability density. Individual heights can exceed 1 when bins are narrow; only areas, not heights, must behave like probabilities.
A cumulative distribution plot can verify histogram readings without bins. At a value $x$, the empirical cumulative proportion is the fraction of observations at or below $x$. It reveals medians and percentiles precisely for the sample and makes group comparisons less sensitive to arbitrary bin boundaries, though it displays local density less directly.
Related topics
Try it yourself
Hints are part of learning. Open one whenever it makes the next step feel possible.
A class of width 5 contains 20 observations. What is its count density?
- Density is $f/w$.
- $20/5=4$.
End of lesson
Nice work making it this far.
Understanding grows through return visits. Save this lesson, try the practice, or continue when you are ready.
