Repository navigation
Commit 7e7f42a
New lecture: Fitting Distributions to Data
Adds the third lecture in the probability sequence, after prob_dist and
observed_distributions, on choosing a distribution to describe a data set.
It sits immediately after observed_distributions in the TOC.
The lecture is organized around the two halves of that question — which
family, and which parameters:
- The method of moments, stated generally and applied to the normal,
lognormal and gamma families. This names what prob_dist already did once
when it fitted a normal to the height data.
- Q-Q plots, defined from first principles rather than introduced through a
library call: the i-th order statistic estimates the quantile of order
(i-0.5)/n, so plotting it against the fitted quantile should give the 45
degree line. Then how to read a departure — curvature means skew, an
S-shape means heavier tails than the fit allows.
- The Kolmogorov-Smirnov statistic as the largest vertical gap between the
ECDF and the fitted CDF, drawn on the figure. It stops short of the test,
and says why: that needs the null distribution of D, and our parameters
came from the same data.
- Choosing between families by fitting each and taking the smallest D. For
the house prices this ranks lognormal (0.053) over gamma (0.070) over
normal (0.123), agreeing with the near-zero skewness of the log prices
found in the previous lecture.
- Count data, fitting a Poisson to goals per football match, where the fit is
good and the reason it should be is worth stating.
- A section on failure: the normal fit to Amazon returns has an unremarkable
D but an obviously S-shaped Q-Q plot, because D looks at the middle of the
distribution and the trouble is in the tails. That hands off to heavy_tails.
Three warnings accompany the model-selection method: D does not charge a
family for having more parameters, it is insensitive in the tails, and the
winner is only the best of the candidates tried.
Exercises fit an exponential to the times between Japanese earthquakes, which
fails because aftershocks cluster and the arrivals are therefore not
independent, and a normal to the Japanese age-at-death data, which fails
through left skew and also exposes the "100 and over" recording cap as a flat
segment in the Q-Q plot.
Data comes from QuantEcon/data-lectures#31.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>1 parent fd3241e commit 7e7f42a
2 files changed
Lines changed: 712 additions & 0 deletions
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
39 | 39 | | |
40 | 40 | | |
41 | 41 | | |
| 42 | + | |
42 | 43 | | |
43 | 44 | | |
44 | 45 | | |
| |||
0 commit comments