Skip to content

Commit 7e7f42a

Browse files
jstacclaude
andcommitted
New lecture: Fitting Distributions to Data
Adds the third lecture in the probability sequence, after prob_dist and observed_distributions, on choosing a distribution to describe a data set. It sits immediately after observed_distributions in the TOC. The lecture is organized around the two halves of that question — which family, and which parameters: - The method of moments, stated generally and applied to the normal, lognormal and gamma families. This names what prob_dist already did once when it fitted a normal to the height data. - Q-Q plots, defined from first principles rather than introduced through a library call: the i-th order statistic estimates the quantile of order (i-0.5)/n, so plotting it against the fitted quantile should give the 45 degree line. Then how to read a departure — curvature means skew, an S-shape means heavier tails than the fit allows. - The Kolmogorov-Smirnov statistic as the largest vertical gap between the ECDF and the fitted CDF, drawn on the figure. It stops short of the test, and says why: that needs the null distribution of D, and our parameters came from the same data. - Choosing between families by fitting each and taking the smallest D. For the house prices this ranks lognormal (0.053) over gamma (0.070) over normal (0.123), agreeing with the near-zero skewness of the log prices found in the previous lecture. - Count data, fitting a Poisson to goals per football match, where the fit is good and the reason it should be is worth stating. - A section on failure: the normal fit to Amazon returns has an unremarkable D but an obviously S-shaped Q-Q plot, because D looks at the middle of the distribution and the trouble is in the tails. That hands off to heavy_tails. Three warnings accompany the model-selection method: D does not charge a family for having more parameters, it is insensitive in the tails, and the winner is only the best of the candidates tried. Exercises fit an exponential to the times between Japanese earthquakes, which fails because aftershocks cluster and the arrivals are therefore not independent, and a normal to the Japanese age-at-death data, which fails through left skew and also exposes the "100 and over" recording cap as a flat segment in the Q-Q plot. Data comes from QuantEcon/data-lectures#31. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
1 parent fd3241e commit 7e7f42a

2 files changed

Lines changed: 712 additions & 0 deletions

File tree

‎lectures/_toc.yml‎

Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -39,6 +39,7 @@ parts:
3939
chapters:
4040
- file: prob_dist
4141
- file: observed_distributions
42+
- file: fitting_distributions
4243
- file: lln_clt
4344
- file: monte_carlo
4445
- file: heavy_tails

0 commit comments

Comments
 (0)