Probability and Distributions Overview
Chapters 2 to 5 were about objects and operations that are certain. A matrix has a rank; a gradient has a value. Chapter 6 is the first one where the answer is a distribution rather than a number, and that changes what “computing” means.
The book’s framing, from the opening line: “Probability, loosely speaking, concerns the study of uncertainty.” And it is quick to say where that uncertainty comes from in machine learning — “uncertainty in the data, uncertainty in the machine learning model, and uncertainty in the predictions produced by the model.” Three different things, all handled by the same machinery.
The one thing to take from this chapter
Section titled “The one thing to take from this chapter”There are only two rules.
| sum rule | Equation 6.20 | |
| product rule | Equation 6.22 |
Bayes’ theorem is not a third rule — it is the product rule written in both orderings and solved for one conditional. Every distribution, every summary statistic, every inference procedure in Chapters 8 through 12 is built from those two lines.
What the rest of the chapter adds is computability. The rules are easy to write and expensive to evaluate: marginalising binary variables costs terms, which is at . So the chapter spends its second half on the two structures that make the cost go away — the Gaussian, which is closed under every operation you need, and the exponential family, which is provably the only class whose parameter count stays fixed as data arrives.
The chapter’s own map
Section titled “The chapter’s own map”The book’s Figure 6.1, as a diagram:
flowchart TD RV["random variable & distribution
Section 6.1, 6.2"] RV -->|"two rules"| SP["sum rule, product rule
Eq 6.20, 6.22"] SP --> BAY["Bayes' theorem
Eq 6.23"] RV --> SS["summary statistics
Section 6.4"] SS --> MV["mean, variance"] SS --> IND["independence"] SS --> IP["inner product
correlation as an angle"] RV --> TR["transformations
Section 6.7"] RV --> G["Gaussian
Section 6.5"] RV --> BER["Bernoulli, Binomial"] BER --> BETA["Beta"] BETA -->|"conjugate"| BER SS --> SUF["sufficient statistics"] SUF --> EF["exponential family
Section 6.6"] BAY -.->|"used in"| C9["Ch 9 Regression"] G -.->|"used in"| C9 EF -.->|"used in"| C10["Ch 10 Dimensionality reduction"] G -.->|"used in"| C11["Ch 11 Density estimation"] TR -.->|"used in"| C11
The through-line, and what each page settles
Section titled “The through-line, and what each page settles”| # | Page | The question it answers |
|---|---|---|
| 6.1 | Construction of a Probability Space | What are the three objects everyone conflates, and why is a random variable a function? |
| 6.2 | Discrete and Continuous Probabilities | What is the difference between a pmf, a pdf and a cdf — and which of them is ever a probability? |
| 6.3 | Sum Rule, Product Rule, and Bayes Theorem | The only two rules, the theorem that follows, and why the sum rule is where the cost lives. |
| 6.4 | Summary Statistics and Independence | What do mean, variance and covariance actually tell you — and what do they hide? |
| 6.5 | Gaussian Distribution | Why is this density everywhere? Because it is closed under marginalising, conditioning, multiplying and affine maps. |
| 6.6 | Conjugacy and the Exponential Family | Is the Gaussian’s convenience luck? No — and here is the family that shares it. |
| 6.7 | Change of Variables and the Inverse Transform | What happens to a density when you transform the variable, and why the Jacobian is not optional. |
| — | Chapter 6 Exercises and Solutions | All thirteen exercises, including the two that derive the Kalman filter and Bayesian linear regression. |
| — | Chapter 6 Formula Sheet | Every definition, theorem and measured number on one page. |
Three reading paths
Section titled “Three reading paths”If you are here for Bayesian inference: §6.1 → §6.3 → §6.6. That is the probability space, the two rules with Bayes, and conjugacy. You can defer §6.4 and §6.5.
If you are here for the Gaussian machinery (Chapter 9’s regression, Kalman filters, Gaussian processes): §6.2 → §6.4 → §6.5, then Exercises 6.5 and 6.12, which build the Kalman filter and the Gaussian-linear-model posterior from §6.5’s closure rules alone.
If you are here for generative models (flows, VAEs): §6.2’s density-versus- probability distinction, then §6.7. The Jacobian in Equation 6.143 is the entire content of a normalising flow, and Chapter 5’s Exercise 5.9 gradient is the reparameterisation trick.
What the chapter looks like once it has been measured
Section titled “What the chapter looks like once it has been measured”Every number below is measured on the page it belongs to.
| Claim | Measured |
|---|---|
| “The probability is 40%” | Example 6.1’s pmf is — and the middle one is a sum over two outcomes, not one |
| “Probability is a long-run frequency” | the error falls as : at , at |
| A density is a probability | the uniform on has density ; a Gaussian at peaks at |
| for continuous | exactly zero — a million samples gave zero exact hits |
| “The test is 99% accurate” | at prevalence — overstated |
| A prior is just a starting guess | a zero prior stays exactly zero after confirming observations |
| Zero covariance means unrelated | Example 6.5: exactly, covariance , total variation from independence |
| Variance has one formula | Equation 6.44 returns for a variance of once the data is offset by |
| Per-feature statistics describe a dataset | two clouds with identical means and marginal variances, correlations and |
| Conditioning is just filtering | it provably shrinks a Gaussian’s covariance — eigenvalues |
| Averaging Gaussians gives a Gaussian | a mixture has excess kurtosis and two modes at identical first two moments |
| Bayesian updating needs integration | with a conjugate prior it is two numbers, still two after observations |
| A model sees all of the data | two visibly different datasets, log-likelihoods equal to |
| Transforming a variable relabels its density | dropping the Jacobian gives a shape error of that renormalising cannot fix |
| The most likely value is a property of the variable | the mode of is , not |
Prerequisites
Section titled “Prerequisites”| You need | From |
|---|---|
| sums, products, set notation | Chapter 0 — Sums, Products and Set Notation |
| integrals, and what a density integrates to | Chapter 0 — Single-Variable Calculus Refresher |
| matrices, inverses, determinants | Chapter 2 — Linear Algebra |
| inner products, orthogonality, norms | Chapter 3 — Analytic Geometry |
| symmetric positive definite matrices, Cholesky | Chapter 4 — Cholesky Decomposition |
| the Jacobian and its determinant | Chapter 5 — Gradients of Vector-Valued Functions |
Two of those are load-bearing rather than background. §6.5’s sampling needs Cholesky to exist, which needs covariance matrices to be symmetric positive definite — Chapter 4’s result. And §6.7’s change of variables is Chapter 5’s Jacobian determinant, used for the purpose it was built for.
Where each section is used later
Section titled “Where each section is used later”| Section | Used in |
|---|---|
| §6.1 probability space | the vocabulary for the rest of the book |
| §6.2 pmf, pdf, cdf | every likelihood in Chapters 9 to 12 |
| §6.3 sum, product, Bayes | Chapter 8’s probabilistic modelling and model selection |
| §6.4 mean and covariance | Chapter 10’s PCA — the covariance matrix is the method |
| §6.5 Gaussian | Chapter 9’s linear regression, Chapter 11’s mixtures, Kalman filters, Gaussian processes |
| §6.6 conjugacy, exponential family | Chapter 9’s Bayesian regression, Chapter 12’s loss functions |
| §6.7 change of variables | normalising flows, reparameterised gradients |
-
How many fundamental rules does this chapter rest on?
Everything else in the chapter is either a named distribution, a summary statistic, or a rule for transforming distributions. The two rules are the generative core.
pch.quizShowAnswer
B — Two. Bayes' theorem is the product rule written in both orderings and solved for one conditional — Equations 6.24 to 6.26 are its whole derivation — Everything else in the chapter is either a named distribution, a summary statistic, or a rule for transforming distributions. The two rules are the generative core.
-
Why does the chapter spend so long on the Gaussian specifically?
Exercises 6.5 and 6.12 are the payoff: between them they derive the Kalman filter and Bayesian linear regression using nothing but those closure rules, with no new mathematics.
pch.quizShowAnswer
B — Because it is CLOSED under marginalising, conditioning, multiplying, adding and affine mapping — each by a closed-form formula in the mean and covariance, so whole inference pipelines become matrix algebra — Exercises 6.5 and 6.12 are the payoff: between them they derive the Kalman filter and Bayesian linear regression using nothing but those closure rules, with no new mathematics.
-
The two rules are easy to write. What makes them hard to use?
The book flags it in a Remark in Section 6.3. Conjugacy (Section 6.6) is the escape hatch: the posterior stays in the prior's family and the integral never has to be done.
pch.quizShowAnswer
B — The sum rule. Marginalising D binary variables costs 2^D terms — 1.27e30 at D = 100 — and there is no known exact polynomial-time algorithm. That is why the chapter's second half is about structures that avoid the cost — The book flags it in a Remark in Section 6.3. Conjugacy (Section 6.6) is the escape hatch: the posterior stays in the prior's family and the integral never has to be done.
-
Which two earlier chapters does Chapter 6 genuinely depend on, rather than merely reference?
Both are load-bearing. Section 6.5.4 names Cholesky explicitly and notes the condition it needs, and Theorem 6.16 is Chapter 5's Section 5.3 applied to densities.
pch.quizShowAnswer
B — Chapter 4, because sampling a Gaussian needs the Cholesky factorisation to exist for a symmetric positive definite covariance; and Chapter 5, because Section 6.7's change of variables IS the Jacobian determinant — Both are load-bearing. Section 6.5.4 names Cholesky explicitly and notes the condition it needs, and Theorem 6.16 is Chapter 5's Section 5.3 applied to densities.
Where §6.8 points next
Section titled “Where §6.8 points next”Section 6.8 opens by conceding that “this chapter is rather terse at times”, and then spends most of its length on the measure theory it deliberately skipped.
| §6.8’s pointer | and where it lands in this module |
|---|---|
| more relaxed presentations for self-study | this module’s own Construction of a Probability Space takes the slower route |
| an overview of exponential families | Conjugacy and the Exponential Family |
| “how to use probability distributions to model machine learning tasks” | When Models Meet Data Overview, the whole of Chapter 8 |
| normalizing flows “rely on change of variables for transforming random variables” | Change of Variables and the Inverse Transform |
| variational inference applied to neural networks | the bound of The Latent Variable Perspective is the same construction |
| measure-theoretic questions “side stepped” | acknowledged, not repaired — see the caution below |
| “an alternative way to approach probability is to start with the concept of expectation” | Summary Statistics and Independence |
Recall card
Section titled “Recall card”- Two rules generate the chapter. The sum rule (Eq 6.20) and the product rule (Eq 6.22). Bayes’ theorem (Eq 6.23) is those two rearranged, not a third axiom.
- The book’s Figure 6.1 map: a random variable and its distribution feed the two rules and Bayes on one side, and summary statistics, independence and the inner-product view on the other; sufficient statistics lead to the exponential family; the Gaussian, Bernoulli, Binomial and Beta are the named examples.
- Three objects get conflated and should not: the sample space, the event space, and the probability measure — plus the random variable, which is a FUNCTION from the first to a target space.
- A density is not a probability, and for a continuous variable P(X = x) = 0 exactly. Only the cdf is a probability at every point.
- The sum rule is the bottleneck: 2^D terms for D binary variables, 1.27e30 at D = 100, with no known exact polynomial-time algorithm.
- The Gaussian earns its place by CLOSURE, not by ubiquity: marginalise, condition, multiply, add, map affinely — all closed form in the mean and covariance.
- The exponential family is provably the only class with finite-dimensional sufficient statistics under repeated sampling (Pitman, Darmois, Koopman, 1935-36), which is what keeps the parameter count fixed as data grows.
- Section 6.7’s change of variables is Chapter 5’s Jacobian, and Section 6.5.4’s sampler is Chapter 4’s Cholesky. Those two dependencies are load-bearing, not decorative.
- Three reading paths: 6.1-6.3-6.6 for Bayesian inference; 6.2-6.4-6.5 plus Exercises 6.5 and 6.12 for the Gaussian machinery; 6.2 and 6.7 for generative models.
Start with: Construction of a Probability Space — the three objects, and the function that is neither random nor a variable.
pch.coffeeTagline
pch.coffeeCtapch.feedbackHeading
pch.feedbackSubheading