Skip to content

Probability and Distributions Overview

Chapters 2 to 5 were about objects and operations that are certain. A matrix has a rank; a gradient has a value. Chapter 6 is the first one where the answer is a distribution rather than a number, and that changes what “computing” means.

The book’s framing, from the opening line: “Probability, loosely speaking, concerns the study of uncertainty.” And it is quick to say where that uncertainty comes from in machine learning — “uncertainty in the data, uncertainty in the machine learning model, and uncertainty in the predictions produced by the model.” Three different things, all handled by the same machinery.

There are only two rules.

sum rulep(x)=∑yp(x,y)p(\mathbf{x}) = \sum_\mathbf{y} p(\mathbf{x},\mathbf{y})Equation 6.20
product rulep(x,y)=p(y∣x)p(x)p(\mathbf{x},\mathbf{y}) = p(\mathbf{y}\mid\mathbf{x})p(\mathbf{x})Equation 6.22

Bayes’ theorem is not a third rule — it is the product rule written in both orderings and solved for one conditional. Every distribution, every summary statistic, every inference procedure in Chapters 8 through 12 is built from those two lines.

What the rest of the chapter adds is computability. The rules are easy to write and expensive to evaluate: marginalising DD binary variables costs 2D2^D terms, which is 1.27×10301.27\times10^{30} at D=100D=100. So the chapter spends its second half on the two structures that make the cost go away — the Gaussian, which is closed under every operation you need, and the exponential family, which is provably the only class whose parameter count stays fixed as data arrives.

The book’s Figure 6.1, as a diagram:

diagram Diagram mermaid

The through-line, and what each page settles

Section titled “The through-line, and what each page settles”
#PageThe question it answers
6.1Construction of a Probability SpaceWhat are the three objects everyone conflates, and why is a random variable a function?
6.2Discrete and Continuous ProbabilitiesWhat is the difference between a pmf, a pdf and a cdf — and which of them is ever a probability?
6.3Sum Rule, Product Rule, and Bayes TheoremThe only two rules, the theorem that follows, and why the sum rule is where the cost lives.
6.4Summary Statistics and IndependenceWhat do mean, variance and covariance actually tell you — and what do they hide?
6.5Gaussian DistributionWhy is this density everywhere? Because it is closed under marginalising, conditioning, multiplying and affine maps.
6.6Conjugacy and the Exponential FamilyIs the Gaussian’s convenience luck? No — and here is the family that shares it.
6.7Change of Variables and the Inverse TransformWhat happens to a density when you transform the variable, and why the Jacobian is not optional.
—Chapter 6 Exercises and SolutionsAll thirteen exercises, including the two that derive the Kalman filter and Bayesian linear regression.
—Chapter 6 Formula SheetEvery definition, theorem and measured number on one page.

If you are here for Bayesian inference: §6.1 → §6.3 → §6.6. That is the probability space, the two rules with Bayes, and conjugacy. You can defer §6.4 and §6.5.

If you are here for the Gaussian machinery (Chapter 9’s regression, Kalman filters, Gaussian processes): §6.2 → §6.4 → §6.5, then Exercises 6.5 and 6.12, which build the Kalman filter and the Gaussian-linear-model posterior from §6.5’s closure rules alone.

If you are here for generative models (flows, VAEs): §6.2’s density-versus- probability distinction, then §6.7. The Jacobian in Equation 6.143 is the entire content of a normalising flow, and Chapter 5’s Exercise 5.9 gradient is the reparameterisation trick.

What the chapter looks like once it has been measured

Section titled “What the chapter looks like once it has been measured”

Every number below is measured on the page it belongs to.

ClaimMeasured
“The probability is 40%”Example 6.1’s pmf is 0.49, 0.42, 0.090.49,\ 0.42,\ 0.09 — and the middle one is a sum over two outcomes, not one
“Probability is a long-run frequency”the error falls as 1/N1/\sqrt N: 0.120.12 at N=10N=10, 0.00080.0008 at N=106N=10^6
A density is a probabilitythe uniform on [0.9,1.6][0.9,1.6] has density 1.42861.4286; a Gaussian at σ=0.01\sigma=0.01 peaks at 39.8939.89
P(X=x)P(X = x) for continuous XXexactly zero — a million samples gave zero exact hits
“The test is 99% accurate”P(disease∣+)=0.0902P(\text{disease}\mid+) = 0.0902 at prevalence 0.0010.001 — overstated 10.98×10.98\times
A prior is just a starting guessa zero prior stays exactly zero after 400400 confirming observations
Zero covariance means unrelatedExample 6.5: y=x2y=x^2 exactly, covariance ≈10−4\approx 10^{-4}, total variation from independence 0.4060.406
Variance has one formulaEquation 6.44 returns 00 for a variance of 44 once the data is offset by 10910^9
Per-feature statistics describe a datasettwo clouds with identical means and marginal variances, correlations −0.8-0.8 and +0.8+0.8
Conditioning is just filteringit provably shrinks a Gaussian’s covariance — eigenvalues [0, 0.186, 1.847][0,\ 0.186,\ 1.847]
Averaging Gaussians gives a Gaussiana mixture has excess kurtosis −1.47-1.47 and two modes at identical first two moments
Bayesian updating needs integrationwith a conjugate prior it is two numbers, still two after 100 000100\,000 observations
A model sees all of the datatwo visibly different datasets, log-likelihoods equal to 2.3×10−132.3\times10^{-13}
Transforming a variable relabels its densitydropping the Jacobian gives a shape error of 0.52750.5275 that renormalising cannot fix
The most likely value is a property of the variablethe mode of exp⁡(X)\exp(X) is 0.36790.3679, not exp⁡(0)=1\exp(0)=1
You needFrom
sums, products, set notationChapter 0 — Sums, Products and Set Notation
integrals, and what a density integrates toChapter 0 — Single-Variable Calculus Refresher
matrices, inverses, determinantsChapter 2 — Linear Algebra
inner products, orthogonality, normsChapter 3 — Analytic Geometry
symmetric positive definite matrices, CholeskyChapter 4 — Cholesky Decomposition
the Jacobian and its determinantChapter 5 — Gradients of Vector-Valued Functions

Two of those are load-bearing rather than background. §6.5’s sampling needs Cholesky to exist, which needs covariance matrices to be symmetric positive definite — Chapter 4’s result. And §6.7’s change of variables is Chapter 5’s Jacobian determinant, used for the purpose it was built for.

SectionUsed in
§6.1 probability spacethe vocabulary for the rest of the book
§6.2 pmf, pdf, cdfevery likelihood in Chapters 9 to 12
§6.3 sum, product, BayesChapter 8’s probabilistic modelling and model selection
§6.4 mean and covarianceChapter 10’s PCA — the covariance matrix is the method
§6.5 GaussianChapter 9’s linear regression, Chapter 11’s mixtures, Kalman filters, Gaussian processes
§6.6 conjugacy, exponential familyChapter 9’s Bayesian regression, Chapter 12’s loss functions
§6.7 change of variablesnormalising flows, reparameterised gradients
pch.quizTag Before you start
  1. How many fundamental rules does this chapter rest on?

    pch.quizShowAnswer

    B — Two. Bayes' theorem is the product rule written in both orderings and solved for one conditional — Equations 6.24 to 6.26 are its whole derivation — Everything else in the chapter is either a named distribution, a summary statistic, or a rule for transforming distributions. The two rules are the generative core.

  2. Why does the chapter spend so long on the Gaussian specifically?

    pch.quizShowAnswer

    B — Because it is CLOSED under marginalising, conditioning, multiplying, adding and affine mapping — each by a closed-form formula in the mean and covariance, so whole inference pipelines become matrix algebra — Exercises 6.5 and 6.12 are the payoff: between them they derive the Kalman filter and Bayesian linear regression using nothing but those closure rules, with no new mathematics.

  3. The two rules are easy to write. What makes them hard to use?

    pch.quizShowAnswer

    B — The sum rule. Marginalising D binary variables costs 2^D terms — 1.27e30 at D = 100 — and there is no known exact polynomial-time algorithm. That is why the chapter's second half is about structures that avoid the cost — The book flags it in a Remark in Section 6.3. Conjugacy (Section 6.6) is the escape hatch: the posterior stays in the prior's family and the integral never has to be done.

  4. Which two earlier chapters does Chapter 6 genuinely depend on, rather than merely reference?

    pch.quizShowAnswer

    B — Chapter 4, because sampling a Gaussian needs the Cholesky factorisation to exist for a symmetric positive definite covariance; and Chapter 5, because Section 6.7's change of variables IS the Jacobian determinant — Both are load-bearing. Section 6.5.4 names Cholesky explicitly and notes the condition it needs, and Theorem 6.16 is Chapter 5's Section 5.3 applied to densities.

Section 6.8 opens by conceding that “this chapter is rather terse at times”, and then spends most of its length on the measure theory it deliberately skipped.

§6.8’s pointerand where it lands in this module
more relaxed presentations for self-studythis module’s own Construction of a Probability Space takes the slower route
an overview of exponential familiesConjugacy and the Exponential Family
“how to use probability distributions to model machine learning tasks”When Models Meet Data Overview, the whole of Chapter 8
normalizing flows “rely on change of variables for transforming random variables”Change of Variables and the Inverse Transform
variational inference applied to neural networksthe bound of The Latent Variable Perspective is the same construction
measure-theoretic questions “side stepped”acknowledged, not repaired — see the caution below
“an alternative way to approach probability is to start with the concept of expectation”Summary Statistics and Independence
  • Two rules generate the chapter. The sum rule (Eq 6.20) and the product rule (Eq 6.22). Bayes’ theorem (Eq 6.23) is those two rearranged, not a third axiom.
  • The book’s Figure 6.1 map: a random variable and its distribution feed the two rules and Bayes on one side, and summary statistics, independence and the inner-product view on the other; sufficient statistics lead to the exponential family; the Gaussian, Bernoulli, Binomial and Beta are the named examples.
  • Three objects get conflated and should not: the sample space, the event space, and the probability measure — plus the random variable, which is a FUNCTION from the first to a target space.
  • A density is not a probability, and for a continuous variable P(X = x) = 0 exactly. Only the cdf is a probability at every point.
  • The sum rule is the bottleneck: 2^D terms for D binary variables, 1.27e30 at D = 100, with no known exact polynomial-time algorithm.
  • The Gaussian earns its place by CLOSURE, not by ubiquity: marginalise, condition, multiply, add, map affinely — all closed form in the mean and covariance.
  • The exponential family is provably the only class with finite-dimensional sufficient statistics under repeated sampling (Pitman, Darmois, Koopman, 1935-36), which is what keeps the parameter count fixed as data grows.
  • Section 6.7’s change of variables is Chapter 5’s Jacobian, and Section 6.5.4’s sampler is Chapter 4’s Cholesky. Those two dependencies are load-bearing, not decorative.
  • Three reading paths: 6.1-6.3-6.6 for Bayesian inference; 6.2-6.4-6.5 plus Exercises 6.5 and 6.12 for the Gaussian machinery; 6.2 and 6.7 for generative models.

Start with: Construction of a Probability Space — the three objects, and the function that is neither random nor a variable.

pch.coffeeTagline

pch.coffeeCta

pch.feedbackHeading

pch.feedbackSubheading