Change of Variables and the Inverse Transform
The book opens with the honest observation that the named distributions run out fast:
It may seem that there are very many known distributions, but in reality the set of distributions for which we have names is quite limited.
So you need to know what happens when you transform one. If , what is the distribution of ? Of ? §6.4.4 gave the mean and variance under affine maps, but not the functional form, and said nothing about nonlinear maps.
This section gives two techniques, and the second one is where Chapter 5’s Jacobian finally gets used for what it was built for.
What you’ll learn
Section titled “What you’ll learn”- Equation 6.125: the discrete case, which needs no Jacobian at all — and why.
- §6.7.1, the distribution function technique: find the cdf, differentiate it. Worked on the book’s Example 6.16.
- Theorem 6.15, the probability integral transform: is uniform, whatever was — measured on four distributions including a Cauchy.
- Run backwards, that theorem is inverse-transform sampling.
- §6.7.2, the change-of-variables technique: Equation 6.143 and its factor — measured to show it is not optional.
- Theorem 6.16: the multivariate version, with of the Jacobian, and Example 6.17 recovering §6.5’s Gaussian result.
- The trap that follows: a density’s mode is not preserved by a nonlinear reparameterisation, while its quantiles are.
Intuition: mass versus density
Section titled “Intuition: mass versus density”For a discrete variable, transforming is trivial. A pmf assigns mass to points; apply an invertible and the points move while their masses ride along unchanged. Equation 6.125 is the whole story.
For a continuous variable, it is not, and the reason is §6.2’s distinction: a density is not a probability. Probability lives in area, so a density is mass per unit length. Transform the variable and you stretch or squash the axis — so the same mass now occupies a different length, and the density must be rescaled to compensate.
That compensation factor is the Jacobian. Squeeze an interval to half its length and the density there must double. This is why the continuous case has an extra term the discrete case does not, and dropping it does not produce a slightly wrong density — it produces something that is not a density at all.
There is a second, less obvious consequence. Because the Jacobian reweights the density unevenly, the location of the peak can move. A “most likely value” is not a property of the random variable; it is a property of the coordinate system you wrote it in.
flowchart TD
Q{"discrete or continuous?"}
Q -->|"discrete"| D["Eq 6.125: P(Y=y) = P(X = U-inverse(y))
masses ride along, no Jacobian"]
Q -->|"continuous"| C["a density is mass PER UNIT LENGTH"]
C --> T1["technique 1, Section 6.7.1
find the cdf, then differentiate"]
C --> T2["technique 2, Section 6.7.2
Eq 6.143, the recipe"]
T1 --> TH["Theorem 6.15
F_X(X) is UNIFORM"]
TH --> SAMP["inverse-transform sampling
F-inverse(U) has the target law"]
TH --> COP["hypothesis testing, copulas"]
T2 --> J["the factor |d U-inverse / dy|"]
J --> MV["Theorem 6.16: |det| of the Jacobian
Chapter 5's Section 5.3"]
J --> TRAP["the density is REWEIGHTED
so the MODE moves"]
TRAP --> MAP["a MAP estimate is not
reparameterisation-invariant"]
The math
Section titled “The math”Equation 6.125: the discrete case
Section titled “Equation 6.125: the discrete case”For a discrete with pmf and an invertible , set :
“For discrete random variables, transformations directly change the individual events.” The probabilities are carried over untouched.
§6.7.1: the distribution function technique
Section titled “§6.7.1: the distribution function technique”Go back to first principles. For :
- Find the cdf, — Equation 6.126.
- Differentiate it to get the pdf, — Equation 6.127.
And keep track of the domain, which the transformation may have changed.
Example 6.16. Let on , and let . Since is increasing on that interval, also lies in :
so and
Theorem 6.15: the probability integral transform
Section titled “Theorem 6.15: the probability integral transform”The technique has a striking special case: use itself as the transformation.
Theorem 6.15. Let be a continuous random variable with a strictly monotonic cdf . Then has a uniform distribution.
Whatever was. A Gaussian, an exponential, a Cauchy with no mean at all — push each sample through its own cdf and the output is uniform on .
Three uses, all named in the book:
- Sampling. Run it backwards: draw and return . That is inverse-transform sampling, and it is why every library only needs one good uniform generator.
- Hypothesis testing whether a sample came from a particular distribution — if it did, the transformed values must look uniform.
- Copulas, which are built on exactly this observation.
§6.7.2: change of variables
Section titled “§6.7.2: change of variables”The distribution-function technique works but has to be redone from scratch each time. The change-of-variables technique is a recipe.
The derivation, compressed. An invertible on an interval is strictly increasing or strictly decreasing; take increasing. Then
Differentiate with respect to , substituting via Equation 6.133 so the integration variable matches, and invoke the fundamental theorem of calculus:
For a decreasing the same derivation produces a minus sign, so take the absolute value and both cases are one formula:
That factor “measures how much a unit volume changes when applying ” — the same quantity §5.3 called the Jacobian.
Theorem 6.16: the multivariate case
Section titled “Theorem 6.16: the multivariate case”The absolute value does not generalise, so it becomes a determinant.
Theorem 6.16. If is differentiable and invertible on the domain of , then
The recipe: work out the inverse transform, substitute it into the density of , then multiply by the absolute determinant of the Jacobian. The determinant appears “because our differentials (cubes of volume) are transformed into parallelepipeds by the Jacobian” — §4.1’s reading of the determinant, used literally.
Example 6.17. Take a standard bivariate normal,
and a linear map with . The inverse transform is the matrix inverse:
Substituting gives (Equation 6.148). The Jacobian of a linear map is the matrix itself (, Equation 6.149), and the determinant of an inverse is the inverse of the determinant. So the extra factor is , and the result is — exactly what §6.5’s Equation 6.88 said, now derived from the general recipe instead of asserted.
Worked example by hand
Section titled “Worked example by hand”Example 6.16, checked
Section titled “Example 6.16, checked”on , , so .
Sanity check first: does it integrate to one? . ✓
Now sample it. , so by inverse transform for . Square four million of those and compare:
| analytic | empirical | |
|---|---|---|
Agreement to about , which is sampling noise at this .
Note the shape flipped. increases on ; also increases but is concave rather than convex, and is avoided only because the exponent is positive. Squaring compressed the region near zero and stretched the region near one, and the density had to be reweighted accordingly.
The Jacobian, dropped
Section titled “The Jacobian, dropped”Take and , so and . Equation 6.143 gives the lognormal
Now compare against what you get by forgetting the :
| integrates to | |
|---|---|
| with the Jacobian | |
| without |
The second is not a density. And renormalising it does not rescue it, because the shape is wrong too:
| sampled histogram | with Jacobian | renormalised, no Jacobian | |
|---|---|---|---|
Worst error with the Jacobian: . Without: — a factor of over a hundred, and qualitatively wrong at both ends.
Theorem 6.15, on four distributions
Section titled “Theorem 6.15, on four distributions”Push each sample through its own cdf. A uniform has mean and variance :
| drawn from | mean of | variance |
|---|---|---|
| Exponential(rate 2) | ||
| Beta | ||
| standard Cauchy |
All four land on the uniform — including the Cauchy, whose own mean does not exist. The theorem does not care about moments; it only needs a strictly monotonic cdf.
Example 6.17, checked against §6.5
Section titled “Example 6.17, checked against §6.5”With , , so the Jacobian factor is . Evaluating Equation 6.144 against the density directly:
| point | Equation 6.144 | gap | |
|---|---|---|---|
The general recipe and the Gaussian shortcut are the same thing.
And the trap: the mode moves
Section titled “And the trap: the mode moves”has its mode at . So where is the mode of ?
The tempting answer is . Measured: , which is . The Jacobian is larger for small , so it tilts the density leftward and drags the peak with it.
| statistic of | image under | actual statistic of |
|---|---|---|
| mode | ✗ | |
| median | ✓ | |
| mean | ✗ |
And it depends on the transform, not on the data:
| transform | mode in -space | image of ‘s mode |
|---|---|---|
| ✓ |
Only the affine map leaves it alone, because its Jacobian is constant and so cannot tilt anything. The median survives every monotone map, because quantiles are defined by mass rather than density.
See it move
Section titled “See it move”From scratch
Section titled “From scratch”"""Section 6.7 — transforming a random variable, and the Jacobian that makes the
result a density rather than a relabelling."""
import numpy as np
np.set_printoptions(precision=6, suppress=True, linewidth=150)
rng = np.random.default_rng(7)
print("########## discrete_needs_no_jacobian")
# Eq 6.125: for a discrete variable, transforming just relabels the events.
states = np.array([1, 2, 3, 4, 5])
pmf = np.array([0.10, 0.15, 0.30, 0.25, 0.20])
U = lambda k: k ** 2 # invertible on these states
print(f" X on {states.tolist()} with pmf {pmf}")
print(f" Y = X^2 on {U(states).tolist()} with pmf {pmf}")
print(f" the probabilities are CARRIED OVER unchanged; both sum to "
f"{pmf.sum():.10f} and {pmf.sum():.10f}")
print(" no Jacobian appears, because a pmf assigns mass to points and the points")
print(" simply move. Eq 6.125b is the whole story for the discrete case.")
print()
print("########## example_6_16")
# Eq 6.128 to 6.131: f(x) = 3x^2 on [0,1], Y = X^2, so f(y) = 1.5 sqrt(y).
print(" f(x) = 3x^2 on [0,1], Y = X^2")
print(" Eq 6.130 F_Y(y) = y^(3/2) Eq 6.131 f(y) = (3/2) y^(1/2)")
# Sample X by inverse cdf: F_X(x) = x^3, so X = U^(1/3).
u = rng.random(4_000_000)
xs = u ** (1 / 3)
ys = xs ** 2
grid = np.linspace(1e-9, 1, 2001)
analytic = 1.5 * np.sqrt(grid)
print(f" analytic f(y) integrates to "
f"{np.trapezoid(analytic, grid):.10f}")
# Compare against a histogram.
hist, edges = np.histogram(ys, bins=60, range=(0, 1), density=True)
centres = 0.5 * (edges[1:] + edges[:-1])
pred = 1.5 * np.sqrt(centres)
print(f" histogram vs analytic: worst gap {np.abs(hist - pred).max():.4f}"
f" mean |gap| {np.abs(hist - pred).mean():.5f}")
print(f" and the cdf at a few points:")
for yv in (0.1, 0.25, 0.5, 0.81):
print(f" F_Y({yv:.2f}) analytic {yv**1.5:.6f} empirical {float((ys <= yv).mean()):.6f}")
print()
print("########## the_jacobian_is_not_optional")
# Eq 6.143. Transform a standard normal by U(x) = exp(x): the lognormal.
# f(y) = f_x(log y) * |d/dy log y| = f_x(log y) / y.
def phi(t):
return np.exp(-0.5 * t ** 2) / np.sqrt(2 * np.pi)
g2 = np.geomspace(1e-4, 60, 400_001)
with_jac = phi(np.log(g2)) / g2 # correct, Eq 6.143
without = phi(np.log(g2)) # the Jacobian dropped
print(" X ~ N(0,1), Y = exp(X). Eq 6.143 gives f(y) = phi(log y) * |1/y|")
print(f" with the Jacobian, integral = {np.trapezoid(with_jac, g2):.10f}")
print(f" without the Jacobian, integral = {np.trapezoid(without, g2):.10f}")
print(" the second is not a density at all -- it does not integrate to 1, and no")
print(" amount of renormalising would give the right SHAPE either:")
sn = rng.standard_normal(4_000_000)
yy = np.exp(sn)
h2, e2 = np.histogram(yy, bins=np.geomspace(0.02, 30, 61), density=True)
c2 = np.sqrt(e2[1:] * e2[:-1])
ok = phi(np.log(c2)) / c2
bad = phi(np.log(c2))
bad = bad / np.trapezoid(phi(np.log(g2)), g2) # even after renormalising
print(f" {'y':>8} {'histogram':>11} {'with Jacobian':>14} {'renormalised, no J':>19}")
for i in (5, 20, 35, 50):
print(f" {c2[i]:>8.4f} {h2[i]:>11.6f} {ok[i]:>14.6f} {bad[i]:>19.6f}")
print(f" worst gap, with Jacobian: {np.abs(h2 - ok).max():.5f}")
print(f" worst gap, without: {np.abs(h2 - bad).max():.5f}")
print()
print("########## theorem_6_15_probability_integral_transform")
# Y := F_X(X) is UNIFORM, whatever X is.
print(" push each sample through its OWN cdf and the result is uniform:")
print(f" {'distribution':>22} {'mean':>9} {'variance':>9} {'max |F_emp - u|':>16}")
cases = []
n = 2_000_000
# Normal
z = rng.standard_normal(n)
from math import erf
cases.append(("N(0,1)", 0.5 * (1 + np.vectorize(erf)(z / np.sqrt(2)))))
# Exponential(2)
e = rng.exponential(1 / 2.0, n)
cases.append(("Exponential(rate 2)", 1 - np.exp(-2 * e)))
# Beta(2,5)
bt = rng.beta(2, 5, n)
from math import lgamma
def beta_cdf(v, a, b, m=4000):
t = np.linspace(0, 1, m)
lc = lgamma(a + b) - lgamma(a) - lgamma(b)
d = np.exp(lc + (a - 1) * np.log(np.clip(t, 1e-12, None))
+ (b - 1) * np.log1p(-np.clip(t, None, 1 - 1e-12)))
cdf = np.concatenate([[0.0], np.cumsum((d[1:] + d[:-1]) / 2 * np.diff(t))])
cdf /= cdf[-1]
return np.interp(v, t, cdf)
cases.append(("Beta(2,5)", beta_cdf(bt, 2, 5)))
# Cauchy, which has no mean at all
c = rng.standard_cauchy(n)
cases.append(("standard Cauchy", 0.5 + np.arctan(c) / np.pi))
for name, v in cases:
srt = np.sort(v[:200_000])
ks = float(np.abs(srt - np.linspace(0, 1, srt.size)).max())
print(f" {name:>22} {v.mean():>9.6f} {v.var():>9.6f} {ks:>16.6f}")
print(" a uniform has mean 0.5 and variance 1/12 = 0.083333. Every row matches,")
print(" including the Cauchy, whose own mean does not exist.")
print()
print("########## inverse_transform_sampling")
# The other direction: F^-1(U) has the target distribution.
uu = rng.random(2_000_000)
print(" and running Theorem 6.15 backwards is how sampling works:")
targets = [
("Exponential(rate 2)", -np.log1p(-uu) / 2.0,
lambda t: 1 - np.exp(-2 * t), (0.0, 3.0)),
("f(x) = 3x^2 on [0,1]", uu ** (1 / 3), lambda t: t ** 3, (0.0, 1.0)),
]
for name, smp, cdf, (lo, hi) in targets:
qs = np.linspace(lo + 1e-6, hi, 7)
worst = max(abs(float((smp <= q).mean()) - cdf(q)) for q in qs)
print(f" {name:>22} worst |empirical cdf - target cdf| over 7 points: {worst:.6f}")
print()
print("########## example_6_17_linear_map")
# Theorem 6.16 with a linear U. det of the Jacobian is 1/|det A|.
A = np.array([[2.0, 1.0], [-0.5, 1.5]])
detA = float(np.linalg.det(A))
print(f" A = {A.tolist()} det A = {detA:.6f}")
print(f" Eq 6.149 d/dy (A^-1 y) = A^-1, and det(A^-1) = 1/det(A) = {1/detA:.6f}")
S_pred = A @ A.T
print(f" so Y = A X with X ~ N(0, I) has covariance A A^T =\n{S_pred}")
X2 = rng.standard_normal((3_000_000, 2))
Y2 = X2 @ A.T
print(f" measured covariance gap: {np.abs(np.cov(Y2.T, bias=True) - S_pred).max():.5f}")
# Now check the DENSITY at a few points, via Eq 6.144.
Ainv = np.linalg.inv(A)
pts = np.array([[0.0, 0.0], [1.0, 0.5], [-2.0, 1.0], [3.0, -1.5]])
lhs = [] # Eq 6.144
rhs = [] # the N(0, A A^T) density directly
for pt in pts:
xpre = Ainv @ pt
f_x = np.exp(-0.5 * xpre @ xpre) / (2 * np.pi)
lhs.append(f_x * abs(1 / detA))
sign, logdet = np.linalg.slogdet(S_pred)
rhs.append(float(np.exp(-0.5 * pt @ np.linalg.solve(S_pred, pt))
/ (2 * np.pi * np.exp(0.5 * logdet))))
lhs, rhs = np.array(lhs), np.array(rhs)
print(f" {'point':>16} {'Eq 6.144':>14} {'N(0, A A^T)':>14} {'gap':>9}")
for pt, l, r in zip(pts, lhs, rhs):
print(f" {str(pt.tolist()):>16} {l:>14.10f} {r:>14.10f} {abs(l-r):>9.1e}")
print(f" worst gap {np.abs(lhs - rhs).max():.1e} -- the change-of-variables recipe")
print(" reproduces the Gaussian result of Section 6.5 exactly.")
print()
print("########## the_mode_is_not_invariant")
# A density's ARGMAX moves under reparameterisation, because the Jacobian
# reweights it. The mean does not have this problem.
print(" X ~ N(0,1), Y = exp(X). Where is the mode of each?")
gy = np.geomspace(1e-3, 30, 2_000_001)
f_y = phi(np.log(gy)) / gy
mode_y = float(gy[int(np.argmax(f_y))])
# The mean is far more tail-sensitive than the mode: y f(y) still carries mass
# well past y = 30, so it needs a wider grid than the peak does.
gy_wide = np.geomspace(1e-4, 5_000, 4_000_001)
mean_y = float(np.trapezoid(gy_wide * (phi(np.log(gy_wide)) / gy_wide), gy_wide))
med_y = 1.0 # exp(median of X) = exp(0)
print(f" mode of X in x-space: 0.000000 -> exp(0) = 1.000000")
print(f" mode of Y in y-space: {mode_y:.6f} (theory exp(-1) = {np.exp(-1):.6f})")
print(f" so the 'most likely' point MOVED: 1.000000 -> {mode_y:.6f}")
print(f" mean of Y {mean_y:.6f} (theory exp(1/2) = {np.exp(0.5):.6f})")
print(f" median of Y {med_y:.6f} -- the median DOES transform correctly")
print(" the mode is not reparameterisation-invariant: the Jacobian |1/y| tilts the")
print(" density and drags the peak. Quantiles are invariant under a monotone map,")
print(" and so is the median; a MAP estimate is not.")
# Show it depends on the transform chosen, not on the data.
for name, fwd, inv, dinv in (
("Y = exp(X)", np.exp, np.log, lambda y: 1 / y),
("Y = X^3", lambda t: t ** 3, np.cbrt, lambda y: np.abs(np.cbrt(y) ** -2) / 3),
("Y = 2X + 1", lambda t: 2 * t + 1, lambda y: (y - 1) / 2, lambda y: 0.5 + 0 * y)):
if name == "Y = exp(X)":
gg = np.geomspace(1e-3, 30, 400_001)
elif name == "Y = X^3":
gg = np.linspace(-40, 40, 400_001)
gg = gg[np.abs(gg) > 1e-6]
else:
gg = np.linspace(-12, 14, 400_001)
dens = phi(inv(gg)) * np.abs(dinv(gg))
peak = float(gg[int(np.argmax(dens))])
print(f" {name:<12} mode in y-space {peak:>10.6f} "
f"image of X's mode {float(fwd(np.array(0.0))):>10.6f}")
print(" only the affine map leaves the mode where the image of the old mode is,")
print(" because its Jacobian is constant.")########## discrete_needs_no_jacobian
X on [1, 2, 3, 4, 5] with pmf [0.1 0.15 0.3 0.25 0.2 ]
Y = X^2 on [1, 4, 9, 16, 25] with pmf [0.1 0.15 0.3 0.25 0.2 ]
the probabilities are CARRIED OVER unchanged; both sum to 1.0000000000 and 1.0000000000
no Jacobian appears, because a pmf assigns mass to points and the points
simply move. Eq 6.125b is the whole story for the discrete case.
########## example_6_16
f(x) = 3x^2 on [0,1], Y = X^2
Eq 6.130 F_Y(y) = y^(3/2) Eq 6.131 f(y) = (3/2) y^(1/2)
analytic f(y) integrates to 0.9999965411
histogram vs analytic: worst gap 0.0083 mean |gap| 0.00322
and the cdf at a few points:
F_Y(0.10) analytic 0.031623 empirical 0.031569
F_Y(0.25) analytic 0.125000 empirical 0.125161
F_Y(0.50) analytic 0.353553 empirical 0.353669
F_Y(0.81) analytic 0.729000 empirical 0.729223
########## the_jacobian_is_not_optional
X ~ N(0,1), Y = exp(X). Eq 6.143 gives f(y) = phi(log y) * |1/y|
with the Jacobian, integral = 0.9999788320
without the Jacobian, integral = 1.6470952340
the second is not a density at all -- it does not integrate to 1, and no
amount of renormalising would give the right SHAPE either:
y histogram with Jacobian renormalised, no J
0.0391 0.053651 0.053321 0.001266
0.2433 0.604508 0.603890 0.089214
1.5143 0.241621 0.241713 0.222228
9.4241 0.003439 0.003419 0.019564
worst gap, with Jacobian: 0.00474
worst gap, without: 0.52751
########## theorem_6_15_probability_integral_transform
push each sample through its OWN cdf and the result is uniform:
distribution mean variance max |F_emp - u|
N(0,1) 0.500185 0.083247 0.001802
Exponential(rate 2) 0.499882 0.083364 0.002249
Beta(2,5) 0.499806 0.083257 0.001523
standard Cauchy 0.499835 0.083290 0.002241
a uniform has mean 0.5 and variance 1/12 = 0.083333. Every row matches,
including the Cauchy, whose own mean does not exist.
########## inverse_transform_sampling
and running Theorem 6.15 backwards is how sampling works:
Exponential(rate 2) worst |empirical cdf - target cdf| over 7 points: 0.000220
f(x) = 3x^2 on [0,1] worst |empirical cdf - target cdf| over 7 points: 0.000481
########## example_6_17_linear_map
A = [[2.0, 1.0], [-0.5, 1.5]] det A = 3.500000
Eq 6.149 d/dy (A^-1 y) = A^-1, and det(A^-1) = 1/det(A) = 0.285714
so Y = A X with X ~ N(0, I) has covariance A A^T =
[[5. 0.5]
[0.5 2.5]]
measured covariance gap: 0.00196
point Eq 6.144 N(0, A A^T) gap
[0.0, 0.0] 0.0454728409 0.0454728409 0.0e+00
[1.0, 0.5] 0.0398236988 0.0398236988 6.9e-18
[-2.0, 1.0] 0.0227198205 0.0227198205 3.5e-18
[3.0, -1.5] 0.0095437907 0.0095437907 5.2e-18
worst gap 6.9e-18 -- the change-of-variables recipe
reproduces the Gaussian result of Section 6.5 exactly.
########## the_mode_is_not_invariant
X ~ N(0,1), Y = exp(X). Where is the mode of each?
mode of X in x-space: 0.000000 -> exp(0) = 1.000000
mode of Y in y-space: 0.367880 (theory exp(-1) = 0.367879)
so the 'most likely' point MOVED: 1.000000 -> 0.367880
mean of Y 1.648721 (theory exp(1/2) = 1.648721)
median of Y 1.000000 -- the median DOES transform correctly
the mode is not reparameterisation-invariant: the Jacobian |1/y| tilts the
density and drags the peak. Quantiles are invariant under a monotone map,
and so is the median; a MAP estimate is not.
Y = exp(X) mode in y-space 0.367878 image of X's mode 1.000000
Y = X^3 mode in y-space -0.000200 image of X's mode 0.000000
Y = 2X + 1 mode in y-space 1.000000 image of X's mode 1.000000
only the affine map leaves the mode where the image of the old mode is,
because its Jacobian is constant.Five things worth stopping on.
The discrete case really needs nothing. on mapped to keeps its pmf exactly; both sum to .
Example 6.16 checks out to about on the cdf and on the histogram.
Dropping the Jacobian is not a small error. The “density” integrates to , and after renormalising the worst pointwise error is against for the correct formula.
Theorem 6.15 is indifferent to the distribution. All four transformed samples have mean and variance — including the Cauchy.
Equation 6.144 reproduces §6.5 exactly, worst gap .
And the finding that matters most downstream: the mode of is , not . The median transforms correctly; the mode does not; and only an affine map — constant Jacobian — leaves it alone.
On real data
Section titled “On real data”Reading the plot
Section titled “Reading the plot”From the first figure. The middle panel’s dashed curve is the interesting failure. It has been renormalised, so its total mass is right — and it is still badly wrong, missing almost all the mass near and putting five times too much at . A missing Jacobian is not a scaling bug you can absorb into a constant; it is a shape error, because the factor varies across the domain.
From the second figure. The left panel is four different distributions drawn on top of each other and you cannot tell them apart, which is the whole point. The right panel is the same fact monetised: because the transform works in both directions, sampling from any distribution with an invertible cdf reduces to sampling a uniform.
From the third figure. Compare the three vertical lines in the right panel. Two of the standard summaries of “where the distribution is” — the mode and the mean — moved somewhere the old summary did not predict; the median did not. If you report a MAP estimate, you are reporting something that depends on whether you parameterised by a variance or a log-variance, a probability or a log-odds.
Pitfalls
Section titled “Pitfalls”Compare
Section titled “Compare”| discrete | continuous | |
|---|---|---|
| Object transformed | pmf, mass at points | pdf, mass per unit length |
| Formula | , Eq 6.125 | Eq 6.143, with a Jacobian |
| Extra factor | none | |
| Why | points move, masses ride along | the axis stretches, so the ratio changes |
| the pmf value | , always |
| Technique | How | Good for |
|---|---|---|
| distribution function (§6.7.1) | find , differentiate | first principles, one-offs |
| change of variables (§6.7.2) | Eq 6.143 / 6.144 | a reusable recipe, any dimension |
| probability integral transform (Thm 6.15) | use as the map | sampling, testing, copulas |
| Statistic | Survives a monotone reparameterisation? |
|---|---|
| median, and any quantile | yes |
| mode (MAP) | no — measured |
| mean | no — measured |
| support | yes, mapped through |
-
Why does the continuous change-of-variables formula have a Jacobian factor when the discrete one does not?
The book's Remark makes the same point from the other side: P(Y = y) = 0 for all y in the continuous case, so f(y) has no description as the probability of an event, and cannot simply be carried over.
pch.quizShowAnswer
B — Because a density is mass PER UNIT LENGTH. Transforming stretches the axis, so the same mass occupies a different length and the density must be rescaled. A pmf assigns mass to points, which simply move — The book's Remark makes the same point from the other side: P(Y = y) = 0 for all y in the continuous case, so f(y) has no description as the probability of an event, and cannot simply be carried over.
-
You transform Y = exp(X) but forget the factor |1/y|, then renormalise so the result integrates to 1. Is that good enough?
A missing Jacobian is a shape error, not a scaling error. The one exception is an affine map, whose Jacobian is constant — there, and only there, renormalising happens to work.
pch.quizShowAnswer
B — No. Measured: the worst pointwise error is still 0.5275 against 0.0047 for the correct density — far too little mass at small y and five times too much in the tail. The factor varies across the domain, so no constant repairs it — A missing Jacobian is a shape error, not a scaling error. The one exception is an affine map, whose Jacobian is constant — there, and only there, renormalising happens to work.
-
Theorem 6.15 says F_X(X) is uniform. What does it require of X?
Strict monotonicity is the real requirement. A distribution with an atom has a flat stretch in its cdf, and then the transform is not uniform — which is why the theorem is stated for continuous random variables.
pch.quizShowAnswer
B — Only a strictly monotonic cdf — measured on a standard Cauchy, whose mean does not exist, and it comes out uniform to the same accuracy as a Gaussian does — Strict monotonicity is the real requirement. A distribution with an atom has a flat stretch in its cdf, and then the transform is not uniform — which is why the theorem is stated for continuous random variables.
-
X is standard normal with mode 0. Where is the mode of Y = exp(X)?
The median IS invariant, since a monotone map preserves quantiles: it stays at exactly 1. So a MAP estimate depends on whether you parameterise by sigma or log sigma, while a posterior median does not.
pch.quizShowAnswer
B — At exp(-1) = 0.3679. The Jacobian 1/y is larger for small y, so it tilts the density leftward and drags the peak — the mode is not reparameterisation-invariant — The median IS invariant, since a monotone map preserves quantiles: it stays at exactly 1. So a MAP estimate depends on whether you parameterise by sigma or log sigma, while a posterior median does not.
-
Example 6.17 applies Theorem 6.16 to y = Ax on a standard bivariate normal. What comes out?
The Jacobian of a linear map is the matrix itself, and the determinant of an inverse is the inverse of the determinant. Section 6.5 asserted this closure property; Section 6.7 derives it from a rule that works for any invertible transform.
pch.quizShowAnswer
B — N(0, A A-transpose), with the Jacobian contributing a factor of one over |det A| — matching Section 6.5's Equation 6.88 to 6.9e-18, so the general recipe and the Gaussian shortcut agree — The Jacobian of a linear map is the matrix itself, and the determinant of an inverse is the inverse of the determinant. Section 6.5 asserted this closure property; Section 6.7 derives it from a rule that works for any invertible transform.
🧪 Try It Yourself
Section titled “🧪 Try It Yourself”Exercise 1 – Example 6.16
Section titled “Exercise 1 – Example 6.16”Exercise 2 – Drop the Jacobian and watch
Section titled “Exercise 2 – Drop the Jacobian and watch”Exercise 3 – The probability integral transform
Section titled “Exercise 3 – The probability integral transform”Exercise 4 – Theorem 6.16 on a linear map
Section titled “Exercise 4 – Theorem 6.16 on a linear map”Exercise 5 – The mode moves, the median does not
Section titled “Exercise 5 – The mode moves, the median does not”Recall card
Section titled “Recall card”- Discrete transforms need no Jacobian. Eq 6.125: P(Y=y) = P(X = U-inverse(y)). A pmf puts mass at points, and the points just move.
- Continuous transforms do, because a density is mass PER UNIT LENGTH and the transformation changes the length.
- Technique 1, Section 6.7.1: find the cdf F_Y(y), the probability that Y is at most y, then differentiate it. Example 6.16: f(x) = 3x^2 with Y = X^2 gives F_Y = y^(3/2) and f(y) = 1.5 sqrt(y).
- Eq 6.143 is the recipe: f(y) = f_x(U-inverse(y)) times |d/dy U-inverse(y)|. The absolute value covers both increasing and decreasing U.
- Dropping the Jacobian does not give a density. Measured on Y = exp(X): the integral is 1.6470952, and even after renormalising the worst pointwise error is 0.5275 against 0.0047. The factor varies across the domain, so no constant fixes it.
- Theorem 6.16 is the multivariate version, with |det| of the Jacobian in place of the absolute derivative — the determinant because differentials of volume become parallelepipeds.
- Example 6.17 recovers Section 6.5’s result. y = Ax on a standard bivariate normal gives N(0, A A-transpose), with the Jacobian supplying 1/|det A| — matched to 6.9e-18.
- Theorem 6.15, the probability integral transform: F_X(X) is UNIFORM for any continuous X with a strictly monotonic cdf. Measured on a Gaussian, exponential, Beta and Cauchy: means near 0.4998, variances near 0.0833 = 1/12.
- It needs monotonicity, not moments. The Cauchy has no mean and transforms just as cleanly. A distribution with an atom has a flat cdf stretch and does NOT.
- Run it backwards and it is inverse-transform sampling: F-inverse(U) has the target law. That is why one uniform generator suffices for a whole library.
- A density’s MODE is not reparameterisation-invariant. X ~ N(0,1) has mode 0, but Y = exp(X) has mode exp(-1) = 0.3679, not exp(0) = 1. The Jacobian tilts the density.
- Quantiles, and hence the median, ARE invariant under a monotone map — measured at exactly 1.0 for the same example. So a MAP estimate depends on your parameterisation and a posterior median does not.
- Only affine maps leave the mode alone, because their Jacobian is constant. Measured: exp and x-cubed move it, 2x+1 does not.
- Check invertibility and the new domain. Y = X^2 works on [0,1] but is not invertible on [-1,1]; there you must split the domain and sum the branches.
Next: Chapter 6 Exercises and Solutions — every exercise from the end of the chapter, worked and checked.
pch.coffeeTagline
pch.coffeeCtapch.feedbackHeading
pch.feedbackSubheading