| variance of Equation 11.5, within / between | 0.95 / 6.84 (87.80%) of 7.79 | 1101 |
| skewness, excess kurtosis of Equation 11.5 | 0.427185, −1.294278 | 1101 |
| KL to the closest single Gaussian | 0.330585 nats | 1101 |
| two equal Gaussians merge at | d=1.0000000 (2σ apart); 1.657251 at weight 0.9 | 1101 |
| a GMM’s edge over one Gaussian | 1.388610 nats per point, for 12 extra parameters | 1101 |
| the book’s starting negative log-likelihood | 28.325536 — its 28.3 | 1102 |
| Equation 11.9’s product underflows at | about 200 data points | 1102 |
| K=1 closed form vs an optimiser | 6×10−8 | 1102 |
| analytic gradient vs central differences | 3×10−10 | 1102 |
| feeding Equation 11.20 back in | moves the means 4.295713, then 0.130775, then 0.018414 | 1102 |
| 400 restarts on 7 points | 17 distinct optima; 218 (54.5%) collapse | 1102 |
| the unbounded likelihood | +16.377668 at σ2=10−30 | 1102 |
| the book’s basin vs the best honest one | −13.973323 vs −13.906162 | 1102 |
| Equation 11.19’s disputed entry (x=0) | printed 0.001; computed 0.000150 | 1103 |
| responsibilities as a softmax of −E | 1.1×10−16 | 1103 |
| mean row entropy at convergence | 0.014 against a possible 1.099 | 1103 |
| Equation 11.17 typed as written | 21 NaNs of 21; log-space agrees to 2.2×10−16 | 1103 |
| K-means from the same start | centres 1.504119 apart, split 3-2-2 | 1103 |
| Example 11.3 reproduced | −2.701230, −0.403411, 3.704287 | 1104 |
| means stay inside the data’s range | 0 excursions over 4,000 starts | 1104 |
| ∑kNkμk=∑nxn | 3.6×10−15 over 2,000 matrices | 1104 |
| component death, Nk underflows at | μ3=73.585070 (and −69.812512) | 1104 |
| the mean update alone | 12.321385 nats, stalling at −15.964728 | 1104 |
| Equation 11.30 with old vs new means | 1.83/0.60/19.98 vs 0.14/0.44/1.53 — factor 13.0878 | 1105 |
| positive semi-definiteness | 0 negatives in 60,000 values, no repair step | 1105 |
| the mixture’s variance after any M-step | = the sample variance, 3.6×10−15 over 20,000 | 1105 |
| that total on the book’s example | 8.336735 from the first M-step onward | 1105 |
| the singularity from the variance side | 2.5×10−4 at a responsibility of 0.99999 | 1105 |
| unconstrained ascent on the weights | the sum reaches 2.54 in six steps | 1106 |
| the Lagrange multiplier | exactly −N, deviation 0 | 1106 |
| Example 11.5 reproduced | 0.293890, 0.287001, 0.419109 | 1106 |
| one full cycle | 28.325536→14.410485 — the book’s 28.3→14.4 | 1106 |
| the split of the 13.92 nats | means 12.32, variances 1.46, weights 0.14 | 1106 |
| “after five iterations” | the 10−6 criterion exactly; 8 at 10−9 | 1107 |
| L vs parameters settling | 5 against 11 iterations at 10−6 | 1107 |
| monotonicity over 180,000 steps | 11,163 negative, none above 8.9×10−15 | 1107 |
| median decrease | 1.776×10−15 — one ULP | 1107 |
| overlap and the convergence rate | 1045 iterations at 53% ambiguous vs 3 at 0% | 1107 |
| K-means initialisation | median 58→19; best optimum 96%→100% | 1107 |
| Equation 11.69 vs Equation 11.17 | 0.0 | 1108 |
| the M-step as the argmax of Q | 5.3×10−15 over 40 restarts | 1108 |
| L=Q+H(r) | 1.1×10−14 | 1108 |
| the step’s rise in L and in Q | 13.915050 against 13.480070 | 1108 |
| parameter recovery, N=200 to 200,000 | errors fall by factors of 10, 20 and 11 | 1108 |
| choosing K: train / held-out / BIC / AIC | 6 / 3 / 3 / 4 | 1109 |
| a variance floor’s price | 21ln10=1.1513 nats per decade | 1109 |
| the honest six-component optimum | −128.9013, against −120.9761 for a spike | 1109 |
| label switching, all 3! relabellings | spread 0.000×100 | 1109 |
| GMM as a classifier, held out | 0.9650 — equal to the labelled fit, Bayes rate 0.9633 | 1109 |
| tying the variances | costs 1.8769 nats held out, saves 8.8856 of BIC | 1109 |
| held-out L recovering K as N grows | 11→19→19→15→14 of 20 | 1109 |
| BIC recovering K as N grows | 10→13→20→20→20 of 20 | 1109 |
| GMM vs the best KDE | 8 numbers, 13.7 nats ahead of 600 points | 1109 |