| ⟨w,xa−xb⟩ on the hyperplane | 0.000×100, 200,000 pairs | 1201 |
| separating hyperplanes sampled | 9,628 of 400,000 | 1201 |
| their held-out accuracy | worst 0.8757, best 0.9877, spread 0.1119 | 1201 |
| correlation, margin against accuracy | 0.590999 | 1201 |
| max-margin accuracy, and its rank | 0.9853; beats 91.76% | 1201 |
| the Bayes rate for that generator | 0.9876 at x=2.25 | 1201 |
| worst accuracy per margin bin | 0.8757→0.9777 | 1201 |
| r=1/∥w∥ | 3.553×10−15, 20,000 draws | 1202, 1203 |
| isotropic rescale: boundary | 2.000000 at every c | 1202 |
| margin as an intruder enters | min{1,(3−t)/2}; infeasible at t=3 | 1202 |
| 12.10 vs 12.21, convergence failures | 251/2000 against 0/2000 | 1202 |
| squaring the norm: change in argmin | 4.441×10−16 | 1203 |
| Theorem 12.1 over 300 datasets | 7.492×10−6 degrees; 1 non-convergence | 1203 |
| Equation 12.24 | 1.137×10−13, 50,000 draws | 1203 |
| ∥w∥ vs 21∥w∥2, 500 starts | 3 failures, 29 iterations against 0 and 5 | 1203 |
| soft margin recovers hard margin | exactly by C=1; ∣b−b∗∣=0 | 1204 |
| C→0 | ∥w∥∝C; margin 800 at C=10−4 | 1204 |
| translation invariance, b free | 1.993×10−15 | 1204 |
| translation invariance, b penalised | objective moves 10.48 | 1204 |
| the knife edge at small C | labels decided by f=±1.2×10−15 | 1204 |
| the exact zero-one optimum | 1 error at w=(1,0), b=−3 | 1205 |
| 12.31 against 12.26 | 4.45×10−13 or better | 1205 |
| hinge boundary with a distant outlier | 2.000000, unchanged at every distance | 1205 |
| squared-loss boundary, same outlier | 2.33→3.97; 2 examples misclassified | 1205 |
| the dual multipliers | (41,41,0,0,21,0,0,0) to 1.075×10−8 | 1206 |
| ∥w∥2=∑αn | 4.782×10−8, 200 datasets | 1206 |
| adding 500 easy examples | still 3 support vectors; 8.289×10−12 of mass | 1206 |
| duality gap | ≤1.8×10−13 across five C | 1206 |
| the margin note at C=2 | on-margin examples have αn=C | 1206 |
| closest hull points | c=(3,0), d=(1,0), distance 2 | 1207 |
| hull-to-SVM conversion, 200 datasets | 9.185×10−8, 4.282×10−8, 5.051×10−11 | 1207 |
| hull weights vs multipliers | 5.780×10−8 | 1207 |
| the reduced hull | ∥c−d∥ 2.00→4.40; class means 4.5 apart | 1207 |
| polynomial feature dimension at D=1000, p=5 | 8,459,043,543,951 | 1208 |
| Gram rank saturation | 6, 10, 21 at every N | 1208 |
| RBF Gram rank | 50→171 (γ=0.5); 50→597 (γ=5) | 1208 |
| tanh smallest eigenvalue | −55.44, 10 of 120 negative | 1208 |
| four kernels on two rings | 0.6417, 1.0000, 1.0000, 1.0000 | 1208 |
| support vectors, linear against degree 2 | 120 against 6 | 1208 |
| subgradient descent, 105 steps | gap 4.091×10−4 | 1209 |
| both standard-form QPs | 7.105×10−15, 2.132×10−14 | 1209 |
| generic solver against LIBSVM | 89×, 12,349×, 49,548× | 1209 |
| max-margin percentile over 192 datasets | 84.11%; best on 6 | 1210 |
| a change of units | 0.2638 of held-out accuracy | 1210 |
| standardising first | spread falls to 0.0029 | 1210 |
| leave-one-out bound | 0.0333 against 0.0000; 0.5500 against 0.2333 | 1210 |
| C‘s good plateau | four orders of magnitude | 1210 |
| γ=0.01 across C | 0.7143→0.9965 | 1210 |
| γ=100 across C | never above 0.8375 | 1210 |
| SVM scores outside [−1,1] | 74.40% | 1210 |
| deleting non-support vectors | SVM 0.000000; logistic offset 0.537167 | 1210 |
| 10 positives against 200 negatives | recall 0.0000 | 1210 |