Convexity and Training Difficulty
Why does the shape of a loss decide whether training is easy?
Vishnu Boddeti
Example 1
Convex probe training is predictable; end-to-end tuning changes both the representation and the decoder.
Example 2
Robustness claims are meaningful only relative to a specified norm, radius, knowledge, and attack objective.
Example 3
Merging works best when checkpoints occupy compatible regions of parameter space and share a common initialization.
Example 4
Even plausible endpoints do not guarantee every mixture satisfies nonlinear safety or permission constraints.
Each scenario hides a mathematical structure from this lecture. Identify the structure, justify it, and work through the resulting problem.
A classifier is certified against Euclidean perturbations, but the deployment contract specifies a maximum per-feature change. The certificate must be translated without rerunning verification.
The input has dimension $d=64$, and the classifier is known to be safe whenever $\|\delta\|_2<0.24$.
A robust optimizer declares perturbations acceptable when $2|\delta_1|+|\delta_2|\le1$. Engineers want to report the smallest scaling needed to bring any perturbation into this safe set as its “risk magnitude.”
$C=\{\delta\in\mathbb{R}^2:2|\delta_1|+|\delta_2|\le1\}$ and $\gamma_C(x)=\inf\{t>0:x\in tC\}$.
A student model matches a teacher's predictions on $[0,1]$, but deployment also depends on sensitivity to input changes. The evaluator must choose a distance that detects both output and slope mismatch.
$f_n(x)=x+\frac{1}{n}\sin(nx)$ and the teacher is $f(x)=x$.