Convexity and Training Difficulty


Why does the shape of a loss decide whether training is easy?

Vishnu Boddeti

Example 1

Linear probes test what a frozen representation exposes

Why is fitting a linear probe predictable while end-to-end transformer tuning is not?

linear probe end-to-end tuning same rule, different destinations
  • A convex linear or logistic probe on frozen activations has no competing local minima.
  • End-to-end transformer tuning is non-convex and depends on initialization, optimizer state, and schedule.
  • What property of the loss landscape explains the difference?

Convex losses have no competing local minima

convex chord stays above non-convex curve rises above chord
  • For a convex $f$, every chord lies on or above the graph: $f(\lambda x+(1-\lambda)y)\le\lambda f(x)+(1-\lambda)f(y)$.
  • The upward-facing bowl has no competing local minima: every local minimum is global.
  • A non-convex loss can have several basins, so different starting points may converge to different minima.

Choose two starting points and compare where they converge

Convex — $f(x) = a x^2 + b x$
Non-convex — $g(x) = a x^4 - b x^2$
Click once on each landscape. The dotted trail shows every gradient-descent step.

Example 2

An adversarial threat model defines allowed changes

How does an adversarial threat set define the size of a perturbation?

0 δ boundary of K K smallest t with δ ∈ tK = γK(δ)
  • The threat set $K$ contains perturbations allowed at unit budget.
  • Scale $K$ until it first reaches $\delta$; that scale is $\gamma_K(\delta)=\inf\{t>0:\delta\in tK\}$.
  • If $K$ is convex, symmetric, bounded, and contains $0$ in its interior, this gauge is a norm.

A symmetric convex set defines a gauge

0 θ boundary point x outside K gauge = how far to shrink x into K
  • Let $K$ be convex, bounded, symmetric about the origin ($K=-K$), and contain $0$ in its interior.
  • Along the ray at angle $\theta$, compare $|x|$ with the boundary radius $R_K(\theta)$: $\|x\|_K=|x|/R_K(\theta)$.
  • Convexity gives the triangle inequality; symmetry gives $\|-x\|_K=\|x\|_K$.
  • These properties make $\|\cdot\|_K$ a norm whose unit ball is exactly $K$.

Reshape $K$, then measure a point with its gauge

Interpreting $K$ as a threat set: gauge below 1 means the perturbation is allowed, 1 is the budget boundary, and above 1 is outside the threat model.

Example 3

Model merging interpolates learned solutions

When can interpolating two model checkpoints produce a useful merged model?

  • Model merging often starts with a line segment between parameter vectors.
  • Endpoints can each perform well, yet the interpolation may cross a loss barrier unless the checkpoints occupy a compatible basin or are first aligned.

Move along a checkpoint interpolation and expose a possible loss barrier

Move the control to compare the two normalized responses.
Parameter averaging is easy; preserving capability along the path is the nonconvex part that must be measured.

Example 4

Safe policy mixtures need geometric guardrails

When can an agent blend tool policies without leaving a safe operating region?

  • If each candidate routing policy satisfies convex resource and risk constraints, a probability mixture remains feasible.
  • Nonconvex requirements—such as mutually exclusive permissions or discrete workflow invariants—do not inherit that guarantee.

Increase policy blending and compare capability coverage with boundary risk

Move the control to compare the two normalized responses.
Convex combinations preserve convex constraints, but agent workflows often contain discrete rules that must be checked separately.

Practice problems

Each scenario hides a mathematical structure from this lecture. Identify the structure, justify it, and work through the resulting problem.

Problem 1 Certified robustness

Translate a certificate between two threat models

A classifier is certified against Euclidean perturbations, but the deployment contract specifies a maximum per-feature change. The certificate must be translated without rerunning verification.

Model evidence

The input has dimension $d=64$, and the classifier is known to be safe whenever $\|\delta\|_2<0.24$.

  1. Find a sufficient $\ell_\infty$ radius that guarantees the existing certificate applies.
  2. Explain why a dimension-dependent constant is unavoidable.
Problem 2 Distributional robustness

Does an uncertainty set define a valid notion of size?

A robust optimizer declares perturbations acceptable when $2|\delta_1|+|\delta_2|\le1$. Engineers want to report the smallest scaling needed to bring any perturbation into this safe set as its “risk magnitude.”

Model evidence

$C=\{\delta\in\mathbb{R}^2:2|\delta_1|+|\delta_2|\le1\}$ and $\gamma_C(x)=\inf\{t>0:x\in tC\}$.

  1. Compute $\gamma_C(x)$ explicitly.
  2. Verify the properties needed for this quantity to be a norm.
  3. Interpret the shape of its unit ball as unequal feature costs.
Problem 3 Function approximation

When are two learned functions close enough?

A student model matches a teacher's predictions on $[0,1]$, but deployment also depends on sensitivity to input changes. The evaluator must choose a distance that detects both output and slope mismatch.

Model evidence

$f_n(x)=x+\frac{1}{n}\sin(nx)$ and the teacher is $f(x)=x$.

  1. Determine whether the predictions converge uniformly to the teacher.
  2. Determine whether the derivatives converge uniformly.
  3. Compare the conclusions under the sup norm on functions and the norm $\|g\|=\|g\|_\infty+\|g'\|_\infty$.