Symmetric Matrices and Positive Definiteness


Why do some optimizers care about curvature, not just slope?

Vishnu Boddeti

Example 1

Curvature changes what counts as a safe optimization step

Why do some optimizers care about curvature, not just slope?

same gradient: ∇f = 0 minimum curves upward maximum curves downward saddle up in one direction down in another
  • A zero gradient identifies a stationary point, but not whether it is a minimum, maximum, or saddle.
  • The missing information is how the objective curves in every direction.

The Hessian and positive definiteness

symmetric Hessian H = a b b d H = Hᵀ positive definite: all λ > 0 bowl indefinite: mixed signs saddle
  • The Hessian $H$ collects second partial derivatives and is symmetric: $H = H^\top$.
  • Near a critical point, $f(x) \approx f(0) + \tfrac12 x^\top Hx$.
  • A strict local minimum requires $x^\top Hx > 0$ for every $x \neq 0$—equivalently, every eigenvalue is positive.
  • Mixed eigenvalue signs produce a saddle.

Change $H$ and classify the critical point

A strict local minimum needs positive curvature in every direction, so both eigenvalues must be positive.

Example 2

Power iteration isolates the strongest curvature direction

What does power iteration have to do with the top Hessian eigenvector?

each step multiplies by H, then normalizes dominant eigendirection v₀ v₁ v₂ v₃ direction converges; normalization controls length
  • Repeated multiplication amplifies the component associated with the eigenvalue of largest magnitude.
  • Normalizing after every multiplication keeps only the evolving direction.

Power iteration targets the largest eigenvalue magnitude

current: shallow direction vk multiply by H after H: rotated + stretched Hvk normalize after normalize: same angle vk+1 = Hvk / ‖Hvk‖
  • The target is the eigenvector whose eigenvalue has the largest magnitude, not necessarily the largest value.
  • Convergence requires a unique largest magnitude and a nonzero starting component in that eigendirection.
  • A magnitude tie has no unique target; a missing component cannot be recovered.

Apply power iteration and track alignment with the target

Repeated multiplication amplifies the component along the unique dominant-magnitude eigendirection.

Example 3

Covariance and PSD curvature encode nonnegative energy

Why must covariance and local-minimum curvature be positive semidefinite?

1direction v → 2energy vᵀAv → 3variance orcurvature vᵀAv ≥ 0 for every direction v
  • For a covariance matrix, every projected variance must be nonnegative.
  • At a local minimum, positive-semidefinite curvature means no direction bends downward.

Increase movement along a direction and compare energy with the stability margin

Move the control to compare the two normalized responses.
Large curvature raises quadratic energy quickly and reduces the displacement that remains locally stable.

Example 4

Trust regions bound risky agent-policy changes

How can a positive-definite metric limit a risky agent-policy update?

1candidate updateΔθ → 2quadratic costΔθᵀFΔθ → 3accept only ifcost ≤ ε the budget forms an ellipsoid in update space
  • A trust region measures policy change using a local quadratic metric rather than raw parameter distance.
  • Positive definiteness gives every nonzero update a positive cost and turns the allowed set into an ellipsoid.

Tighten the trust region and compare protection with update freedom

Move the control to compare the two normalized responses.
The metric penalizes directions according to behavioral sensitivity, not merely the number of changed parameters.

Practice problems

Each scenario hides a mathematical structure from this lecture. Identify the structure, justify it, and work through the resulting problem.

Problem 1 Principal component analysis

Choose the most informative one-dimensional embedding

A two-feature representation must be compressed to one scalar before transmission. The retained direction should capture as much sample variation as possible.

Model evidence

The centered embedding covariance is $S=\begin{bmatrix}4&2\\2&1\end{bmatrix}$. A unit direction $u$ retains variance $u^\top Su$.

  1. Find the direction that maximizes retained variance without searching over angles.
  2. Compute the retained and discarded variances.
  3. Explain why symmetry of the covariance matrix is essential to the argument.
Problem 2 Kernel methods

Is a proposed similarity table a valid kernel?

A custom similarity function produces a Gram matrix for two training examples. A kernel classifier requires every quadratic energy $c^\top Kc$ to be nonnegative.

Model evidence

$K_\rho=\begin{bmatrix}1&\rho\\\rho&1\end{bmatrix}$.

  1. Find every value of $\rho$ for which this table can be used as a positive-semidefinite Gram matrix.
  2. At the boundary values, identify the collapsed coefficient direction.
  3. Interpret $|\rho|>1$ as an impossible similarity geometry.
Problem 3 Model diagnostics

Find the most amplified failure mode

A symmetric sensitivity matrix scores how strongly a unit perturbation excites a model failure mode. The red-team budget permits any unit-norm direction.

Model evidence

$H=\begin{bmatrix}2&1\\1&2\end{bmatrix}$ and the failure score is $r(v)=v^\top Hv$ for $\|v\|_2=1$.

  1. Find the perturbation direction with the largest score and compute that score.
  2. Find the safest unit direction and its score.
  3. Prove that no other unit direction can exceed the reported maximum.