Singular Value Decomposition and the Pseudo-Inverse


How does one factorization explain image compression, low-rank structure, and least squares?

Vishnu Boddeti

Example 1

SVD compresses images and transformer weight matrices

How does SVD explain compressing an image — or a transformer weight matrix?

a huge matrix A ≈ few dominant directions σ₁ σ₂ σ₃ … discarded
  • A matrix can be huge, but most of the "signal" often lives in a handful of directions.
  • Is there a way to peel off exactly those directions, in order of importance, and throw the rest away?

The singular value decomposition

A = U × Σ × Vᵀ keep top k terms → best rank-k fit
  • Every matrix factors as $A=U\Sigma V^\top$.
  • $U$ and $V$ are orthogonal; the diagonal entries $\sigma_1\ge\sigma_2\ge\cdots\ge0$ are singular values.
  • Keeping the first $k$ terms gives the best rank-$k$ approximation $A_k=\sum_{i=1}^k\sigma_i u_i v_i^\top$ in Frobenius norm.

Adjust the rank and compare reconstruction quality

original $16\times16$ matrix $A$
Watch how the largest singular values recover the main structure before the smaller details.

Example 2

A probe decodes features without changing the model

How does a linear probe decode a feature from frozen model activations?

? activation versus attribute — which linear decoder fits best?
  • A linear probe fits a small decoder on many frozen activation examples.
  • The resulting system is overdetermined and noisy: no decoder matches every label exactly.
  • What does "best" mean, and how does the pseudo-inverse compute it?

Least squares via the pseudo-inverse

col(A) 0 Ax* b residual b − Ax*
  • An overdetermined system $Ax=b$ usually has no exact solution.
  • $Ax^\star$ is the orthogonal projection of $b$ onto $\operatorname{col}(A)$, so the residual is perpendicular to the column space.
  • The pseudo-inverse gives the minimizer directly: $x^\star=A^+b$.
  • For a line $y=mx+c$, the two parameters can also be found from the $2\times2$ normal equations $A^\top A x=A^\top b$.

Drag an activation example and inspect the linear probe

activation / attribute pairs (draggable) least-squares probe
The probe minimizes squared prediction residuals. A good fit shows that a feature is linearly decodable; it does not by itself prove the model uses that feature causally.

Example 3

Spectral directions organize weight sensitivity

How does truncated SVD reveal which weight directions are most sensitive?

  • A truncated SVD separates high-energy operator directions from a low-energy tail.
  • Compression or quantization can allocate more precision to dominant singular directions and tolerate larger error where the operator has less energy.

Retain more singular modes and compare signal recovery with sensitivity

Move the control to compare the two normalized responses.
Spectral structure provides a principled way to decide where approximation error is most likely to matter.

Example 4

Redundant tools create an underdetermined allocation problem

How can a pseudoinverse allocate work across redundant tools?

  • An agent may have several tools whose effects overlap.
  • A linearized planner can represent desired outcome changes as Au≈b; the pseudoinverse returns the minimum-norm tool-intensity vector when many allocations fit equally well.

Add usable tool directions and compare task fit with coordination ambiguity

Move the control to compare the two normalized responses.
The pseudoinverse resolves redundancy by choosing the smallest joint action among equally good linearized plans.

Practice problems

Each scenario hides a mathematical structure from this lecture. Identify the structure, justify it, and work through the resulting problem.

Problem 1 Model compression

Compress an embedding table under a rank budget

An embedding table is too large to ship to an edge device, and there is room for only one latent direction. Which rank-one table keeps as much of the original structure as possible?

Model evidence

$M=\begin{bmatrix}3&0\\0&2\\0&0\end{bmatrix}$.

  1. What is the optimal rank-one approximation?
  2. What Frobenius reconstruction error does it incur?
  3. Which semantic direction is discarded, and why is no other rank-one matrix better?
Problem 2 Linear probing

Fit a probe when the equations disagree

Three labeled examples disagree slightly about the parameters of a two-number probe. There is no exact solution, so the probe must make the smallest collective compromise without inventing unnecessary parameter magnitude.

Model evidence

$X=\begin{bmatrix}1&0\\0&1\\1&1\end{bmatrix}$ and $y=(1,2,2)^\top$.

  1. What minimum-norm least-squares parameter does the pseudoinverse produce?
  2. What are the resulting predictions and residuals?
  3. Which orthogonality condition confirms that the fit is optimal?
Problem 3 Robustness auditing

Separate worst-case amplification from average matrix size

Two researchers say a layer has a large “norm,” but they mean different things. One worries about the single input that is amplified most; the other is summarizing the energy of every matrix entry. How different are their answers here?

Model evidence

$W=\begin{bmatrix}2&0\\0&1/2\end{bmatrix}$.

  1. What quantity controls $\max_{\|x\|_2=1}\|Wx\|_2$?
  2. What is the total squared-entry measure, and how does it compare with the worst-case amplification?
  3. Which input direction realizes the worst case?