Singular Value Decomposition and the Pseudo-Inverse
How does one factorization explain image compression, low-rank structure, and least squares?
Vishnu Boddeti
Example 1
Truncated SVD gives an optimal matrix approximation, but downstream model quality also depends on activation statistics, truncation error, and layer sensitivity.
Example 2
Probe accuracy indicates that information is linearly accessible, not necessarily that the model causally uses it.
Example 3
Truncated SVD separates dominant transformations from directions that can often be approximated or removed.
Example 4
Minimum-norm solutions avoid arbitrary extreme assignments when many tool mixtures achieve the same result.
Each scenario hides a mathematical structure from this lecture. Identify the structure, justify it, and work through the resulting problem.
An embedding table is too large to ship to an edge device, and there is room for only one latent direction. Which rank-one table keeps as much of the original structure as possible?
$M=\begin{bmatrix}3&0\\0&2\\0&0\end{bmatrix}$.
Three labeled examples disagree slightly about the parameters of a two-number probe. There is no exact solution, so the probe must make the smallest collective compromise without inventing unnecessary parameter magnitude.
$X=\begin{bmatrix}1&0\\0&1\\1&1\end{bmatrix}$ and $y=(1,2,2)^\top$.
Two researchers say a layer has a large “norm,” but they mean different things. One worries about the single input that is amplified most; the other is summarizing the energy of every matrix entry. How different are their answers here?
$W=\begin{bmatrix}2&0\\0&1/2\end{bmatrix}$.