p-norms, Sparsity, Whitening
How do sparsity penalties and representation covariance shape modern models?
Vishnu Boddeti
Example 1
Hardware benefits when sparsity removes structured blocks rather than isolated weights that remain expensive to index.
Example 2
Diagonal terms encourage invariance while off-diagonal penalties discourage redundant coordinates.
Example 3
Controlling scale improves optimization while preserving mean information and avoiding the centering operation.
Example 4
Sparse selection reduces latency, prompt size, attack surface, and opportunities for an irrelevant action.
Each scenario hides a mathematical structure from this lecture. Identify the structure, justify it, and work through the resulting problem.
A tied-weight block applies the same transformation 100 times. Direct multiplication is slow and obscures which representation directions survive.
$A=\begin{bmatrix}3&1\\0&2\end{bmatrix}$ and the block output is $A^{100}x$.
Two coordinates of a learned embedding share a strong common mode. Before sending the representation to a retrieval head, the pipeline should rotate the coordinates and give both resulting directions unit variance.
The centered embedding covariance is $\Sigma=\begin{bmatrix}5&4\\4&5\end{bmatrix}$, and one centered embedding is $x=(3,1)^\top$.
A causal feature mixer is upper triangular because token $i$ may depend only on itself and earlier internal features. The team wants a quick stability test and an efficient way to invert the layer.
$T=\begin{bmatrix}0.9&2&-1\\0&0.5&3\\0&0&1.1\end{bmatrix}$.