p-norms, Sparsity, Whitening


How do sparsity penalties and representation covariance shape modern models?

Vishnu Boddeti

Example 1

Structured pruning removes coherent model components

How can structured sparsity remove whole heads, channels, or experts?

group penalty corners → groups off dense shrinkage smooth → all active
  • Current compression and conditional-compute systems often want an entire attention head, channel, or expert block to become inactive.
  • Group-sparsity penalties apply an L2 norm within each block and an L1-style sum across blocks, so the geometry can drive whole groups—not isolated scalars—to exactly zero.

p-norm and quasi-norm balls

solution touches corner p = 2 (circle) p = 1 (diamond)
  • For $p \ge 1$, $\|x\|_p = \left(\sum_i |x_i|^p\right)^{1/p}$ is a p-norm.
  • For $0 < p < 1$, the same formula defines a quasi-norm, because it does not satisfy the triangle inequality.
  • The $p=2$ norm ball is a smooth circle, while the $p=1$ norm ball is a diamond with corners on the coordinate axes.
  • Smaller values such as $p=0.5$ produce a nonconvex quasi-norm ball with even sharper axis points.

Change $p$ and compare the boundary near each axis

Interpret each coordinate as a group norm: axis-aligned corners explain why a structured penalty can deactivate a whole head, channel, or expert.

Example 2

Feature covariance exposes redundant representations

Why do representation-learning systems monitor and regularize feature covariance?

redundant features → decorrelated
  • Embedding dimensions can become redundant or collapse onto a few directions.
  • Current representation-learning objectives therefore monitor off-diagonal covariance
  • and some pipelines decorrelate or whiten embeddings before retrieval or a downstream head.
  • What transformation separates correlation from scale?

Whitening rotates, then rescales

v₁ (λ₁) v₂ (λ₂) A = QΛQᵀ
  • A positive definite covariance matrix $\Sigma$ diagonalizes as $\Sigma = Q\Lambda Q^\top$.
  • Multiplying centered data by $Q^\top$ rotates it into the eigenbasis, so the coordinates are decorrelated but retain variances $\lambda_i$.
  • Full whitening also rescales each coordinate by $1/\sqrt{\lambda_i}$.
  • The complete transform is $z = \Lambda^{-1/2}Q^\top x$, and the covariance of $z$ is the identity matrix.

Step through rotation and rescaling of the data

Covariance regularization discourages redundant directions. Full whitening goes further: rotation removes correlation, then Λ⁻¹ᐟ² gives every direction unit variance.

Example 3

RMSNorm stabilizes activation scale

How does RMSNorm control activation scale without centering features?

  • RMSNorm rescales a token representation using its root-mean-square magnitude.
  • Unlike whitening, it does not rotate features or subtract their mean; it stabilizes scale while preserving the direction of the vector.

Increase input magnitude and compare raw with normalized activation scale

Move the control to compare the two normalized responses.
Scaling every coordinate changes raw magnitude, while normalization keeps the layer input on a controlled scale.

Example 4

Agents should activate only useful tools

How should an agent choose a sparse tool set under a cost budget?

  • A broad tool catalog increases capability but also routing ambiguity, latency, and permissions exposure.
  • A structured sparsity penalty can deactivate whole tool groups while preserving a compact set that covers the task distribution.

Activate more tools and compare task coverage with execution cost

Move the control to compare the two normalized responses.
The useful optimum is sparse by groups: enough tools to cover tasks, few enough to route and authorize reliably.

Practice problems

Each scenario hides a mathematical structure from this lecture. Identify the structure, justify it, and work through the resulting problem.

Problem 1 Deep linear networks

Compute a deep linear block without multiplying 100 matrices

A tied-weight block applies the same transformation 100 times. Direct multiplication is slow and obscures which representation directions survive.

Model evidence

$A=\begin{bmatrix}3&1\\0&2\end{bmatrix}$ and the block output is $A^{100}x$.

  1. Decide whether a change of coordinates can make repeated application scalar-by-scalar.
  2. Construct such coordinates and give a closed expression for $A^{100}x$.
  3. Explain which component dominates numerically and why.
Problem 2 Representation learning

Whiten correlated embedding features before retrieval

Two coordinates of a learned embedding share a strong common mode. Before sending the representation to a retrieval head, the pipeline should rotate the coordinates and give both resulting directions unit variance.

Model evidence

The centered embedding covariance is $\Sigma=\begin{bmatrix}5&4\\4&5\end{bmatrix}$, and one centered embedding is $x=(3,1)^\top$.

  1. Find an orthonormal eigenbasis of $\Sigma$ and its eigenvalues.
  2. Construct the whitening map $W=\Lambda^{-1/2}Q^\top$ and verify that $W\Sigma W^\top=I$.
  3. Compute the whitened embedding $z=Wx$ and identify which coordinate measures shared variation and which measures contrast variation.
Problem 3 Autoregressive models

Debug a causal mixer from its diagonal

A causal feature mixer is upper triangular because token $i$ may depend only on itself and earlier internal features. The team wants a quick stability test and an efficient way to invert the layer.

Model evidence

$T=\begin{bmatrix}0.9&2&-1\\0&0.5&3\\0&0&1.1\end{bmatrix}$.

  1. Read off the spectral values governing repeated application without computing a characteristic determinant.
  2. Decide whether every hidden state decays under repeated application.
  3. Explain how the triangular structure makes solving $Tx=y$ sequential.