Rank, Determinant, LoRA Intuition
How much information does a matrix actually carry?
Vishnu Boddeti
Example 1
Low rank reduces trainable parameters and memory while retaining useful adaptation capacity.
Example 2
Exact likelihood requires accounting for local expansion and contraction; flow matching can train the velocity field without that full computation.
Example 3
The rotation has determinant one and inverse equal to its transpose, so it preserves norms and area while relative position changes attention.
Example 4
A rank-deficient evaluator cannot observe every independent failure direction. Collision-pair tests expose blind spots hidden by aggregate success rates.
Each scenario hides a mathematical structure from this lecture. Identify the structure, justify your choice, and then solve the resulting problem.
A frozen projection sends an activation vector to a smaller feature vector. During adapter training, the optimizer receives the gradient with respect to projected features and needs the gradient with respect to the original activation.
$z=Wx$, with $W=\begin{bmatrix}1&2&0\\-1&0&3\end{bmatrix}$, and the downstream gradient is $g_z=(4,-2)^\top$.
Two teams use different bases for the same two-dimensional latent subspace. A checkpoint stores coordinates in Team A's basis, while Team B's decoder expects its own coordinates.
In ambient coordinates, $P=\begin{bmatrix}1&1\\0&1\end{bmatrix}$ contains Team A's basis vectors and $Q=\begin{bmatrix}1&0\\1&1\end{bmatrix}$ contains Team B's. The checkpoint coordinate is $c_A=(2,-1)^\top$.
A fusion layer mixes three learned features. A hyperparameter controls how strongly the third output combines the first two, and the layer must not collapse a latent direction.
$A_\alpha=\begin{bmatrix}1&0&1\\0&1&1\\1&1&\alpha\end{bmatrix}$.