Inner Products and Orthogonal Matrices
How does attention decide what to "attend" to?
Vishnu Boddeti
Example 1
Dot products become normalized weights that decide which token information reaches the next representation.
Example 2
Orthogonal coordinates turn inner-product geometry into a practical control mechanism: one steering direction should have zero component along protected directions.
Example 3
Cosine similarity ignores vector magnitude and compares semantic direction in the learned space.
Example 4
Semantic closeness is not permission; retrieval must enforce access control before memories enter the model context.
Each scenario hides a mathematical structure from this lecture. Identify the structure, justify it, and work through the resulting problem.
A probe identifies a unit latent direction $b$ associated with a nuisance attribute. Before retrieval, the system should remove exactly the component aligned with $b$ while changing the embedding as little as possible.
$x=(2,1,3)^\top$ and $b=(1,1,0)^\top/\sqrt2$.
A frozen representation permits a prediction vector to lie only in the span of two feature templates. The target cannot be matched exactly, so the head should minimize squared error.
$u_1=(1,0,1)^\top$, $u_2=(0,1,1)^\top$, and target $y=(2,1,0)^\top$.
A privacy layer rotates every database and query embedding using the same matrix. Retrieval uses dot products and Euclidean distances, so the team needs a condition guaranteeing rankings do not change.
The transformed embedding is $x'=Qx$ for a square real matrix $Q$.