Inner Products and Orthogonal Matrices


How does attention decide what to "attend" to?

Vishnu Boddeti

Example 1

Attention performs content-based information routing

How does attention decide what to "attend" to?

θ query q key k alignment score compare their directions
  • A transformer computes a score between a query vector and a key vector for every token pair, then uses those scores to decide how much each token should "attend" to every other token.
  • What operation turns two vectors into a single number that captures how well they align?

Attention score as an inner product

q k ||k||cosθ ⟨q,k⟩ = ||q|| ||k|| cosθ
  • The attention score between a query $q$ and a key $k$ is $\propto \langle q, k \rangle = q^\top k = \sum_i q_i k_i$.
  • Vectors pointing in the same direction (parallel, small angle) score high; vectors at right angles (orthogonal) score zero; vectors pointing opposite ways score negative.
  • The dot product is exactly $\|q\|\|k\|\cos\theta$, so it measures alignment scaled by magnitude.

Drag $q$ and $k$ to compare alignment with attention score

query $q$ key $k$
Parallel vectors have a positive score, orthogonal vectors score zero, and opposite vectors have a negative score.

Example 2

LatentCompass isolates attributes in orthogonal spaces

How can steering one generative attribute avoid changing another?

steer protect orthogonal attributes generic W leakage coupled attributes
  • LatentCompass constructs low-dimensional concept coordinates from exemplars.
  • Orthogonal directions separate intended steering from protected attributes.
  • A generic matrix can stretch or mix those directions and reintroduce leakage.

Orthogonal matrices preserve norms and angles

unit circle W x Wx same length, rotated W⁠ᵀW = I
  • A matrix $W$ is orthogonal iff $W^\top W = I$, equivalently its columns are unit vectors that are mutually perpendicular.
  • Orthogonal maps preserve inner products, $\langle Wx, Wy\rangle = \langle x,y\rangle$, so they preserve norms ($\|Wx\| = \|x\|$) and angles
  • no vector gets stretched, shrunk, or sheared.
  • Geometrically, the unit circle maps to itself exactly when $W$ is orthogonal; otherwise it maps to an ellipse.

Change $W$ and test whether concept geometry is preserved

An orthogonal steering map preserves lengths and angles between concept directions. A generic matrix can shear the space and couple attributes that were separated.

Example 3

Dense retrieval turns meaning into nearest-neighbor search

Why does cosine similarity retrieve semantically related embeddings?

  • Retrieval systems commonly normalize query and item embeddings, turning the inner product into cosine similarity.
  • Direction then carries semantic agreement while vector magnitude no longer dominates ranking.

Rotate two unit embeddings and watch similarity fall with angle

Move the control to compare the two normalized responses.
After normalization, alignment—not raw magnitude—determines which item receives the larger similarity score.

Example 4

Agent memory retrieval needs authorization as well as relevance

How should an agent retrieve relevant memory without exposing unrelated records?

  • A memory retriever ranks normalized embeddings by inner product, then applies metadata and permission filters.
  • Raising the similarity threshold improves precision but can hide useful context; isolation filters must remain hard constraints rather than soft similarity preferences.

Raise the retrieval threshold and compare precision with memory recall

Move the control to compare the two normalized responses.
Semantic similarity decides relevance only after authorization determines which records are eligible.

Practice problems

Each scenario hides a mathematical structure from this lecture. Identify the structure, justify it, and work through the resulting problem.

Problem 1 Representation debiasing

Remove a nuisance direction from an embedding

A probe identifies a unit latent direction $b$ associated with a nuisance attribute. Before retrieval, the system should remove exactly the component aligned with $b$ while changing the embedding as little as possible.

Model evidence

$x=(2,1,3)^\top$ and $b=(1,1,0)^\top/\sqrt2$.

  1. Compute the nuisance component and the cleaned embedding.
  2. Prove that the cleaned embedding contains no component along $b$.
  3. Explain why this update is the minimum-Euclidean-change vector satisfying that constraint.
Problem 2 Linear probing

Fit a linear head using only a subspace

A frozen representation permits a prediction vector to lie only in the span of two feature templates. The target cannot be matched exactly, so the head should minimize squared error.

Model evidence

$u_1=(1,0,1)^\top$, $u_2=(0,1,1)^\top$, and target $y=(2,1,0)^\top$.

  1. Find the coefficients producing the closest allowable prediction.
  2. Verify that the residual is orthogonal to every allowable prediction direction.
  3. Interpret the orthogonality condition as the optimality certificate for least squares.
Problem 3 Embedding alignment

Does a latent rotation preserve retrieval scores?

A privacy layer rotates every database and query embedding using the same matrix. Retrieval uses dot products and Euclidean distances, so the team needs a condition guaranteeing rankings do not change.

Model evidence

The transformed embedding is $x'=Qx$ for a square real matrix $Q$.

  1. Derive the matrix condition under which every pairwise dot product is preserved.
  2. Prove that the same condition preserves all Euclidean distances.
  3. Test $Q=\frac1{\sqrt2}\begin{bmatrix}1&-1\\1&1\end{bmatrix}$.