Self-Attention Explained: How Transformers Read Context
Self-attention lets each token in a sequence weigh every other token when building its representation, which is how transformers understand context.
Self-attention is the mechanism that lets a model weigh every token in a sequence against every other token, producing a representation of each word that’s informed by its full context rather than just its own meaning in isolation. It’s the core operation inside a transformer, and it’s the reason transformers replaced recurrent architectures for language tasks — earlier models processed a sequence one token at a time, while self-attention lets every token look at every other token directly, all at once.
The problem it solves
Consider the sentence “the trophy didn’t fit in the suitcase because it was too big.” What does “it” refer to — the trophy or the suitcase? A human resolves this instantly from context, but a model that processes words one at a time, carrying forward only a compressed summary of everything it’s seen so far, has to guess based on how well that summary preserved the relevant detail. Earlier recurrent architectures suffered from exactly this: information from early in a long sequence tends to get diluted by the time the model reaches a later token that needs it, since it has only a fixed-size running summary to carry that information forward.
Self-attention sidesteps the bottleneck by giving every token a direct path to every other token, regardless of distance. When the model builds a representation for “it,” it can directly weigh how relevant “trophy” and “suitcase” each are, rather than relying on that information having survived a long chain of intermediate steps.
Queries, keys, and values
Mechanically, self-attention represents each token with three learned vectors, all derived from its embedding via separate weight matrices:
- Query (Q) — what this token is “looking for” in the rest of the sequence.
- Key (K) — what this token “advertises” about itself, for other tokens to match against.
- Value (V) — the actual content this token contributes if another token attends to it.
For each token, its query is compared against every other token’s key (via a dot product), producing a relevance score for each pair. Those scores are scaled and passed through a softmax, turning them into a set of weights that sum to 1 — how much attention this token should pay to each other token. The token’s new representation is then a weighted sum of every token’s value vector, using those weights.
In the trophy example, when the model processes “it,” its query vector ends up matching the key vector for “trophy” more strongly than “suitcase” (learned from patterns during training), so “it” absorbs more of the trophy’s value vector into its own representation — effectively resolving the reference.
Why “self”-attention
The mechanism is called self-attention because the queries, keys, and values all come from the same sequence — a token attends to other tokens within the same input. This is distinct from the cross-attention used in encoder-decoder architectures, where the queries come from one sequence (say, a partially generated translation) and the keys and values come from another (the source text being translated). Self-attention is what lets a model understand a single passage internally; cross-attention is what lets it relate two different sequences to each other.
Multi-head attention
A single attention computation only learns one kind of relationship between tokens. In practice, transformers run several attention computations in parallel — “heads” — each with its own learned Q, K, and V projections, so each head can specialize in a different kind of relationship: one head might track subject-verb agreement, another might track which adjective modifies which noun, another might track longer-range topical relevance. The outputs of all the heads are concatenated and combined, giving the model several independent “views” of the same sequence rather than just one.
The cost: quadratic scaling
Self-attention’s biggest practical drawback is that every token compares itself against every other token, so the computation grows quadratically with sequence length — doubling the input length roughly quadruples the attention computation. This is a major reason context windows have historically been limited, and why techniques like mixture-of-experts and various sparse or approximate attention variants exist: they aim to keep most of self-attention’s benefit while avoiding the full quadratic cost at very long sequence lengths.
Self-attention vs recurrence
| Self-attention (transformers) | Recurrence (RNN/LSTM) | |
|---|---|---|
| Access to earlier tokens | Direct, regardless of distance | Indirect, through a carried-forward state |
| Parallelizable across a sequence | Yes — all positions computed at once | No — must process tokens in order |
| Cost as sequence grows | Quadratic in sequence length | Linear, but slow and sequential |
| Long-range dependencies | Handled directly | Prone to degrading over long distances |
The parallelizability turned out to matter as much as the modeling improvement: because every token’s attention computation is independent of the others, training can process an entire sequence at once on modern hardware, rather than being forced through a slow, token-by-token loop.
The takeaway
Self-attention lets every token in a sequence directly weigh every other token when building its own representation, using learned query, key, and value vectors and a softmax-weighted combination of values. That direct access to full context — computed in parallel across the whole sequence rather than carried forward step by step — is what makes transformers both better at long-range context and dramatically faster to train than the recurrent architectures they replaced, at the cost of a computation that grows quadratically with how long the input gets.
Tagged
Keep reading
Chisato · · 5 min read Constrained Decoding: How LLMs Output Guaranteed JSON
Constrained decoding masks invalid tokens at each step so an LLM can only emit output matching a grammar, regex, or JSON Schema. How it works and its limits.
Chisato · · 5 min read LLM Logprobs Explained: Token Probabilities in Practice
Logprobs are the log probabilities an LLM assigns to each token it generates. What they mean, how to read them, and practical uses like classification.
Chisato · · 4 min read Precision vs Recall, Explained
Precision measures how many of a model's positive predictions were correct; recall measures how many actual positives it found. Why you can't max both.