Articles

Self-Attention Explained: How Transformers Read Context

Self-attention lets each token in a sequence weigh every other token when building its representation, which is how transformers understand context.

Chisato Chisato · · 4 min read
Abstract illustration of interconnected neural pathways

Self-attention is the mechanism that lets a model weigh every token in a sequence against every other token, producing a representation of each word that’s informed by its full context rather than just its own meaning in isolation. It’s the core operation inside a transformer, and it’s the reason transformers replaced recurrent architectures for language tasks — earlier models processed a sequence one token at a time, while self-attention lets every token look at every other token directly, all at once.

The problem it solves

Consider the sentence “the trophy didn’t fit in the suitcase because it was too big.” What does “it” refer to — the trophy or the suitcase? A human resolves this instantly from context, but a model that processes words one at a time, carrying forward only a compressed summary of everything it’s seen so far, has to guess based on how well that summary preserved the relevant detail. Earlier recurrent architectures suffered from exactly this: information from early in a long sequence tends to get diluted by the time the model reaches a later token that needs it, since it has only a fixed-size running summary to carry that information forward.

Self-attention sidesteps the bottleneck by giving every token a direct path to every other token, regardless of distance. When the model builds a representation for “it,” it can directly weigh how relevant “trophy” and “suitcase” each are, rather than relying on that information having survived a long chain of intermediate steps.

Queries, keys, and values

Mechanically, self-attention represents each token with three learned vectors, all derived from its embedding via separate weight matrices:

  • Query (Q) — what this token is “looking for” in the rest of the sequence.
  • Key (K) — what this token “advertises” about itself, for other tokens to match against.
  • Value (V) — the actual content this token contributes if another token attends to it.

For each token, its query is compared against every other token’s key (via a dot product), producing a relevance score for each pair. Those scores are scaled and passed through a softmax, turning them into a set of weights that sum to 1 — how much attention this token should pay to each other token. The token’s new representation is then a weighted sum of every token’s value vector, using those weights.

In the trophy example, when the model processes “it,” its query vector ends up matching the key vector for “trophy” more strongly than “suitcase” (learned from patterns during training), so “it” absorbs more of the trophy’s value vector into its own representation — effectively resolving the reference.

Why “self”-attention

The mechanism is called self-attention because the queries, keys, and values all come from the same sequence — a token attends to other tokens within the same input. This is distinct from the cross-attention used in encoder-decoder architectures, where the queries come from one sequence (say, a partially generated translation) and the keys and values come from another (the source text being translated). Self-attention is what lets a model understand a single passage internally; cross-attention is what lets it relate two different sequences to each other.

Multi-head attention

A single attention computation only learns one kind of relationship between tokens. In practice, transformers run several attention computations in parallel — “heads” — each with its own learned Q, K, and V projections, so each head can specialize in a different kind of relationship: one head might track subject-verb agreement, another might track which adjective modifies which noun, another might track longer-range topical relevance. The outputs of all the heads are concatenated and combined, giving the model several independent “views” of the same sequence rather than just one.

The cost: quadratic scaling

Self-attention’s biggest practical drawback is that every token compares itself against every other token, so the computation grows quadratically with sequence length — doubling the input length roughly quadruples the attention computation. This is a major reason context windows have historically been limited, and why techniques like mixture-of-experts and various sparse or approximate attention variants exist: they aim to keep most of self-attention’s benefit while avoiding the full quadratic cost at very long sequence lengths.

Self-attention vs recurrence

Self-attention (transformers)Recurrence (RNN/LSTM)
Access to earlier tokensDirect, regardless of distanceIndirect, through a carried-forward state
Parallelizable across a sequenceYes — all positions computed at onceNo — must process tokens in order
Cost as sequence growsQuadratic in sequence lengthLinear, but slow and sequential
Long-range dependenciesHandled directlyProne to degrading over long distances

The parallelizability turned out to matter as much as the modeling improvement: because every token’s attention computation is independent of the others, training can process an entire sequence at once on modern hardware, rather than being forced through a slow, token-by-token loop.

The takeaway

Self-attention lets every token in a sequence directly weigh every other token when building its own representation, using learned query, key, and value vectors and a softmax-weighted combination of values. That direct access to full context — computed in parallel across the whole sequence rather than carried forward step by step — is what makes transformers both better at long-range context and dramatically faster to train than the recurrent architectures they replaced, at the cost of a computation that grows quadratically with how long the input gets.

Chisato Chisato · · 4 min read

Precision vs Recall, Explained

Precision measures how many of a model's positive predictions were correct; recall measures how many actual positives it found. Why you can't max both.

#AI #Machine Learning #LLMs