What Is Attention? Transformer Models

What Is Attention? Transformer Models

What Is Attention? Transformer Models 150 150 Creangel Portal.

Introduction to the Attention Mechanism in Transformers

What Is Attention?

 

Attention is like when you read a sentence and your brain decides which words matter most for understanding its meaning.

Introduction to the attention mechanism in Transformers

Introduction to the attention mechanism in Transformers

Based on.

 

 

Scaled Dot-Product Attention

It is a mathematical formula that works like this:

  1. Compare: It takes one word (the query) and compares it with all the other words (the keys)
  2. Compute weights: It gives more importance to the most closely related words
  3. Combine: It blends all the information according to those weights

Formula: Attention(Q,K,V) = softmax(QK^T/√dk)V

Multi-Head Attention

Instead of using a single attention “head,” it uses several (8 in this case):

  • Each head focuses on different aspects of the relationships between words
  • It is like having 8 different perspectives on the same text
  • At the end, all the perspectives are combined

Three Uses in the Transformer:

  1. Encoder-Decoder: The decoder looks at the entire input
  2. Encoder self-attention: Each word looks at all the others in its layer
  3. Decoder self-attention: Each word looks only at the previous ones (it cannot “cheat” by seeing the future)

In short: It is a mechanism that helps the model understand which parts of the text are most relevant to each word.

 

In the formula Attention(Q,K,V) = softmax(QK^T/√dk)V, the T stands for transpose.

What Is the Transpose?

The transpose of a matrix means flipping the matrix by swapping rows and columns.

Example:

Original matrix K:    Transpose K^T:
[1  2  3]            [1  4]
[4  5  6]            [2  5]
                     [3  6]

Why Is K^T Used?

To be able to perform matrix multiplication between Q and K:

  • Q has dimensions: (number of queries × dk)
  • K has dimensions: (number of keys × dk)
  • K^T has dimensions: (dk × number of keys)

Therefore:

  • Q × K^T = (number of queries × dk) × (dk × number of keys)
  • Result = (number of queries × number of keys)

This gives us a matrix in which each element represents how compatible each query is with each key.

In Python code:

python
scores = torch.matmul(query, key.transpose(-2, -1))

The .transpose(-2, -1) is equivalent to applying the T (transpose).

 

dk is the dimension of the keys and queries.

Simple Definition

dk = the number of elements (features) in each key and query vector.

In the Context of the Transformer

According to the paper:

  • dk = dv = dmodel/h = 64

Where:

  • dmodel = total model dimension (for example, 512)
  • h = number of attention heads (8 in this case)
  • dk = 512 ÷ 8 = 64

Why Is It Important?

It is used in two key places:

  1. In the division by √dk:
   scores = QK^T / √dk
  • If dk = 64, then √dk = 8
  • This normalizes the scores to keep them from becoming too large
  1. In the matrix dimensions:
    • Each query vector (Q) has 64 elements
    • Each key vector (K) has 64 elements
    • Each value vector (V) also has 64 elements

Why Divide by √dk?

As the paper explains: when dk is large, the QK^T dot products become very large, causing the softmax function to have extremely small gradients. Dividing by √dk stabilizes training.

Example: If dk = 64, we divide by √64 = 8 to keep the values within a reasonable range.