多头注意力
term_id: multi_head_attention
Category: basic_concepts
Definition
多头注意力通过并行运行多次标准注意力机制(使用不同的学习到的线性投影)来扩展标准注意力机制。这使得模型能够联合关注来自不同位置的不同表示子空间的信息。
Summary
Transformer模型中的一种机制,允许模型同时关注来自不同表示子空间的信息。
Key Concepts
- 自注意力
- 线性投影
- 拼接
Use Cases
- 自然语言处理 (NLP)
- 机器翻译
- 使用Vision Transformer进行图像分类
Code Example
import torch.nn as nn
class MultiHeadAttention(nn.Module):
def __init__(self, d_model, num_heads):
super().__init__()
self.num_heads = num_heads
self.d_k = d_model // num_heads
self.W_q = nn.Linear(d_model, d_model)
self.W_k = nn.Linear(d_model, d_model)
self.W_v = nn.Linear(d_model, d_model)
self.W_o = nn.Linear(d_model, d_model)
def forward(self, x):
# Simplified forward pass logic
pass