Deep Learning & Natural Language Processing · NeurIPS 2017
Attention Is All You Need (Transformer Architecture)
Authors: Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, Illia Polosukhin (Google Brain & Google Research) · arXiv: 1706.03762
Core Methodological Innovation
Replaced sequential recurrence (RNN/LSTM) with parallelizable Multi-Head Scaled Dot-Product Attention, reducing maximum path length between any two token positions to O(1) while enabling massive GPU training throughput.
Key Quantitative & Theoretical Takeaway: Scaling the dot product by 1/sqrt(d_k) prevents large inner products in high dimensions from pushing the softmax function into regions with vanishingly small gradients.
Abstract
The dominant sequence transduction models are based on complex recurrent or convolutional neural networks that include an encoder and a decoder. We propose a new simple network architecture, the Transformer, based solely on attention mechanisms, dispensing with recurrence and convolutions entirely.
Step-by-Step Equation & Methodology Breakdown
Why does Scaled Dot-Product Attention divide by sqrt(d_k)?
If components of q and k are independent random variables with mean 0 and variance 1, their dot product has mean 0 and variance d_k. Dividing by sqrt(d_k) restores unit variance so softmax does not saturate into vanishing gradients.
What is the computational complexity per layer of Self-Attention vs Recurrent layers?
Self-Attention has O(n^2 * d) complexity per layer with O(1) sequential operations and O(1) maximum path length, whereas Recurrent layers require O(n * d^2) complexity with O(n) sequential operations.
Read Any arXiv Paper Side-by-Side in Nujan
Replace arxiv.org with nujan.app on any arXiv URL (for example, https://nujan.app/abs/2006.11239) to open the PDF and AI research partner side-by-side.
Launch Interactive Split Reader for ATTENTION