Mastering Transformers

[video player]
0:00
Welcome back to the course on transformers.
0:15
The attention mechanism was introduced in 2017 by Vaswani and colleagues.
1:02
Multi-head attention lets the model attend to different representation subspaces at once.
1:12:34
And that concludes our deep dive into positional encoding.