Kyle Harrison
research-paper

Attention Is All You Need

Ashish Vaswani et al. 2017 View original ↗

TL;DR — The 2017 Google paper that introduced the Transformer, the architecture behind modern large language models.

How much weight it carries: One of the most cited and consequential papers in machine learning.

Where this came from

15 pages. A copy is archived locally against link rot; the header links the original source.