Series

LLM Architecture: GPT-2

The fundamentals every model since has inherited.

Chapters

4 of 8
  1. 01Tokenization2 min
  2. 02From Raw Text to LLM Training Data6 min
  3. 03Layer normalization
  4. 04Attention, Step by Step9 min
  5. 05Dropout4 min
  6. 06Shortcut connections
  7. 07The feed-forward network
  8. 08From vectors to logits

Search

Search pages, articles, and resources