Series
LLM Architecture: GPT-2
The fundamentals every model since has inherited.
Chapters
4 of 8- 01Tokenization2 min
- 02From Raw Text to LLM Training Data6 min
- 03Layer normalization—
- 04Attention, Step by Step9 min
- 05Dropout4 min
- 06Shortcut connections—
- 07The feed-forward network—
- 08From vectors to logits—