XFollowListPeople, ideas and original work on X
Sebastian Raschka’s avatar

LLM research engineer and author

Sebastian Raschka

@rasbt on X

Explains language models through small implementations, connecting tokens and attention to the code for building, training and adapting a GPT-style model.

Why read Sebastian Raschka?

Read Raschka when you want to connect a model diagram to calculations you can inspect. His attention tutorial starts with a six-word sentence, turns the words into vectors, then constructs queries, keys, values and a context vector. You can follow how information from other words changes a word's representation.

Continue with the LLMs from Scratch repository to see that mechanism inside a GPT-style model, followed by pretraining and fine-tuning. The chapter map gives you a sequence to follow rather than a collection of unrelated examples. Python knowledge is useful; the repository points readers new to PyTorch to its introductory appendix.

Start with the original

Selected work

  1. Tutorial ·

    Understanding and Coding the Self-Attention Mechanism of Large Language Models From Scratch

    How does attention give a word context? Raschka builds a small PyTorch example from word vectors to attention weights and a context vector, then explains multiple heads and cross-attention.

  2. Code

    Build a Large Language Model (From Scratch): companion code

    How do the pieces become a model? Follow the chapter map from text data to a GPT implementation, pretraining and fine-tuning. Start with the introductory PyTorch appendix if tensor operations are unfamiliar.

Projects & roles

  • LLMs from Scratch

    The official companion code for Raschka's book, including a small GPT-style model and training examples.

  • Ahead of AI

    His publication for longer explanations of model architectures, training methods and evaluation.

Community rating

82

Vote once every 24 hours. No account needed.

Sources & further reading