Most explainers stop where things get interesting, and most textbooks assume you already read the previous one. This course covers the whole distance: what a token is, why attention is shaped the way it is, and how to build something with it.
Because the existing ones stop too early or start too late.
Introductory material tends to explain what attention does and leave it there. Academic material explains why, but assumes you can already read the notation. The gap between those two is where most people give up.
This course is written to close that gap. Every section comes in two registers: a plain-language version that uses no notation at all, and a technical version that works through the maths properly. Same concept, same page, and you can switch between them whenever one stops making sense.
There are no videos. Reading is faster, and a diagram you can drag and poke at teaches more than watching someone else point at one. Every section ends with a quiz so you find out what didn't stick while it's still fresh.
modules, grouped into 5 units
sections, every one written
interactive diagrams and visualisations
quiz and exam questions
This is an independent project and still a work in progress. Sections get rewritten as better explanations turn up. If something reads badly or is flat out wrong, reporting it is welcome and it gets fixed.
Six things you can't do before you start.
Each one is a thing you can go and do afterwards, not a topic you've read about.
Tokens, attention, embeddings, training. The things people use as buzzwords, explained until they stop being buzzwords.
Few-shot, chain of thought, structured output, and why each one works rather than a list of tricks to copy.
Your first request, streaming responses, conversation state, tool use, and keeping the bill under control.
Embeddings, chunking, vector search, and the parts that usually go wrong when you connect a model to your own documents.
The curriculum follows the published objectives for NVIDIA's generative AI exam, so studying here covers that ground.
Hallucination, bias, prompt injection, evaluation. Where these systems break and what to do about it.
Three steps, and then you're just learning.
Sign-up asks about your background and what you want out of this. If you've never written code, you get the Explorer path. If you have, you get the extra code and maths. You can switch later.
Explorer, Builder or ResearcherEvery section has the written explanation, a diagram you can interact with, usually a code example you can run, and a short quiz. There's a tutor in the corner that knows which section you're on if you get stuck.
Theory, visuals, code, quizEach module ends with a timed exam, around 20 questions, 70% to pass. It's there so you find out what didn't stick while you can still go back and reread it.
Timed, 70% to pass, retakeable16 modules, in dependency order.
Each module assumes only what came before it. Start at Module 1 if any of this is new, or skip ahead if it isn't.
No prior knowledge assumed. What the word intelligence means once you apply it to a machine, how the field got from hand-written rules to ChatGPT, and where the AI you already use fits in.
What training actually does. Datasets, the three kinds of learning, how a model measures its own error, why it can memorise instead of learn, and what a neural network is underneath.
Why depth turned out to matter. How backpropagation lets a network correct itself, how convolutional networks read images, and the vanishing gradient problem that nearly ended the field.
The architecture behind ChatGPT, Gemini and Claude. Where earlier sequence models fell short, what attention fixed, and how a Transformer is assembled from self-attention, positional encoding and multiple heads.
How a model trained only to predict the next word ends up useful. Tokenization, pretraining at scale, context windows, what separates BERT from GPT, and what RLHF changed.
The models behind image, audio and video generation. How diffusion works, what GANs and VAEs each do differently, and how multimodal systems handle more than one kind of input at once.
About the NVIDIA certification
The curriculum is structured around the published exam objectives for the NVIDIA-Certified Associate: Generative AI LLMs exam, so working through these modules covers that ground. To be explicit about what this is not: an independent project with no affiliation with or endorsement from NVIDIA. Completing it does not certify you. The exam is sat with NVIDIA directly.
Read the official exam page →Nothing, for now at least.
There is no payment page, because there is nothing to pay for. Every module, every exam and the tutor are all included.
An account just saves your progress. No card, ever.
The honest answer.
Writing and hosting cost time rather than money. The one real expense is the tutor, which calls a paid API every time somebody uses it. If that bill becomes uncomfortable, tutor messages get a daily limit before anything gets a price.
If part of this is ever charged for, it will not be material that was free the day before, and accounts that already exist keep the access they have. Any change gets announced here rather than applied quietly.
As it stands there is no billing code in this project at all, which is a more convincing version of that promise than a sentence about it.