A modern generative model is a gigantic function that maps a sequence of symbols to a probability for every possible next symbol. Training tunes that function so likely next symbols in the data become high-probability predictions. That one trick — next-piece prediction at scale — now writes, codes, composes, animates, and sculpts.
What is inside
Parameters — billions of numbers (weights) that define the function.
Architecture — usually a transformer: stacked layers of attention + feed-forward nets.
Tokenizer — a dictionary that chops text, pixels, audio, or 3D into discrete tokens.
Objective — minimize surprise (cross-entropy / diffusion loss) on training data.
What is not inside
No searchable encyclopedia of facts. Recalling “Paris is the capital of France” is a side effect of compression.
No inner monologue required. Chain-of-thought is more generated tokens, not a separate mind.
No guarantee of truth. High probability ≠ correct. Models hallucinate fluent gaps.
No senses unless you attach cameras, mics, tools, or retrieval.
The training loop in one breath
Take a batch of real examples (sentences, spectrograms, video latents, meshes).
Hide the future: mask the next token, or add noise, or drop a patch.
Ask the net to restore it. Measure error with a loss.
Backpropagate: nudge every weight a tiny step that would have reduced that error.
Repeat on trillions of tokens until the model is a sharp statistical model of the world-as-data.
Chapter 2
Prediction is a softmax over a dictionary
Language models do not pick “the answer.” They emit a score (a logit) for every token in a vocabulary of tens or hundreds of thousands of pieces. Softmax turns those scores into probabilities that sum to 1. Sampling then draws the next token. Repeat. That is generation.
Temperature flattens or sharpens the distribution. At 0.1 the model is nearly greedy. At 1.5 it becomes chaotic. Top-k keeps only the k likeliest tokens before sampling — a simple way to cut off nonsense tails.
Next-token distribution
pi = exp(zi/T) / Σ exp(zj/T)
Four sampling recipes
Greedy — always take argmax. Repetitive, safe, dull.
Temperature — divide logits by T before softmax.
Top-k — zero out everything outside the k best, renormalize.
Nucleus (top-p) — keep the smallest set whose mass ≥ p (often 0.9).
This lab’s honest cheat
Your browser is not running a 70B transformer. It runs a compact interpolating n-gram over a built-in essay about AI, with a unigram backoff for unknown words. The interface of logits → softmax → sample is the same one a frontier model uses. The world model is tiny on purpose so you can see the gears.
Chapter 3
Attention is a soft lookup table
A transformer does not read left-to-right like a human. Each token builds three vectors — Query, Key, and Value — then asks: “which other tokens should I listen to?” Compatibility is a dot product. Softmax turns it into weights. The output is a blend of Values. Stack that 96 times with residual paths and you get GPT-class behavior.
Tap a query token. Cells show a toy attention pattern (content + position), not weights from a trained 7B model — but the algebra is identical: softmax(QKᵀ / √d) V.
Embeddings
Each token id becomes a vector. Nearby meanings often sit nearby in that space after training — “king” minus “man” plus “woman” was the old slogan. Real geometry is messier, but the idea holds: meaning is direction.
Residual stream
Each layer adds a delta to a running residual. Early layers do syntax; later layers do facts, style, and long-range binding. You can literally read some facts out of that stream with a linear probe.
Context window
Attention is O(n²) in sequence length unless you approximate it. That is why context is expensive, why long videos are hard, and why engineers invent sliding windows, state-space models, and sparse patterns.
Chapter 4
The Codex: a dictionary the model can speak
Every generative system needs a finite alphabet. In papers this alphabet is often called a codebook or codex: a numbered list of reusable pieces. Text uses subword tokens. Images, audio, and 3D often use a learned vector-quantized codebook. Code models are the same next-token engine aimed at source code. Three ideas, one spine.
Real tokenizers (BPE, WordPiece, Unigram) learn merge rules from a huge corpus. This demo uses a frozen, readable merge table so you can see “at”+“ten”+“tion” collapse into fewer, more meaningful pieces. Fewer tokens = more meaning per step, cheaper context.
Tokens
1. Text codex
GPT-style vocabs hold ~32k–200k subwords. “unbelievable” might be un + believ + able. Spaces and punctuation are tokens too. The model never sees “characters” unless you train it that way.
2. VQ codebook
A VQ-VAE or residual vector quantizer learns K prototype vectors. An encoder maps a spectrogram patch or image patch to the nearest prototype. Generation becomes “predict codebook indices,” i.e. language modeling on a visual or audio alphabet. EnCodec, SoundStream, and many image tokenizers work this way.
3. Code models
Codex-class models are LLMs trained heavily on public repositories. Same softmax. The difference is the data: syntax is locally rigid, names repeat, and unit tests are a verifiable reward. That is why fill-in-the-middle and tool-use (run the compiler) help so much.
Why a codebook is the secret handshake across media
Once sound, pictures, video clips, or 3D parts are discrete tokens, you can drop them into the same transformer that already speaks English. Multimodal models are often several tokenizers feeding one predictor — a shared codex of the senses.
Chapter 5
Song models predict air, not lyrics only
Audio is a pressure wave at 44,100 samples per second. No transformer wants to softmax 44k times a second over 16-bit integers. So music AIs compress the wave into a slow sequence of codes, then generate those codes — or they diffuse in a spectrogram / latent space and decode back to a waveform.
A toy pentatonic Markov melody: each note is a token. This is how early symbolic models (and MIDI transformers) work. Neural codecs do the same idea on learned audio tokens instead of piano-roll notes.
Two families
Autoregressive codecs — encode audio with EnCodec-like RVs, then an LM predicts the next codebook index (MusicGen-style). Great local coherence, slow sequential decode.
Diffusion / flow — start from noise in spectrogram or latent space, iteratively denoise conditioned on text or melody (Stable Audio-style). Parallel, high fidelity, needs many steps or a distilled sampler.
Hybrids — language model for structure, diffusion for timbre.
Why lyrics and groove disagree
Text tokens are ~4 characters. Audio tokens might represent 10–20 ms. A three-minute song is thousands of codes. Long-range form (verse / chorus) is the same long-context problem as novels. That is why models get help from beat grids, chord charts, or a separate “plan” token stream.
Chapter 6
Video is time-shaped noise becoming physics
A video model must invent pixels that agree with each other across space and across time. The dominant recipe: compress frames into a latent grid (a spacetime patch tokenizer), then either denoise all patches together or predict the next latent token. Motion is not bolted on later — it is whatever the loss forced to be consistent.
Press denoise. Gaussian noise is iteratively pulled toward a bouncing-ball clip. Real systems denoise in a compressed latent, with a text embedding injected at every step, often using extra temporal attention layers.
Step 0 — pure noise in pixel space (real models use latents).
Spacetime patches
Instead of “a token per pixel,” models pack a 16×16×temporal tube into one vector. The transformer then attends over tubes. That is how a few seconds of 24 fps video become a tractable sequence.
Temporal attention
Each patch may attend to the same location in other frames (motion) and to neighbors in-frame (shape). Factorized attention — spatial then temporal — is a common engineering compromise.
Consistency hacks
Optical-flow guidance, overlapping windows, image-to-video conditioning, and overlapping latent decoding all fight flicker. Humans notice a 1-pixel wobble that a still-image FID score will happily ignore.
Chapter 7
3D models must pick a representation
A song is a 1D wave. A video is a 2D image plus time. 3D has no single native tensor. Generators therefore invent an object in some internal dialect — voxels, meshes, implicit fields, or clouds of Gaussians — then render it. The interactive globe below is a stand-in: noise points that snap onto a platonic solid, the way a model commits to a surface.
WebGL unavailable — showing a 2D schematic instead.
Native 3D generators
Train on ShapeNet-style assets: predict occupancy, SDF, mesh tokens, or splat parameters directly. Fast at inference, limited by the bias of the 3D dataset (lots of chairs, not lots of wet streets).
Distill from 2D
Score Distillation Sampling (DreamFusion-style) optimizes a 3D scene so that random camera renders look like a frozen 2D text-to-image model’s denoising gradient. You borrow the 2D model’s visual prior when you lack 3D captions.
Chapter 8
Check the gears
Eight questions. Your best score is stored on this device, and on your signed-in Popchat account when available.