Field guide

The Prediction Engine

Hand-drawn atlas of glowing nodes linked like a constellation

Chapter 1

AIs do not know. They continue.

A modern generative model is a gigantic function that maps a sequence of symbols to a probability for every possible next symbol. Training tunes that function so likely next symbols in the data become high-probability predictions. That one trick — next-piece prediction at scale — now writes, codes, composes, animates, and sculpts.

What is inside

  • Parameters — billions of numbers (weights) that define the function.
  • Architecture — usually a transformer: stacked layers of attention + feed-forward nets.
  • Tokenizer — a dictionary that chops text, pixels, audio, or 3D into discrete tokens.
  • Objective — minimize surprise (cross-entropy / diffusion loss) on training data.

What is not inside

  • No searchable encyclopedia of facts. Recalling “Paris is the capital of France” is a side effect of compression.
  • No inner monologue required. Chain-of-thought is more generated tokens, not a separate mind.
  • No guarantee of truth. High probability ≠ correct. Models hallucinate fluent gaps.
  • No senses unless you attach cameras, mics, tools, or retrieval.

The training loop in one breath

  1. Take a batch of real examples (sentences, spectrograms, video latents, meshes).
  2. Hide the future: mask the next token, or add noise, or drop a patch.
  3. Ask the net to restore it. Measure error with a loss.
  4. Backpropagate: nudge every weight a tiny step that would have reduced that error.
  5. Repeat on trillions of tokens until the model is a sharp statistical model of the world-as-data.