ESC
Type to search guides, tutorials, and reference documentation.
← AI and Machine Learning
🧠

Generative AI and LLMs

Models that learn a distribution and sample from it: prefill and decode, the adaptation ladder, evaluation, and why confident wrong answers are the objective working as designed.

Generative AI covers models that learn a distribution over data and can draw new samples from it, rather than only mapping an input to a label. A large language model is the most visible case: it is trained to predict the next token given the preceding context, and text generation is simply that prediction applied repeatedly to its own output. Understanding generative systems mostly means understanding what that objective does and does not guarantee.

Model families

  • Autoregressive models factorize a sequence into a product of conditional probabilities and generate one element at a time. This is the basis of modern language models and of several audio and image models.
  • Diffusion models learn to reverse a gradual noising process, starting from noise and denoising in steps toward a sample. They dominate high-quality image and video synthesis, and trade compute-per-sample for quality through the number of denoising steps.
  • Variational autoencoders learn a compressed latent space with a probabilistic structure you can sample from; they are often used as a component inside larger systems rather than as the end product.
  • Generative adversarial networks pit a generator against a discriminator. Fast to sample from, historically difficult to train stably.

How an LLM actually runs

Text is first split into tokens by a subword tokenizer, so the model never sees characters or words directly — which is why character-level tasks such as counting letters or reversing a string are awkward for it. Tokens are embedded, passed through a stack of transformer blocks, and turned into a probability distribution over the vocabulary for the next position.

Generation has two distinct phases with very different cost profiles. Prefill processes the whole prompt in parallel and is compute-bound. Decode emits one token at a time and is bound by memory bandwidth, because each step must read the model weights and the accumulated KV cache — the stored keys and values for every previous position. That cache grows with context length and with the number of concurrent requests, and it is usually what limits how many sessions a given amount of accelerator memory can hold.

Decoding parameters control how the distribution is turned into a choice. Greedy decoding always takes the most likely token. Temperature flattens or sharpens the distribution; top-k and top-p (nucleus) sampling restrict the candidate set before sampling. Sampling is not a defect to be tuned away — it is the mechanism by which the model produces varied, non-degenerate text — but it does mean identical inputs need not produce identical outputs.

The adaptation ladder

There is a rough ordering of ways to make a general model do your task, from cheapest and most reversible to most expensive and most committed:

  1. Prompting. Instructions, examples, output schemas. Instant to change, no training infrastructure, but the behavior lives in a string that is easy to break and hard to regression-test.
  2. Retrieval augmentation (RAG). Fetch relevant documents at request time and place them in the context. The right answer when the problem is missing knowledge — private data, recent facts, long tail — because you can update the corpus without touching the model.
  3. Parameter-efficient fine-tuning. Train a small set of additional parameters while freezing the base weights. Suits format, tone and task-shape adaptation.
  4. Full fine-tuning. Updates all weights. Most capable, most expensive, and it pins you to a base model version.
  5. Pretraining. Rarely the right call outside organizations whose product is the model itself.

The common mistake is reaching for fine-tuning to inject facts. Fine-tuning is good at teaching behavior and poor at reliably storing knowledge; retrieval is the reverse. Diagnose which one you are missing before choosing.

Evaluation is the hard part

Generative output has no single correct answer, so there is no accuracy number to read off. What works is unglamorous: build a golden set of representative inputs with expected properties, and check properties rather than exact strings — does the output parse as valid JSON, does it cite a retrieved document, does it refuse when it should. Pairwise comparison between two candidate systems is often more reliable than absolute scoring. Using a model as a judge scales evaluation but carries known biases, including preferences for longer answers and for its own style, so calibrate it against human labels on a subset before trusting it.

Treat prompts and retrieval configurations as versioned artifacts with a regression suite. Without one, every prompt edit is an unmeasured change to production behavior.

Failure modes

  • Confabulation. The model is trained to produce probable continuations, not true ones. Fluent, confident, wrong output is the objective working as designed — mitigate with retrieval, citation requirements and verification, not by asking the model to be accurate.
  • Prompt injection. Any untrusted text that reaches the context — a fetched web page, a user file, a tool result — can carry instructions. Treat retrieved content as data, never as instructions, and gate every consequential tool call on an authorization check outside the model.
  • Long-context degradation. Filling a large window does not mean the model attends evenly across it; material in the middle is easiest to lose. More context is not free in quality or in latency.
  • Silent version drift. When a hosted model is updated behind a stable name, your prompts inherit the change. Pin versions where you can and keep the regression suite running.
  • Evaluation contamination. Public benchmarks may appear in pretraining data, which makes published scores a weak predictor of performance on your task.
  • Cost and latency surprises. Cost scales with tokens, and long prompts are paid on every call. Caching, shorter contexts and smaller models for easy requests are the usual levers.

When not to use it

Do not use a generative model where a deterministic transformation will do — parsing a known format, applying a fixed business rule, or performing arithmetic that a calculator does exactly. Avoid putting one in an unsupervised path to an irreversible action: payments, deletions, outbound communication. And where the requirement is an auditable, stable decision, a conventional model with a fixed threshold is easier to defend than a sampled one. Related material: deep learning for the underlying architectures, model compression for serving them affordably, and enterprise RAG pipelines for retrieval in production.

RELATED GUIDES