Contents
Map

01 · LLM Foundations

Appendix - Summary & Key Terms

View as:

Appendix - LLM Foundations

What We Learned

  • Tokenization turns text into IDs; vocabulary choices affect length, cost and failure modes.
  • A transformer combines embeddings, position, attention and feed-forward layers to predict tokens.
  • Causal masking supports parallel training; generation still proceeds token by token.
  • Encoder, decoder, MoE and hybrid designs trade capability against compute and memory.
  • Choose models using task evaluations; test hallucination, context loss and sycophancy.

Key Acronyms, Concepts & Jargon

TermShort meaning
LLMLarge Language Model: a model that learns patterns in token sequences.
BPEByte Pair Encoding: builds tokens by repeatedly merging frequent adjacent pairs.
Token / context windowA model input unit / the token budget available to a request.
Temperature / top-pControls sampling sharpness / limits sampling to a probability mass.
Q/K/VQuery, Key, Value: attention matches queries to keys and combines values.
RoPERotary Position Embeddings: rotations that encode token position in attention.
RMSNorm / FFNRoot Mean Square Normalization / Feed-Forward Network: normalize activations / transform each token.
MHA / MQA / GQAMulti-Head / Multi-Query / Grouped-Query Attention: independent / shared / grouped K/V heads.
FlashAttentionExact attention computed with less intermediate memory and data movement.
MoEMixture of Experts: routes tokens to a subset of expert networks.
MLA / SSMMulti-head Latent Attention / State-Space Model: compressed attention state / recurrent sequence state.
Hallucination / sycophancyUnsupported output / agreeing with the user over correctness.
Lost in the middleReduced use of relevant information buried inside a long context.

Back to section overview

⚡AI-assisted content - always verify, always explore multiple perspectives·