Contents
Map

01 · LLM Foundations

Model Architecture Types

View as:

Model Architecture Types

Transformers come in three shapes - encoder-only, decoder-only and encoder-decoder - distinguished by the attention mask and the training objective, and each shape can be made sparse with mixture-of-experts layers. This chapter explains what each shape can and cannot do, why decoder-only models took over general-purpose AI, and how to choose one for a task.

Learning objectives 45 min
By the end of this page you will be able to:
  • Distinguish encoder-only, decoder-only and encoder-decoder models by attention mask, training objective and what they can output
  • Explain why decoder-only models became the default for general-purpose assistants, and where encoders and encoder-decoders still win
  • Explain mixture-of-experts: total vs active parameters, routing, load balancing, and why it saves compute but not memory
  • Choose an architecture (and model size class) for a task and justify it

Encoder-Only Architecture

Concept

Encoder-only models use bidirectional self-attention - each token can attend to all other tokens simultaneously. There is no causal mask; the model sees the full context in both directions.

Training objective - Masked Language Modeling (MLM):

  • Randomly mask 15% of tokens in the input
  • Train the model to predict the masked tokens from context
  • This forces the model to understand bidirectional context

Architecture specifics:

  • Input sequence → bidirectional attention → contextual representations per token
  • No autoregressive generation - the model produces a fixed-size representation for each token, not new tokens
  • Classification and span-extraction heads are added on top of the final hidden states

When to use encoder-only:

  • Text classification (sentiment analysis, intent detection): take the [CLS] token embedding → linear head
  • Named Entity Recognition (NER): classify each token's hidden state → per-token labels
  • Semantic similarity / embeddings: mean pool or [CLS] embedding → similarity search
  • Extractive QA (SQuAD): predict start/end token positions within the context

Key models:

ModelParametersContextKey innovation
BERT-base110M512Bidirectional MLM + NSP
BERT-large340M512Larger BERT
RoBERTa125M/355M512Removed NSP, more data, better MLM
DeBERTa-v3183M512Disentangled attention, ELECTRA pretraining
DistilBERT66M51240% smaller, 60% faster, 97% of BERT quality
BGE-M3~560M8KMulti-function embedding model

Tricky Q: Can a BERT-family model generate text?
No - BERT was never trained to autoregressively predict the next token. Its output is a contextual representation, not a probability distribution over the next token. You can build a seq2seq encoder-decoder on top of a BERT encoder, but the encoder alone cannot generate.


Decoder-Only Architecture

Concept

Decoder-only models use causal (left-to-right) self-attention - each token can only attend to itself and all preceding tokens. The model is trained to predict the next token.

Training objective - Causal Language Modeling (CLM):

  • Given tokens t₁, t₂, ..., tₙ₋₁, predict tₙ
  • Loss is cross-entropy on the predicted next-token distribution vs the actual next token
  • The simplicity of this objective is a key reason for the architecture's dominance

Why decoder-only won:

  1. Unified objective: CLM pretraining + instruction fine-tuning (SFT) + RLHF all use the same token prediction framework - no architectural changes between stages
  2. Natural generation: The architecture is inherently designed to generate; no adapter or cross-attention bridge needed
  3. Emergent reasoning: At scale, decoder-only models develop chain-of-thought reasoning by learning to produce intermediate "thinking" tokens
  4. In-context learning: Few-shot prompting works naturally by prepending examples before the query
  5. Scalability: Simple objective → easy to scale to trillions of training tokens

Architecture flow:

flowchart LR
    A["📝 Input: 'The cat sat on the'"] --> B["✂️ Tokenize"] --> C["🔢 Embed"]
    C --> D["🔁 Transformer Block × N<br/>with causal mask"] --> E["🧮 Final LN"] --> F["📐 Linear"] --> G["🎯 Softmax"]
    G --> H["P(next token)"] --> I["🎲 Sample"] --> J["➕ Append"]
    J -.->|repeat| B

Key models:

ModelParamsContextKey innovation
GPT-21.5B1024First large-scale CLM demo
GPT-3175B2048In-context learning at scale
GPT-4 (2023)Undisclosed8K / 32KMultimodal input, top-tier reasoning at release
LLaMA-27B–70B4KOpen weights, commercial license
LLaMA-3.18B–405B128KGQA, strong open weights
Gemma 22B–27B8KGoogle open weights, interleaved local (sliding-window) and global attention
Mistral 7B7B8K (v0.1) / 32K (v0.2+)GQA + sliding window attention (v0.1)
Phi-3-mini3.8B128KHigh quality from small size
Falcon7B–180B2K–8KMulti-Query Attention
DeepSeek-V3671B MoE (37B active)128KMLA, fine-grained MoE, FP8 training

Encoder-Decoder Architecture

Concept

Encoder-decoder models have two components:

  1. Encoder: reads the input with bidirectional attention → produces context representations
  2. Decoder: generates output tokens autoregressively, attending to both its own output (causal self-attention) and the encoder output (cross-attention)

Cross-attention:

Q = decoder hidden states (what the decoder wants to know)
K = V = encoder output    (what the input provides)
cross_attention_output = softmax(Q·Kᵀ / √d_k) · V

This allows the decoder to directly "look up" relevant parts of the input at each generation step.

Training objectives:

  • T5 (Span Corruption): Mask contiguous spans of input tokens (replace with single mask token), train decoder to reconstruct the spans
  • BART (Denoising): Corrupt input via masking, shuffling, deletion; train decoder to reconstruct the clean sequence

When to use encoder-decoder:

  • Machine translation: long input → long output with reordering
  • Abstractive summarization: input document → shorter summary (not extractive)
  • Document-to-structured output: parse a document into a structured format
  • Question generation: answer + context → question

Key models:

ModelParametersKey innovation
T5-base/large220M/770MText-to-text unified framework
FLAN-T580M–11BInstruction fine-tuned T5
BART140M/400MDenoising pretraining
mT5300M–13BMultilingual T5
mBART610MMultilingual denoising

Tricky Q: Why did decoder-only models overtake encoder-decoder for summarization and translation?

Two reasons: (1) At sufficient scale, decoder-only models learn to produce appropriate-length summaries and translations via instruction fine-tuning, without the architectural inductive bias of cross-attention. (2) A single decoder-only model can handle many tasks via prompting, while encoder-decoder models require task-specific fine-tuning to perform well. Operational simplicity wins.


Architecture Comparison

Concept

DimensionEncoder-OnlyDecoder-OnlyEncoder-Decoder
Attention directionBidirectionalCausal (left-to-right)Encoder: bi; Decoder: causal + cross
Training objectiveMLM / ELECTRACLM (next token)Denoising / span corruption
Can generate text?NoYesYes
Best forClassification, NER, embeddingsGeneration, reasoning, chatSeq2seq tasks
In-context learningPoorExcellentModerate
Instruction fine-tuningAwkwardNaturalPossible
Production dominanceEmbedding modelsGeneral LLMsNiche seq2seq

Mixture of Experts (MoE)

Concept

MoE is an architectural modification to the FFN layer, not a new architecture class. Instead of one dense FFN per layer, the model has N "expert" FFNs and a learned router that selects top-K experts per token.

flowchart LR
    subgraph STD["Standard FFN"]
        T1["token"] --> F1["single FFN"] --> O1["output"]
    end
    subgraph MOE["MoE FFN"]
        T2["token"] --> R["router<br/>(softmax over N experts)"]
        R --> K["select top-K experts<br/>by router score"]
        K --> W["weighted sum of<br/>selected expert outputs"]
    end

    style STD fill:#e8f4fd,stroke:#4a9eca
    style MOE fill:#d4edda,stroke:#28a745

Key MoE concept: sparse activation

  • Total parameters: N × (d_model × d_ffn) - much larger than a dense model
  • Active parameters per token: only K experts are used - much smaller computation
  • Example: Mixtral-8x7B has 8 expert FFNs per layer and routes each token to 2 of them. Only the FFN blocks are replicated - attention and embeddings are shared - so the total is ~47B parameters (not 8 × 7B = 56B), with ~13B active per token

MoE advantages:

  • Scale model capacity without proportional compute increase
  • Different experts can specialize in different domains/languages/task types
  • Efficient at inference for large models (only K/N of FFN is computed per token)

MoE challenges:

  • Load balancing: without constraints, router collapses all tokens to the same 1–2 experts
  • Communication overhead in distributed training (all-to-all for expert routing)
  • MoE saves compute, not memory: all experts must be resident in GPU memory (47B for Mixtral), and the KV cache is unchanged because attention is not sparse

Models:

  • Mixtral-8x7B: 8 experts, top-2 routing, 47B total / 13B active
  • DeepSeek-V3, Qwen3-235B-A22B, Llama 4, gpt-oss, Kimi K2: fine-grained MoE with many small experts - see Modern Architectures
  • Switch Transformer (Google): pioneered MoE at scale with top-1 routing

Model Family Comparison Table

Concept

A reference table of influential model families and the architectural idea each is known for. It is deliberately historical - for the current lineup see the Model Landscape. Closed-model internals are listed as undisclosed unless the vendor published them.

Model FamilyOrgTypeContextArchitecture innovations
GPT-2OpenAIDecoder1KFirst demo of large CLM
GPT-3OpenAIDecoder2K175B in-context learning
GPT-4 / 4oOpenAIUndisclosed128KMultimodal, top-tier at release
GPT-4.1OpenAIUndisclosed1MExtended context, code + general
LLaMA-2MetaDecoder4KOpen weights, GQA (70B)
LLaMA-3 / 3.1MetaDecoder8K / 128KGQA all sizes, 405B
LLaMA-4 ScoutMetaMoE Transformer10MUltra-long context, open weights
LLaMA-4 MaverickMetaMoE Transformer1MBalanced capability/cost
GemmaGoogleDecoder8KMulti-query attention (2B), open weights
Gemma 2GoogleDecoder8KGQA + local+global attn
Gemini 2.5 ProGoogle DeepMindUndisclosed1MComplex reasoning, multimodal
Gemma 3GoogleDecoder128K5:1 local:global attention, QK-norm
Mistral 7BMistralDecoder8K / 32KGQA + sliding window (v0.1)
Mixtral 8x7BMistralDecoder (MoE)32KTop-2 MoE, 47B/13B active
Phi-3-miniMicrosoftDecoder128K3.8B, textbook-quality data
Qwen3Alibaba CloudDense + MoE32K native (128K with YaRN)Hybrid thinking mode, Apache-2.0
DeepSeek-V3 / R1DeepSeekMoE (671B / 37B active)128KMLA, aux-loss-free balancing, open reasoning (R1)
gpt-ossOpenAIMoE (117B / 5.1B active)128KOpen weights, MXFP4 experts, banded attention
Command ACohereDecoder256KTuned for RAG, tool use and enterprise search
BERT-baseGoogleEncoder512MLM + NSP, bidirectional
RoBERTaMetaEncoder512Better BERT training
DeBERTa-v3MicrosoftEncoder512Disentangled attention
T5 / FLAN-T5GoogleEnc-Dec512–2KText-to-text, instruction
BARTMetaEnc-Dec1KDenoising pretraining

LLM Types by Modality

Concept

Beyond architecture type, LLMs are also categorized by the modalities they handle. This affects model selection, embedding strategy, and system design.

TypeDescriptionExamplesUse Cases
Text-OnlyTrained on text corpora onlyGPT-3, LLaMA, BERT, FalconGeneral NLP, summarization, generation
MultilingualTrained on multilingual corporaXLM-R, mT5, BLOOM, GPT-4Translation, cross-lingual search
MultimodalInput: image/video/audio + textGPT-4o, Gemini, Qwen-VL, LLaVAImage captioning, audio Q&A, OCR, document understanding
CodeSpecialized in programming languagesCodex, CodeLLaMA, CodeGemma, StarCoderCode generation, completion, refactoring
Speech-TextIntegrate speech recognition and synthesisWhisper, SeamlessM4TTranscription, speech translation
Image/VisionImage understanding and classificationViT, ConvNeXT, DETR, Mask2FormerObject detection, segmentation, depth estimation

Modality in system design: When building a multi-modal pipeline (e.g., processing scanned PDFs + structured data + free text), you must choose whether to use a single multi-modal foundation model (simplicity, one context) or multiple specialized models (higher per-task accuracy, more complex orchestration). For RAG over images, multimodal embeddings (CLIP, Gemini Embeddings) are required - text-only embeddings cannot capture visual content.

Models by specific task (quick reference):

ModelTask
Wav2Vec2Audio classification, automatic speech recognition (ASR)
Vision Transformer (ViT), ConvNeXTImage classification
DETRObject detection
Mask2FormerImage segmentation
GLPNDepth estimation
BERTText classification, token classification, question answering
GPT-2, LLaMAText generation
BART, T5Summarization and translation
Codex, CodeLLaMACode generation
WhisperSpeech-to-text transcription

When to Choose Which Architecture

Concept

Use encoder-only (BERT/DeBERTa/BGE) when:

  • Task is purely classification, NER, or span extraction
  • You need high-quality fixed-size embeddings for search/RAG
  • Inference must be very fast (smaller models, no generation overhead)
  • Labels are sentence-level or token-level - not generative

Use decoder-only (LLaMA/Gemma/GPT) when:

  • Task requires free-form text generation
  • You want a single model for multiple tasks via prompting
  • You need in-context learning (few-shot examples in prompt)
  • You plan to instruction fine-tune for custom tasks

Use encoder-decoder (T5/FLAN-T5/BART) when:

  • Task is explicitly seq2seq with fixed input and output schemas
  • You have limited compute - smaller fine-tuned encoder-decoder can beat large decoder-only for specific structured tasks
  • Translation or summarization at scale where encoder-decoder efficiency matters

Practical rule of thumb: Default to decoder-only for new projects. Switch to encoder-only if you need embeddings or fast token classification. Encoder-decoder is a niche choice for legacy systems or very specific seq2seq workloads.

Code

from transformers import (
    AutoTokenizer, AutoModelForSequenceClassification,  # Encoder-only
    AutoModelForCausalLM,                               # Decoder-only
    AutoModelForSeq2SeqLM,                             # Encoder-Decoder
    pipeline
)

# --- Encoder-only: text classification ---
classifier = pipeline(
    "text-classification",
    model="distilbert-base-uncased-finetuned-sst-2-english"
)
result = classifier("This movie was absolutely fantastic!")
print(f"Encoder-only classification: {result}")

# --- Decoder-only: text generation ---
generator = pipeline(
    "text-generation",
    model="gpt2",  # small, runs locally
    max_new_tokens=50,
    temperature=0.7
)
result = generator("The transformer architecture revolutionized AI because")
print(f"\nDecoder-only generation: {result[0]['generated_text']}")

# --- Encoder-Decoder: summarization ---
summarizer = pipeline(
    "summarization",
    model="facebook/bart-large-cnn",
    max_length=60,
    min_length=20
)
article = """
The transformer architecture, introduced in the paper "Attention Is All You Need" by Vaswani 
et al. in 2017, revolutionized natural language processing. It replaced recurrent neural 
networks with a self-attention mechanism that could process all tokens simultaneously, 
enabling much faster training and better handling of long-range dependencies.
"""
result = summarizer(article)
print(f"\nEncoder-Decoder summary: {result[0]['summary_text']}")

# --- Checking architecture type programmatically ---
from transformers import AutoConfig

for model_name in ["bert-base-uncased", "gpt2", "facebook/bart-large"]:
    config = AutoConfig.from_pretrained(model_name)
    print(f"{model_name}: {config.model_type} - {config.architectures}")

Functional Model Categories

Concept

Beyond the three architecture shapes, people describe models by what they are trained to do. These are overlapping labels, not separate architectures - one model can be an MoE, a reasoning model and a vision-language model at once.

CategoryWhat it isExamplesTypical use
Reasoning modelA decoder-only LLM post-trained, largely with RL on verifiable tasks, to produce a long chain of thought before answering (see Post-Training & Reasoning)DeepSeek-R1, OpenAI o-series, Qwen3 in thinking modeMaths, code, multi-step planning
Mixture of expertsSparse FFN layers: a router sends each token to a few expertsMixtral, DeepSeek-V3, Qwen3-235B-A22B, gpt-ossLarge capacity at lower compute per token
Vision-language modelA vision encoder (often a ViT) whose features are mapped into the LM's embedding space by a projector, trained on image-text data, or a natively multimodal modelLLaVA, Qwen-VL, Gemma 3, GPT-4oDocuments, screenshots, image Q&A, agent perception
Small language modelModels of roughly 1-8B parameters, trained (often with distillation from a larger model) to run on a single GPU or devicePhi, Gemma, Qwen and Llama small sizesOn-device use, latency-critical sub-tasks, fine-tuning
Computer-use / action modelAn LLM trained or tuned to output actions - tool calls, clicks, keystrokes - from screenshots or UI treesOpenAI's Computer-Using Agent, Anthropic's computer-use capability, Gemini computer-use modelsBrowser and desktop agents (see Computer-Use and Browser Agents)
Embedding modelUsually an encoder (or a decoder adapted) trained with contrastive objectives to output a fixed vectorBGE, E5, Gemini and OpenAI embedding modelsSearch, RAG, clustering
Large concept model (research)Models that predict the next sentence-level embedding ("concept") instead of the next tokenMeta's LCM (2024)Research into language-agnostic reasoning

Distinctions worth knowing:

  • Reasoning model vs chat model: same architecture; the difference is post-training that rewards correct final answers and lets the model spend tokens thinking first.
  • Vision-language training typically goes: pretrained vision encoder and LM → train the projector on image-caption pairs → instruction-tune on visual tasks (LLaVA's recipe); newer models are trained multimodally from the start.
  • Small model vs quantized large model: a small model has fewer parameters; quantization keeps the parameter count and lowers the precision. They trade off differently (see Quantized Inference).
  • Action models vs tool use: any tool-calling LLM can act through tools; computer-use models are additionally trained on screen-and-action data so they can operate interfaces that have no API.

Study Notes

Must-know for interviews:

  • Three architecture families: encoder-only (bidirectional, MLM, classification), decoder-only (causal, CLM, generation), encoder-decoder (cross-attention, seq2seq)
  • Decoder-only dominates production for general LLMs - simple CLM objective + instruction fine-tuning = powerful general purpose model
  • BERT uses bidirectional attention - cannot generate; best for embeddings and classification
  • MoE: sparse routing selects top-K experts per token → large total params, small active params
  • LLaMA-3, Gemma, Mistral all use GQA for efficient KV caching
  • When in doubt: reach for a decoder-only model (e.g. Llama, Qwen, Gemma)

Check Yourself

Check yourself
0 / 5 answered
  1. Why can't BERT generate free text?
  2. Mixtral 8x7B has about 47B parameters, not 56B. Why?
  3. What does a mixture-of-experts model save compared with a dense model of the same total size?
  4. What distinguishes a reasoning model from an ordinary chat model with the same architecture?
  5. You need fast, cheap intent classification for 50 million support messages a day. Which architecture class do you start with, and why?

Exercises

Exercise - Choose the architecture

For each task pick encoder-only, decoder-only or encoder-decoder and a size class, and justify it: (a) deduplicating 20M product titles by semantic similarity; (b) a customer-facing assistant that answers policy questions; (c) translating 1M short UI strings into 30 languages on a fixed budget; (d) tagging personal data in documents at the token level.

Solution

(a) Encoder embedding model - fixed vectors plus approximate nearest-neighbour search; no generation. (b) Decoder-only instruct model with RAG - open-ended generation grounded in documents. (c) Either a small encoder-decoder translation model (cheap, strong for a fixed task) or a small decoder-only model; measure quality per dollar on a sample. (d) Encoder token classifier (NER-style) - per-token labels with bidirectional context, fast and cheap.

Exercise - Total vs active parameters

A MoE layer has 64 experts, each an FFN with d_model = 4096 and d_ffn = 1408 (SwiGLU), and routes each token to 6 experts plus 2 always-on shared experts. Compute the FFN parameters per layer in total and active per token.

Solution

Each expert: 3 × 4096 × 1408 ≈ 17.3M parameters. Total per layer: (64 + 2) × 17.3M ≈ 1.14B. Active per token: (6 + 2) × 17.3M ≈ 138M - about 12% of the layer's FFN parameters, which is the compute saving; memory must still hold all 1.14B.

References

Last reviewed: 2026-09

⚡AI-assisted content - always verify, always explore multiple perspectives·