Contents
Map

01 ยท LLM Foundations

Model Landscape

View as:

Model Landscape

New frontier models now ship every few weeks, so any list of "the best models" is stale within a quarter. This page gives you a durable way to read the landscape - the axes that actually separate models - plus a dated snapshot with a link to each vendor's own model page. It is the only page in the course that names current model versions; everywhere else, notes refer to model families so they don't go out of date.

Learning objectives 30 min
By the end of this page you will be able to:
  • Place any model on the axes that matter - access and license, dense vs MoE, reasoning mode, modality, context, cost
  • Explain the main 2025-2026 trends (sparse MoE, adaptive "thinking", 1M-token context, the open/closed gap) with concrete examples
  • Choose a shortlist of models for a use case, and design the evaluation that makes the final call
  • Keep this knowledge current from primary sources rather than leaderboard screenshots
Prerequisites

The Axes That Matter

mindmap
  root((๐Ÿ—บ๏ธ An LLM))
    ๐Ÿ” Access
      Closed API
      Open weights
      License terms
    ๐Ÿงฎ Shape
      Dense
      Mixture of experts
      Total vs active params
    ๐Ÿง  Reasoning mode
      Non-thinking
      Always thinking
      Hybrid / adaptive with effort control
    ๐Ÿ‘๏ธ Modality
      Text
      Vision in
      Audio in / out
      Image out
    ๐Ÿ“ Context
      128K
      1M+
    ๐Ÿ’ฐ Cost and speed
      Price per million tokens
      Tokens per second
      Time to first token
AxisWhat to checkWhy it matters
Access and licenseClosed API only? Open weights? Apache-2.0 / MIT, or a custom license with usage limits (e.g. Llama's community license)?Decides whether you can self-host, fine-tune, run on-prem, or ship in a product
ShapeDense or MoE; total and active parametersActive params drive cost per token; total params drive the GPU memory you need to self-host
Reasoning modeDoes the model think before answering? Can you switch it off or set an effort level?Thinking improves hard math, code and multi-step tasks, and costs latency and output tokens
ModalityWhich inputs and outputs are nativeNative audio/vision avoids a pipeline of separate models
ContextAdvertised window, and measured quality near the end of itAdvertised length โ‰  usable length - see Failure Modes
Cost and speedInput/output price, cached-input price, throughput, latencyOften the deciding factor once two models are "good enough"

1. Reasoning became a dial, not a separate model. In 2024 reasoning models (OpenAI o1, then DeepSeek-R1) were separate products. By late 2025 most flagships were hybrid: DeepSeek-V3.1 switches between thinking and non-thinking modes, Qwen3 exposes the same toggle, and API models expose an effort setting. Current Claude models, for example, use adaptive thinking - the model decides how much to think, bounded by an effort level from low to max - rather than a fixed thinking-token budget. The training recipe behind this is covered in Post-Training & Reasoning.

2. Sparse MoE became the default for large open models. DeepSeek-V3/R1 (671B total, 37B active), Qwen3-235B-A22B, Llama 4 Maverick (400B / 17B), gpt-oss-120b (117B / 5.1B) and Kimi K2 (~1T / 32B) are all mixtures of experts. Dense models remain common below ~30B parameters, where they are simpler to serve.

3. Million-token context became standard at the frontier. Current Claude, GPT and Gemini flagships advertise ~1M-token windows; Llama 4 Scout advertises 10M. Measured quality near the end of the window still lags the advertised number, and long prompts are expensive without prompt caching.

4. The open-weight gap narrowed to months. DeepSeek-R1 (January 2025, MIT license) matched the then-leading closed reasoning model on math and code benchmarks; through 2025-2026, open releases from DeepSeek, Qwen, Moonshot (Kimi) and Z.ai (GLM) repeatedly landed within a few points of closed frontier models on coding and agentic benchmarks. OpenAI's return to open weights with gpt-oss (August 2025, Apache-2.0) is part of the same shift.

5. Small models got much better, mostly through distillation. Sub-10B models trained on the outputs of large reasoning models (e.g. the DeepSeek-R1 distilled Qwen and Llama checkpoints, Qwen3's small dense models) now handle tasks that needed a frontier model two years earlier - which makes on-device and low-cost serving realistic.


Snapshot: Major Families (as of 2026-09-26)

Names and versions change quickly - treat this table as a starting point and follow the source link for the current lineup.

DeveloperCurrent flagship family (per source)AccessNotesSource
OpenAIGPT-6 (Astra, Sol, Luna tiers)Closed API~1M-token context; open-weight gpt-oss-120b / 20b (Aug 2025, Apache-2.0)OpenAI models
AnthropicClaude Fable 5.1 (most capable), Claude Opus 5 / Opus 5.5, Sonnet 5, Haiku 4.5Closed API; also on AWS, Google Cloud, Microsoft Foundry1M-token context (Haiku 4.5: 200K); adaptive thinking with effort levelsClaude models
GoogleGemini 3.x (e.g. 3.8 Flash stable, 3.1 Pro preview)Closed API; open Gemma familyNative audio/vision; separate Live (voice) and image modelsGemini models
DeepSeekDeepSeek-V4 (deepseek-v4-pro, deepseek-flash)API + open weightsMLA + fine-grained MoE lineage from V2/V3; R1 (Jan 2025) popularized open reasoningDeepSeek API docs
Alibaba (Qwen)Qwen3 family and successorsOpen weights (mostly Apache-2.0) + APIDense 0.6B-32B and MoE up to 235B-A22B at the Qwen3 launch (Apr 2025); hybrid thinkingQwen on Hugging Face
MetaLlama 4 (Scout, Maverick)Open weights, Llama 4 Community LicenseNatively multimodal MoE; Scout advertises 10M contextLlama
Moonshot AIKimi K2 lineOpen weights (modified MIT) + API~1T-parameter MoE, 32B active, MLA; trained with the MuonClip optimizerKimi K2 paper
Z.ai (Zhipu)GLM-4.5 / 4.6 lineOpen weights (MIT for GLM-4.5) + API355B MoE (and a 106B "Air" variant) aimed at agentic codingGLM-4.5 on Hugging Face

Reference Open-Weight Architectures

These are stable, published facts - good anchors for the architecture notes and for self-hosting estimates.

ModelReleasedTotal / active paramsAttentionContextLicense
DeepSeek-V3 / R1Dec 2024 / Jan 2025671B / 37BMLA128KMIT (R1); V3 weights under a model license
Qwen3-235B-A22BApr 2025235B / 22BGQA + QK-norm32K native, 128K with YaRNApache-2.0
Qwen3-32B (dense)Apr 202532B / 32BGQA + QK-norm32K native, 128K with YaRNApache-2.0
Llama 4 MaverickApr 2025400B / 17BGQA, iRoPE1MLlama 4 Community License
Llama 4 ScoutApr 2025109B / 17BGQA, iRoPE10MLlama 4 Community License
Gemma 3 27BMar 202527B denseGQA, 5:1 local:global128KGemma terms of use
gpt-oss-120bAug 2025117B / 5.1BGQA, alternating banded window128KApache-2.0
gpt-oss-20bAug 202521B / 3.6BGQA, alternating banded window128KApache-2.0
Kimi K2Jul 2025~1T / 32BMLA128KModified MIT
GLM-4.5Jul 2025355B / 32BGQA128KMIT

Choosing a Model

flowchart TD
    Q1{"๐Ÿ” Must data stay on your<br/>infrastructure, or must you fine-tune weights?"}
    Q1 -->|Yes| OW["๐Ÿ“ฆ Shortlist open-weight models<br/>that fit your GPUs"]
    Q1 -->|No| Q2{"๐Ÿง  Hard reasoning, long agentic tasks,<br/>or complex code?"}
    Q2 -->|Yes| FR["๐Ÿš€ Frontier tier with thinking on<br/>tune effort per route"]
    Q2 -->|No| Q3{"โšก High volume or<br/>latency-critical?"}
    Q3 -->|Yes| SM["๐Ÿ’จ Small / fast tier<br/>(or distilled open model)"]
    Q3 -->|No| MID["โš–๏ธ Mid tier"]
    OW --> EV["๐Ÿงช Evaluate the shortlist<br/>on YOUR task set"]
    FR --> EV
    SM --> EV
    MID --> EV
    EV --> DEC["โœ… Pick on quality, cost per<br/>completed task, and latency"]

    style Q1 fill:#e8e0d4,stroke:#c8b89a
    style Q2 fill:#e8e0d4,stroke:#c8b89a
    style Q3 fill:#e8e0d4,stroke:#c8b89a
    style EV fill:#dde4dc,stroke:#b0c4b0
    style DEC fill:#d8dfe8,stroke:#b0bac8

Three habits separate good model choices from benchmark-chasing:

  1. Measure cost per completed task, not per token. A cheaper model that needs three retries, or a thinking model that solves it in one call, changes the maths.
  2. Evaluate on your own data. Public benchmarks are saturated or contaminated for many tasks - see Evaluation & Benchmarks.
  3. Try the strong model at low effort before building a cascade. A frontier model at reduced reasoning effort is often cheaper and simpler than routing between several models.

Staying Current

SourceUse it for
Vendor model pages (linked in the snapshot)Exact model IDs, context limits, prices, deprecation dates
Technical reports and model cards (arXiv, Hugging Face)Architecture, training data, evaluation methodology
Artificial AnalysisIndependent speed, price and aggregate-intelligence comparisons
LMArenaHuman-preference rankings (style-sensitive; see the Evaluation module)
Epoch AITraining-compute trends and benchmark tracking

Check Yourself

Check yourself
0 / 3 answered
  1. Two models score within a point of each other on your evaluation. Model A costs half as much per token but needs, on average, 2.5 attempts to finish a task; Model B finishes in one. Which is cheaper per completed task?
  2. You must self-host on 8ร— 80 GB GPUs. Which fact about an MoE model decides whether it fits?
  3. Why does this course keep current model names on one page instead of in every note?

Exercises

Exercise - Build a shortlist

You are choosing a model for an internal code-review assistant: ~5,000 pull requests a day, diffs up to 30K tokens, source code must not leave the company's cloud account.

  1. Which axes rule models in or out before any benchmark?
  2. Produce a shortlist of three (at least one open-weight), using the vendor pages linked above.
  3. Describe the evaluation you would run to choose between them.
Solution
  1. Data residency rules out consumer APIs but not cloud-hosted ones (Bedrock, Vertex AI and Microsoft Foundry serve closed models inside your cloud account) and allows self-hosted open weights. Context must comfortably exceed 30K tokens; volume makes price and cached-input pricing important.
  2. Any reasonable mix: a frontier closed model via your cloud provider, a mid-tier closed model, and an open-weight coding-strong MoE sized to your GPUs.
  3. Build a set of ~100-200 real historical PRs with known issues (bugs caught later, reviewer comments), score each model's findings for precision and recall with a rubric (and spot-check with humans), and compare cost and latency per review at the effort level you would run in production.

Study Notes

Must-know:

  • Read models on six axes: access/license, shape (dense vs MoE, total vs active), reasoning mode, modality, context, cost/speed
  • 2025-2026 trends: reasoning as an effort dial, sparse MoE for large open models, ~1M-token context at the frontier, open weights within months of closed, strong distilled small models
  • Self-hosting memory is driven by total parameters; per-token compute by active parameters
  • Choose on cost per completed task and on your own evaluation, not on leaderboard rank

References

Last reviewed: 2026-09-26

โšกAI-assisted content - always verify, always explore multiple perspectivesยท