| SFT | Supervised Fine-Tuning: training on examples of desired responses. |
| PEFT | Parameter-Efficient Fine-Tuning: adapting a model with few trainable parameters. |
| LoRA / QLoRA | Low-Rank Adaptation / Quantized LoRA: train low-rank updates / do so on a quantized base. |
| RL / RLHF | Reinforcement Learning / RL from Human Feedback: learn from rewards / human-derived preferences. |
| Reward model | Predicts preference or quality to supply a training signal. |
| DPO | Direct Preference Optimization: learns from preferred/rejected pairs without an explicit RL rollout loop. |
| PPO | Proximal Policy Optimization: constrains policy updates during reward optimization. |
| GRPO | Group Relative Policy Optimization: uses group-relative rewards without a separate value critic. |
| RLVR | Reinforcement Learning with Verifiable Rewards: trains against checkable outcomes. |
| KL penalty | Discourages a policy from drifting away from a reference distribution. |
| CoT / test-time compute | Chain of Thought / inference work spent reasoning, searching or checking. |
| Distillation | Trains a smaller or cheaper model using a stronger model's outputs. |
| Reward hacking / forgetting | Exploiting the reward without solving the task / losing earlier capabilities. |