Contents
Map

02 · Prog Langs

Git Workflows for ML

View as:

Git Workflows for ML

Git is how a team keeps code, configuration and prompts consistent, reviewed and recoverable. ML and LLM projects add problems ordinary software doesn't have: large files that don't belong in Git, notebooks that produce noisy diffs, results that are only meaningful if you can say exactly which code, data and settings produced them, and prompts that change behaviour as much as code does. This note covers a working Git setup for that: what to version where, branching and reviews, pre-commit checks, large files with LFS, DVC and the Hugging Face Hub, reproducible runs, and recovering from leaked secrets.

Learning objectives 40 min
By the end of this page you will be able to:
  • Decide what belongs in Git, Git LFS, DVC or a model registry, and set up .gitignore accordingly
  • Run a trunk-based workflow with short-lived branches, pull requests and CI checks, including eval gates for prompt changes
  • Configure pre-commit hooks for formatting, notebooks, large files and secrets
  • Tie every experiment to a commit, data version, config and environment so it can be reproduced
  • Respond correctly to a secret committed to a repository, and use git bisect to find a regression
Prerequisites
  • Basic Git: clone, add, commit, push, branch

What Goes Where

ArtifactWhereWhy
Code, configs, prompts, small eval sets, lockfiles, Dockerfiles, IaCGitSmall, text, needs review and diffs
NotebooksGit, outputs strippedOutputs bloat diffs and can leak data
Medium binary files (images, small models) in a code repoGit LFSKeeps pointers in Git, contents in LFS storage
Datasets and pipeline outputsDVC (or a data lake with versioned snapshots)Versions large data in object storage, tied to Git commits
Model weights and checkpointsModel registry or the Hugging Face HubLarge, with lifecycle stages and access control (Model Lifecycle & Rollout)
Secrets (API keys, tokens)Secret manager - never GitGit history is permanent and copied to every clone

A starting .gitignore for an LLM project excludes .venv/, __pycache__/, .env, data/ (tracked by DVC instead), checkpoints/, wandb/ and mlruns/, and large model files such as *.safetensors and *.gguf.


Branches, Pull Requests and CI

gitGraph
    commit id: "main: eval baseline"
    branch prompt-v7
    checkout prompt-v7
    commit id: "tighten refusal rules"
    commit id: "add 12 eval cases"
    checkout main
    branch retriever-rerank
    commit id: "add reranker"
    checkout main
    merge prompt-v7 id: "PR #41: eval gate passed"
    checkout retriever-rerank
    commit id: "tune top-k"
    checkout main
    merge retriever-rerank id: "PR #42"
    commit id: "tag v1.4.0" tag: "v1.4.0"

Trunk-based development works well for ML teams: short-lived branches (hours to a few days) merged into main through pull requests, with main always deployable. Long-lived experiment branches drift and become unmergeable; run experiments through configuration and flags on main instead.

Pull request checks for an LLM application:

  • Lint, format, type check and unit tests on every PR (Python for AI Engineering)
  • A fast eval subset whenever prompts, model versions, retrieval settings or tool definitions change - prompts are code, and a prompt edit needs the same review and gate as a code edit (Prompts in Production)
  • Infrastructure plans posted to the PR for IaC changes (Infrastructure as Code)

Commits: small, one logical change each, with messages that say why. Squash-merging PRs keeps main's history readable, and tags mark releases.


Pre-commit Hooks

The pre-commit framework runs checks on staged files before each commit, and the same config runs in CI:

# .pre-commit-config.yaml
repos:
  - repo: https://github.com/astral-sh/ruff-pre-commit
    rev: v0.16.10
    hooks:
      - id: ruff-check
        args: [--fix]
      - id: ruff-format
  - repo: https://github.com/pre-commit/pre-commit-hooks
    rev: v6.0.0
    hooks:
      - id: check-added-large-files    # blocks accidental multi-MB commits
        args: [--maxkb=1000]
      - id: detect-private-key
      - id: check-merge-conflict
      - id: check-yaml
  - repo: https://github.com/kynan/nbstripout
    rev: 0.9.1
    hooks:
      - id: nbstripout                 # strip notebook outputs and execution counts
  - repo: https://github.com/gitleaks/gitleaks
    rev: v8.30.1
    hooks:
      - id: gitleaks                   # scan staged changes for API keys and tokens

Install once per clone with pre-commit install; run on everything with pre-commit run --all-files. Pin rev values (the ones above were current in October 2026) and update them deliberately with pre-commit autoupdate.


Large Files: LFS, DVC and the Hub

  • Git LFS replaces large files in Git with small pointer files and stores the contents on an LFS server. Simple, but every clone that needs the files downloads them, and LFS storage on hosted Git is metered.
  • DVC tracks datasets and pipeline outputs: the data lives in object storage (S3, GCS, Azure Blob), and small .dvc files and dvc.lock in Git record exactly which version each commit used. dvc repro reruns only the pipeline stages whose inputs changed.
  • The Hugging Face Hub hosts models and datasets in Git-based repositories, with large files stored on its Xet storage backend (which deduplicates chunks across versions). It is the default distribution channel for open weights; pin a specific revision (commit hash) when you download, not just the repository name.

Reproducible Runs

A result is reproducible only if you can recover everything that produced it:

IngredientRecord
CodeGit commit SHA - and refuse to launch official runs from a dirty working tree
DataDVC version, dataset revision or snapshot id
ConfigurationThe full resolved config (not just overrides), saved with the run
EnvironmentLockfile hash, container image digest, CUDA and driver versions
RandomnessSeeds - knowing that GPU kernels can still be non-deterministic
ModelsBase model id and revision; adapter or checkpoint ids

Experiment trackers record most of this automatically - MLflow, for example, tags each run with the Git commit of the source it ran from. Make the tracker's run id the link between a number in a report and everything behind it.


Finding Regressions with git bisect

When quality or latency dropped "somewhere in the last 40 commits", git bisect binary-searches the history - about 6 test runs for 40 commits instead of 40:

git bisect start
git bisect bad HEAD                 # current version fails
git bisect good v1.3.0              # last known good release
git bisect run ./scripts/eval_gate.sh   # exit 0 = good, 1 = bad, 125 = skip this commit
git bisect reset

With a script that runs a small deterministic eval and exits non-zero below a threshold, bisect finds the commit that broke the behaviour unattended. Keep that eval subset fast and stable for exactly this reason.


When a Secret Is Committed

  1. Revoke and rotate the credential immediately. Assume it is compromised the moment it was pushed - public repositories are scanned for keys within minutes.
  2. Then remove it from history with git filter-repo (the tool the Git documentation recommends over filter-branch), force-push, and ask collaborators to re-clone.
  3. Add a secret-scanning hook (gitleaks, above) and enable your Git host's push protection so it doesn't happen again.

Rewriting history without rotating the key fixes nothing: forks, clones and caches still contain it.


Check Yourself

Check yourself
0 / 5 answered
  1. Where should a 40 GB training dataset used by several experiments be versioned?
  2. An API key was pushed to a public repository an hour ago. What is the first thing to do?
  3. Why should prompt changes go through pull requests with an eval gate?
  4. Quality dropped somewhere in the last 64 commits. Roughly how many eval runs does git bisect need to find the culprit?
  5. Why strip notebook outputs before committing?

Exercises

Exercise - Make a run reproducible

A teammate reports "fine-tuned model scored 71.4% on our eval" in a chat message. List what must be recorded for that number to be reproducible, and how you would make recording it automatic.

Solution

Record: Git commit SHA of the training and eval code (clean tree); base model id and revision; dataset version (DVC or snapshot id) for training and eval sets; the full resolved training and eval configs; lockfile hash and container image digest; hardware, CUDA and driver versions; seeds; the eval harness version and settings (prompt template, shots, decoding); and the checkpoint id. Make it automatic by launching runs through a script that refuses a dirty tree, logs all of this to the experiment tracker, and prints the tracker run id - then report "71.4% (run abc123)" instead of a bare number.

Exercise - Write the bisect script

Write the outline of scripts/eval_gate.sh for git bisect: it should skip commits that don't build, run a 50-item deterministic eval, and fail below 85% accuracy.

Solution
#!/usr/bin/env bash
set -u
uv sync --frozen >/dev/null 2>&1 || exit 125      # can't build this commit: tell bisect to skip it
score=$(uv run python -m evals.run --subset smoke50 --temperature 0 --print-accuracy) || exit 125
python3 -c "import sys; sys.exit(0 if float('$score') >= 0.85 else 1)"

Exit code 125 means "skip"; 0 means good; any other non-zero means bad. Use temperature 0 and a pinned model version so the eval is as deterministic as possible.

Study Notes

Must-know:

  • Git for code, configs, prompts, small eval sets, lockfiles; DVC or snapshots for data; registry or Hub for weights; secret manager for secrets
  • Trunk-based: short-lived branches, PRs, CI checks; eval gates on prompt, model and retrieval changes
  • pre-commit: ruff, large-file guard, private-key and secret scanning (gitleaks), nbstripout; same config in CI; pin and update revs deliberately
  • Reproducibility = commit + data version + resolved config + environment + seeds + model revisions, logged with a tracker run id
  • git bisect with an eval script finds regressions in log2(n) runs
  • Leaked secret: rotate first, then git filter-repo, then prevention

References

Last reviewed: 2026-10

⚡AI-assisted content - always verify, always explore multiple perspectives·