# digest-finetune

View on GitHub

# Postmortem (2026-08-25): dead end, not pursuing further

GRPO on digest-sft3 regressed rather than improved it (0.6844 vs sft3’s 0.8700, matched sampled eval — see Weights) even after the reward had been fuzzed, calibrated, and hardened against every exploit found across two separate RL runs. Converting sft3 to GGUF and running best-of-N against it with reward.py as a verifier, wired into git-digest, still wasn’t enough on repos outside the training set.

That’s the actual finding: a rising reward score was never proof the task was solvable at this size. Writing a digest from commits + diffs requires inferring what the rest of the repo does — what it’s for, how the pieces relate — from a truncated diff and a commit message. A 135M model can match the vocabulary and structure of the ~100 repos it trained on; it has no spare capacity to generalize that inference to a repo it has never seen, and produces fluent, ungrounded confabulation instead. sft3 is a very good fit to this project’s reward function. That was never the same thing as a model that can write a truthful digest for an arbitrary repo, and no amount of RL or dialing in the reward further was going to close that gap.

Real fix: don’t ask a language model to infer facts it was never shown. git-digest is getting a --static mode (commits/diffs → extracted facts → summary templates, no model, no network beyond GitHub, no hallucination). It generalizes to any repo for free, which the finetuning route structurally cannot.

This repo is frozen as reference, not deleted: the reward-hardening methodology (fuzzing, teacher calibration, grounding-exploit fixes) and the GRPO divergence/abort-guard work below generalize past this specific dead end even though the model they were built around doesn’t belong in git-digest’s release path. The goal was always automated codeblog generation, not finetuning this specific model — static analysis solves that goal, distillation didn’t.


Distill the Claude-written daily digests from git-digest into SmolLM2-135M-Instruct so digest writing runs offline on CPU — eventually bundled into git-digest itself.

# Goals

  1. SFT SmolLM2-135M on ~100 real (commits.json → digest.md) pairs to learn format + voice.
  2. GRPO/RLVR pass using verifiable rewards only (no reward model): format, repo-grounding, coverage, penalties — implemented in scripts/reward.py.
  3. Release model (Apache-2.0 weights) + dataset + MIT training code to HF.
  4. Stretch: quantize to GGUF and embed in git-digest as a local backend; retire cloud path.

# Progress

  • Dataset — build_dataset.py parses ~/Documents/portfolio/digests into chat-format jsonl: 84 real train / 10 eval. Prompts now mirror main.go’s format exactly, including file stats and truncated patches (see Known issues). augment_dataset.py adds synthetic examples — real days merged into combined high-commit-count windows, completions generated by the actual teacher (claude -p --model claude-haiku-4-5-20251001, same backend git-digest uses), scored with reward.py and spot-checked before merging. Regenerated once diffs were backfilled (the prior synthetic batch had stale message-only-informed completions); of 18 generated, 3 violated the “one ### repo section per repo” instruction on very high-repo-count merge days (12-13 repos) — two silently dropped trivial single-commit repos, one consolidated 7 of them into a non-conforming ### Others section — dropped. Train set: 99 examples (84 real + 15 verified synthetic). Covered by tests (pytest tests/).
  • SFT — full fp32 fine-tune, TRL SFTTrainer, 6 epochs on 6 CPU threads, on the 101-example set (checkpoints/sft2). Per-step checkpointing (--save-steps 1 --save-only-model, ~514MB/checkpoint instead of 1.6GB) so every step survives for reward-based selection. Loss 3.90 → 1.99 (min at step 41), then spiked to 2.47 at the final step — same instability pattern as the first run; last-step checkpoints should not be trusted blindly.
  • Reward module — reward.py: parse + score digests on format/coverage/grounding with banned-phrase and duplicate-section penalties. 25 tests green, including adversarial cases for known gaming patterns (see Known issues). Grounding now draws from file stats + diffs, not just commit messages.
  • Checkpoint selection — evaluate.py is expensive (~5-8 min/generation × 10 eval days per checkpoint), so a full sweep over all 42 saved steps isn’t viable. Filtered to loss ≤2.2 (11 candidates), thinned to every 3rd (steps 23/31/35/40), evaluated those. checkpoint-40 won (mean reward 0.499) and is now promoted to checkpoints/sft2 directly; the other 41 checkpoints were deleted after selection.
  • Eval result (checkpoint-40, greedy + rep-penalty 1.08): mean total 0.499 over 10 held-out days, up from the 0.304 baseline (checkpoint-33, pre-augmentation). The original failure mode — total generation collapse, EOS after ~9 tokens on high-commit-count days — is fixed; those days now produce full, fluent prose.

# Known issues

  • Checkpoint selection must use best mean reward on a filtered/thinned shortlist, not lowest loss or last step — the final-step loss spike recurred on this run too.
  • Greedy decode loops without repetition penalty ≥1.05.
  • Collapse is fixed, but a narrower failure remains: on complex (15+ commit) and sparse (1-commit) days the model sometimes skips the required ### repo structure and fills in plausible-sounding but ungrounded detail instead of failing outright. Format + grounding failure, not collapse — this is the primary GRPO target now.
  • First GRPO run (60 steps, T4 Colab) collapsed: final checkpoint looped identical commit-sha sections verbatim instead of improving. Cause: num_generations=4 made group-relative advantage noisy enough that a degenerate completion could still rank “best of group”, and format_score’s binary all-or-nothing gate combined with max_completion_length=450 truncating 50-88% of rollouts gave a spiky, unreliable reward signal. Fixed in scripts/reward.py (graded format score, duplicate-section penalty) and scripts/grpo.py (gens=8, beta=0.05, max_completion=768, mask_truncated_completions=True) — retry pending.
  • checkpoints/ and data/ are gitignored — moving to another machine (e.g. the homelab LXC for GRPO) needs an explicit rsync/scp of both, git clone alone won’t carry them.
  • Root cause found for the GRPO collapse (bigger than the reward-formula fixes above): the dataset’s reconstructed prompts were commit-message-only, but production git-digest sends file stats for every commit and truncated patches on sparse-commit days (main.go: sparseCommitThreshold=5, maxPatchLines=5). The historical *-commits.json archives never persisted that data (writeCommits only wrote sha/message/url), so training labels (written by the real teacher, which did see diffs) contained real, verifiable detail the reconstructed prompt never showed the model — a genuine train/inference mismatch, not hallucination. Confirmed directly: an “init”-only commit’s teacher completion named Svelte 5, three.js, and a hand-rolled raycaster — all pulled verbatim from the initial commit’s README diff. scripts/backfill_diffs.py replays git-digest’s own GET /repos/{repo}/commits/{sha} call to recover this for every commit already in the dataset; build_prompt() and reward.py’s source_tokens now mirror production’s stat-line/patch-inclusion logic exactly (scripts/digest_format.py holds the shared constants so they can’t drift apart). Rescoring the real dataset after backfill: mean reward 0.918 → 0.979, zero rows below 0.8 (previously 8 rows below 0.7, one at 0.30 — all sparse/“init”/merge-commit days with no lexical ground truth). sft2/grpo checkpoints predate this fix and were trained on message-only prompts — do not resume GRPO from them; rerun SFT on the diff-aware dataset first (colab/digest_sft_colab.ipynb).
  • Second bug this surfaced: sft.py was training on the whole prompt, not just the completion. It flattened prompt+completion into one "text" field, which TRL’s SFTTrainer treats as plain language modeling — loss over the entire sequence, no masking. That was mostly harmless when prompts were short (message-only), but with diff-aware prompts running up to ~3.5k tokens against ~300-600 token completions, most of every gradient step went to predicting diff/JSON syntax instead of digest text. Caught by comparing a live Colab run’s loss curve against sft2.log: it tracked better than the old run through epoch ~2, then plateaued around 2.7-2.9 while the old run kept dropping to ~2.1-2.4 by epoch 3.3-3.6 — consistent with the model quickly fitting the easy, now-dominant prompt tokens and further completion-quality gains getting diluted into invisibility in the aggregate loss. Fixed: sft.py now passes native TRL "prompt"/"completion" columns (prompt as a conversational message list, so the chat template + completion_only_loss auto-enable) instead of a flat "text" field. Verified directly on a real example: 900/1196 tokens (75%) were prompt and are now correctly masked from the loss. This affected every prior SFT run, including the one behind sft2/checkpoint-40 — its loss curve likely understates how well it actually fit the completions, though its message-only prompts were short enough that the effect was much smaller.
  • Fixed one real bug the backfill surfaced: a commit touching a minified dist/ bundle produced a single ~60k-character diff line, which line-count truncation (maxPatchLines) doesn’t bound. digest_format.is_low_signal_file() now excludes generated/build-output paths and lockfiles from patch content entirely (kept for reward.py too, so grounding isn’t gamed by naming generic lockfile/license vocabulary), plus a per-line character cap as defense in depth.
  • Reward hardening this session (adversarial testing against synthetic exploits, not just the real dataset): closed empty-section-body, nonsense/SHA-copy filler, and single-real-word padding exploits in coverage_score (now requires min_tokens=3 content words and min_overlap=2 traceable to the repo’s own commits/diff, up from a bare presence check). Added MANIFEST_STOPWORDS (name, version, private, license, …) so generic package.json/ license boilerplate pulled in by diff-aware grounding can’t be named to fake relevance. Accepted residual risk (deliberate, not fixed): a “2 real keywords + 1 filler word” pattern can still score close to honest prose (~0.85 vs ~0.9-1.0) since coverage is a binary per-repo gate rather than continuous — watch the GRPO spotcheck’s per-component grounding score for drift toward short/sparse bodies as an early warning.
  • GRPO on a T4 is bounded by sequence length, not batch size or step count (2026-08-24 session, three OOMs). SDPA falls back to the math path under GRPO’s completion mask, so it materialises the full (batch, heads, seq, seq) score matrix. Gradient checkpointing does not help — it bounds how many layers are held, not how big one layer is. The fatal run died in backward() at step 30 asking for exactly 8 x 9 x 3654^2 x 4 bytes = 3.58 GiB on the 3014-token prompt. PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True did not help and the traceback proves why: only 157 MiB was reserved-but-unallocated, so there was nothing to defragment. grpo.py now runs a CPU pre-flight (--dry-run) that computes this spike from the tokenised dataset, drops rows over the cap, and refuses to start if that exceeds a quarter of them. At batch 8 / max_completion=768 the cap is 1596 prompt tokens and keeps 89/99 rows; max_completion=512 keeps 94/99. Batch 16 (--prompts-per-step 2) is rejected outright — it OOMed on the very first step.
  • save_strategy="no" cost 29 steps of good training. The crash killed the process before trainer.save_model() ran, so a run that had trained cleanly for 30 steps produced nothing. Periodic saving is back as --save-steps (default 10, weights only, save_total_limit=2). This does not revive reward-based checkpoint selection, which was removed for good reasons — it is only about surviving a crash inside a fixed GPU window.
  • lr=1e-5 does not move this policy in a run that fits the window. kl sat at 0.0031-0.0041 across 29 steps — the policy stayed essentially identical to digest-sft3, so even a clean 40-step run could not have produced a measurable delta. The conservative LR was chosen because the first GRPO run collapsed at a higher rate, but that run had 4 generations, a binary format gate and 50-88% truncation; none of that describes the current setup (reward_std 0.23, zero truncation, beta=0.1). Note the direction: kl went 0.0041 -> 0.0028, it declined. A too-small LR would show KL climbing slowly; flat-to-declining KL is the signature of the KL penalty holding an equilibrium, so beta is as likely the binding constraint as lr. The notebook raises LR to 3e-5 first because beta=0.1 still bounds the damage — if kl pins near 0.003 at that rate too, beta is confirmed and comes down next. One variable at a time.
  • The reward saturates rather than being gamed. fuzz_reward.py passes clean (317 adversarial inputs, 0 violations) and the teacher’s own labels average 0.978, so 1.000 spotchecks are the ceiling being reachable, not the scorer being exploited. But frac_reward_zero_std hit 1 on 2 of the first 10 steps: on easy days all 8 rollouts score exactly 1.000 and the group contributes no gradient. Raising the ceiling (uncapped grounding, activity-weighted coverage, or teacher-relative targets) is worth more than making the scorer harsher.
  • Rollout spread on the training set is healthy and was never the problem: probing the five hardest rows at GRPO’s own sampling settings gave mean within-group std 0.247 with 0/5 flat groups, against 0.109 and 1/3 on the held-out rows (scripts/probe_rollouts.py).

# Reward hardening (pre-GRPO session)

Before the next training run the reward went through adversarial fuzzing (scripts/fuzz_reward.py) plus teacher calibration (scripts/calibrate_reward.py). Four real holes found by mutation-testing the scorer against all 124 dataset rows:

  • Summary was a free fabrication channel — penalties never scanned it and format only counted words. A hallucinated filename in the summary cost nothing (1.000 → 1.000); swapping the whole summary for unrelated prose cost nothing (0.875 → 0.875). Fixed: summary must now trace ≥2 content tokens to the prompt’s own material, and fabricated-file scanning includes it.
  • Orphan prose under ## Per-Repo Activity (between the header and the first ###) was invisible to every scorer — parse_sections now emits it as an unnamed pseudo-section so it fails header precision, trips the excess-header penalty, and gets fabricated-file-scanned.
  • Coverage was a binary per-repo gate — the documented residual risk (“2 real keywords + filler”) fired in the wild: randomized soup hit “cleanup”+“scroll” (both real dominion-tracker commit words) and banked full coverage credit for a total of 0.75. Coverage credit is now continuous, saturating at ~4 traced tokens, scaled by body length so tight honest paraphrases keep full credit.
  • Grounding averaged over present sections only — dropping your weakest repo section raised the total (+0.038 on a synthetic 6-repo day): zero-credit sections cost nothing to omit while dragging the mean down. Grounding now averages over every expected repo (missing section = 0), so omission can’t pay. Also fixed a crash on degenerate activity dicts (commits: [{}]) found by garbage-input property checks.

Fuzzer contract: MUST_COST mutations (fabrication, duplication, truncation, hollowing) must each lose ≥0.02 on every row; MUST_NOT_PAY mutations (dropping/shuffling sections) must never gain; format-valid ungrounded soup digests must stay below a ceiling (max observed 0.525 vs teacher mean 0.98); garbage inputs never crash and stay in [0,1].

Calibration after the fixes: teacher means train 0.978 / eval 0.975 / synthetic 0.939, zero rows failing format; two known thin-source days (2026-06-03, 2026-08-15 — merge/scaffold commits with little lexical ground truth) sit at 0.75–0.81 and are honest scores, not miscalibration. Reward totals from before this session are not comparable (formula changed) — sft2’s 0.499 baseline and older eval logs don’t carry over; re-baseline any model against the new scorer before comparing.

# Run guards (post-hardening session)

Fuzzing the reward found four real holes, but a fifth check — a minimal-effort digest with 3 content tokens per section and 2 traceable — came back already covered, scoring 0.60 against a teacher mean of 0.978 and a fuzzer-built soup ceiling of 0.525. That margin is wide enough to train against, so the wasted runs were traced to the training scripts instead, and the effort moved there.

Periodic checkpointing and reward-based checkpoint selection are both gone. Every sweep so far cost 10 generations per candidate and landed on a checkpoint indistinguishable from the last step — the selection machinery consumed more GPU time than the training it was selecting from. sft.py and grpo.py now write only the final model (save_strategy="no"), and both notebooks compare exactly two things: the untrained base (or SFT baseline) against the trained result.

That also deleted the sft6 footgun rather than guarding it. sft.py used to auto-resume whenever --out happened to contain a checkpoint-* dir, which is what ruined that run: re-running the Colab train cell after an interrupt picked up a stale checkpoint, replayed the LR warmup at epoch 8.08 (0 → 2.7e-5 → 5.3e-5 → 8e-5), and mixed the new args with the old trainer_state.json (TRL warned save_steps: 25 (from args) != 3). The cosine schedule never decayed below 6.2e-5 and loss sat flat at ~2.6 with token accuracy ~0.53 for four epochs. With no periodic checkpoints there is nothing to resume from and nothing to resume into, so the whole hazard is gone. Trade-off accepted: an interrupted run restarts from scratch.

grpo.py can now stop itself. abort_reason() judges the last --abort-patience spotchecks (default 3) and halts the run on any of:

  • training reward gaining ≥0.05 while the held-out spotcheck total drops by ≥0.10 — fitting the scorer rather than the task, which is the failure the reward cannot self-detect
  • mean completion length dropping below half the reference digest’s — measured against each day’s own label, since a quiet day is legitimately short and a cross-day peak conflated the two
  • more than half of rollouts truncating at --max-completion

The guard’s own false-positive rate is measured, not assumed. grpo3 died at step 30 on spotchecks of 1.0, 1.0, 0.05 — three different days at one rollout each. Pinning the day and averaging --spotcheck-gens rollouts fixed the input; the rule still read “held-out failed to gain” as divergence, which a day the policy already scores ~0.85 on can never satisfy. Simulating a flat policy from sft3’s own day-0 rollouts (draws 0.7/1.0) against grpo3’s per-step training rewards: <= 0 aborted 59% of 6-tick runs (66% at one rollout per tick), the ≥0.10 band 16%, while a genuine pinned-day collapse (0.95 → 0.30) is caught 99.5% of the time either way. The band costs no detection. Note the remaining 16% is dominated by window count — --spotcheck-every 10 over 60 steps gives four overlapping windows; at 20 the same guard sits at 3%, but with --abort-patience 3 its only window closes at the final step, which disables it rather than fixing it.

An aborted run still saves its policy (save_model runs after train() returns), so the base-vs-result comparison works either way. --beta also moved 0.05 → 0.1: against a lexical, gameable reward the risk is drift from the SFT policy, and tighter KL is cheaper insurance than another penalty term in reward.py.

Both Colab notebooks now run pytest, fuzz_reward.py, and calibrate_reward.py as a CPU pre-flight before touching the GPU, and carry their hyperparameters as named constants so the notebook records what actually ran.

# Weights

  • usr-wwelsh/digest-sft2 — SFT checkpoint, mean reward 0.499 on held-out eval. Stale: trained on message-only prompts, predates the diff-aware dataset fix (see Known issues). Superseded by digest-sft3.
  • usr-wwelsh/digest-sft3 — current release candidate. Diff-aware prompts, hardened reward. 0.8700 mean reward sampled (temperature=0.8, matched 10-day held-out set, current reward.py).
  • usr-wwelsh/digest-grpo — dropped, not in the release path. main is a GRPO pass on top of digest-sft3 that advertised +0.0125 over its baseline; re-evaled under the fixed reward (sampled, matched settings) it scores 0.6844 against sft3’s 0.8700 — a regression, not a gain, once the reward hole it was partly exploiting got closed. Kept published for the record, not for use. The earlier revision 1cb58df (GRPO on top of the now-superseded digest-sft2, which collapsed into looping identical commit-sha sections) is unrelated and also not in use.

# Next steps

  • grpo.py plumbing + laptop smoke test (2 prompts × 2 rollouts × 3 steps) — hit and fixed a crash (save_safetensors invalid for installed trl==1.10.0
  • Backfill file stats + patches into the dataset, rebuild prompts to match production, make reward.py
  • Fuzz the reward (mutation + property tiers) and calibrate against teacher completions;
  • Run colab/digest_sft_colab.ipynb — SFT on the diff-aware dataset, reward-based checkpoint selection against both the untrained base and old sft2, pushed as
  • Point colab/digest_grpo_colab.ipynb’s BASE_MODEL at digest-sft3, rerun GRPO, pushed as digest-grpo (+0.0125 over the sft3 baseline, greedy eval) — see open
  • Closed a grounding exploit: overlap was scored against a repo’s whole commit pool, so real vocabulary from unrelated commits could be recombined into an unsupported claim on busy days and still clear the bar. Now scored against the best single
  • digest_live.py (replays production prompt-building against a real day’s commits) and compare_teacher.py (scores local checkpoints against the live teacher on
  • GRPO dropped. Re-evaled sft3 vs grpo under the fixed reward, sampled (temperature=0.8) and matched on the full 10-day held-out set: sft3 0.8700 vs grpo 0.6844 — GRPO scores worse than the checkpoint it trained on top of, with 0% truncation on both sides so it isn’t a decoding artifact. The +0.0125 it advertised was measured against the reward hole above; closing that hole exposed a real regression, not noise. Not rerunning GRPO against the fixed reward either: this is the second run on this project to score well against a reward that had a real hole in it, which points at RL finding shortcuts faster
  • Ship digest-sft3 (temperature=0.8) + best-of-N sampling with reward.py as verifier — pipe into git-digest

# Usage

uv venv --python 3.12 .venv
uv pip install --python .venv/bin/python torch --index-url https://download.pytorch.org/whl/cpu
uv pip install --python .venv/bin/python transformers trl datasets accelerate pytest

.venv/bin/python scripts/build_dataset.py          # rebuild data/ from portfolio/digests
.venv/bin/python scripts/augment_dataset.py --n 18 # synthetic long-window examples (real teacher calls)
.venv/bin/python -m pytest tests/ -q               # unit tests (parser + rewards)
.venv/bin/python scripts/fuzz_reward.py            # adversarial fuzz: degradations must drop the score, garbage must not crash
.venv/bin/python scripts/calibrate_reward.py       # teacher completions must sit near the reward ceiling
# SFT locally on CPU: ~12s/example at batch 1 on an i7-6700HQ (8 threads), so ~20min/epoch
nohup .venv/bin/python scripts/sft.py --epochs 6 --lr 8e-5 --out checkpoints/sft3 > sft3.log 2>&1 &
.venv/bin/python scripts/grpo.py --model checkpoints/sft3 --spotcheck-every 10 --abort-patience 3
.venv/bin/python scripts/eval_reward.py --model checkpoints/sft3 --tokenizer HuggingFaceTB/SmolLM2-135M-Instruct --n 10
.venv/bin/python scripts/generate.py --model checkpoints/sft2 --file DATE-1d.md --repeat-penalty 1.08   # eyeball one day
.venv/bin/python scripts/generate.py --model checkpoints/sft2 --repeat-penalty 1.08 --prompt "..."       # freeform prompt
.venv/bin/python scripts/evaluate.py --model checkpoints/sft2      # scored sweep over eval set

🔒 This site's search assistant runs a local AI model in your browser — no cloud AI, no tracking. See the ℹ️ in the chat panel for details.