# digest-finetune
# Postmortem (2026-08-25): dead end, not pursuing further
GRPO on digest-sft3 regressed rather than improved it (0.6844 vs sft3’s 0.8700,
matched sampled eval — see Weights) even after the reward had been fuzzed, calibrated, and
hardened against every exploit found across two separate RL runs. Converting sft3 to GGUF
and running best-of-N against it with reward.py as a verifier, wired into git-digest, still
wasn’t enough on repos outside the training set.
That’s the actual finding: a rising reward score was never proof the task was solvable at this
size. Writing a digest from commits + diffs requires inferring what the rest of the repo does —
what it’s for, how the pieces relate — from a truncated diff and a commit message. A 135M model
can match the vocabulary and structure of the ~100 repos it trained on; it has no spare capacity
to generalize that inference to a repo it has never seen, and produces fluent, ungrounded
confabulation instead. sft3 is a very good fit to this project’s reward function. That was
never the same thing as a model that can write a truthful digest for an arbitrary repo, and no
amount of RL or dialing in the reward further was going to close that gap.
Real fix: don’t ask a language model to infer facts it was never shown. git-digest is getting a
--static mode (commits/diffs → extracted facts → summary templates, no model, no network
beyond GitHub, no hallucination). It generalizes to any repo for free, which the finetuning
route structurally cannot.
This repo is frozen as reference, not deleted: the reward-hardening methodology (fuzzing, teacher calibration, grounding-exploit fixes) and the GRPO divergence/abort-guard work below generalize past this specific dead end even though the model they were built around doesn’t belong in git-digest’s release path. The goal was always automated codeblog generation, not finetuning this specific model — static analysis solves that goal, distillation didn’t.
Distill the Claude-written daily digests from git-digest into SmolLM2-135M-Instruct so digest writing runs offline on CPU — eventually bundled into git-digest itself.
# Goals
- SFT SmolLM2-135M on ~100 real
(commits.json → digest.md)pairs to learn format + voice. - GRPO/RLVR pass using verifiable rewards only (no reward model): format, repo-grounding,
coverage, penalties — implemented in
scripts/reward.py. - Release model (Apache-2.0 weights) + dataset + MIT training code to HF.
- Stretch: quantize to GGUF and embed in git-digest as a local backend; retire cloud path.
# Progress
- Dataset —
build_dataset.pyparses~/Documents/portfolio/digestsinto chat-format jsonl: 84 real train / 10 eval. Prompts now mirror main.go’s format exactly, including file stats and truncated patches (see Known issues).augment_dataset.pyadds synthetic examples — real days merged into combined high-commit-count windows, completions generated by the actual teacher (claude -p --model claude-haiku-4-5-20251001, same backend git-digest uses), scored withreward.pyand spot-checked before merging. Regenerated once diffs were backfilled (the prior synthetic batch had stale message-only-informed completions); of 18 generated, 3 violated the “one### reposection per repo” instruction on very high-repo-count merge days (12-13 repos) — two silently dropped trivial single-commit repos, one consolidated 7 of them into a non-conforming### Otherssection — dropped. Train set: 99 examples (84 real + 15 verified synthetic). Covered by tests (pytest tests/). - SFT — full fp32 fine-tune, TRL
SFTTrainer, 6 epochs on 6 CPU threads, on the 101-example set (checkpoints/sft2). Per-step checkpointing (--save-steps 1 --save-only-model, ~514MB/checkpoint instead of 1.6GB) so every step survives for reward-based selection. Loss 3.90 → 1.99 (min at step 41), then spiked to 2.47 at the final step — same instability pattern as the first run; last-step checkpoints should not be trusted blindly. - Reward module —
reward.py: parse + score digests on format/coverage/grounding with banned-phrase and duplicate-section penalties. 25 tests green, including adversarial cases for known gaming patterns (see Known issues). Grounding now draws from file stats + diffs, not just commit messages. - Checkpoint selection —
evaluate.pyis expensive (~5-8 min/generation × 10 eval days per checkpoint), so a full sweep over all 42 saved steps isn’t viable. Filtered to loss ≤2.2 (11 candidates), thinned to every 3rd (steps 23/31/35/40), evaluated those. checkpoint-40 won (mean reward 0.499) and is now promoted tocheckpoints/sft2directly; the other 41 checkpoints were deleted after selection. - Eval result (checkpoint-40, greedy + rep-penalty 1.08): mean total 0.499 over 10 held-out days, up from the 0.304 baseline (checkpoint-33, pre-augmentation). The original failure mode — total generation collapse, EOS after ~9 tokens on high-commit-count days — is fixed; those days now produce full, fluent prose.
# Known issues
- Checkpoint selection must use best mean reward on a filtered/thinned shortlist, not lowest loss or last step — the final-step loss spike recurred on this run too.
- Greedy decode loops without repetition penalty ≥1.05.
- Collapse is fixed, but a narrower failure remains: on complex (15+ commit) and
sparse (1-commit) days the model sometimes skips the required
### repostructure and fills in plausible-sounding but ungrounded detail instead of failing outright. Format + grounding failure, not collapse — this is the primary GRPO target now. - First GRPO run (60 steps, T4 Colab) collapsed: final checkpoint looped identical
commit-sha sections verbatim instead of improving. Cause:
num_generations=4made group-relative advantage noisy enough that a degenerate completion could still rank “best of group”, andformat_score’s binary all-or-nothing gate combined withmax_completion_length=450truncating 50-88% of rollouts gave a spiky, unreliable reward signal. Fixed inscripts/reward.py(graded format score, duplicate-section penalty) andscripts/grpo.py(gens=8,beta=0.05,max_completion=768,mask_truncated_completions=True) — retry pending. checkpoints/anddata/are gitignored — moving to another machine (e.g. the homelab LXC for GRPO) needs an explicitrsync/scpof both, git clone alone won’t carry them.- Root cause found for the GRPO collapse (bigger than the reward-formula fixes above):
the dataset’s reconstructed prompts were commit-message-only, but production git-digest
sends file stats for every commit and truncated patches on sparse-commit days (
main.go:sparseCommitThreshold=5,maxPatchLines=5). The historical*-commits.jsonarchives never persisted that data (writeCommitsonly wrote sha/message/url), so training labels (written by the real teacher, which did see diffs) contained real, verifiable detail the reconstructed prompt never showed the model — a genuine train/inference mismatch, not hallucination. Confirmed directly: an “init”-only commit’s teacher completion named Svelte 5, three.js, and a hand-rolled raycaster — all pulled verbatim from the initial commit’s README diff.scripts/backfill_diffs.pyreplays git-digest’s ownGET /repos/{repo}/commits/{sha}call to recover this for every commit already in the dataset;build_prompt()andreward.py’ssource_tokensnow mirror production’s stat-line/patch-inclusion logic exactly (scripts/digest_format.pyholds the shared constants so they can’t drift apart). Rescoring the real dataset after backfill: mean reward 0.918 → 0.979, zero rows below 0.8 (previously 8 rows below 0.7, one at 0.30 — all sparse/“init”/merge-commit days with no lexical ground truth).sft2/grpocheckpoints predate this fix and were trained on message-only prompts — do not resume GRPO from them; rerun SFT on the diff-aware dataset first (colab/digest_sft_colab.ipynb). - Second bug this surfaced:
sft.pywas training on the whole prompt, not just the completion. It flattened prompt+completion into one"text"field, which TRL’s SFTTrainer treats as plain language modeling — loss over the entire sequence, no masking. That was mostly harmless when prompts were short (message-only), but with diff-aware prompts running up to ~3.5k tokens against ~300-600 token completions, most of every gradient step went to predicting diff/JSON syntax instead of digest text. Caught by comparing a live Colab run’s loss curve againstsft2.log: it tracked better than the old run through epoch ~2, then plateaued around 2.7-2.9 while the old run kept dropping to ~2.1-2.4 by epoch 3.3-3.6 — consistent with the model quickly fitting the easy, now-dominant prompt tokens and further completion-quality gains getting diluted into invisibility in the aggregate loss. Fixed:sft.pynow passes native TRL"prompt"/"completion"columns (prompt as a conversational message list, so the chat template +completion_only_lossauto-enable) instead of a flat"text"field. Verified directly on a real example: 900/1196 tokens (75%) were prompt and are now correctly masked from the loss. This affected every prior SFT run, including the one behindsft2/checkpoint-40— its loss curve likely understates how well it actually fit the completions, though its message-only prompts were short enough that the effect was much smaller. - Fixed one real bug the backfill surfaced: a commit touching a minified
dist/bundle produced a single ~60k-character diff line, which line-count truncation (maxPatchLines) doesn’t bound.digest_format.is_low_signal_file()now excludes generated/build-output paths and lockfiles from patch content entirely (kept forreward.pytoo, so grounding isn’t gamed by naming generic lockfile/license vocabulary), plus a per-line character cap as defense in depth. - Reward hardening this session (adversarial testing against synthetic exploits, not just the
real dataset): closed empty-section-body, nonsense/SHA-copy filler, and single-real-word
padding exploits in
coverage_score(now requiresmin_tokens=3content words andmin_overlap=2traceable to the repo’s own commits/diff, up from a bare presence check). AddedMANIFEST_STOPWORDS(name, version, private, license, …) so generic package.json/ license boilerplate pulled in by diff-aware grounding can’t be named to fake relevance. Accepted residual risk (deliberate, not fixed): a “2 real keywords + 1 filler word” pattern can still score close to honest prose (~0.85 vs ~0.9-1.0) sincecoverageis a binary per-repo gate rather than continuous — watch the GRPO spotcheck’s per-component grounding score for drift toward short/sparse bodies as an early warning. - GRPO on a T4 is bounded by sequence length, not batch size or step count (2026-08-24
session, three OOMs). SDPA falls back to the math path under GRPO’s completion mask, so it
materialises the full
(batch, heads, seq, seq)score matrix. Gradient checkpointing does not help — it bounds how many layers are held, not how big one layer is. The fatal run died inbackward()at step 30 asking for exactly8 x 9 x 3654^2 x 4 bytes = 3.58 GiBon the 3014-token prompt.PYTORCH_CUDA_ALLOC_CONF=expandable_segments:Truedid not help and the traceback proves why: only 157 MiB was reserved-but-unallocated, so there was nothing to defragment.grpo.pynow runs a CPU pre-flight (--dry-run) that computes this spike from the tokenised dataset, drops rows over the cap, and refuses to start if that exceeds a quarter of them. At batch 8 /max_completion=768the cap is 1596 prompt tokens and keeps 89/99 rows;max_completion=512keeps 94/99. Batch 16 (--prompts-per-step 2) is rejected outright — it OOMed on the very first step. save_strategy="no"cost 29 steps of good training. The crash killed the process beforetrainer.save_model()ran, so a run that had trained cleanly for 30 steps produced nothing. Periodic saving is back as--save-steps(default 10, weights only,save_total_limit=2). This does not revive reward-based checkpoint selection, which was removed for good reasons — it is only about surviving a crash inside a fixed GPU window.lr=1e-5does not move this policy in a run that fits the window.klsat at 0.0031-0.0041 across 29 steps — the policy stayed essentially identical todigest-sft3, so even a clean 40-step run could not have produced a measurable delta. The conservative LR was chosen because the first GRPO run collapsed at a higher rate, but that run had 4 generations, a binary format gate and 50-88% truncation; none of that describes the current setup (reward_std0.23, zero truncation,beta=0.1). Note the direction:klwent 0.0041 -> 0.0028, it declined. A too-small LR would show KL climbing slowly; flat-to-declining KL is the signature of the KL penalty holding an equilibrium, sobetais as likely the binding constraint aslr. The notebook raisesLRto 3e-5 first becausebeta=0.1still bounds the damage — ifklpins near 0.003 at that rate too, beta is confirmed and comes down next. One variable at a time.- The reward saturates rather than being gamed.
fuzz_reward.pypasses clean (317 adversarial inputs, 0 violations) and the teacher’s own labels average 0.978, so 1.000 spotchecks are the ceiling being reachable, not the scorer being exploited. Butfrac_reward_zero_stdhit 1 on 2 of the first 10 steps: on easy days all 8 rollouts score exactly 1.000 and the group contributes no gradient. Raising the ceiling (uncapped grounding, activity-weighted coverage, or teacher-relative targets) is worth more than making the scorer harsher. - Rollout spread on the training set is healthy and was never the problem: probing the five
hardest rows at GRPO’s own sampling settings gave mean within-group std
0.247with0/5flat groups, against0.109and1/3on the held-out rows (scripts/probe_rollouts.py).
# Reward hardening (pre-GRPO session)
Before the next training run the reward went through adversarial fuzzing
(scripts/fuzz_reward.py) plus teacher calibration (scripts/calibrate_reward.py).
Four real holes found by mutation-testing the scorer against all 124 dataset rows:
- Summary was a free fabrication channel — penalties never scanned it and format only counted words. A hallucinated filename in the summary cost nothing (1.000 → 1.000); swapping the whole summary for unrelated prose cost nothing (0.875 → 0.875). Fixed: summary must now trace ≥2 content tokens to the prompt’s own material, and fabricated-file scanning includes it.
- Orphan prose under
## Per-Repo Activity(between the header and the first###) was invisible to every scorer — parse_sections now emits it as an unnamed pseudo-section so it fails header precision, trips the excess-header penalty, and gets fabricated-file-scanned. - Coverage was a binary per-repo gate — the documented residual risk (“2 real keywords + filler”) fired in the wild: randomized soup hit “cleanup”+“scroll” (both real dominion-tracker commit words) and banked full coverage credit for a total of 0.75. Coverage credit is now continuous, saturating at ~4 traced tokens, scaled by body length so tight honest paraphrases keep full credit.
- Grounding averaged over present sections only — dropping your weakest repo
section raised the total (+0.038 on a synthetic 6-repo day): zero-credit sections
cost nothing to omit while dragging the mean down. Grounding now averages over every
expected repo (missing section = 0), so omission can’t pay. Also fixed a crash on
degenerate activity dicts (
commits: [{}]) found by garbage-input property checks.
Fuzzer contract: MUST_COST mutations (fabrication, duplication, truncation, hollowing) must each lose ≥0.02 on every row; MUST_NOT_PAY mutations (dropping/shuffling sections) must never gain; format-valid ungrounded soup digests must stay below a ceiling (max observed 0.525 vs teacher mean 0.98); garbage inputs never crash and stay in [0,1].
Calibration after the fixes: teacher means train 0.978 / eval 0.975 / synthetic 0.939, zero rows failing format; two known thin-source days (2026-06-03, 2026-08-15 — merge/scaffold commits with little lexical ground truth) sit at 0.75–0.81 and are honest scores, not miscalibration. Reward totals from before this session are not comparable (formula changed) — sft2’s 0.499 baseline and older eval logs don’t carry over; re-baseline any model against the new scorer before comparing.
# Run guards (post-hardening session)
Fuzzing the reward found four real holes, but a fifth check — a minimal-effort digest with 3 content tokens per section and 2 traceable — came back already covered, scoring 0.60 against a teacher mean of 0.978 and a fuzzer-built soup ceiling of 0.525. That margin is wide enough to train against, so the wasted runs were traced to the training scripts instead, and the effort moved there.
Periodic checkpointing and reward-based checkpoint selection are both gone. Every sweep
so far cost 10 generations per candidate and landed on a checkpoint indistinguishable from
the last step — the selection machinery consumed more GPU time than the training it was
selecting from. sft.py and grpo.py now write only the final model (save_strategy="no"),
and both notebooks compare exactly two things: the untrained base (or SFT baseline) against
the trained result.
That also deleted the sft6 footgun rather than guarding it. sft.py used to auto-resume
whenever --out happened to contain a checkpoint-* dir, which is what ruined that run:
re-running the Colab train cell after an interrupt picked up a stale checkpoint, replayed the
LR warmup at epoch 8.08 (0 → 2.7e-5 → 5.3e-5 → 8e-5), and mixed the new args with the old
trainer_state.json (TRL warned save_steps: 25 (from args) != 3). The cosine schedule never
decayed below 6.2e-5 and loss sat flat at ~2.6 with token accuracy ~0.53 for four epochs. With
no periodic checkpoints there is nothing to resume from and nothing to resume into, so the
whole hazard is gone. Trade-off accepted: an interrupted run restarts from scratch.
grpo.py can now stop itself. abort_reason() judges the last --abort-patience
spotchecks (default 3) and halts the run on any of:
- training reward gaining ≥0.05 while the held-out spotcheck total drops by ≥0.10 — fitting the scorer rather than the task, which is the failure the reward cannot self-detect
- mean completion length dropping below half the reference digest’s — measured against each day’s own label, since a quiet day is legitimately short and a cross-day peak conflated the two
- more than half of rollouts truncating at
--max-completion
The guard’s own false-positive rate is measured, not assumed. grpo3 died at step 30 on
spotchecks of 1.0, 1.0, 0.05 — three different days at one rollout each. Pinning the day and
averaging --spotcheck-gens rollouts fixed the input; the rule still read “held-out failed to
gain” as divergence, which a day the policy already scores ~0.85 on can never satisfy. Simulating
a flat policy from sft3’s own day-0 rollouts (draws 0.7/1.0) against grpo3’s per-step training
rewards: <= 0 aborted 59% of 6-tick runs (66% at one rollout per tick), the ≥0.10 band 16%,
while a genuine pinned-day collapse (0.95 → 0.30) is caught 99.5% of the time either way. The
band costs no detection. Note the remaining 16% is dominated by window count — --spotcheck-every 10
over 60 steps gives four overlapping windows; at 20 the same guard sits at 3%, but with
--abort-patience 3 its only window closes at the final step, which disables it rather than
fixing it.
An aborted run still saves its policy (save_model runs after train() returns), so the
base-vs-result comparison works either way. --beta also moved 0.05 → 0.1: against a lexical, gameable reward the risk is drift from the
SFT policy, and tighter KL is cheaper insurance than another penalty term in reward.py.
Both Colab notebooks now run pytest, fuzz_reward.py, and calibrate_reward.py as a CPU
pre-flight before touching the GPU, and carry their hyperparameters as named constants so the
notebook records what actually ran.
# Weights
- usr-wwelsh/digest-sft2 — SFT checkpoint, mean reward
0.499 on held-out eval. Stale: trained on message-only prompts, predates the diff-aware
dataset fix (see Known issues). Superseded by
digest-sft3. - usr-wwelsh/digest-sft3 — current release
candidate. Diff-aware prompts, hardened reward.
0.8700mean reward sampled (temperature=0.8, matched 10-day held-out set, currentreward.py). - usr-wwelsh/digest-grpo — dropped, not
in the release path.
mainis a GRPO pass on top ofdigest-sft3that advertised+0.0125over its baseline; re-evaled under the fixed reward (sampled, matched settings) it scores0.6844againstsft3’s0.8700— a regression, not a gain, once the reward hole it was partly exploiting got closed. Kept published for the record, not for use. The earlier revision1cb58df(GRPO on top of the now-supersededdigest-sft2, which collapsed into looping identical commit-sha sections) is unrelated and also not in use.
# Next steps
-
grpo.pyplumbing + laptop smoke test (2 prompts × 2 rollouts × 3 steps) — hit and fixed a crash (save_safetensorsinvalid for installedtrl==1.10.0 - Backfill file stats + patches into the dataset, rebuild prompts to match production,
make
reward.py - Fuzz the reward (mutation + property tiers) and calibrate against teacher completions;
- Run
colab/digest_sft_colab.ipynb— SFT on the diff-aware dataset, reward-based checkpoint selection against both the untrained base and oldsft2, pushed as - Point
colab/digest_grpo_colab.ipynb’sBASE_MODELatdigest-sft3, rerun GRPO, pushed asdigest-grpo(+0.0125 over the sft3 baseline, greedy eval) — see open - Closed a grounding exploit: overlap was scored against a repo’s whole commit pool, so real vocabulary from unrelated commits could be recombined into an unsupported claim on busy days and still clear the bar. Now scored against the best single
-
digest_live.py(replays production prompt-building against a real day’s commits) andcompare_teacher.py(scores local checkpoints against the live teacher on - GRPO dropped. Re-evaled
sft3vsgrpounder the fixed reward, sampled (temperature=0.8) and matched on the full 10-day held-out set:sft30.8700 vsgrpo0.6844 — GRPO scores worse than the checkpoint it trained on top of, with 0% truncation on both sides so it isn’t a decoding artifact. The+0.0125it advertised was measured against the reward hole above; closing that hole exposed a real regression, not noise. Not rerunning GRPO against the fixed reward either: this is the second run on this project to score well against a reward that had a real hole in it, which points at RL finding shortcuts faster - Ship
digest-sft3(temperature=0.8) + best-of-N sampling withreward.pyas verifier — pipe intogit-digest
# Usage
uv venv --python 3.12 .venv
uv pip install --python .venv/bin/python torch --index-url https://download.pytorch.org/whl/cpu
uv pip install --python .venv/bin/python transformers trl datasets accelerate pytest
.venv/bin/python scripts/build_dataset.py # rebuild data/ from portfolio/digests
.venv/bin/python scripts/augment_dataset.py --n 18 # synthetic long-window examples (real teacher calls)
.venv/bin/python -m pytest tests/ -q # unit tests (parser + rewards)
.venv/bin/python scripts/fuzz_reward.py # adversarial fuzz: degradations must drop the score, garbage must not crash
.venv/bin/python scripts/calibrate_reward.py # teacher completions must sit near the reward ceiling
# SFT locally on CPU: ~12s/example at batch 1 on an i7-6700HQ (8 threads), so ~20min/epoch
nohup .venv/bin/python scripts/sft.py --epochs 6 --lr 8e-5 --out checkpoints/sft3 > sft3.log 2>&1 &
.venv/bin/python scripts/grpo.py --model checkpoints/sft3 --spotcheck-every 10 --abort-patience 3
.venv/bin/python scripts/eval_reward.py --model checkpoints/sft3 --tokenizer HuggingFaceTB/SmolLM2-135M-Instruct --n 10
.venv/bin/python scripts/generate.py --model checkpoints/sft2 --file DATE-1d.md --repeat-penalty 1.08 # eyeball one day
.venv/bin/python scripts/generate.py --model checkpoints/sft2 --repeat-penalty 1.08 --prompt "..." # freeform prompt
.venv/bin/python scripts/evaluate.py --model checkpoints/sft2 # scored sweep over eval set