MIP-0025 tasks¶
Written by hand, following .claude/skills/mip-tasks/SKILL.md's Step 1 process. That skill has
disable-model-invocation: true, so this file was produced without invoking it as a slash command;
the process it documents was followed manually instead.
Before task 1: the MIP's own blocking dependency is already partially satisfied, undocumented.
MIP-0025's own Depends on row says Tier 2 (QLoRA) needs to complete at least one real run before
the rest is worth building. As of this session, it has: a real tiny-preset run (SmolLM2-360M,
CPU, 3 epochs, finetune/README.md's existing small-model ladder) completed end to end:
finetune-dataset → finetune-train preset=tiny --no-4bit → convert_lora_to_gguf.py →
ollama create → verified live via just run -- --summarize producing a real (if rough-quality)
model reply. Nothing from that run is committed yet (finetune/out/* is gitignored by design, and
the run wasn't done through this stacked-task process). Task 1 below formalizes that evidence as a
real, committed artifact before layers 1-3 build on it.
| # | slug | delivers | tests (must exist before the PR) | depends on |
|---|---|---|---|---|
| 1 | tier2-baseline-evidence | ✅ merged (PR #192): finetune/README.md's Tier 2 status line updated from "written, not run" to "run, tiny preset verified" with the real command sequence and eval-loss numbers (3.032→2.866→2.799) from this session's run; MIP-0025's own Depends on/Cost so far rows updated to reflect it |
none (docs-only); this is the "doc travels with the task that changes behaviour" exception the mip-tasks skill allows for a status correction | — |
| 2 | dataset-scale-layer1 | finetune/build_dataset.py extended per MIP-0025 §4.3: synthetic Q/A pairs added to the existing knowledge-corpus extraction, reaching thousands of examples instead of dozens (Layer 1, Marine Corpus Domain) |
build_dataset.py's existing self-test extended: dataset size floor assertion, every synthetic example has a real knowledge/*.md source line it was derived from (no invented facts, per the mip skill's own "no unsourced text" rule applied to training data) |
1 |
| 3 | tool-call-sft-layer2 | A new dataset slice in build_dataset.py teaching the real MCP tool-call shape for SwimConditionsMcpServer.scala's four tools (find_nearby_beaches, get_swim_recommendation, get_water_quality, ask_ocean_question), Layer 2, genuinely new work, nothing in finetune/ does this today |
A fixture-driven test asserting the generated examples cover all four real tool names with syntactically valid call shapes (confirmed against the actual server source, not guessed) | 1 |
| 4 | dpo-preference-data-layer3 | New finetune/build_dpo_dataset.py: preference pairs (chosen/rejected) derived from core/llm/Reviewer.scala's own historical reject/revise decisions (rejected = the flawed draft, chosen = the reviewer's correction), Layer 3's data half |
A test against a fixture set of Reviewer decisions asserting one preference pair is emitted per reject/revise event, none invented when the fixture has none |
1 |
| 5 | dpo-training-layer3 | New finetune/train_dpo.py (MIP-0025 §5's Layer 3 training half) consuming task 4's dataset; a combined SFT (tasks 2+3) + DPO (task 5) run on the tiny preset at minimum, extending task 1's evidence |
just benchmark run against the resulting model, compared to the best kept run and the plain baseline per AGENTS.md's "re-run and compare before adopting" rule (same gate scripts/benchmark_gate.py already applies to marola-local); "done" = the comparison is recorded in docs/benchmarks/, not just eyeballed |
2, 3, 4 |
| 6 | hf-publish | MIP-0025 §5.1's export chain: peft merge → convert_hf_to_gguf.py → quantize (Q4_K_M + Q8_0) → CHECKSUMS → Hugging Face upload_folder with a model card (base model, training-data description, docs/benchmarks/ numbers, the IMPRÓPRIA safety note from §6 verbatim); naming per §5.1(3) (a Llama-based model needs the Llama- name prefix; tiny/small presets use ungated bases with no such constraint — confirm which base task 5 actually used before naming) |
ollama run hf.co/<user>/<repo> pulls and runs from a machine that never had the model locally — the real acceptance test from MIP-0025 §7 |
5 |
Sequencing notes¶
- Tasks 2, 3, and 4 are independent of each other (different dataset slices/scripts) and can be
worked in parallel once task 1 lands, no shared file, per the parallelism check
scripts/mip_graph.py --parallelwould run once these exist as their own MIP-numbered entries (they don't; this is one MIP's task list, not separate MIPs, so that tool doesn't apply here; noted for consistency with this session's other MIP-hardening work, not as a gap in this file). - Task 5 is the real risk concentration: it's the first task that trains anything beyond the
tinypreset's baseline, and MIP-0025's own Risk row (a tuned model hallucinating a specific number with more apparent authority) is squarely here.Reviewerstaying in the loop (unchanged by this MIP, per §5's own "what stays deterministic" note) is the existing mitigation, not something task 5 needs to rebuild. - Task 6 is genuinely optional for a
tiny/small-preset run under MIP-0033 (Release 0)'s "if affordable" framing — publishing a roughtinymodel is honest labelling (PHILOSOPHY.md's new Pillar 3 already says so), but the maintainer may prefer to wait for asmallorbaserun before spending the naming/model-card effort. Not decided here; flag before starting task 6.