Skip to content

MIP-0025 tasks

Written by hand, following .claude/skills/mip-tasks/SKILL.md's Step 1 process. That skill has disable-model-invocation: true, so this file was produced without invoking it as a slash command; the process it documents was followed manually instead.

Before task 1: the MIP's own blocking dependency is already partially satisfied, undocumented. MIP-0025's own Depends on row says Tier 2 (QLoRA) needs to complete at least one real run before the rest is worth building. As of this session, it has: a real tiny-preset run (SmolLM2-360M, CPU, 3 epochs, finetune/README.md's existing small-model ladder) completed end to end: finetune-dataset → finetune-train preset=tiny --no-4bit → convert_lora_to_gguf.py → ollama create → verified live via just run -- --summarize producing a real (if rough-quality) model reply. Nothing from that run is committed yet (finetune/out/* is gitignored by design, and the run wasn't done through this stacked-task process). Task 1 below formalizes that evidence as a real, committed artifact before layers 1-3 build on it.

# slug delivers tests (must exist before the PR) depends on
1 tier2-baseline-evidence ✅ merged (PR #192): finetune/README.md's Tier 2 status line updated from "written, not run" to "run, tiny preset verified" with the real command sequence and eval-loss numbers (3.032→2.866→2.799) from this session's run; MIP-0025's own Depends on/Cost so far rows updated to reflect it none (docs-only); this is the "doc travels with the task that changes behaviour" exception the mip-tasks skill allows for a status correction —
2 dataset-scale-layer1 finetune/build_dataset.py extended per MIP-0025 §4.3: synthetic Q/A pairs added to the existing knowledge-corpus extraction, reaching thousands of examples instead of dozens (Layer 1, Marine Corpus Domain) build_dataset.py's existing self-test extended: dataset size floor assertion, every synthetic example has a real knowledge/*.md source line it was derived from (no invented facts, per the mip skill's own "no unsourced text" rule applied to training data) 1
3 tool-call-sft-layer2 A new dataset slice in build_dataset.py teaching the real MCP tool-call shape for SwimConditionsMcpServer.scala's four tools (find_nearby_beaches, get_swim_recommendation, get_water_quality, ask_ocean_question), Layer 2, genuinely new work, nothing in finetune/ does this today A fixture-driven test asserting the generated examples cover all four real tool names with syntactically valid call shapes (confirmed against the actual server source, not guessed) 1
4 dpo-preference-data-layer3 New finetune/build_dpo_dataset.py: preference pairs (chosen/rejected) derived from core/llm/Reviewer.scala's own historical reject/revise decisions (rejected = the flawed draft, chosen = the reviewer's correction), Layer 3's data half A test against a fixture set of Reviewer decisions asserting one preference pair is emitted per reject/revise event, none invented when the fixture has none 1
5 dpo-training-layer3 New finetune/train_dpo.py (MIP-0025 §5's Layer 3 training half) consuming task 4's dataset; a combined SFT (tasks 2+3) + DPO (task 5) run on the tiny preset at minimum, extending task 1's evidence just benchmark run against the resulting model, compared to the best kept run and the plain baseline per AGENTS.md's "re-run and compare before adopting" rule (same gate scripts/benchmark_gate.py already applies to marola-local); "done" = the comparison is recorded in docs/benchmarks/, not just eyeballed 2, 3, 4
6 hf-publish MIP-0025 §5.1's export chain: peft merge → convert_hf_to_gguf.py → quantize (Q4_K_M + Q8_0) → CHECKSUMS → Hugging Face upload_folder with a model card (base model, training-data description, docs/benchmarks/ numbers, the IMPRÓPRIA safety note from §6 verbatim); naming per §5.1(3) (a Llama-based model needs the Llama- name prefix; tiny/small presets use ungated bases with no such constraint — confirm which base task 5 actually used before naming) ollama run hf.co/<user>/<repo> pulls and runs from a machine that never had the model locally — the real acceptance test from MIP-0025 §7 5

Sequencing notes

  • Tasks 2, 3, and 4 are independent of each other (different dataset slices/scripts) and can be worked in parallel once task 1 lands, no shared file, per the parallelism check scripts/mip_graph.py --parallel would run once these exist as their own MIP-numbered entries (they don't; this is one MIP's task list, not separate MIPs, so that tool doesn't apply here; noted for consistency with this session's other MIP-hardening work, not as a gap in this file).
  • Task 5 is the real risk concentration: it's the first task that trains anything beyond the tiny preset's baseline, and MIP-0025's own Risk row (a tuned model hallucinating a specific number with more apparent authority) is squarely here. Reviewer staying in the loop (unchanged by this MIP, per §5's own "what stays deterministic" note) is the existing mitigation, not something task 5 needs to rebuild.
  • Task 6 is genuinely optional for a tiny/small-preset run under MIP-0033 (Release 0)'s "if affordable" framing — publishing a rough tiny model is honest labelling (PHILOSOPHY.md's new Pillar 3 already says so), but the maintainer may prefer to wait for a small or base run before spending the naming/model-card effort. Not decided here; flag before starting task 6.