MIP-0048: Scaling marola-sea — which model, which checkpoint, which hardware, and the data ceiling¶
| Status | Draft |
| Author | Claude Opus 5, for M. Hoffmann (session of 2026-09-07, consolidating findings from PR #277) |
| Created | 2026-09-07 |
| Phase | 0 (local-only; no earlier-phase prerequisite missing) |
| Related | MIP-0025 (marola-sea itself), MIP-0032 (model × strategy benchmark matrix), MIP-0007 (a different kind of model entirely), FUTURE-WORK.md §9.1 |
| Effort | M — no new module: preset tables, training flags and a preflight estimator already exist on experiment/qwen-presets-and-training-perf (PR #277). The work left is corpus growth and one real GPU run, not new machinery |
| Gain | user value (a domain model that answers ocean questions usefully rather than proving a pipeline); infra/dev-loop (a preflight that says whether a run fits before four hours are spent finding out) |
| Effort vs Gain | do next for the 4B/7B step; park 27B — §4.4 shows it needs instruction data marola does not have and should not write |
| Depends on | MIP-0025 (the pipeline this scales — merged tasks 1-5, publish tooling in place). No hardware prerequisite — the host's RTX 4090 is healthy (driver 595.84, CUDA 13.2). Coordinates with MIP-0032, which would benchmark whatever this produces; neither blocks the other. No Phase 1 gate, no paid resource, so AGENTS.md's cost rule does not apply |
| Blocked by | 0025 |
| Risk | Spending a GPU day on a bigger model that answers no better, because the ceiling is the corpus (168 unique facts) and not the parameter count — §4.5 |
| Cost so far | — |
1. Summary¶
marola-sea today is SmolLM2-360M fine-tuned on marola's corpus: a proof that the pipeline works end to end, explicitly not a quality bar. This MIP records what it would take to make it good, and, more usefully, which of the obvious upgrades are worth it. The short answer is that changing the base model is cheap and helps a little; growing the corpus is the only thing that raises the ceiling; and the 27B model that prompted this research is the wrong choice for reasons that have nothing to do with hardware.
2. Motivation¶
finetune/README.md is honest that the tiny preset "was to validate the pipeline end to end on
hardware anyone has, which it did", and that "quality itself is rough at this scale". MIP-0025 §7
asks for a benchmark against the plain base model before trusting the tuned one, and
docs/benchmarks/ still has no tiny run. So the state is: a working pipeline, an unmeasured
model, and no written basis for choosing what to run next.
The concrete trigger was a request to train Qwen3.8-27B on an RTX 4090. Researching that turned up three findings that are worth more than the answer to the original question, and are recorded in §4 rather than lost in a chat log.
3. User-visible change¶
None directly: this MIP chooses a model, it does not ship a feature. The downstream effect is on
--ask and the chat widget, where today a 360M model produces answers like this real output from
the merged Q4_K_M GGUF through Ollama (2026-09-07):
$ ollama run marola-sea:test "In one sentence: is it safe to swim when jellyfish risk is high?"
Most stings are harmless, but a sting in any part of the body can cause redness, itching, and pain.
Source: en.wikipedia.org/wiki/Jellyfish_(marine_animal)
The trained format is right (it cites a source, which is exactly what build_dataset.py teaches),
and the content dodges the question that was asked. That gap between format learned and
question answered is what a bigger model and a bigger corpus each address differently; §4.5
argues only one of them is the binding constraint.
4. Data sources and dependencies reviewed¶
4.1 The Qwen family (candidate bases)¶
Fetched from huggingface.co/api/models?author=Qwen on 2026-09-07 — 60 models by download count.
Sizes and licences below are from that response, not from memory.
| model | params | licence | note |
|---|---|---|---|
| Qwen3.5-0.8B / Qwen3-1.7B / Qwen3.5-2B | 0.8-2B | apache-2.0 | CPU-viable |
| Qwen2.5-1.5B-Instruct | 1.5B | apache-2.0 | Instruct |
| Qwen2.5-3B-Instruct | 3B | other |
the one exception in the family — excluded |
| Qwen3-4B-Instruct-2507 | 4B | apache-2.0 | best quality-per-hour candidate |
| Qwen2.5-7B-Instruct | 7.6B | apache-2.0 | the safe default |
| Qwen2.5-14B-Instruct | 14.8B | apache-2.0 | comfortable QLoRA on 24 GB |
| Qwen2.5-32B-Instruct / Qwen3-32B / Qwen3.6-27B | 27-32B | apache-2.0 | VRAM-tight, see §4.3 |
| Qwen3.8-27B | 27B | apache-2.0 | base, not Instruct — §4.4 |
| Qwen3.5-35B-A3B, Qwen3.5-122B-A10B | 35-122B | apache-2.0 | MoE — all experts must be resident; out of reach |
| Qwen3-Embedding-0.6B / 4B / 8B, Qwen3-Reranker-4B | 0.6-8B | apache-2.0 | a different lever — §4.6 |
Apache-2.0 throughout (bar the two noted) is the decisive property: unlike the small/base
Llama presets, a Qwen derivative carries no naming obligation. Llama 3.2's Community Licence
§1.b.i requires "Llama" at the beginning of any such AI model name plus a bundled agreement and a
"Built with Llama" notice, verified against the model card on 2026-09-07 and quoted in
finetune/README.md.
4.2 The reference machine¶
Measured on 2026-09-07, not assumed:
| resource | measured | binding? |
|---|---|---|
| GPU | RTX 4090, 24 GB — healthy: driver 595.84, CUDA 13.2, 34 °C idle, 933 MiB used (host nvidia-smi, 2026-09-07) |
no |
| RAM | 188 GB total, ~158 GB available | no |
| CPU | 32 threads, 50 °C at load 14 (high = 80, crit = 100) | no |
| disk | 7,030 GB free on the repo's filesystem | no |
Both of the machine facts above were first measured from inside a sandboxed agent session, and both were wrong about the machine, a methodological note worth recording since it was learned twice:
df /homein the sandbox reported 95 GB;os.statvfson the repo reported 7,030 GB. The first number produced a confident "27B cannot finish here" that survived two messages.nvidia-smiin the sandbox reports "couldn't communicate with the NVIDIA driver" because the jail maps no/dev/nvidia*nodes. On the host the same command reports driver 595.84, CUDA 13.2, a healthy RTX 4090. An earlier draft of this MIP recorded the GPU as broken and built recommendations on it.
The rule this MIP adopts: a measurement taken inside the sandbox describes the sandbox, not the
machine. Anything about hardware gets confirmed on the host before it is written down, and
finetune/preflight.py is designed to be run there rather than by an agent.
4.3 What 24 GB of VRAM actually allows¶
QLoRA holds the base in 4-bit (~0.55 GB per billion parameters with double quantization) plus adapter weights, gradients and optimizer state; the adapter is under 1% of parameters, so activations dominate the rest.
| params | 4-bit weights | QLoRA train | full export disk | verdict on this machine |
|---|---|---|---|---|
| 4B | ~2.2 GB | ~4.7 GB | ~23 GB | comfortable |
| 7.6B | ~4.2 GB | ~6.7 GB | ~40 GB | comfortable |
| 14.8B | ~8.1 GB | ~11 GB | ~80 GB | comfortable |
| 27B | ~15 GB | ~17 GB | ~154 GB | fits — VRAM tight, disk irrelevant at 7 TB |
| 35B MoE | all experts resident | — | — | does not fit |
Hardware does not rule out 27B on this machine; §4.4 does.
4.4 Instruct vs base — the finding that decides it¶
Qwen3.8-27B is a base model. The HF listing shows no Instruct variant for it, unlike
Qwen2.5-7B-Instruct or Qwen2.5-14B-Instruct.
That matters more than parameter count. A base model has never been taught to follow an instruction, respect a system prompt, refuse, or hold a turn. Fine-tuning one on marola's 2,774 domain rows produces a model that completes marola-shaped prompts and does nothing else: no assistant behaviour, because nothing in the data teaches it.
Closing that gap means general instruction data, mixed with the domain rows. Published reference points: LIMA reported ~1,000 carefully curated, highly diverse examples sufficient to align a 65B base; Dolly used 15k; Alpaca 52k. Marola would need somewhere in that range of general instruction data it does not have and has no business authoring.
An Instruct checkpoint has already had that work done by its publisher, for free. This is the
argument for qwen-14b over qwen-27b: same 24 GB card, already aligned, Apache-2.0.
4.5 The data ceiling — the finding that matters most¶
Measured from the built dataset on 2026-09-07:
dataset rows: 2,774
unique answers: 168
corpus: 7 files, 2,535 words
The 2,774 rows are 168 distinct facts asked roughly sixteen ways each. build_dataset.py's
augmentation (every chunk and sentence asked with several paraphrased templates) inflates the row
count, not the information. The corpus is about five pages of text.
This reframes every other question in this MIP. A 14B model trained on 168 facts still knows 168
facts. Raising the augmentation multiplier changes nothing: the model sees the same fact reworded.
The ceiling is knowledge/, and no choice of base model raises it.
Rough shape of the fix: 10-20× more corpus (25k-50k words) would yield ~1,700-3,400 unique facts.
The corpus-doc skill exists to add sourced entries to knowledge/ and is the correct tool.
Tool-call coverage has a smaller version of the same problem: 320 rows across the four real MCP tools, generated from only 2 coordinate pairs and 3 radii. More real locations would generalise better at near-zero cost.
4.6 The lever nobody asked about¶
--ask answers are produced by retrieval then generation. The retriever is nomic-embed-text.
If the right chunk is not retrieved, no amount of model quality recovers it, and marola has never
measured retrieval quality separately from answer quality. Qwen3-Embedding-0.6B/4B and
Qwen3-Reranker-4B (both apache-2.0) are candidates. Not proposed here; noted in §11 as a follow-up.
Pick¶
qwen-7b (Qwen2.5-7B-Instruct) for the next real run, qwen-14b next: the 4090 is healthy, so both
are a matter of hours, not hardware. qwen-27b is parked pending general instruction data that is
out of scope, and the corpus work in §4.5 outranks all of them.
4.7 Growing the corpus with quality, not volume¶
§4.5 says the ceiling is 168 unique facts. Raising it badly is easy and would make the model worse, so this records how the field says to do it well, and what marola already has lying around.
How much. Published 2026 guidance: 500-1,000 examples move the needle on formatting or classification; 3,000-10,000 high-quality examples are the range for adapting a model to a new domain. marola has 168 unique facts behind 2,774 rows, an order of magnitude short on the axis that matters, and already past the point where more paraphrases help.
Real data first. The consistent advice is to prioritise real domain text and, when generating, generate from real seeds rather than from generic prompts. marola already ingests real sources it does not use for training:
- The agency bulletins it already parses.
IneaPdfParser/InemaPdfParserread real INEA/RJ and INEMA/BA bathing-water PDFs. Every historical bulletin is real, sourced, dated domain text. scripts/arxiv_digest.pyalready fetches and caches ocean-forecasting papers. Abstracts are citable domain prose.- Open-Meteo and OSM documentation: the vocabulary of the tool-call layer.
- Tide tables and the existing
sea_lore.json: small, but curated and sourced. - Lifeguard and civil-defence safety material (SALVAMAR/Bombeiros for SC, and the equivalents in RJ/BA): exactly the register the safety footer answers in.
Synthetic, done properly. The 2026 Self-Instruct shape is: 150-200 human-written seed tasks, a stronger teacher model to expand them with a diversity-promoting prompt, then a judge model that re-scores generated rows and discards those below a threshold, commonly sampling 5-10% for audit. CRAFT (arXiv 2409.02098) is the retrieval-flavoured variant: pull real corpus passages, then augment around them, which fits marola better than free generation because every row stays anchored to a citable source.
Two constraints marola must add to that recipe, from its own rules:
AGENTS.md's "sourced or clearly labelled, never invented" already forbids unsourced facts reaching a user. Training data deserves the same bar:build_dataset.py --self-testalready asserts every generated fact appears verbatim in theknowledge/*.mdit cites, and any synthetic expansion must keep that property or the assertion becomes theatre.- The teacher's licence follows the output. This is the same clause that made a Llama teacher a problem for a SmolLM2 student (see MIP-0025's licence note): Llama 3.2 §1.b.i reaches "any outputs or results of the Llama Materials" used to train a model. An Apache-2.0 teacher (a larger Qwen) has no such term; a commercial API's terms of service need reading before its output enters a published model's training set.
Capability collapse is the failure mode to watch: narrow domain fine-tuning on a small set can destroy general ability (Dial-insight, arXiv 2403.09167). The defence is mixing in general instruction data, which is also §4.4's argument, arriving from the other direction.
Precedents from other domains worth copying rather than inventing: REx86 (arXiv 2510.20975) builds a local domain model for x86 reverse engineering from a modest curated corpus; both it and CRAFT are the same shape marola needs: narrow domain, small real corpus, careful augmentation, a local model at the end.
5. Design¶
Nothing new is required; the machinery landed in PR #277 on
experiment/qwen-presets-and-training-perf:
finetune/train_lora.py:PRESETSgainsqwen-4b,qwen-7b,qwen-14b,qwen-27b, each carryingparams_bandlicence;--device auto|cuda|cpu|hybrid;--batch/--grad-accum;--no-packing/--no-grad-checkpointing.finetune/preflight.py: VRAM/RAM/disk/ETA before a run, measuring the real filesystem.finetune/merge_export.py: unchanged; already merges any adapter into any base.
Deterministic vs LLM is unchanged and worth restating, since this MIP is about a model: nothing
here touches the map. SiteBuilder contains no LLM reference and Swimability.score is a pure
function (ARCHITECTURE.md §3b). A better marola-sea improves --ask and the chat widget only.
LoRA rank and target layers. Today r=16, α=32, targeting all seven projections
(q,k,v,o,gate,up,down). Target modules should stay as they are: attention-only LoRA is the common
economy and it costs quality on domain adaptation. Rank is worth raising to 32 at 7B and above:
adapter memory is negligible next to a 4-bit base, and rank is the main capacity knob. DoRA is a
one-flag peft change usually worth a small quality gain at the same rank; untested here.
What to change, in order:
just finetune-preflight preset=qwen-7bon the host, to confirm the numbers where the GPU is actually visible.- Grow
knowledge/per §4.5 and §4.7. This is the highest-value item and needs no GPU. - Train
qwen-7bwith--rank 32(anduse_dora=True, Appendix A), benchmark against the untuned base per MIP-0025 §7. - Only then consider 14B.
Multi-stage and resumable training¶
A long run should not be an all-or-nothing bet on the machine staying up. Three separable mechanisms, in increasing order of how much they change the design:
1. Checkpoint and resume, the direct answer. HuggingFace Trainer (which SFTTrainer
extends) writes a full checkpoint: model/adapter weights, optimizer state, LR-scheduler state, RNG
state and the step counter. trainer.train(resume_from_checkpoint=True) picks up mid-epoch, not
just at an epoch boundary, so a run can be stopped and restarted across days, reboots or a moved
GPU. What marola needs to change to use it:
# today, in train_lora.py
SFTConfig(
save_strategy="epoch", # a checkpoint only every epoch
save_total_limit=1, # keeps the newest only
)
# for a multi-day run
SFTConfig(
save_strategy="steps",
save_steps=200, # tune so a crash costs minutes, not hours
save_total_limit=2, # one to resume from, one as a fallback
)
plus a --resume flag passing resume_from_checkpoint through. Note the interaction with
save_total_limit=1, added in PR #277 to stop a 27B run writing ~100 GB of unread checkpoints:
that is still resumable, but it leaves no fallback if the newest checkpoint is truncated by the
crash that stopped the run. For long runs, 2 is the safer number.
2. Staged training: marola already does this. SFT and DPO are separate scripts producing
separate adapters (out/adapter, out/dpo-adapter), where DPO continues from the SFT adapter. That
is already a two-stage pipeline with a durable artifact between stages, and each stage can be run on
a different day. The natural third stage, if §4.7's corpus grows enough, is continued
pre-training on raw ocean text before the SFT stage: domain knowledge first, instruction
format second, preferences last.
3. Batched corpus work. The expensive part of this MIP is not GPU time, it is §4.7's corpus
growth, which is inherently incremental: every knowledge/*.md file added raises the ceiling a
little, build_dataset.py is deterministic and re-runnable, and its self-test asserts provenance on
every regeneration. There is no reason to wait for a "complete" corpus before training on the
current one: train, benchmark, add documents, retrain, and keep the benchmark rows to see whether
the corpus is actually helping.
The practical shape for a multi-day 14B run: save_steps=200, save_total_limit=2, run under
tmux or a systemd unit so an SSH drop does not kill it, and resume after any interruption. The
GPU only needs to be free while a stage is running, not for the whole calendar span.
6. Scoring / safety impact¶
None. Swimability scoring, the water verdict and the safety footer are deterministic Scala and are
not touched by any model change. A larger model does not gain authority over whether marola tells
someone to swim; that boundary is the point of ARCHITECTURE.md §3b.
7. Verification plan¶
- Already passing (PR #277):
python3 finetune/preflight.py --self-test, 9 assertions, wired intojust quality-other. - Add
docs/benchmarks/entries fortinyand for whichever Qwen preset is run: MIP-0025 §7 asks for the comparison against the untuned base and it has never been done. Without it, "the tuned model is better" is an assumption. - Live check:
just finetune-preflight preset=qwen-7bon the host, where the GPU is visible, then the full chain toollama run hf.co/<user>/<repo>. - Done looks like: a Qwen-based marola-sea published, and a benchmark row showing it beats both the untuned base and the 360M tuned model on marola's own questions. If it does not beat them, that is a finding to record, not to hide.
8. Risks, limitations, and honest caveats¶
- The corpus ceiling makes a bigger model look pointless. 168 unique facts is the real limit; a 14B model may benchmark barely above the 360M one and the effort will look wasted. It will not be, but expectations should be set now, not after the run.
- Nothing in PR #277 has been run. The Qwen presets are wired and preflighted, never trained. Not for want of hardware (the host's RTX 4090 is healthy) but because the agent session that wrote them has no GPU access. The first real run is the maintainer's, on the host.
- The ETA estimator is calibrated on one data point (SmolLM2-360M on CPU) and scaled by parameter count. Treat its numbers as order-of-magnitude.
- CPU training above ~3B is days. The estimator warns; the warning is real.
- A tuned small model can be worse than an untuned larger one. That is exactly what MIP-0032's benchmark matrix exists to find out, and why §7 insists on the comparison.
9. Alternatives considered¶
- Do nothing: keep
tiny. Legitimate: the pipeline is proven and the corpus is the real ceiling. Loses because "rough at this scale" is not a model anyone should publish as v1. - Train Qwen3.8-27B as originally asked. Hardware allows it (§4.3). Loses on §4.4: it is a base model, and aligning it needs instruction data marola should not author.
- Llama-3.2-3B (
basepreset). Works, but carries the naming obligation, the bundled agreement and the "Built with Llama" notice, for no quality advantage over an Apache-2.0 Qwen of similar size. - Grow the corpus and keep the 360M model. Genuinely competitive, and cheaper. Rejected only as an exclusive choice: §4.5's corpus work is proposed alongside, not instead.
- Skip fine-tuning; rely on RAG with a stock model. Arguably the best answer for factual accuracy. It loses the tone and citation format the fine-tune teaches, and MIP-0032 is the right place to settle it with numbers.
11. Open questions¶
- Does a 7B tuned model actually beat the untuned 7B on marola's questions? Unknown, and the whole justification rests on it. MIP-0032's matrix is the instrument.
- Is DoRA worth the flag? Reported to beat LoRA at equal rank; not tested here.
- What is the right corpus size? §4.5's 25k-50k words is arithmetic from the current augmentation ratio, not an empirical finding.
- Follow-up MIP: retrieval quality as a separate, measured axis: marola has never measured
whether
--askfailures are retrieval misses or generation misses, and Qwen3-Embedding / Qwen3-Reranker are candidates to improve the former. It needs the next MIP number; it is out of this MIP's scope, which is the generator. - Human decision: whether to publish a Qwen-based marola-sea as v1 and retire the 360M one, or publish both and let the benchmark decide.
Appendix¶
Appendix A — State-of-the-art training techniques worth adopting¶
Researched 2026-09-07. The user's constraint was explicit: a longer run is acceptable if the result is better. These are ordered by confidence, and none has been run here.
Adopt now, near-free:
- DoRA (
use_dora=Trueinpeft). Decomposes the weight update into magnitude and direction and LoRA-fits only the direction; reported to converge faster and match full fine-tuning at the same rank. 2026 guides describe it as a default-on free upgrade. One flag intrain_lora.py. - All-linear target modules. Current 2026 benchmarks say including q, k, v, o, gate, up and down consistently beats attention-only for minimal VRAM cost. marola already does this, worth recording as validated rather than changed.
- rsLoRA (rank-stabilised LoRA), a scaling factor that makes the output scale invariant to rank. Recommended once r ≥ 32, which is exactly the rank §5 proposes.
Adopt if the tooling proves out:
- Unsloth. Reported ~2× faster training at ~50% less VRAM through fused kernels and PEFT
optimisations, plus MoE support since Feb 2026 with a claimed 7-12× speedup there. Packaging is
not the obstacle it looked like:
nix eval nixpkgs#python3Packages.unslothresolves topython3.14-unsloth-2026.4.5(checked 2026-09-07), so it is one line in.github/nix-ml-env.nixand needs no pip path or overlay. That makes it cheap to try; the speed and VRAM claims are still vendor figures and unmeasured here.
Deliberately not proposed:
- Longer runs / more epochs. The user offered more hours, but with 168 unique facts more epochs buys memorisation, not knowledge. Spend the hours on §4.7's corpus instead: this is the one place where "it can take longer" does not convert into quality.
- Model merging. See Appendix B: it is a legitimate technique, and it is precisely what the Rio project was criticised for presenting as training.
Appendix B — Comparison with Rio de Janeiro's municipal LLM¶
In April 2026 Rio de Janeiro's city government, through IplanRio, launched a family of six models ("Rio 3"), including Rio 3.0 Open at 235 billion parameters and Rio 3.0 Open Mini at 44 billion, described as open source, with a total project cost of R$ 500 mil, claimed as "30 vezes menor" than an off-the-shelf system. A larger Rio 3.5 Open at 397 billion parameters was published with strong benchmark numbers.
Within hours the benchmark claims were contested. The developer of N2 Pro analysed the release and found that the published artifact combined the weights of Qwen and Nex: a raw merge, a mathematical mixture of two existing models with no additional training, rather than the distillation the project described, and without disclosing the use of N2 Pro. The prefecture acknowledged that existing models had been reused, and said the release had been premature and represented an unfinished intermediate version. Its director had earlier said of the Qwen base: "o modelo não tem nada a ver com o Qwen original, mas usamos a estrutura e o treinamento."
Three things are worth stating fairly before the comparison. Merging is a legitimate technique when disclosed; the criticism was about description, not method. A municipal government building public AI capability is a good thing, and R$ 500 mil is genuinely small for the ambition.
| marola-sea today | marola-sea proposed (§4 pick) | Rio 3.5 Open | |
|---|---|---|---|
| Parameters | 360 M | 7.6 B (Qwen2.5-7B-Instruct) | 397 B (claimed) |
| Base | SmolLM2-360M-Instruct | Qwen2.5-7B-Instruct | Qwen + Nex (merged) |
| Method | QLoRA SFT + tool-call SFT + DPO | same, larger base, r=32 | raw weight merge, no training |
| Evidence of training | eval loss 3.032 → 2.866 → 2.799 over 3 epochs; DPO to 0.2594 | to be measured | none — no training performed |
| Weights published | tooling ready, not yet uploaded | same | "open source"; location not confirmed in reporting |
| Licence | Apache-2.0 end to end, deliberately | Apache-2.0 | Qwen is Apache-2.0; Nex component unclear |
| Attribution | base model named in the model card by publish_hf.py |
same | N2 Pro use undisclosed |
| Independent verification | GGUF loaded in Ollama and answered in trained format | benchmark planned (MIP-0025 §7) | claims contested within hours |
| Cost | ~$0 (local CPU) | one GPU-day | R$ 500.000 |
marola-sea is roughly a thousandth the size of Rio's headline number and does strictly more actual
training than a merge does. What separates the two is not scale, but provenance: which base, which
licence, what was run, and what the numbers were. This repo already has that discipline written
down: AGENTS.md's "sourced or clearly labelled, never invented", the Cost: trailer on every PR,
and finetune/README.md's "run, tiny preset verified" versus "written, not run" vocabulary.
MIP-0025 §7 requires a benchmark against the untuned base before claiming the tuned model is
better.
The concrete practice to keep, stated as a rule this MIP adopts: the model card must say what was
actually done. publish_hf.py already generates the base model, the training-data description and
the eval numbers. If a future marola-sea is ever a merge rather than a fine-tune, the card says
merge. That is the whole difference between the two columns above.
Appendix C — Curated lists worth watching¶
The fine-tuning landscape moves faster than a MIP can be revised, so these are the maintained indexes to re-read before acting on anything in Appendix A rather than trusting this document's snapshot of 2026-09-07. Listed as pointers, not endorsements: none was audited here beyond confirming it exists and is on topic.
| List | Why it is relevant to marola |
|---|---|
| Hannibal046/Awesome-LLM | The general index — models, papers, tooling. The first place a new base model shows up. |
| Curated-Awesome-Lists/awesome-llms-fine-tuning | Tutorials, papers and tools specifically for fine-tuning; the closest match to §5's decisions. |
| pdaicode/awesome-LLMs-finetuning | Second fine-tuning collection; useful as a cross-check when two lists disagree. |
| horseee/Awesome-Efficient-LLM | Efficiency-focused, with a dedicated tuning.md — the right index for DoRA/rsLoRA-class techniques on one GPU. |
| rafska/Awesome-local-LLM | Running models locally: the axis marola actually cares about, since Ollama is the deployment target. |
| ethicals7s/awesome-local-ai | Local-only tooling, no cloud or API keys — the same constraint as ARCHITECTURE.md §5's local default. |
| mlabonne/llm-datasets | Post-training datasets. Directly relevant to §4.4's "general instruction data marola does not have" and §4.7's corpus work. |
| onejune2018/Awesome-LLM-Eval | Evaluation tooling and benchmarks — the gap MIP-0025 §7 and MIP-0032 both point at. |
Two frameworks surfaced repeatedly across these lists and are worth naming next to Unsloth in
Appendix A: Axolotl (LoRA/QLoRA/DeepSpeed/PEFT, multi-GPU) and xtuner (explicitly supports
Qwen among others). Neither was evaluated here; both are alternatives to hand-rolling
train_lora.py further if its flag surface keeps growing.
Checked live¶
https://huggingface.co/api/models?author=Qwen&sort=downloads&limit=60, checked 2026-09-07. Returned 60 models; sizes and licence fields in §4.1 are from this response. ConfirmedQwen2.5-3B-Instructreportsotherwhile its siblings reportapache-2.0, and thatQwen3.8-27Bappears with no Instruct variant.https://huggingface.co/HuggingFaceTB/SmolLM2-360M-Instruct, checked 2026-09-07. Licenceapache-2.0; no naming requirement on derivatives.https://huggingface.co/meta-llama/Llama-3.2-3B-Instruct, checked 2026-09-07. Llama 3.2 Community Licence §1.b.i quoted verbatim in §4.1.- Reference machine, 2026-09-07, on the host:
nvidia-smireports driver 595.84, CUDA 13.2,NVIDIA GeForce RTX 4090, 933 MiB / 24564 MiB used, 34 °C, P8. Inside the sandboxed agent session the same command fails and/dev/nvidia*does not exist (see §4.2).free -g188 GB;nproc32;sensorsPackage id 0 +50.0 °C (high +80, crit +100);os.statvfsonfinetune/7,030 GB free againstdf /home95 GB. - Built dataset, 2026-09-07: 2,774 rows across
train.jsonl/eval.jsonl, 168 unique assistant answers;knowledge/7 files, 2,535 words. https://www.mobiletime.com.br/noticias/02/04/2026/prefeitura-do-rio-3-llms/, checked 2026-09-07. Rio 3.0 Open 235B, Mini 44B, six models, R$ 500 mil, "30 vezes menor", Qwen-derived, described as open source; the article does not say where the weights are published. Director quote in Appendix B is from this page.- Web search on the Rio controversy, 2026-09-07, results from Canaltech, Tecnoblog, Baguete, TMC: Rio 3.5 Open at 397B, contested benchmarks, the N2 Pro developer's finding that the artifact was a raw Qwen+Nex merge rather than the claimed distillation, and the prefecture's response that the release was premature and intermediate.
- Web search on 2026 fine-tuning practice, 2026-09-07: DoRA as a default-on upgrade, rsLoRA for r ≥ 32, all-linear target modules beating attention-only, Unsloth's ~2×/50% claims and Feb 2026 MoE support.
- Web search for curated GitHub lists (Appendix C), 2026-09-07, restricted to github.com. Returned the eight repositories linked there plus mentions of Axolotl and xtuner. Existence and topic confirmed from the search result titles/descriptions; none of the repositories was opened, audited, or its recommendations verified.
- Host
nvidia-smi, 2026-09-07, pasted by the maintainer: driver 595.84, CUDA 13.2, RTX 4090, 933 MiB / 24564 MiB, 34 °C. This corrected an earlier draft of this MIP that recorded the GPU as unusable based on a sandbox-side failure (see §4.2). - Web search on domain dataset practice, 2026-09-07: 500-1,000 examples for formatting tasks, 3,000-10,000 for domain adaptation; Self-Instruct with a judge filtering 5-10%; prefer real domain data and generate synthetic from real seeds. Papers surfaced: CRAFT (2409.02098), Dial-insight (2403.09167), REx86 (2510.20975).
- Merged model artifacts produced this session:
merged/model.safetensors723 MB,marola-sea-tiny-f16.gguf725 MB,Q4_K_M270 MB,Q8_0386 MB; the Q4_K_M loaded into Ollama and answered in the trained "Source:" format (§3).
Not checked¶
- Every ETA in §4.3 and the estimator. Scaled from one measured run (SmolLM2-360M on CPU); no Qwen model has been trained on this machine.
- VRAM figures are arithmetic from ~0.55 GB/B for 4-bit plus an activation allowance, not
measured with
nvidia-smiduring a real run. - LIMA / Dolly / Alpaca dataset sizes in §4.4 are recalled from the literature, not fetched from the papers this session. The qualitative claim (a base model needs general instruction data) is not in doubt; the specific counts should be re-checked before anyone plans against them.
- DoRA's reported quality advantage: from memory, not verified.
- Qwen3-Embedding / Qwen3-Reranker quality: listed in the HF response, never evaluated here.
- Rio's published weights. Not located or inspected. The merge finding in Appendix B is reported by the N2 Pro developer via the press coverage above, not independently reproduced here: the comparison table's "Rio 3.5 Open" column is therefore as reported, not as verified. The Nex component's licence in particular was not established.
- Every SOTA claim in Appendix A: DoRA/rsLoRA/Unsloth figures are from 2026 guides and vendor claims found in search, not benchmarked here and not read from the primary papers.
- Unsloth's runtime behaviour: the package is present in nixpkgs (verified), but it has not been imported, run, or benchmarked here; the ~2×/50% figures remain vendor claims.