MIP-0010: MLflow as marola's experiment ledger — benchmark runs, prompt compiles and LLM traces, local server first¶
| Status | Implemented — v1 (local server, Phase 0), all 7 tasks merged: PRs #44 → #51 → #52 → #49 → #48 → #47 → #50 (docs/MIPs/MIP-0010.tasks.md); cost ~$16.52 across the 7 PRs (summed Cost: trailers, below) |
| Author | Claude Fable 5.1, for M. Hoffmann (request of 5 Sep 2026: "MIP for adding MLflow") |
| Created | 2026-09-05 |
| Tasks | docs/MIPs/MIP-0010.tasks.md — stacked PRs, one per task |
| Phase | 0 for the local server and the Scala/Python logging (developer tooling, no user-visible change) |
| Related | FUTURE-WORK.md §10 (the Langfuse-shaped tracing gap this closes for the JVM side), FUTURE-WORK.md §4.1 (evaluation harness — the ledger this MIP adds is what a harness writes to), ARCHITECTURE.md §5f (Telemetry.scala, the existing OpenTelemetry plumbing), docs/benchmarks/ and scripts/benchmark_gate.py (today's Markdown ledger), dspy/ and finetune/ (the offline Python steps) |
| Effort | L — 7 stacked PRs across three lanes (ledger, tracing, dspy), a new REST client, an OTel split |
| Gain | infra/dev-loop (replaces a hand-pasted Markdown ledger with a queryable one) |
| Effort vs Gain | cheap win, delivered — developer-only, no user-facing risk |
| Depends on | none (local only, delivered) |
| Risk | MLflow's OTLP ingest and REST surface are both young (server 3.16.0 vs. a lagging Java client) — a version bump could break the REST contract silently |
| Cost so far | ~$16.52 across the 7 implementation PRs (#44, #51, #52, #49, #48, #47, #50); the MIP's own drafting cost is bundled into a shared ~$9.65 session total with MIP-0011 and other PRs (commit 3fdcd05), not separately split |
1. Summary¶
marola already measures itself: just benchmark scores three answer modes on 22 questions,
the DSPy compile step optimises two prompts against a metric, and the fine-tune tiers produce model
variants. But every result is a Markdown file under data/ or docs/benchmarks/, compared by
hand and parsed back by a regex (scripts/benchmark_gate.py). This MIP adds MLflow as the
ledger those steps write to: one experiment per kind of run, params (model, embedder,
corpus hash, git SHA), metrics (coverage per arm, cited %, latency; the DSPy metric scores),
artifacts (the Markdown report, the compiled prompt JSONs), and, because the MLflow server
accepts OpenTelemetry traces over OTLP/HTTP from any language, LLM-call traces from the Scala
pipeline itself (summariser + reviewer spans with model, latency and token counts). Local default:
mlflow server on SQLite via a compose profile or just mlflow-up, no account.
2. Motivation¶
- The benchmark ledger is Markdown.
docs/benchmarks/2026-09-05.mdholds three runs as hand-pasted tables;benchmark_gate.pyfinds "the result that matters" (rag-generalcoverage, 0.84) by regex over| arm | coverage ... |rows. It works, and it is exactly what an experiment tracker exists to replace: runs with params, comparable in a UI, queryable by the gate. - Prompt compiles leave no record but their output.
compile_recommendation_prompt.pywritesrecommendation_prompt.json/review_prompt.json; which model compiled them, against which trainset, with what metric score, is in nobody's notes. Langfuse tracing is optional there today (_init_langfuse_tracing), Python-only, and needs a hosted account or a second server. - The JVM side has OpenTelemetry but nothing LLM-shaped.
Telemetry.scalawraps one span around the pipeline and had no local export target (ARCHITECTURE.md§5f, "unverified against a live resource").FUTURE-WORK.md§10 names "Langfuse-shaped tracing … a scoped, real improvement" as the actionable gap; Langfuse has no JVM SDK. MLflow's OTLP endpoint makes the existing OpenTelemetry plumbing enough.
3. User-visible change¶
None for the swimmer. For the developer:
$ just mlflow-up # http://127.0.0.1:5000, SQLite + local artifacts, no account
$ MAROLA_MLFLOW_TRACKING_URI=http://127.0.0.1:5000 just benchmark
marola :: benchmark — 22 questions × 3 arms on llama3.2
...
mlflow: experiment "marola/benchmark", run 3f9c… — params model=llama3.2 embed=llama3.2 min_score=0
corpus=knowledge@a1b2c3 git=abe1ba4; metrics rag_general.coverage_all=0.83 …; artifact
data/benchmark-20260905-1550.md → http://127.0.0.1:5000/#/experiments/1/runs/3f9c…
$ MAROLA_MLFLOW_TRACKING_URI=http://127.0.0.1:5000 MAROLA_TRACES=mlflow just run -- --summarize
... (unchanged output)
traces: 1 trace, 3 spans (marola.recommend, llm.summarize 2.1 s, llm.review 1.8 s) → experiment "marola/traces"
Without MAROLA_MLFLOW_TRACKING_URI nothing changes: the Markdown report is still written, the
gate still reads it, no network call is made.
4. Data sources and dependencies reviewed¶
4.1 MLflow (the server and the concept)¶
Apache-2.0; latest release v3.16.0 on 2026-09-04 (GitHub API, fetched 2026-09-05). Container
image ghcr.io/mlflow/mlflow exists (GHCR tag list fetched 2026-09-05; the first page reached
v3.9.0rc0; the newest tag was not confirmed, pin whatever just mlflow-up first uses). Local
server: mlflow server --backend-store-uri sqlite:///… --artifacts-destination … --serve-artifacts
--host 127.0.0.1 --port 5000; --app-name basic-auth adds basic authentication (CLI reference,
raw page fetched 2026-09-05; --host default 127.0.0.1, --port default 5000).
4.2 Writing runs from Scala — REST, not the Java client¶
- REST API (fetched 2026-09-05):
POST /api/2.0/mlflow/experiments/create,runs/create,runs/update,runs/search,runs/log-batch(the last capped at 1000 items per call, ≤ 1000 metrics, ≤ 100 params, ≤ 100 tags, 1 MB payload, 250-character keys and param/tag values). Artifact upload goes throughartifacts/presigned-upload-urlor the server's artifact proxy when started with--serve-artifacts. - Java client
org.mlflow:mlflow-client: latest 3.11.1, Maven Central metadata last updated 2026-04-08 (fetched 2026-09-05), five minor versions behind the server;MlflowClientofferscreateExperiment,createRun,logParam/logMetric/setTag/logBatch,logArtifact(s),searchRuns,setTerminated(Javadoc, fetched 2026-09-05). No tracing API. - Pick: REST via the existing
Http/Jsonhelpers. Four endpoints, the same shape as every other client incore/, no new dependency tree, no version lag. The Java client is the fallback if artifact upload over REST proves awkward (§11).
4.3 Traces from the JVM — OTLP into the MLflow server¶
MLflow ≥ 3.6.0 exposes POST /v1/traces (OTLP/HTTP only, no gRPC), requires the header
x-mlflow-experiment-id: <id>, requires a SQL backend store (SQLite is fine), supports gzip
from 3.7.0; the documented SDK env vars are OTEL_EXPORTER_OTLP_TRACES_ENDPOINT=http://localhost:5000/v1/traces
and OTEL_EXPORTER_OTLP_TRACES_HEADERS=x-mlflow-experiment-id=123 (ingest page, fetched
2026-09-05). MLflow states GenAI semantic-convention support (gen_ai.request.model,
gen_ai.usage.input_tokens, …) for ingestion. On the JVM side io.opentelemetry:opentelemetry-exporter-otlp
is at 1.65.0 (Maven Central, 2026-08-07); the repo pins opentelemetry-sdk-extension-autoconfigure
1.49.0; bump to one consistent
OpenTelemetry BOM in the implementation.
4.4 Writing runs from the Python steps¶
mlflow (Python) is the reference client: mlflow.start_run, log_params/log_metrics/
log_artifact. MLflow's GenAI evaluation (mlflow.genai.evaluate, scorers, LLM judges) and the
Prompt Registry (mlflow.genai.register_prompt) are Python-only APIs (docs fetched
2026-09-05; the prompt registry's REST surface for non-Python clients was not documented on the
page fetched); usable in dspy/ and finetune/, not from Scala. Whether mlflow.dspy.autolog()
exists at the pinned DSPy/MLflow versions was not checked (§11).
5. Design¶
A new pluggable integration, same shape as the six others (ARCHITECTURE.md §5): a trait in
core/, a no-op default, an implementation in local/ (plain REST, no SDK), env-var selection in AppConfig.
// core/src/main/scala/marola/ledger/RunLedger.scala
trait RunLedger:
def start(experiment: String, name: String, params: Map[String, String]): RunHandle < Sync
def metrics(run: RunHandle, values: Map[String, Double], step: Int = 0): Unit < Sync
def artifact(run: RunHandle, path: java.nio.file.Path): Unit < Sync
def end(run: RunHandle, ok: Boolean): Unit < Sync
object RunLedger:
val Noop: RunLedger // every method returns unit; the default
final case class RunHandle(experimentId: String, runId: String, url: Option[String])
// local/src/main/scala/marola/ledger/MlflowRunLedger.scala — REST over Http/Json:
// experiments/get-by-name → create; runs/create; runs/log-batch (chunked to the §4.2 caps,
// keys truncated to 250 with a warning); artifact upload via the server's artifact proxy;
// runs/update RUNNING→FINISHED/FAILED.
Where runs are written:
OceanBenchmark.run→ experimentmarola/benchmark: paramsmodel,embed_model,min_score,corpus_sha(hash ofknowledge/*.md),git_sha,questions; metrics per arm (<arm>.coverage_in_corpus,.coverage_general,.coverage_all,.cited_pct,.abstained_pct,.mean_ms); artifact: the Markdown report. The Markdown report stays the canonical gate input in v1:benchmark_gate.pyis unchanged; reading the best kept run from MLflow is v2.dspy/compile_recommendation_prompt.py→ experimentmarola/prompt-compile: params (model, optimiser, trainset size), metrics (the metric's score per compiled program), artifacts (both JSONs). Replaces the optional Langfuse hook or sits beside it (§11).finetune/train_lora.py→ experimentmarola/finetune(later task; "written, not run" today).- Traces:
Telemetrysplits into acoretrait (Tracing.withSpan,Tracing.llmSpan) with two backends:local/MlflowTracing(OTLP/HTTP exporter to<tracking uri>/v1/traces, header fromMAROLA_MLFLOW_EXPERIMENT) and a no-op. ATracedLlmClient(inner, tracing)decorator wrapsLocalLlmClientand emits one span percompletewithgen_ai.request.model, latency, and token counts when the response carriesusage; prompt and completion text are attached only whenMAROLA_TRACE_CONTENT=1(locations are personal data).
Env vars (AppConfig): MAROLA_MLFLOW_TRACKING_URI (unset = Noop), MAROLA_MLFLOW_EXPERIMENT
(default marola), MAROLA_TRACES=off|mlflow. Recipes: just mlflow-up
(compose profile mlflow on ghcr.io/mlflow/mlflow, SQLite + .tmp/mlflow/ artifacts, bound to
127.0.0.1) and just mlflow-ui. Deterministic: everything above. Through the LLM: nothing.
6. Scoring / safety impact¶
None. No scoring code is touched; Swimability is unchanged.
7. Verification plan¶
Unit tests (all offline, munit): MlflowRunLedgerSpec, the exact JSON of runs/create and
runs/log-batch for a Report, chunking at 1000 metrics / 100 params, 250-character key
truncation, FAILED status on end(ok = false), via a scripted Http.Transport;
NoopRunLedgerSpec (no transport call ever); TracedLlmClientSpec (span name, gen_ai.*
attributes and the content-off default, using opentelemetry-sdk-testing's in-memory exporter);
BenchmarkLedgerSpec (the params/metrics map derived from an OceanBenchmark.Report).
Live checks: just mlflow-up && MAROLA_MLFLOW_TRACKING_URI=… just benchmark shows the run and
its artifact at 127.0.0.1:5000; MAROLA_TRACES=mlflow just run -- --summarize shows one trace
with the summariser and reviewer spans; dspy compile logs a run with two artifacts. Done: the
next benchmark PR's comparison paragraph links a run URL instead of pasting a table.
8. Risks, limitations, and honest caveats¶
- MLflow 3 moves fast (server 3.16.0 vs Java client 3.11.1); REST is the stable surface, and OTLP
ingest only exists from 3.6.0, so pin the image tag and say so in
RUN-LOCALLY.md. - Traces can carry prompts, which carry a user's location. Content off by default; the local
server binds to loopback; the MLflow UI has no auth unless
--app-name basic-auth. - MLflow is a heavy Python dependency: it lives in a compose profile /
uvx, never in the runtime image (Dockerfile) or inflake.nix's default shell unless the nixpkgs package proves light (§11). - Two ledgers during the transition (Markdown + MLflow). Mitigated by keeping Markdown canonical for the gate until the MLflow query path is tested.
9. Alternatives considered¶
- Do nothing: the Markdown ledger and regex gate work today; they do not scale past one benchmark and record nothing about prompt compiles or traces.
- Langfuse: already optional in
dspy/; no JVM SDK (FUTURE-WORK.md§10); whether its server accepts OTLP from other languages was not checked. One server for runs and traces favoured MLflow. - Weights & Biases / hosted trackers: an account and a key for a personal project's ledger.
- Java client instead of REST: see §4.2; kept as the fallback.
11. Open questions¶
- Is
python3Packages.mlflowin nixpkgs light enough for the dev shell, or is the compose profile /uvx mlflowthe only sane local path? mlflow.dspy.autolog()at the pinned DSPy version: exists? worth it, or log the two compile metrics by hand?- Keep the Langfuse hook in
dspy/alongside MLflow, or remove it once MLflow logs the compile? - Artifact upload over REST vs the Java client's
logArtifact: decide after the first spike. - Experiment naming: one
marolaexperiment with tags, ormarola/<kind>as sketched here.
Appendix¶
- MLflow release:
https://api.github.com/repos/mlflow/mlflow/releases/latest→v3.16.0, 2026-09-04; licence Apache-2.0 (2026-09-05). - Java client:
https://repo1.maven.org/maven2/org/mlflow/mlflow-client/maven-metadata.xml→ 3.11.1, lastUpdated 20260408 (2026-09-05). - OTLP ingest:
https://mlflow.org/docs/latest/genai/tracing/opentelemetry/ingest/:/v1/traces, OTLP/HTTP only,x-mlflow-experiment-id, MLflow ≥ 3.6.0, SQL store required (2026-09-05). - REST:
https://mlflow.org/docs/latest/api_reference/rest-api.html: log-batch caps (2026-09-05). - OpenTelemetry Java:
io.opentelemetry:opentelemetry-exporter-otlp1.65.0 (Maven Central, 2026-08-07).