MIP-0012: llm4s as marola's Scala-native LLM/agent layer — and the deprecation of the Python DSPy step¶
| Status | Draft |
| Author | Claude Fable 5.1, for M. Hoffmann (request of 5 Sep 2026: "implement llm4s in this repo, and deprecate DSPy — comprehensive, all llm4s capabilities, all marola ambitions, see PHILOSOPHY.md") |
| Created | 2026-09-05 |
| Phase | 0 for the Scala prompt compiler, the DSPy removal and the llm4s/ module behind marola's own traits (developer tooling; the swimmer sees a re-compiled prompt, nothing else). The agent loop waits on Phase 1 (MIP-0002, still Draft); the HTTP MCP server for a hosted agent is Phase 2/3 and behind the cost gate |
| Related | PHILOSOPHY.md ("Why Scala 3 on the JVM" — the Python paragraph this MIP acts on), FUTURE-WORK.md §10 (ds4s, "DSPy stays a Python subprocess indefinitely" — revisited here) and §4.1 (held-out eval), AGENT-FRAMEWORKS-SURVEY.md §1.2 (llm4s row), ARCHITECTURE.md §5a (the DSPy step, CompiledPrompt) and §5c ("HTTP/SSE MCP transport — not built"), MIP-0010 (ledger + traces this step should write to; its task 7 becomes moot), MIP-0008 (native image — dependency weight), MIP-0002 (the first real agent surface) |
| Effort | XL — a new module, a from-scratch Scala prompt compiler replacing DSPy, a dependency-boundary CI check, an agent/MCP layer gated on Phase 1 |
| Gain | infra/dev-loop (removes the Python/DSPy toolchain and its training/serving skew) |
| Effort vs Gain | do when X lands — tasks 1-3 (the prompt compiler) are dependency-free and could go now; tasks 4-7 (the agent) wait on MIP-0002 |
| Depends on | MIP-0010 (the ledger the compiler logs to); MIP-0002/Phase 1 (the agent tasks); no cloud resource |
| Risk | llm4s is pre-1.0 with 154 transitive jars and one coordinate rename already — real churn for what would otherwise be a 30-line adapter |
| Cost so far | n/a — the merged commit's own Cost: line was left as an unfilled placeholder ("~unmeasured in-agent · fill from just claude-cost before PR") |
1. Summary¶
Two changes, deliberately separable. First, the offline DSPy step (dspy/, Python) is replaced by
a marola-owned Scala compiler in core/prompt/: the same metric-gated few-shot bootstrap marola
actually uses today (three examples, max_bootstrapped_demos=3, a deterministic metric), writing the
same JSON artifact CompiledPrompt already loads, so the prompt is compiled and replayed through
one code path, and the gate PHILOSOPHY.md wants ("there by construction") covers the last piece of
marola that only had it by discipline. This part needs no new dependency: it runs over marola's
existing LlmClient. Second, llm4s (org.llm4s:llm4s-core
0.4.1, MIT, Scala 3, verified consumable from marola's 3.9.0 build, one live completion against
the local Ollama) is adopted as an opt-in fifth module llm4s/ behind marola's own traits, for
what marola lacks and llm4s has: an agent loop with tools, handoffs and guardrails,
structured output, an MCP client, and an HTTP/Streamable-HTTP MCP server (the §5c gap). It
cannot live in local/: llm4s-core drags 154 jars transitively (§4.1), and local/'s small
dependency set is load-bearing.
2. Motivation¶
- PHILOSOPHY.md's own argument, applied to the one place it isn't. "Marola's offline steps
(
dspy/,finetune/) are Python, chosen because the libraries only exist there." Fordspy/that premise is now too strong: what marola uses of DSPy isBootstrapFewShotover three hand-written examples with a keyword metric (compile_recommendation_prompt.py). That is ~150 lines of Scala, not a DSPy port.finetune/(peft/trl, a GPU) keeps the argument;dspy/no longer does. - Training/serving skew, admitted in the code.
CompiledPromptis "a good-faith replication of DSPy's ownChatAdapterformat … not a byte-identical replay" (ARCHITECTURE.md§5a). A Scala compiler that bootstraps demos throughCompiledPrompt.buildMessagesitself removes the skew by construction, not by care. - Two toolchains for one artifact. A venv,
dspy>=3.3,<4,langfuse, anLD_LIBRARY_PATHworkaround fortokenizers(dspy/README.md), a Python lint lane inci.yml, and MIP-0010 task 7 planning a second logging client just for this script. The JVM side already has the client, the ledger (RunLedger, MIP-0010 tasks 1-2) and the tests. FUTURE-WORK.md§4.1 never got built: "rundspy.Evaluateafter every compile" was the plan, but it lived on the Python side nobody ran. In Scala a held-outEvaluateis one function next to the compiler and one row in the ledger.- The agent ambitions have no substrate. marola wants summarize / critique / escalate as
addressable agents with a named topology;
ARCHITECTURE.md§5c says a hosted agent needs an HTTP MCP transport marola doesn't have;AGENT-FRAMEWORKS-SURVEY.mdfound "nothing in this list is a Scala-native multi-agent framework" and listed llm4s as "pre-production per its own roadmap". Checked against the jar (§4), llm4s has an agent loop, handoffs, guardrails, an MCP client with three transports and an MCP server with two. It is the closest Scala-native fit.
3. User-visible change¶
None for the swimmer, except that both prompts get re-compiled against the runtime default model
(llama3.2) by the new step, so the summary wording may shift. For the developer:
$ just prompt-compile # Scala, over the configured LlmClient — no venv
marola :: prompt-compile — SummarizeSwimConditions, 3 examples, metric jellyfishAndWhale
example Arpoador/07:00 metric 1.00 → demo (augmented)
example Praia Vermelha/15:00 metric 1.00 → demo (augmented)
example Praia do Leme/18:00 metric 1.00 → demo (augmented)
held-out eval (2 examples): 1.00
wrote core/src/main/resources/recommendation_prompt.json (3 demos)
marola :: prompt-compile — ReviewSwimSummary, 3 examples, metric reviewJsonWellFormed
...
mlflow: experiment "marola/prompt-compile", run 7c1a… — params model=llama3.2 optimiser=bootstrap-few-shot
trainset=3 max_demos=3; metrics summarize.train=1.0 summarize.dev=1.0 review.train=0.9 …;
artifacts recommendation_prompt.json review_prompt.json (only with MAROLA_MLFLOW_TRACKING_URI)
$ MAROLA_LLM_PROVIDER=llm4s-ollama just run -- --summarize # same output shape as today
just run -- --summarize with no provider set is unchanged (marola's own LocalLlmClient).
4. Data sources and dependencies reviewed¶
4.1 llm4s — what was verified (2026-09-05) and how¶
Everything below was checked against the published jar (javap, class listing, cs resolve) or the
source file named in Appendix C, not the README; the detailed notes are Appendix D. Marola use is
the ambition each capability serves, with the phase it lands in (ARCHITECTURE.md §11).
| llm4s capability (0.4.1) | Verified how | marola use | Phase |
|---|---|---|---|
LLMClient.complete/streamComplete: Either[LLMError, Completion], token usage |
javap; live: Scala 3.9.0 probe, one completion on local Ollama (Appendix A) | Llm4sLlmClient adapter behind marola.llm.LlmClient (§5.2) |
0 |
Providers: OpenAI (vendor Java SDK, key only), Ollama (native /api/chat, drops tool messages, no JSON mode), Anthropic, Gemini, DeepSeek, Cohere, Mistral, OpenRouter, Vertex, Z.ai |
javap + OpenAIClient.scala/OllamaClient.scala |
llm4s-ollama/llm4s-openai providers, opt-in |
0 |
Structured output: ResponseFormat.Json/JsonSchema(schema, name, strict) — OpenAI path only |
javap + ResponseFormat.scala |
Reviewer verdict as a strict JSON schema (§5.3) |
0 |
Tools: ToolFunction, ToolBuilder, Schema DSL, ToolRegistry (ujson) |
jar + tool-calling-api-design.md; ScalaMeta code-gen claimed, not found |
one Tools.scala source for MCP SDK + llm4s (§5.5) |
1 |
Agent loop: Agent.run(query, tools, guardrails, handoffs, maxSteps=50), runStep, events; Handoff.to; AgentStatus |
javap + /guide/agents/, /handoffs |
the bot's /ask agent; escalation agent via handoff (§5.4) |
1 → 4 |
Guardrails: InputGuardrail/OutputGuardrail, ValidationMode.Block/Warn/Log, 18 built-ins |
jar package + /guardrails |
output checks on /ask; never the human gate (§5.4) |
1 |
Orchestration: Plan (DAG, topological order, parallel batches), TypedAgent[I,O], PlanRunner — Future-based |
javap | candidate for the named summarize/critique/escalate topology; Future sits awkwardly next to Kyo |
4 |
MCP client: stdio / SSE / Streamable HTTP, getTools: Either[String, Seq[ToolFunction]] |
javap | consume external servers from the agent — noted, not scoped | 4 |
MCP server: HTTP+SSE (2024-11-05) + Streamable HTTP (2025-06-18) on com.sun.net.httpserver, bearer auth, loopback-only without a key; no stdio |
javap + MCPServer.scala |
just mcp-http; the remote-tool transport a hosted agent needs (§5.5) |
2/3, cost gate |
Tracing: Tracing trait, modes langfuse/otel/console/collector/noop; OTel module = gRPC exporter, gen_ai.* attrs |
javap + OpenTelemetryTracing.scala |
not adopted — MLflow is OTLP/HTTP; TracedLlmClient wraps the adapter (MIP-0010) |
— |
Embeddings (OllamaEmbeddingProvider), rag + sqlite/pgvector stores, reranker, RAGAS-style eval |
jar packages | not needed: knowledge/ + OllamaEmbedder already live-verified (MIP-0001) |
— |
eval.dataset: Dataset, Example[I,O], ExampleSelector, JsonlCodec |
javap | a dataset model only; core/prompt/Example stays marola's (§5.1) |
— |
Prompt optimisation / Signature / BootstrapFewShot / MIPROv2 |
jar class search: absent; README/roadmap silent | none — marola owns the compiler (§5.1) | — |
ReliableClient (retry, circuit breaker), middleware, memory, context pruning, image, speech (vosk), knowledge graph, workspace runner |
jar packages | not needed | — |
| Java/Kotlin/Spring interop | roadmap phase 3 | not needed | — |
Coordinates and terms. org.llm4s:llm4s-core_3:0.4.1 (Maven Central via coursier; v0.4.0 and
v0.4.1 both 2026-08-29, the former "rename-only"; v0.3.0 2026-02-21 added OTel, MCPServer,
Result). MIT (LICENSE, Rory Graves, 2025). Only llm4s-core_3, llm4s-observability-otel_3, the
two workspace artifacts and knowledgegraph-neo4j_3 are published; llm4s-mcp/-rag/-memory
ship inside llm4s-core ("the latest release (v0.4.1) still ships as a single artifact", 1.0 scope).
Built with Scala 3.7.1 ("Scala 3 only"; the homepage's "2.13 fully supported" contradicts the
roadmap: treat as stale); JDK 21 CI. Verified live from Scala 3.9.0 with marola's strict flags.
Pre-1.0: "API stabilizing", MiMa and "deprecate before removing" only after 1.0 and only for Frozen
modules; the pending module split will move coordinates again.
Dependency closure, the decisive number. cs resolve llm4s-core_3:0.4.1: 154 jars, ~150 MB
versus ~32 MB for marola's Kyo + MCP SDK + logback today; includes a vendor OpenAI Java SDK
1.0.0-beta.16 (+ its core library, Netty, Reactor), anthropic-java (+ OkHttp, Kotlin), AWS S3/STS,
PDFBox, POI, Tika, PostgreSQL, sqlite-jdbc, HikariCP, vosk, Prometheus, cats-core, upickle/ujson,
Jackson 2.19 (coursier flags Jackson 2.17→2.19 and JNA 5.7→5.19 conflicts).
Not checked: the Streamable-HTTP server against a real MCP client; llm4s's OpenAIClient
pointed at Ollama's /v1 (would give tools + response_format locally, plausible, not run); a
GraalVM native image with llm4s on the classpath (MIP-0008); llm4s's MCPClient against marola's
stdio server.
4.2 What marola uses of DSPy today (dspy/compile_recommendation_prompt.py)¶
Two dspy.Signatures (8 and 9 string fields), two three-example trainsets, two deterministic
metrics (keyword presence weighted 0.7/0.3; JSON well-formedness 0.3/0.7 + verdict match 0.3),
dspy.teleprompt.BootstrapFewShot(metric, max_bootstrapped_demos=3), compiled.save(path).
Artifact schema (confirmed against dspy==3.3.1 output, CompiledPrompt's doc comment): demos[]
(input fields + output field + "augmented": true), signature.instructions, signature.fields[]
(prefix, description). No MIPROv2, no dspy.Evaluate, no dev split. Optional Langfuse tracing
(never verified end to end, dspy/README.md).
4.3 Ollama's OpenAI-compatible endpoint — response_format¶
Verified live 2026-09-05: POST /v1/chat/completions with "response_format":{"type":"json_object"}
on marola-llama3.2 returned {"score":35,"verdict":"revise"} and a usage block. Relevant because
marola's own LocalLlmClient can use it and llm4s's OllamaClient (native /api/chat) does not.
4.4 Alternatives for the client/agent layer (from AGENT-FRAMEWORKS-SURVEY.md, not re-verified)¶
sttp-ai (Ollama + OpenAI-compatible, structured output, tools; needs an fs2/ZIO/Ox backend), Kyo's own AI modules (typed tools, MCP: "check what's in the pinned RC5 jar"), agent4s (cats-effect, v0.1.0), LangChain4j (Java). None is Scala-native and carries an agent loop with handoffs.
Pick¶
llm4s, as an opt-in module behind marola's traits, for the agent/MCP/guardrail/structured-output layer; marola's own code for the prompt compiler (llm4s has nothing there).
5. Design¶
5.1 The prompt compiler — core/prompt/, no new dependency¶
Pure Scala over the existing LlmClient; only the complete calls are effects.
package marola.prompt
final case class Field(name: String, description: String) // snake_case names, as today
final case class Signature(instructions: String, inputs: List[Field], output: Field)
final case class Example(inputs: Map[String, String], output: String)
type Metric = (Example, String) => Double // pure, in core/prompt/PromptMetrics
object BootstrapFewShot:
/** Runs the zero-shot program on each trainset row, keeps the rows whose prediction scores
* ≥ threshold as demos (marked augmented), stops at maxDemos: what DSPy's BootstrapFewShot does
* for marola's trainset size; not a port of DSPy. */
def compile(client: LlmClient, sig: Signature, trainset: List[Example], metric: Metric,
maxDemos: Int = 3, threshold: Double = 1.0): CompiledPrompt < Sync
object Evaluate:
def run(client: LlmClient, prompt: CompiledPrompt, devset: List[Example], metric: Metric): Double < Sync
extension (p: CompiledPrompt) def render(sig: Signature): String // the JSON CompiledPrompt.loadFromString reads
PromptMetrics.jellyfishAndWhale and PromptMetrics.reviewJsonWellFormed port the two Python
metrics with the same weights. The two signatures and trainsets move to core/prompt/Programs.scala
(the instructions text verbatim). cli/prompt/CompilePrompts (just prompt-compile) wires
AppConfig.llmClient, runs both compiles plus Evaluate on a held-out split (FUTURE-WORK.md
§4.1: the dev rows are new examples, added in the same task), writes the two resources, and logs
params/metrics/artifacts to AppConfig.runLedger (MIP-0010) under marola/prompt-compile. The
Scala side already replays via CompiledPrompt.buildMessages; the compiler bootstraps through the
very same method, so compile and serve cannot drift.
5.2 The llm4s/ module — marola-llm4s, depends on core, used by cli¶
build.sbt gains lazy val llm4s = project.dependsOn(core) with "org.llm4s" %% "llm4s-core" %
"0.4.1"; cli depends on it. core/ and local/ are untouched, and just quality gains a
dependency-boundary check (sbt local/dependencyList core/dependencyList must contain no
org.llm4s) so the rule is enforced, not remembered. Why a module and not a replacement of
LlmClient: (a) pre-1.0 coordinate churn stops at the adapter; (b) the local default keeps its
30-line client and its native image.
package marola.llm4s
enum LlmFailure: // marola's view of org.llm4s.error.LLMError
case Auth(msg: String); case RateLimited(msg: String); case Network(msg: String)
case Service(msg: String); case Invalid(msg: String); case Other(msg: String)
final class Llm4sLlmClient(client: org.llm4s.llmconnect.LLMClient) extends LlmClient:
def complete(messages: List[ChatMessage]): String < Sync = // v1 keeps the trait's shape
Sync.defer(client.complete(toConversation(messages))).map {
case Right(c) => c.message.content
case Left(err) => throw Llm4sLlmClient.Failed(LlmFailure.from(err)) // OQ5: Abort[LlmFailure] instead
}
object Llm4sLlmClient:
def ollama(baseUrl: String, model: String): LlmClient // OllamaConfig → LLMConnect.getClient
def openAiCompatible(baseUrl: String, model: String, apiKey: String): LlmClient // Ollama /v1 too (OQ2)
AppConfig: MAROLA_LLM_PROVIDER=local | llm4s-ollama | llm4s-openai; default stays
local until the parity gate in §7 passes (OQ1). TracedLlmClient (MIP-0010 task 6) wraps this
client like any other, so traces arrive without llm4s's tracing.
5.3 Structured output — Reviewer¶
When the provider path supports it (llm4s-openai), Reviewer asks for
ResponseFormat.JsonSchema(score:int, verdict:enum, final_summary:string, strict = true) and
drops the extractJsonObject fallback for that path; the fallback stays for local and
llm4s-ollama. Parsing still lands in the existing ReviewResult.
5.4 The agent — after MIP-0002, Phase 1¶
The summarize→review pipeline stays two plain calls: a single completion gains nothing from an agent
loop. Where the loop earns its place is the bot's free-form /ask (MIP-0002): one llm4s Agent
over marola's four tools as ToolFunctions (find_nearby_beaches, get_swim_recommendation,
get_water_quality, ask_ocean_question), maxSteps = 4, output guardrails JSONValidator /
GroundingGuardrail in ValidationMode.Block, and a marola budget on model calls per request
(survey §2). Handoffs (Handoff.to(escalationAgent, …)) are how the third agent becomes
addressable — the human-confirmation gate before any proactive alert stays marola code; a
Block guardrail is not a human gate and never replaces one. Caveat
that shapes this: through llm4s's OllamaClient tool messages are dropped, so on the local default
the loop must go through openAiCompatible at Ollama's /v1 (OQ2) or wait.
5.5 MCP — the HTTP server, and a client¶
org.llm4s.mcp.MCPServer over the same tool definitions gives the HTTP+SSE / Streamable-HTTP
transport ARCHITECTURE.md §5c lacks, bound to loopback locally (just mcp-http), and the piece a
hosted agent's remote MCP tool needs once Phase 3 deploys it (cost gate, human go-ahead). The
stdio server stays on the official Java SDK (llm4s has none). Tool definitions get one source
(cli/agent/Tools.scala) rendering both the SDK's McpSchema.Tool and llm4s's ToolFunction
(OQ4). MCPClient is what a future marola agent would use to consume external servers, noted,
not scoped.
5.6 What is deterministic, what goes through a model¶
Unchanged: everything in scoring/, the water veto, lore. Through a model: the two compiled
prompts (as today), the compile step's bootstrap calls (offline), and, later, the /ask agent's
tool selection, whose outputs are the same deterministic tools the CLI runs.
5.7 Deprecating dspy/ — the steps, as tasks (mip-tasks will number them)¶
core/prompt/compiler, metrics,render, round-trip golden test; no llm4s, no behaviour change.just prompt-compile; re-compile both artifacts againstllama3.2;just e2eandjust benchmarkbefore/after; commit the new artifacts with the ledger run id in the message.- Delete
dspy/and itsrequirements.txt;ci.yml/justfileruff paths dropdspy;flake.nixcomments; README badge (prompts-compiled in Scala) and §"Offline prompt optimization";ARCHITECTURE.md§5a rewritten aroundcore/prompt/;AGENTS.mdmodule list;FUTURE-WORK.md§10 (marola's need is met;ds4sas a library stays a non-marola idea) and §4.1 (built);docs/index.md; MIP-0010 task 7 marked superseded by task 2 above.finetune/build_dataset.pyreads the demos from the JSON artifacts; unchanged. llm4s/module,Llm4sLlmClient,AppConfigproviders, dependency-boundary check.ReviewerJSON schema on the OpenAI-compatible path.Tools.scalasingle source + HTTP MCP server behindjust mcp-http(loopback only).- The
/askagent: blocked on MIP-0002.
Tasks 1-3 do not depend on 4-7 and can merge first; that is the honest shape of "deprecate DSPy".
6. Scoring / safety impact¶
None. Swimability.score, waterVerdict, thresholds and notes are untouched; the reviewer keeps its
role; guardrails and structured output only make the model's output more checkable, never a
substitute for the deterministic layer.
7. Verification plan¶
Unit (all deterministic, no network): PromptMetricsSpec (exact 0.0/0.3/0.7/1.0 values on the six
trainset rows, matching the Python metrics), BootstrapFewShotSpec (scripted LlmClient:
metric-gated selection, maxDemos cap, augmented flag, threshold), CompiledPromptRoundTripSpec
(render then loadFromString equals the loaded checked-in artifacts, both), EvaluateSpec,
Llm4sLlmClientSpec (message mapping; Left → failure), AppConfigProviderSpec (every
MAROLA_LLM_PROVIDER value), ToolsSpec (SDK and llm4s schemas from one definition agree),
McpHttpServerSpec (tools/list over a loopback client). Gate: the dependency-boundary check in
just quality. Live: just prompt-compile against Ollama with MAROLA_MLFLOW_TRACKING_URI set
shows one run with two artifacts; just e2e; just benchmark compared with docs/benchmarks/;
MAROLA_LLM_PROVIDER=llm4s-ollama just run -- --summarize; just native-image still builds and its
size delta is recorded in the PR. Done = dspy/ gone, both artifacts regenerated by Scala with
a ledger run, all gates green, MIP flipped to Implemented with PR numbers and summed Cost:.
8. Risks, limitations, and honest caveats¶
- Weight. 154 jars / ~150 MB for a client marola uses 30 lines of; cloud SDKs, PDF, Office and
speech libraries on the
cliclasspath; Jackson version conflicts; native-image reachability unknown (MIP-0008). Mitigation: the module boundary, the size measurement in §7, and OQ6. - Pre-1.0 churn. One rename already (v0.4.0); the module split will move coordinates again; compatibility promises start after 1.0 and only for Frozen modules. The adapter is the blast wall.
- Scala versions. 3.7.1 TASTy reads fine from 3.9.0 today; the reverse never will: if llm4s moves to a Scala newer than marola's pin, marola must bump first.
- The local provider is the weakest path in llm4s: no tools, no JSON mode through
OllamaClient. marola's ownLocalLlmClientcan do both via/v1(§4.3). llm4s is not an upgrade for the default path; it is the agent/MCP layer. - Tracing. llm4s's OTel exporter is gRPC; MLflow wants HTTP. Use marola's
TracedLlmClient. - No optimiser in llm4s. marola owns the compiler;
MIPROv2is not ported (never used); Langfuse's DSPy instrumentation goes away (never verified anyway); the ledger replaces it. - Re-compiled prompts change wording. Small local models follow instructions imperfectly
(
ARCHITECTURE.md§5a); the reviewer and the benchmark are the check, as today. - Homepage vs roadmap disagree on 2.13: rely on neither; marola is Scala 3 anyway.
9. Alternatives considered¶
- Do nothing. Keeps two toolchains and the admitted skew for a three-example bootstrap. Lost.
- Build
ds4sas a library first (FUTURE-WORK.md§10). Right scope for a general port, wrong for marola's need; §5.1 is what marola needs and it is small.ds4sstays a non-marola idea. - Replace
LlmClientwith llm4s everywhere. Lost on churn and weight on the local default. - Adopt llm4s only for the compile step. Pointless: the compiler needs only
LlmClient. - Kyo's AI modules / sttp-ai instead of llm4s. Cheaper effect fit; neither has an agent loop with handoffs and guardrails. Revisit when marola bumps past RC5 (survey §1.2).
- Make
llm4s-ollamathe default. Rejected until tools/JSON mode reach the local path (OQ1/OQ2).
11. Open questions¶
- Default flip. Criteria for
MAROLA_LLM_PROVIDERdefaulting to an llm4s path, if ever:just benchmarkparity andjust e2egreen for N runs? Human decision. OpenAIClient→ Ollama/v1for tools +response_formaton the local path: run it.- Native image with llm4s (MIP-0008): measure; if it breaks,
clisplits intocli(native, no llm4s) andcli-agent. - One tool-definition source for the Java MCP SDK and llm4s: shape of
Tools.scala. LlmClienterror channel: keep throwing, orString < (Sync & Abort[LlmFailure])across all three clients in task 4 (theenumrule says the latter; the fan-out isMain/Reviewer/MCP).- Wait for llm4s's module split? An
llm4s-ollama/llm4s-agentwithout the cloud SDKs would cut most of §8's weight. Track the roadmap; do not block tasks 1-3 on it. - Which model compiles the artifacts in task 2:
llama3.2(runtime default) or the 8x7b that produced today's demos. Human decision; the ledger records it either way.
Appendix¶
A. The probe (Scala 3.9.0, strict flags) — compiled and run 2026-09-05¶
//> using scala 3.9.0
//> using dep org.llm4s::llm4s-core:0.4.1
//> using options -language:strictEquality -Wvalue-discard -Wnonunit-statement -deprecation
import org.llm4s.llmconnect.{LLMClient, LLMConnect}
import org.llm4s.llmconnect.config.OllamaConfig
import org.llm4s.llmconnect.model.*
import org.llm4s.model.ModelRegistryService
val result = for
registry <- ModelRegistryService.default()
client <- { given ModelRegistryService = registry
LLMConnect.getClient(OllamaConfig("marola-llama3.2", "http://localhost:11434", 8192, 1024)) }
text <- client.complete(Conversation(Seq(SystemMessage("Answer in one short sentence."),
UserMessage("Is 0.4 m swell calm for swimming?")))).map(_.message.content)
yield text
// → OK via llm4s OllamaClient (marola-llama3.2): Yes, a 0.4 meter wave is generally considered very calm …
No warnings under the strict flags. ModelRegistryService is a using parameter of getClient.
B. Signatures confirmed by javap on llm4s-core_3-0.4.1.jar¶
LLMClient.complete(Conversation, CompletionOptions): Either[LLMError, Completion] ·
CompletionOptions(temperature, topP, maxTokens, presencePenalty, frequencyPenalty, tools:
Seq[ToolFunction[?, ?]], reasoning: Option[ReasoningEffort], budgetTokens, responseFormat:
Option[ResponseFormat]) · OllamaConfig(model, baseUrl, contextWindow, reserveCompletion) ·
Agent(client).run(String, ToolRegistry,
Seq[InputGuardrail], Seq[OutputGuardrail], Seq[Handoff], Option[Int], Option[String],
CompletionOptions, AgentContext): Either[LLMError, AgentState] · MCPServer(MCPServerOptions(port,
path, name, version, apiKey: Option[String], host), Seq[ToolFunction[?, ?]]).start(): Either[
Exception, Unit] · MCPServerConfig.stdio(name, command, timeout) / sse / streamableHTTP ·
MCPClient.getTools: Either[String, Seq[ToolFunction[?, ?]]] · OpenTelemetryConfig(serviceName,
endpoint, headers: Map[String, String]) · Tracing.traceCompletion(Completion, String):
Either[LLMError, Unit] · TypedAgent[I, O].execute(I)(using ExecutionContext): Future[Either[
LLMError, O]] · Plan.topologicalOrder, PlanRunner.execute(Plan, Map, CancellationToken).
C. Sources fetched 2026-09-05¶
github.com/llm4s/llm4s (README, LICENSE, build.sbt, project/Dependencies.scala, releases,
releases/tag/v0.3.0, modules/ tree, modules/core/.../provider/OllamaClient.scala,
.../provider/OpenAIClient.scala, .../model/ResponseFormat.scala, .../http/Llm4sHttpClient.scala,
modules/mcp/.../MCPServer.scala, modules/trace-opentelemetry/.../OpenTelemetryTracing.scala);
llm4s.org (/, /reference/roadmap, /reference/v1-scope, /migrations/0x-to-1x, /guide/agents/,
/guide/agents/handoffs, /guide/agents/guardrails, /guide/providers, /guide/observability/,
/getting-started/ollama-quickstart, /getting-started/configuration, /llm4s-api-spec,
/tool-calling-api-design, /PRODUCTION_DEPLOYMENT); Maven Central via coursier (cs resolve,
cs complete-dep, cs fetch); local Ollama /api/tags and /v1/chat/completions. Not fetched:
/guide/tools, /guide/structured-output (404 — do not exist under those names).
D. llm4s 0.4.1 — detailed notes behind the §4.1 table¶
- Coordinates, version, licence.
org.llm4s:llm4s-core_3:0.4.1on Maven Central (coursiercomplete-dep; releases page: v0.4.0 29 Aug "rename-only release … artifact prefix standardization", v0.4.1 29 Aug "CI/docs … Maven relocation POMs"; v0.3.0 21 Feb added the OpenTelemetry backend,MCPServer, Gemini/DeepSeek,Resultmigration). MIT (LICENSE, Rory Graves, 2025). Published at 0.4.1:llm4s-core_3,llm4s-observability-otel_3,llm4s-workspace-client_3,llm4s-workspace-shared_3,llm4s-knowledgegraph-neo4j_3— thellm4s-mcp/llm4s-rag/llm4s-memorymodules named in itsbuild.sbtand 1.0 scope are not separately published; their packages ship insidellm4s-core(jar inspection:org.llm4s.mcp32 classes,org.llm4s.rag,org.llm4s.agent.memory). The 1.0-scope page says so itself: "the latest release (v0.4.1) still ships as a single artifact". - Scala/JDK. Built with Scala 3.7.1 ("Scala 3 only (3.7.1)", roadmap and 1.0 scope; CI on JDK
21). The homepage's "Scala 2.13.x fully supported" contradicts both; treat as stale; the
_2.13artifacts on Central carry the pre-rename names. Verified live: a probe compiled with Scala 3.9.0 and marola's flags (-language:strictEquality -Wvalue-discard -Wnonunit-statement) againstllm4s-core_3:0.4.1and completed one chat turn against the local Ollama (marola-llama3.2) through llm4s'sOllamaClient(Appendix A). - Dependency closure (the decisive number).
cs resolveofllm4s-core_3:0.4.1: 154 jars, ~150 MB on disk, versus ~32 MB for marola's current Kyo + MCP SDK + logback set. It includes a vendor OpenAI Java SDK 1.0.0-beta.16 (+ its core library, Netty, Reactor),com.anthropic: anthropic-java:2.42.0(+ OkHttp, Kotlin stdlib), the AWS SDK (S3, STS), PDFBox, POI, Tika, PostgreSQL, sqlite-jdbc, HikariCP,vosk(speech), Prometheus, cats-core, upickle/ujson, Jackson 2.19. Coursier flags conflicts (Jackson 2.17.2→2.19.4 under that SDK's core library, JNA 5.7→5.19). - Client API (javap):
LLMClient.complete(Conversation, CompletionOptions): Either[LLMError, Completion],streamComplete,validate,close(AutoCloseable);type Result[+A] = Either[LLMError, A], a sealedLLMErrorhierarchy of 20+ cases;LLMConnect.getClient(config)( using ModelRegistryService);Completion(id, created, message, usage: Option[TokenUsage]). - Providers (javap + source):
OpenAIClientwraps a vendor OpenAI Java SDK, API key only (KeyCredential, noTokenCredential);OllamaClientposts to Ollama's native/api/chatoverjava.net.http, streams, and drops tool messages ("Tool messages are not supported by Ollama chat API; drop them") and sends noresponseFormat. Also Anthropic, Gemini, DeepSeek, Cohere, Mistral, OpenRouter, Vertex AI, Z.ai. - Structured output.
CompletionOptions.responseFormat: Option[ResponseFormat];ResponseFormat .Json | .JsonSchema(schema: ujson.Value, name, strict);ResponseFormatMapper .toOpenAIResponseFormat, OpenAI-path only. No case-class→schema derivation found in the jar. - Tools.
ToolFunction[T, R: ReadWriter](name, description, schema, handler: SafeParameterExtractor => Either[String, R]),ToolBuilder, aSchemaDSL (string,number,object,array,.withEnum…),ToolRegistry.execute: Either[ToolCallError, ujson.Value]. ScalaMeta code-gen is claimed in the README; not found in the published jar. - Agent.
new Agent(client);run(query, tools, inputGuardrails, outputGuardrails, handoffs, maxSteps (default 50), …): Either[LLMError, AgentState],runStep,runWithEvents,continueConversation;Handoff.to(agent, reason);AgentStatusInProgress / WaitingForTools / Complete / Failed / HandoffRequested.agent.orchestration:Plan(nodes, edges,topologicalOrder,getParallelBatches),TypedAgent[I, O]andPlanRunner,Future-based. - Guardrails.
InputGuardrail/OutputGuardrail,ValidationMode.Block | Warn | Log; 18 built-ins (JSONValidator,RegexValidator,GroundingGuardrail,SourceAttributionGuardrail,PromptInjectionDetector,LLMFactualityGuardrail, …). No human-approval concept. - MCP. Client:
MCPClient.getTools: Either[String, Seq[ToolFunction]]overStdioTransport( command),SSETransport,StreamableHTTPTransport(MCPServerConfig.stdio/sse/streamableHTTP). Server:MCPServer(MCPServerOptions(port, path, name, version, apiKey, host), tools)oncom.sun.net.httpserver; HTTP+SSE (protocol 2024-11-05) and Streamable HTTP (2025-06-18); bearer auth with constant-time compare; "refuses to bind a non-loopback host" without a key. No stdio server. - Tracing.
Tracingtrait (traceEvent,traceCompletion,traceTokenUsage,traceCost,traceToolCall,traceAgentState); modeslangfuse | opentelemetry | console | collector | noop. The OTel module (llm4s-observability-otel) usesOtlpGrpcSpanExporter, one span named "LLM Completion" withgen_ai.request.modelandgen_ai.usage.*. MLflow ingests OTLP/HTTP only (MIP-0010 §4.3), not compatible as-is. - Also present, not needed here: embeddings (
OllamaEmbeddingProvider),rag+ vector stores (sqlite/pgvector) + RAGAS-stylerag.evaluation, reranker, memory + consolidation, context-window pruning,ReliableClient(retry, circuit breaker), middleware, image generation, speech (vosk), knowledge graph, workspace runner.eval.dataset(Dataset,Example[I, O],ExampleSelector,JsonlCodec) is a dataset model, no optimiser. - Prompt optimisation / a DSPy equivalent: none. A class-name search of the jar for
Prompt/Optim/FewShot/Template/Signature finds only
PromptInjectionDetector,ConsolidationPromptsand an imageTemplateName; README and roadmap never mention DSPy. - Roadmap. "Pre-1.0, API stabilizing"; 2026 phases (module boundaries, provider capability
matrix, Java/Kotlin/Spring interop, security hardening); "v1.0 date intentionally not fixed";
binary compatibility (MiMa) and "deprecate before removing" apply to Frozen modules after
1.0. The pending module split (
llm4s-core/llm4s-agent/llm4s-ollama…) means at least one more coordinate change; v0.4.0 was already one. - Effect model. Plain
Either,Futurefor orchestration; no cats-effect/ZIO/Kyo. Fits marola's rule as a boundary call wrapped inSync.defer, theLeftlifted intoAbort.
Not checked: the Streamable-HTTP server against a real MCP client; llm4s's OpenAIClient
pointed at Ollama's /v1 (would give tools + response_format on the local path, plausible, not
run); a GraalVM native image with llm4s on the classpath (MIP-0008, Netty/Jackson/OkHttp
reachability); llm4s's MCPClient against marola's stdio server.