Skills roadmap
A roadmap of skills to actually practice, in order, using this repo as the vehicle.
Each item names the concrete marola artifact that exercises it and whether it's practice-ready
today or needs something built first. Work top-to-bottom within a section; sections are ordered
from single-model work to multi-agent systems.
Stage 1 — AI solution planning
| Skill |
Practice it via |
Ready now? |
| Responsible AI: stating limitations honestly |
ARCHITECTURE.md §8/§9 — write your own one-paragraph "known limitations" section for a feature you add, in that style, before calling it done |
Yes |
Stage 2 — Generative AI implementation
| Skill |
Practice it via |
Ready now? |
| Calling a model through an OpenAI-compatible endpoint |
LocalLlmClient.complete (Ollama's /v1/chat/completions) |
Yes |
| Structured/parsed output from a model |
CompiledPrompt.buildMessages + LlmClient.extractContent (summary field), Reviewer.extractJsonObject (JSON-from-prose fallback) |
Yes |
| Systematic prompt optimization (not hand-tuning) |
dspy/compile_recommendation_prompt.py — run it yourself against llama3.2:1b, inspect the compiled recommendation_prompt.json, then hand-edit the trainset and re-run to see the artifact change |
Yes — see RUN-LOCALLY.md |
| RAG |
Not built. First real exercise: implement core/knowledge/KnowledgeStore per FUTURE-WORK.md §9.1 against a small local corpus |
Design only — build it to practice this |
| Fine-tuning a small open model |
FUTURE-WORK.md §9.1 step 4 (Ollama Modelfile + QLoRA-style adapter) |
Design only |
Stage 3 — Agentic solutions
| Skill |
Practice it via |
Ready now? |
| Exposing app logic as MCP tools |
cli/agent/SwimConditionsMcpServer.scala — read it, then add a new tool (e.g. ask_ocean_question once §9.1 exists) yourself |
Yes |
| Testing an MCP server without a full agent client |
Pipe raw JSON-RPC to the server's stdin yourself (ARCHITECTURE.md §5c's Status note describes how this was verified) — do this once by hand before trusting any higher-level client |
Yes |
| Multi-step agent pipelines (plan → act → critique) |
Recommender.bestPerBeachTomorrow → Reviewer.review — trace one real request through both LLM calls end to end with just run -- --summarize |
Yes |
| Recognizing when orchestration frameworks are and aren't worth adopting |
FUTURE-WORK.md §5 (workflows4s review) — do your own version of this exercise on a framework not yet reviewed here before adding one |
Yes, as a practice exercise |
Stage 4 — Computer vision
| Skill |
Practice it via |
Ready now? |
| Local multimodal model calls |
local/vision/LocalVisionClient via just run -- --analyze-photo <path> |
Yes |
| Closing the loop: vision output feeding a decision |
Not built — SightingStore records vision-analyzed sightings but nothing yet feeds them back into Swimability's heuristics (ARCHITECTURE.md §8). Build the calibration step to practice this |
Design only |
Stage 5 — NLP / text analysis
| Skill |
Practice it via |
Ready now? |
| Structured extraction from unstructured model output |
Reviewer.extractJsonObject |
Yes |
| A first-class text-analysis use case (not just a parsing fallback) |
Not built — the marine-literature ingestion step in FUTURE-WORK.md §9.1 (step 2: pulling structured hazard facts out of prose bulletins) is the concrete gap-closer |
Design only — build it to close this gap for real |
Stage 6 — Scala/engineering craft (implementation quality, exercised on every task above)
marola's own code was reviewed for this directly (September 2026). Treat these as the concrete
skill gaps to close by practicing on this codebase, not abstract advice:
- Typed error channels over
throw+catch-all. Every I/O boundary (Http.scala, Json.scala,
Reviewer.scala, vision clients) defines a real exception type but throws it and catches
it as bare Throwable via Abort.catching[Throwable] at the call site, which discards the type
information Kyo's Abort[E] effect exists to track. Practice: pick one client, change its
signature to ... < (Sync & Abort[HttpError]), and thread the typed error through instead.
- Kyo's own bulk-effect combinators over hand-rolled recursion.
Recommender.traverse/
traverseSingle reimplement what kyo.Async.foreach/collectAll already provide (confirmed
present in the exact pinned kyo-core/kyo-combinators 1.0.0-RC5 jars via javap), kept
hand-rolled specifically because only map/flatMap on < Sync were confirmed at the time.
Practice: verify live whether Async.foreach works over < Sync callers (does the effect type
widen to Sync & Abort[...] correctly?) and replace the hand-rolled version if so. A real,
scoped verification exercise, not a guess.
- Property-based tests where they're an obvious fit.
Swimability.score's 0-100 clamped range
and monotonic threshold behavior is a textbook ScalaCheck property (score is never outside
[0,100]; a strictly worse wave height never increases the score), none exist yet; only
example-based tests do (SwimabilitySpec).
- Test coverage is thin outside the one pure module. ~1900 lines of main source, ~280 lines of
test source, and exactly one deterministic unit-test file (
SwimabilitySpec): the hand-rolled
JSON parser (Json.scala, escape sequences and all), Recommender's grouping/sorting, and
AppConfig.fromEnv's parsing have zero unit tests today. Practice: write JsonSpec first. It's
the highest-value, lowest-effort gap (pure function, pure input/output, no mocking needed).
- Scalafix, not just scalafmt. Formatting is enforced (
scalafmtCheckAll in CI); nothing
enforces the code's own unwritten conventions (no stray var, no bare throw outside an
Abort.catching boundary, no unused imports) mechanically. Practice: add scalafix with
DisableSyntax rules for exactly the conventions this repo already tries to follow by hand.
What's already solid, worth recognizing rather than only listing gaps: pinned exact dependency
versions with reproducible Nix builds; a real, documented pattern of verifying library claims
against decompiled jars/live calls instead of trusting docs (FUTURE-WORK.md throughout,
EFFECTS-MAP.md); strict compiler flags most Scala 3 codebases skip (-Wvalue-discard,
-Wnonunit-statement, -language:strictEquality, promoted to errors); a genuinely clean
pure-core/effectful-shell split (EFFECTS-MAP.md); deliberate, reasoned dependency minimalism
(multiple libraries reviewed and explicitly not adopted, with the reasoning kept, not just silently
skipped).
Stage 7 — Multi-agent architecture
| Skill |
Practice it via |
Ready now? |
| Recognizing an implicit multi-agent system inside a "simple" pipeline |
Name the two roles already in Recommender/Reviewer as agents before building anything new — this recognition step is the actual skill, not just adding code |
Yes |
| Designing a third agent with a genuinely different role, not a variant of the first two |
FUTURE-WORK.md §9.2 — the escalation/hazard-detection agent |
Design only |
| Naming and documenting an orchestration topology |
Write the topology doc before building the third agent, not after |
Yes, as a practice exercise |
Stage 8 — Multi-agent development
| Skill |
Practice it via |
Ready now? |
| MCP as an agent-to-agent capability-exposure mechanism, not just a chat-tool bridge |
Extend SwimConditionsMcpServer with one tool per agent role |
Design only |
| Shared state/memory across agents |
Extend SightingStore's trait-plus-local-backend pattern to a UserPreferencesStore/conversation-state store (FUTURE-WORK.md §1.5) |
Design only |
Stage 9 — Evaluating and monitoring multi-agent systems
| Skill |
Practice it via |
Ready now? |
| An agent evaluating another agent (generator/critic) |
Reviewer.review — already built; study it as the generator/critic pattern it is, not just a marola feature |
Yes |
| A held-out eval set, not just a training set doing double duty |
FUTURE-WORK.md §4.1's gap — build a real dspy.Evaluate loop over held-out examples |
Design only |
| LLM-call-shaped tracing (prompt/completion/cost/eval-score per span), not just infra tracing |
FUTURE-WORK.md §10 flags this as a real JVM/Scala tooling gap (no Langfuse-equivalent) — extend Telemetry.scala's existing OpenTelemetry plumbing yourself |
Design only, scoped and doable |
Stage 10 — Securing, governing, and deploying multi-agent systems
| Skill |
Practice it via |
Ready now? |
| A human-confirmation gate on autonomous (not just responsive) agent behavior |
AGENTS.md's existing cost-safety gate is the template — design the equivalent for the escalation agent's proactive alerts before building it (FUTURE-WORK.md §9.2) |
Design only |
| Content-safety review on agent output meant for the public |
Same escalation agent — its alert text is the first marola output that isn't only shown to the person who asked for it |
Design only |
| Audit trails across agent boundaries |
Extend Telemetry.scala rather than building a new system |
Design only |