MIP-0057: GCP as an opt-in cloud backend¶
| Status | Draft |
| Author | Claude Sonnet 5, from a repo-maintainer brief of 2026-09-15 |
| Created | 2026-09-15 |
| Phase | 0 — this MIP is a design exercise only; if accepted, the first build task (§9's recommended path) is Phase 0 work (no Telegram bot dependency), but any real GCP provisioning is gated by AGENTS.md's cost-safety rule |
| Related | docs/2-Building-marola/ARCHITECTURE.md §5 (the pluggable integrations), MIP-0012 (llm4s — a Scala-native LLM/agent layer; whatever LlmClient shape it lands on is what a gcp/ Vertex AI backend would implement against), MIP-0055 (KnowledgeStore's opt-in backend — same shape question for GCP, not addressed here) |
| Effort | S for this MIP itself (a survey, no code); the thing it recommends building (§9, LLM-only gcp/ module) is M — a new sbt module (one trait implementation, one dependency, no new CI workflow, no new trait) |
| Gain | infra/dev-loop (an opt-in cloud tier past the local default); cost/ops (a free-tier shape to pick from per integration, e.g. Cloud Vision's flat 1,000-units/month free tier) |
| Effort vs Gain | do when X lands — this MIP does not itself justify building anything; the recommended first slice (§9) is a cheap win sized on its own once someone wants a hosted LLM, but nothing here is blocking or blocked-on by product need today |
| Depends on | Not blocked by any MIP. Depends on MIP-0012's LlmClient/agent-layer shape settling first if the LLM slice is built after that lands, since a gcp/ module should implement whatever trait shape exists at build time, not a shape this MIP freezes. No Phase 1 gate (Telegram bot) — this is Phase 0 design work, and even the recommended build slice needs no Telegram bot. A paid GCP resource is gated by AGENTS.md's cost rule: explicit human go-ahead before any gcloud/terraform apply/console provisioning step |
| Blocked by | none |
| Risk | A second path (local / GCP) per integration adds code paths that are "compiles, never run against a live account" — adding GCP without ever provisioning it just adds unverified surface area. The mitigating design choice (§9) is to add only one GCP-backed trait implementation at a time |
| Cost so far | — |
1. Summary¶
marola's docs/2-Building-marola/ARCHITECTURE.md §5 treats several capabilities as pluggable behind traits, each
with a free local default. This MIP asks whether Google Cloud Platform (GCP) should be an opt-in
backend for some or all of those traits, never replacing the local-first default. It names the
closest GCP service for each integration, states what could and couldn't be price/terms-verified
this session, and recommends starting, if at all, with a single, low-risk slice (Vertex AI/Gemini
behind LlmClient) rather than a parallel build across every integration. It commits to nothing:
no code changes, no module, no provisioning.
2. Motivation¶
Every one of marola's pluggable integrations (ARCHITECTURE.md §5) runs locally, with no paid
opt-in tier. That's a deliberate, working pattern. Three concrete reasons to look at GCP
specifically, not "some cloud" in the abstract:
- Model marketplace breadth. Vertex AI's Model Garden serves Google's own Gemini family alongside third-party and open models from one API surface, a natural opt-in for §5a.
- Free-tier shape, not just price. Cloud Vision's free tier is a flat 1,000 units/month across the whole account (confirmed this session, see §4.4). The shape matters as much as the price, and it is worth knowing before any paid tier is turned on.
- Firestore/Cloud SQL/BigQuery for §5d's sighting store, Google Maps Platform for §5b's route distance, and Cloud Logging/Trace for §5f's telemetry are the natural GCP services once the LLM question is on the table. A maintainer choosing an opt-in tier has no survey to make that decision against today.
This MIP does not argue GCP is cheap. It lays out what's actually available and priced, so a future
"turn on the paid tier" decision (still gated by AGENTS.md's cost rule) has real numbers instead
of assumed ones.
3. User-visible change¶
None from this MIP alone, no code changes. If §9's recommended slice is later built and a human opts in through config, the user-visible change is the same CLI/MCP output shape, generated by a different backend. No new output field, no new flag surface beyond the one new config value.
4. Data sources and dependencies reviewed¶
One subsection per integration, matching ARCHITECTURE.md §5's own lettering (5a/5b/5d/5e/5f) plus
§5c (agentic tool access, which needs no cloud service) and §5g/5h (regional/local-only, so GCP
doesn't apply).
4.1 §5a — Query synthesis: Vertex AI (Gemini)¶
- GCP side: Vertex AI serves Gemini models (and Model Garden third-party models) via a REST/ gRPC API. A chat call is a plain HTTP request, no SDK required.
- Pricing (checked live,
ai.google.dev/gemini-api/docs/pricing, 2026-09-15): Gemini 2.5 Flash is $0.30/1M input tokens, $2.50/1M output tokens (text/image/video input; $1.00/1M for audio input). Gemini 2.5 Flash-Lite is cheaper still: $0.10/1M input, $0.40/1M output. Both list a free tier, but it is scoped to Google Search grounding requests specifically (a shared 500 requests/day quota), not general chat completions, so, unlike Ollama's genuinely free local path, there is no meaningfully free way to run Gemini chat completions at marola's actual usage (query synthesis, not grounded search). - Auth: GCP's answer to "never hardcode a key" (
AGENTS.md) is a service account plus Workload Identity Federation (WIF), which lets a workload authenticate without a downloaded JSON key. Not independently verified against a live GCP project this session (none provisioned, per the cost-safety rule). The shape is well-documented, but documented is not confirmed. - Migration effort: low.
LlmClientis already a trait; a second implementation (VertexAiLlmClient) is additive. The real dependency risk is MIP-0012 (llm4s) potentially reshapingLlmClientbefore this is built. Build against whatever trait shape exists at the time, not this MIP's description of it.
4.2 §5b — Real travel distance: Google Maps Platform (Routes API)¶
- GCP side: Google Maps Platform's Routes API (
computeRoutes) returns road distance. marola uses straight-line haversine distance today. - Pricing/free-tier shape (checked live,
mapsplatform.google.com/pricing, 2026-09-15): Google Maps Platform moved off its old $200/month credit model in March 2025 to a per-SKU free-call allotment instead, reported by the fetched page as 10K free calls/SKU/month on the "Essentials" plan, 5K on "Pro," 1K on "Enterprise," with fixed-price subscription tiers ($100/$275/$1,200 per month) as an alternative to pure pay-as-you-go, and volume discounts (20%–80%) at high monthly call counts. What that costs at marola's actual call volume needs a real usage estimate. - Auth: an API key taken from the environment.
- Migration effort: low-to-medium. marola has no routing code today, so a Routes API call is
new code.
computeRoutesis a POST with a JSON body and a field mask, not a query string.
4.3 §5d — Sighting reports: Firestore¶
- GCP side: Firestore is the closest shape match, a managed NoSQL document store that fits
Sighting's access pattern (append + query-by-beach). Cloud SQL (relational) and BigQuery (analytical/columnar) are GCP options too, but neither matches that pattern as directly. They're included only for completeness per the task brief, not as serious contenders. - Pricing/free-tier shape (checked live,
firebase.google.com/docs/firestore/quotas, 2026-09-15): Firestore's free tier is a daily quota, not monthly: 50,000 reads/day, 20,000 writes/day, 20,000 deletes/day, 1 GiB stored. That is effectively free at marola's current, near-zero sighting volume. - Auth: Firestore via a GCP service account (WIF for a deployed workload, same mechanism as §4.1), so a Firestore backend built with WIF from day one needs no key in the environment.
- Migration effort: low-to-medium.
SightingStore'srecord/recentForshape maps directly onto Firestore's collection/document model.
4.4 §5e — Photo analysis: Cloud Vision API¶
- GCP side: Cloud Vision API's label/caption-style detection. marola uses a local Ollama vision model today.
- Pricing/free-tier shape (checked live,
cloud.google.com/vision/pricing, 2026-09-15): first 1,000 units/month free across all features; $1.50 per 1,000 units from 1,001 to 5,000,000 units/month, $1.00/1,000 beyond that (label detection tier; other features like Web Detection are priced differently, e.g. $3.50/1,000). This free tier is a flat per-account monthly cap. - Auth: GCP service account/WIF, same as §4.1.
- Migration effort: low.
VisionClient.describe(imageBytes)is a single-method trait; aCloudVisionClientis additive. One real design fork: Cloud Vision returns structured labels/annotations, not a free-text caption by default.describe's return type (presumably a caption string today) may need to compose Cloud Vision's label list into a sentence itself, a small but real implementation difference, not a drop-in swap.
4.5 §5f — Observability: Cloud Logging/Cloud Trace¶
- GCP side: two separate services: Cloud Logging (log ingestion) and Cloud Trace (distributed
tracing/spans), both under the wider "Google Cloud Observability" (formerly Stackdriver)
umbrella.
Tracing.withSpan/llmSpan(ARCHITECTURE.md§5f) map onto Cloud Trace specifically, not Cloud Logging. - Pricing/free-tier shape (search-result summary, not independently fetched from the source
page this session, see Appendix "Not checked"): commonly reported as 50 GiB/month free log
ingestion (then $0.50/GiB) for Cloud Logging, and 2.5 million free spans/month (then $0.20/million
spans) for Cloud Trace. Flag explicitly: these two numbers came back from a WebSearch summary,
not a direct fetch of
cloud.google.com/logging/pricing/cloud.google.com/trace/pricing, both of which returned unusable/truncated content this session, treat them as plausible, not confirmed, and re-verify against the live pages before using either number in a real decision. - Auth: GCP service account/WIF, consistent with every other GCP integration above.
- Migration effort: medium. The core repo already has a vendor-neutral
Tracingtrait (core/observability/Tracing), so a new backend is additive. OpenTelemetry has an official Google Cloud exporter, so aGoogleCloudTracingimplementation would plug into the same OTLP-shaped patternMlflowTracingalready uses, rather than needing a wholly new export mechanism.
4.6 §5c — Agentic tool access: no GCP service needed¶
§5c's MCP server runs over stdio and needs no cloud account at all. A hosted variant using the HTTP/SSE MCP transport is not built. Vertex AI Agent Builder / Agent Engine would be the GCP-side host if that variant were ever built, but nothing runs today, so there is nothing concrete to design. Noted for completeness, not pursued further here.
4.7 §5g/§5h — no cloud service applies¶
Bathing-water quality (§5g) is sourced from Brazilian regional agencies, not a cloud service
(ARCHITECTURE.md §5g explicitly: "none — regional agencies, not a cloud service"). Ocean
knowledge/RAG (§5h) runs from the local corpus. A Vertex AI Search backend would be a reasonable
follow-up, but is out of scope here.
Pick¶
No pick is made here, see §9 for the recommended path. This section's job was surveying what exists, not choosing a winner.
5. Design¶
No code is proposed to change in this MIP. If a future MIP or task builds §9's recommended slice, the shape would be:
// core/ — no change to the LlmClient trait signature itself (whatever MIP-0012 settles it to)
trait LlmClient:
def complete(messages: List[ChatMessage]): String < (Abort[LlmError] & Sync)
// gcp/ — new sbt module depending on core/
final class VertexAiLlmClient(project: String, location: String, model: String) extends LlmClient:
// REST call to Vertex AI's generateContent endpoint, auth via a service account /
// Workload Identity Federation token — plain REST, not the full agent SDK
def complete(messages: List[ChatMessage]): String < (Abort[LlmError] & Sync) = ???
AppConfig.llmClient would grow a way to select GCP, returning None with a clear message if GCP
is selected but not fully configured. build.sbt would gain a gcp/ project depending on core/
(zero dependency from local/ or core/ back onto it). Nothing here is built by this MIP.
6. Scoring / safety impact¶
None. Every integration this MIP discusses (§5a/§5b/§5d/§5e/§5f) is either non-safety-scoring
(query synthesis is deliberately kept out of Swimability.score, ARCHITECTURE.md §5a) or
infrastructure (storage, distance refinement, telemetry). No GCP backend changes what
Swimability.score computes or how, the pluggable integrations were designed precisely so a
backend swap can never touch scoring logic, and this MIP doesn't change that design.
7. Verification plan¶
This MIP proposes no code, so there is nothing to unit-test yet. If §9's recommended slice is later built as its own task/PR, that PR's own verification plan should require, at minimum:
- A unit test for
AppConfig.llmClient's new GCP selection (config-present/config-absent cases). - A compile-only check of
VertexAiLlmClientagainst the real Vertex AI REST API shape, explicitly labeled unverified-against-a-live-account until a human opts in and provisions a real GCP project. just build && just test && just qualitygreen, same gate as every other change.
"Done" for this MIP is: merged as Draft, indexed in docs/MIPs/README.md, with every §4 claim
either cited live or explicitly flagged unverified, not a working GCP integration.
8. Risks, limitations, and honest caveats¶
- Every price/free-tier figure in §4 was fetched on 2026-09-15 and can drift, GCP's own Maps Platform pricing already changed shape once (the March 2025 move off the $200/month credit, confirmed live this session), so a figure here going stale is not hypothetical.
- §4.5's Cloud Logging/Cloud Trace figures are the weakest sourced numbers in this MIP, they came back from a WebSearch summary after two direct page fetches failed (truncated/404), and are flagged as such rather than presented as confirmed.
- Three GCP price points cited here (Vertex AI Gemini, Cloud Vision) came from an AI-summarized fetch of the vendor's own page, not a raw HTML/text capture, the fetch tool itself does model- assisted extraction, which is a real, if small, risk of paraphrase-introduced error even when the page was genuinely retrieved (as opposed to the cases where the tool visibly failed and fell back to "general knowledge," which are separately flagged in the Appendix).
- No GCP resource has been, or will be, provisioned to verify any of this against a live account, every auth-model claim (service accounts, Workload Identity Federation) is a documented shape, not something exercised against a real GCP project.
- Adding a backend option per integration is a real, if small, maintenance cost even before anything is built: this MIP itself is now a doc another maintainer has to read. Worth stating plainly rather than pretending a survey MIP is free.
9. Alternatives considered¶
- Do nothing (stay local-only). Loses nothing marola uses today, every local default still works with zero cloud account. The cost is optionality: anyone who wants a paid tier has none. Not unreasonable. This MIP's real argument is optionality, not urgency.
- Add GCP as a backend for every integration at once. Rejected: one integration at a time keeps each backend small enough to verify before the next one starts.
- Recommended path: add
gcp/as a fully opt-in backend module behind the existing core traits, starting with LLM only (Vertex AI/Gemini), if and when a maintainer wants a hosted LLM. This is the lowest-risk single slice:LlmClientis already a trait (so adding an implementation changes no interface), query synthesis is explicitly non-safety-scoring (§6), and §4.1's pricing is cited (Gemini 2.5 Flash-Lite at $0.10/$0.40 per 1M tokens). The task brief's own suggested order (LLM first) held up against a read of the actual architecture, nothing in §4/§5's other integrations argued for starting anywhere else. This MIP does not commit to building even that slice, it only says that if one integration is picked first, the evidence points here. - A GCP-specific cost/deployment guard (
guard-gcp.sh), built now, ahead of any GCP code. Rejected as premature: there is no GCP code or command in this repo to guard yet, so building aguard-gcp.shtoday would be guarding an empty directory. Stated instead as an explicit prerequisite (§11), a real task for whoever picks up §9's recommended slice, before anygcloud/Terraform command that provisions anything, not something this MIP builds.
11. Open questions¶
- Which integration, if any, does a maintainer actually want a second vendor for? This MIP argues LLM (§9) is the lowest-risk first slice if one is picked, not that one must be picked at all, that's a product decision, not something this MIP can settle.
- A same-day price check for every §4 service against marola's actual (tiny) usage volume.
- §4.5's Cloud Logging/Cloud Trace numbers need a direct re-fetch of
cloud.google.com/logging/pricingandcloud.google.com/trace/pricing, both returned unusable content (truncated / 404) this session; the figures used came from a search-result summary instead and are flagged as such throughout. - A
guard-gcp.sh-style PreToolUse hook, and the.claude/settings.jsondeny-list entries it would need (literal-prefix denials forgcloud deployment-manager/terraform apply/gcloud projects create) is a named prerequisite for §9's slice, not something to design in more detail until someone actually starts that build, flagged here so it isn't forgotten when that day comes. - Follow-up MIP: a Vertex AI Search backend for §5h, not designed here (§4.7).
Appendix¶
Checked live¶
https://ai.google.dev/gemini-api/docs/pricing, fetched 2026-09-15: returned Gemini 2.5 Flash ($0.30/1M input text/image/video, $1.00/1M input audio, $2.50/1M output) and Gemini 2.5 Flash-Lite ($0.10/1M input text/image/video, $0.30/1M input audio, $0.40/1M output) pricing, plus a free tier scoped to Google Search grounding (shared 500 requests/day), not general chat completions.https://mapsplatform.google.com/pricing/, fetched 2026-09-15: returned the post-March-2025 per-SKU free-call model (10K/5K/1K free calls per SKU per month on Essentials/Pro/Enterprise), fixed subscription tiers ($100/$275/$1,200/month), and volume discounts (20%–80%) at high usage.https://firebase.google.com/docs/firestore/quotas, fetched 2026-09-15: returned Firestore's free-tier daily quota, 50,000 reads/day, 20,000 writes/day, 20,000 deletes/day, 1 GiB stored.https://cloud.google.com/vision/pricing, fetched 2026-09-15: returned Cloud Vision's free tier (first 1,000 units/month across all features) and tiered per-1,000-unit pricing ($1.50 from 1,001–5,000,000 units/month, $1.00 beyond, with other features like Web Detection priced separately at e.g. $3.50/1,000).https://cloud.google.com/vertex-ai/generative-ai/pricing, fetched 2026-09-15: failed, the fetch tool reported truncated content and fell back to "based on general knowledge" figures, which are explicitly not used anywhere in §4 of this MIP (the ai.google.dev fetch above supplied the real, cited Gemini numbers instead).https://cloud.google.com/logging/pricing, fetched 2026-09-15: failed, truncated content, no usable pricing extracted.https://cloud.google.com/trace/pricing, fetched 2026-09-15: failed, the tool reported a 404 at that URL.https://cloud.google.com/firestore/pricing, fetched 2026-09-15: failed, truncated content; superseded by the successfulfirebase.google.com/docs/firestore/quotasfetch above.https://cloud.google.com/stackdriver/pricing, fetched 2026-09-15: failed, truncated content, no usable pricing extracted.
Not checked¶
- Cloud Logging's "50 GiB/month free, then $0.50/GiB" and Cloud Trace's "2.5 million spans/month free, then $0.20/million spans" (§4.5) came from a WebSearch results summary on 2026-09-15, not a direct fetch of either service's own pricing page (both direct fetches failed, see above). Treat both numbers as plausible, not confirmed.
- Workload Identity Federation's exact setup mechanics (how a Cloud Run job or a GitHub Actions workflow attaches to it) were not verified against GCP's own docs this session, described here only at the level of "the WIF concept exists," which is a lower bar than this repo normally holds for a claim that would drive a real auth implementation.
- Vertex AI Agent Builder / Agent Engine (§4.6) is named from background knowledge, not a fetched page, flagged as unverified since §4.6 doesn't recommend building against it anyway.
- Google Maps Platform's Routes API request/response shape (
computeRoutes, POST with a field mask) is stated from background knowledge, not verified against Google's own API reference this session, a real implementation would need that reference checked first.