MIP-0011: Claude Code best practices in this repository — hooks as gates, a shared permission allowlist, path-scoped rules, subagents and skills the dev flow already implies¶
| Status | Implemented — all eleven tasks merged (docs/MIPs/MIP-0011.tasks.md); verified with /code-review ultra (ultrareview) against the full stack diff on 2026-09-06, see §7 |
| Author | Claude Fable 5.1, for M. Hoffmann (request of 5 Sep 2026: "MIP for implementing Claude Code best practices, by Anthropic and community") |
| Created | 2026-09-05 |
| Tasks | docs/MIPs/MIP-0011.tasks.md — stacked PRs, one per task |
| Phase | 0 — developer tooling; nothing a user of marola sees. No earlier-phase prerequisite |
| Related | AGENTS.md (the rules this turns from advisory into enforced), docs/3-Working-on-the-repo/DEV-FLOW.md, docs/3-Working-on-the-repo/AGENT-SKILLS.md §3 (skill candidates), docs/4-Research-and-plans/FABLE_REVIEW.md §3 (jail environment notes), PHILOSOPHY.md (why constraints, not prose), .claude/settings.json, .claude/skills/, justfile (jail-claude, jcf, jcs); MIP-0013 (an OpenCode tryout — most of tasks 1-5, 7 and 9 here have a one-config-key equivalent there); docs/en/interactive-mode, docs/en/routines (background/overnight execution, task 11) |
| Effort | M — ten small, independent config/hook PRs; each individually small, but ten of them |
| Gain | infra/dev-loop (turns advisory AGENTS.md rules into enforced config); cost/ops (fewer permission prompts, a shorter AGENTS.md) |
| Effort vs Gain | do next — cheap and overdue |
| Depends on | none blocking; Phase 0, developer tooling only; MIP-0013's tryout result may show which tasks are Claude-Code-specific |
| Risk | a shell-text permissions.deny is a guard rail, not a boundary — bash -c "gh pr merge" evades it; ai-jail and the human go-ahead remain the real boundary |
| Cost so far | ~$12.86 total across the eleven merged task PRs (Cost: trailers, mostly measured via scripts/cost-split.py; task 6's $3.40 is a diff-size estimate, no session log matched) — not counting the design/MIP-writing session bundled into MIP-0010's shared total (commit 3fdcd05), or the follow-on ultrareview verification (see §7 — a free-tier request this time, at $0 to the user; note the real per-review price here once a paid one is spent, so this line stays honest about what future reviews cost) |
1. Summary¶
marola already follows the parts of Anthropic's guidance that are habits: a CLAUDE.md that imports
AGENTS.md, two in-repo skills, one session per feature, a cost line per PR, the agent in a
sandbox. What it does not have is the part that is configuration the harness enforces: no hooks
(every rule in AGENTS.md is text the model may or may not weigh), no shared permission allowlist
(every session re-approves just test), no path-scoped rules (AGENTS.md is 196 lines loaded into
every session, on the edge of the documented adherence cliff), no subagents (the reviewer
DEV-FLOW.md §5 describes is run by hand), no .mcp.json (marola's own MCP server is not offered
to the agent working on marola). This MIP adopts the documented practices in that order, each as one
small PR, with a measurable "done".
2. Motivation¶
Quoted from the docs fetched 2026-09-05 (code.claude.com/docs/en/…):
- "Unlike CLAUDE.md instructions which are advisory, hooks are deterministic and guarantee the
action happens." (
best-practices).AGENTS.md's cost rule is the highest-stakes rule in the repo and it is advisory. - "target under 200 lines per CLAUDE.md file. Longer files consume more context and reduce
adherence." (
memory).AGENTS.mdis 196 lines and grows with every doc-table row. - "Commit
.claude/settings.jsonso everyone who clones the repository gets the same permissions, hooks, telemetry, and plugins." (settings). Ours holds attribution and one plugin. - "Use
disable-model-invocation: truefor workflows with side effects" (skills).mip-taskscreates branches and PRs and can be auto-invoked. - Local evidence:
docs/4-Research-and-plans/FABLE_REVIEW.md§3 (ghunauthenticated inside the jail,.env.examplereading as empty) is re-discovered per session; auto-memory currently carries it. ASessionStarthook can print it once. - "When a claude.ai usage limit stops Claude mid-task, Claude Code waits in the open session
and continues the task on its own after the limit resets... automatic continue is on by
default." (
interactive-mode, fetched 2026-09-05). This closes what was, until recently, a well-documented gap:anthropics/claude-codeissues #35744, #26775, #18980, #36320, #38263, #62788 all requested exactly this, several as late as March 2026, with a small ecosystem of third-party wrappers (claude-auto-retry,claude-auto-resume,claude-delayed-message) built to work around its absence. What it does and doesn't cover matters before relying on it for an unattended overnight run; see the new risks below.
3. User-visible change¶
None for marola's users. For whoever runs just jcf / just jcs:
$ just jcf
▸ SessionStart: branch mip-0010/1-run-ledger · gh: not logged in (push works, PRs by hand) · 0 uncommitted
> …edits core/src/main/scala/marola/ledger/RunLedger.scala…
▸ PostToolUse(Edit *.scala): scalafmt ok
> (turn ends with .scala changes and no test run)
▸ Stop: 3 Scala files changed since HEAD and `just test` has not run this session — run it or say why.
Permission prompts for just build|test|quality|fmt, sbt …, git status|diff|log, scripts/*.sh
disappear; anything not on the allowlist still asks.
4. Data sources and dependencies reviewed¶
4.1 Anthropic — Claude Code documentation (fetched 2026-09-05)¶
- Best practices (
/docs/en/best-practices;anthropic.com/engineering/claude-code-best-practicesredirects there): give Claude a check it can run (tests, build, a Stop hook, a/goal, a verification subagent); explore → plan → code → commit, plan mode for multi-file or unfamiliar changes; specific prompts; CLAUDE.md short, "would removing this cause a mistake?", run/doctorto trim; permissions allowlist + sandbox; CLI tools (gh) over API calls; hooks "for actions that must happen every time with zero exceptions"; skills in.claude/skills/; subagents in.claude/agents/;/clearbetween tasks,/renamesessions; a fresh-context reviewer ("Writer/Reviewer"); an adversarial review subagent told to report gaps that affect correctness only. - Hooks reference (
/docs/en/hooks): events incl.SessionStart,UserPromptSubmit,PreToolUse(can block),PostToolUse(read-only),Stop(can block; the harness overrides after 8 consecutive blocks),PreCompact,InstructionsLoaded; matchers by tool name or permission-rule syntax ("if": "Bash(git *)"); exit code 2 blocks, JSON output carriespermissionDecision/systemMessage; handler types command/http/mcp_tool/prompt/agent;${CLAUDE_PROJECT_DIR}substitution; documented examples: block destructive shell, auto-format after edits, watch.env. - Memory (
/docs/en/memory):CLAUDE.mdload order;@import;.claude/rules/*.mdwithpaths:frontmatter loading only when matching files are read;CLAUDE.local.mdgitignored; auto-memory limits (200 lines / 25 KB index); hooks, not CLAUDE.md, for "must happen at a point". - Settings (
/docs/en/settings): precedence managed → CLI →.claude/settings.local.json→.claude/settings.json→~/.claude/settings.json;permissions.allow/deny/ask,hooks,env,attribution,enabledPlugins,autoMemoryEnabled; the local file is git-ignored by Claude Code when it creates it; add it to.gitignorewhen created by hand. - Skills (
/docs/en/skills): frontmattername,description,disable-model-invocation,allowed-tools,context: fork,paths,arguments; dynamic context via!`cmd`;/skill-doctorfor token cost and hit rate; theskill-creatorplugin runs evals inevals/evals.json; keep SKILL.md under 500 lines. - Subagents (
/docs/en/sub-agents):.claude/agents/*.mdwithname,description,tools,model,permissionMode,memory: project,maxTurns; read-only tool sets for review; fresh context by design.
4.2 Community (fetched 2026-09-05)¶
hesreallyhim/awesome-claude-code: 53.5k stars, pushed 2026-09-05, the index of hooks, skills, statuslines and CLAUDE.md examples; used as a directory, no single claim taken from it.MuhammadUsmanGM/claude-code-best-practices(wiki, "Last updated May 12, 2026 v1.6"): format-on-write and test-on-stop hooks, pre-commit secret blocking, aPreToolUsehook that validates edits to CLAUDE.md itself, worktrees for parallel tasks, cost budgeting per task type. Its model advice ("Claude 3.5 Sonnet") is stale and ignored.- Several 2026 blog guides (DEV, SmartScope, mcp.directory; search results, not fetched) converge on the same list: lean CLAUDE.md, plan mode, subagents for noisy research, worktrees, hooks as guardrails, verification loop. Nothing beyond the official docs was adopted from them.
4.3 What is already here (read 2026-09-05)¶
CLAUDE.md = @AGENTS.md + two lines; .claude/settings.json = attribution + superpowers
plugin; .claude/skills/mip, mip-tasks (no disable-model-invocation); .githooks/pre-commit
(sbt compile of staged Scala, a git hook, not a Claude hook, keep it); .ai-jail; just
jail-claude|jcf|jcs; scripts/uprd.sh, stack.sh, cost-split.py; auto-memory notes about gh
and Docker in sessions. No .claude/hooks/, .claude/rules/, .claude/agents/, .mcp.json,
CLAUDE.local.md, statusline.
5. Design¶
Each bullet is one task/PR (mip-tasks order), all under .claude/ or the root, each with a
self-test the quality recipe runs (scripts/hooks/*.sh --self-test, like the Python scripts).
.claude/settings.json, permissions.allow:Bash(just build*),Bash(just test*),Bash(just quality*),Bash(just fmt*),Bash(sbt *),Bash(git status*),Bash(git diff*),Bash(git log*),Bash(scripts/stack.sh status*),Read,Grep,Glob;deny:Read(.env),Read(**/*.pem),Read(**/*.key). Derive the allow list from real transcripts with thefewer-permission-promptsskill first; personal extras go tosettings.local.json(add it to.gitignore).- Hooks, cost gate. A
PreToolUsehook onBashthat blocked paid-deploy commands until a human go-ahead. It has since been removed, with the integrations it guarded. ai-jail and the human go-ahead remain. - Hooks, format on write.
PostToolUseonEdit|Writematching*.scala→scalafmton that file via the native binary (coursier launch scalafmtcached in the flake; whether the flake already provides a nativescalafmtwas not checked, §11).*.py→ruff formatwhen present. A slow hook is worse than none: 30 s timeout, single file only. - Hooks, test on stop, once.
Stop→.claude/hooks/stop-gate.sh: ifgit diff --name-only HEADhas.scalafiles and no marker from ajust testrun exists for this session, block once with "runjust testor say why", write the marker, allow the next stop. Never loop (the harness stops blocking after 8; this stops after 1). - Hooks, session start.
SessionStartprints branch,gh auth statusin one line, uncommitted count, and the two jail caveats fromFABLE_REVIEW.md§3. Replaces two auto-memory notes with a fact the harness states. .claude/rules/. Move the path-specific halves ofAGENTS.mdintorules/scala.md(paths: ["**/*.scala", "build.sbt"]: Kyo boundary, strictEquality, enum-over-exceptions, verify-against-the-jar),rules/docs.md(paths: ["docs/**"]: status vocabulary, MIP template pointer).AGENTS.mdkeeps what every session needs (what the repo is, commands, phase discipline, attribution, cost) and links the rules; target ≤ 150 lines; run/doctorfor the trim list. Decision needed (§11):AGENTS.mdis read by non-Claude agents too; rules are excerpts that link back, never the only copy.- Subagents.
.claude/agents/mip-reviewer.md(toolsRead, Grep, Glob, Bash,model: fable,memory: project): reviews a task branch against itsMIP-NNNN.tasks.mdrow, reports gaps that affect correctness or the stated requirement only,DEV-FLOW.md§5's prompt, made a file..claude/agents/jar-verifier.md(toolsBash,model: sonnet):javapa pinned Kyo class and report the real signature;AGENTS.md's rule as a cheap delegate. - Skills.
disable-model-invocation: trueonmip-tasks; the four candidates fromAGENT-SKILLS.md§3 (fixture-refresh,benchmark-compare,corpus-doc,water-provider), one PR each, each withevals/evals.jsonand askill-creatorrun before merge. .mcp.json.marola→just mcp-server(stdio) so a session working on marola can callfind_nearby_beaches/get_swim_recommendation, dogfoodingARCHITECTURE.md§5c.CLAUDE.mdadditions under the import:/clearper feature and/renameto the branch (already inDEV-FLOW.md, restated once); compaction instruction "preserve the list of modified files, the test commands run, and the Cost figure"; aCLAUDE.local.mdmention. Optional: a statusline showing context use and branch (statuslineskill), personal, not committed.- Background/overnight MIP execution.
/goal+/looplocally, or a cloud routine for something that survives the laptop sleeping, to run aMIP-NNNN.tasks.mdrow-by-row unattended, stopping at a stated condition rather than running indefinitely. Same guard rails as a daytime session, not fewer: the branch/PR boundaries tasks 1-2 already set, and thegh pr mergedeny rule from the earlier GitHub-hygiene work; an overnight run still never merges itself.
Not adopted: claude -p in CI (a paid run per PR with no reviewer, revisit with MIP-0010's
ledger measuring it); /batch fan-out (nothing here is a 2 000-file migration); agent teams
(experimental).
6. Scoring / safety impact¶
None. No product code changes.
7. Verification plan¶
scripts/hooks/*.sh --self-testinjust quality:stop-gate.shblocks once then allows;format.shleaves an already-formatted file unchanged (byte-equal).- A session check per PR, recorded in the PR body:
/hookslists the hooks;/contextshowsrules/scala.mdonly after a.scalafile is read;/skill-doctorbefore/after the skills PR. - Measure: permission prompts per session (count from the transcript, before vs after task 1);
AGENTS.mdline count (196 → ≤ 150); Cost trailers on the next three task PRs vs the previous three (the hooks should not raise them; a Stop hook that loops would). - Done: all eleven tasks merged (PRs #102, #108-110, #115-122),
FABLE_REVIEW.md§3's jail caveats are printed by theSessionStarthook;gh-not-logged-in-in-sessions.md(the one matching auto-memory note that actually existed (task 5's PR found only one of the "two" this row originally expected) is deleted. - Ultrareview, 2026-09-06:
/code-review ultrarun against the full stack diff (mip-0011/11-overnight-run-guardrail→main, 22 files, +1118/-134) after all eleven tasks merged, at the user's request. A free-tier request this session (Free ultrareview 1 of 3): $0 to the user; log the real price here the first time a paid one is spent, so this line doesn't silently imply every future review is free. Findings tracked separately, not folded into this MIP's own scope (a design doc doesn't get rewritten post-hoc for review findings; those become new fix commits/PRs againstmain, referencing this MIP for context). - How to refresh the Cost so far figure above, since all eleven PRs are already merged (the
normal
just uprds MIP-0011flow only updates open PR bodies, not this row): git log --oneline main | grep 'mip-0011 task': the eleven merged commit SHAs.for c in <shas>; do git show -s --format=%B "$c" | grep '^Cost:'; done: pull each commit's ownCost:trailer (already measured viascripts/cost-split.pyat PR time, except task 6, still the one diff-sizeest.) and sum the dollar figures by hand; there is no single command that re-sums an already-merged, cross-session stack's trailers today (just cost-split MIP-0011is for an open, unmerged stack in one session's own log).- Edit this row and the matching one in
docs/MIPs/README.mdtogether; they must always agree;docs/3-Working-on-the-repo/DEV-FLOW.mdhas no automation for that agreement, so it's a manual pair-edit. - If task 6's
est.figure is ever superseded by a real measured one (e.g. a later session finds the original log), update this row's total (the merged commit's ownCost:trailer stays as originally written, commit messages are historical, not corrected in place), and note in the doc edit that the figure changed from estimate to measured, so a reader doesn't assume it was always precise.
8. Risks, limitations, and honest caveats¶
- Hooks run shell on every matching event inside the jail: a slow formatter or an sbt-backed hook would tax every edit; native binaries only, per-file, timeouts set.
- A
Stophook that blocks repeatedly burns tokens and trust; the once-per-session design and the self-test exist for that. permissions.denymatches command text;bash -c "gh pr merge"or a script can evade it. It is a guard rail, not a boundary; ai-jail and the human go-ahead remain the boundary.- Auto mode's classifier and the deny list can disagree; deny wins for what it matches, and the rest stays the classifier's call.
- Splitting
AGENTS.mdrisks two copies of one rule drifting; rules link back and CI can grep for the sentinel sentences. - Docs quoted here are a moving target (the hooks page lists events that did not exist months ago); each task re-fetches the page it relies on and notes the date in its PR.
- Auto-continue is reactive, not a throttle. It waits out a hit limit and resumes; nothing
in the docs describes pacing usage to avoid hitting the wall mid-task. A long unattended run
can still burn the 5-hour window in one uncontrolled burst. (The community
heavy-usageplugin claims to "stop safely before the wall", not verified against its source, treat as an unconfirmed community claim, not confirmed Claude Code behavior, until checked.) - It needs a session that stays alive. The doc says it "waits in the open session"; a closed terminal or a sleeping laptop breaks this, the same gap the OS-level community schedulers exist to work around. A routine doesn't depend on the local machine staying awake and is the safer default for a genuine walk-away run; check its own cost model first (§11).
- A 7-day reset is not an overnight wait. Same mechanism, different practical meaning: a
task that exhausts the weekly window stops for up to a week. Check which window
(
rate_limits.five_hourvs..seven_day, per the statusline schema) is actually at risk before planning an overnight run around it.
9. Alternatives considered¶
- Do nothing: the prose works most of the time; "most" is the problem for the cost rule.
- Managed settings (
managed-settings.json): organisation-level; this is a personal repo. - ai-jail alone: filesystem/process containment, no notion of a paid API call or a test gate.
- One giant
settings.jsonPR: harder to attribute a regression (a hook slowing sessions) to a change; ten small PRs matchDEV-FLOW.md.
11. Open questions¶
- Does
flake.nixalready provide a nativescalafmt(coursier launcher orpkgs.scalafmt)? If not, which is lighter for a per-edit hook? AGENTS.mdvs.claude/rules/: excerpts that link back (proposed) or move the text and leave a pointer for other agents? The former duplicates; the latter drops the rule for non-Claude tools.Stophook policy: block once (proposed) or advise only viasystemMessage?- Which allowlist entries, from real transcripts (
fewer-permission-prompts) rather than from the list above. - Statusline: worth committing a project default, or personal only?
- Should
.mcp.jsonalso register a local MLflow (MIP-0010) or Overpass helper, or nothing beyond marola's own server? - Resolved by task 11's spike, see
docs/3-Working-on-the-repo/DEV-FLOW.md§7. Routines vs. local/goal+/loopfor overnight MIP runs: local/goal+/loop(viaCronCreate) was run for real this session to work through this very MIP's own task stack (real pushed branches, realGH_POST_MORTEM.mdentries), and is the chosen default: no extra environment setup, demonstrated working. A cloud routine survives the laptop closing but needs Claude Code on the web/a cloud environment, not confirmed available in every setup; adopt it once that's confirmed, not assumed as a prerequisite. - Resolved by task 11's spike, see
docs/3-Working-on-the-repo/DEV-FLOW.md§7.heavy-usage's "stop before the wall" claim, checked against its actual source (~/.claude/plugins/cache/heavy-usage): real, but soft: aUserPromptSubmitprompt injection at a linear-projection threshold (90% five-hour / 95% weekly), not aPreToolUseblock, so compliance is advisory; its data source (usage-live.json) is populated only while the interactive statusLine renders and can go stale in a genuinely headless run (the plugin's own comment: "never suppress a wind-down" even on stale data). Treat it asmip-solve-perpetual's own backup usage-guard layer, never the sole or primary stop condition, consistent with how the skill already used it before this check.
Appendix¶
https://code.claude.com/docs/en/best-practices(2026-09-05), the redirect target ofhttps://www.anthropic.com/engineering/claude-code-best-practices(HTTP 308).https://code.claude.com/docs/en/hooks,/memory,/settings,/skills,/sub-agents(2026-09-05).https://github.com/hesreallyhim/awesome-claude-code: 53 559 stars, pushed 2026-09-05T18:10Z (GitHub API).https://github.com/MuhammadUsmanGM/claude-code-best-practices: v1.6, 2026-05-12.- Repo state read on 2026-09-05:
AGENTS.md196 lines;.claude/settings.jsonkeysattribution,enabledPlugins; skillsmip,mip-tasks.