cost-split
cost-split — attribute a Claude Code session's real token usage to the commits it produced.
scripts/cost-split.py # commits of the current branch (main..HEAD)
scripts/cost-split.py mip-0005 # every commit of a MIP stack (mip-0005/1-*, /2-*, ...)
scripts/cost-split.py --session 336ebf4f # restrict to one session (prefix of its id)
scripts/cost-split.py --json # machine-readable
scripts/cost-split.py --estimate # also fill in commits with no logged usage
scripts/cost-split.py --estimate --verbose # ... and print the calibration fit
scripts/cost-split.py --estimate-commit <sha> # just one commit's estimated `Cost:` trailer
scripts/cost-split.py --self-test # parser + estimator self-check (just quality-other)
AGENTS.md wants a Cost: trailer per commit and a Cost section per PR. /usage and
just claude-cost (ccusage) give one figure per session; when one session produces several
stacked PRs that figure has to be split. This does it the only honest way available: every
assistant message in the session log (~/.claude/projects/
Subagent usage: an Agent(...) call (.claude/skills, code-review, a plain subagent) writes its
own transcript to <session>/subagents/agent-<id>.jsonl next to the parent session's own
<session>.jsonl — same shape (assistant messages with a usage block), same billing. The parent
log's own record of the call is a tool_result placeholder with no usage of its own (confirmed by
hand: no isSidechain:true line in a top-level .jsonl carries a usage block), so folding these
transcripts in adds real cost that was previously invisible without double-counting anything.
Some commits still have nothing logged at all — a subagent that ran in a worktree whose session
was never re-attached, work done on another machine, or a commit written outside a Claude Code
session. --estimate fills those in from the diff alone: lines changed, files touched, and a
tokens-per-changed-line coefficient calibrated from this repo's own history (every commit on
origin/main that already carries a measured Cost: trailer with a token figure, paired with its
own diff size). The result is always clearly labelled est. — never mistaken for a measurement,
and cache-read share (impossible to guess from a diff) is simply left out of the price.
Prices: LiteLLM's public price table (the same source ccusage uses), cached one day under .tmp/. Offline with no cache, the tokens are still split and the dollar column says n/a. On a subscription the dollars are not a bill — they are the quota proxy AGENTS.md asks for.
1#!/usr/bin/env python3 2"""cost-split — attribute a Claude Code session's real token usage to the commits it produced. 3 4 scripts/cost-split.py # commits of the current branch (main..HEAD) 5 scripts/cost-split.py mip-0005 # every commit of a MIP stack (mip-0005/1-*, /2-*, ...) 6 scripts/cost-split.py --session 336ebf4f # restrict to one session (prefix of its id) 7 scripts/cost-split.py --json # machine-readable 8 scripts/cost-split.py --estimate # also fill in commits with no logged usage 9 scripts/cost-split.py --estimate --verbose # ... and print the calibration fit 10 scripts/cost-split.py --estimate-commit <sha> # just one commit's estimated `Cost:` trailer 11 scripts/cost-split.py --self-test # parser + estimator self-check (just quality-other) 12 13AGENTS.md wants a `Cost:` trailer per commit and a Cost section per PR. `/usage` and 14`just claude-cost` (ccusage) give one figure per *session*; when one session produces several 15stacked PRs that figure has to be split. This does it the only honest way available: every 16assistant message in the session log (~/.claude/projects/<project>/<session>.jsonl) carries its 17timestamp and token usage, so the usage between two commits is the cost of the second commit. 18Messages before the first commit (planning) go to the first commit; messages after the last one 19are reported as "uncommitted (so far)". 20 21Subagent usage: an `Agent(...)` call (`.claude/skills`, `code-review`, a plain subagent) writes its 22own transcript to `<session>/subagents/agent-<id>.jsonl` next to the parent session's own 23`<session>.jsonl` — same shape (assistant messages with a `usage` block), same billing. The parent 24log's own record of the call is a `tool_result` placeholder with no usage of its own (confirmed by 25hand: no `isSidechain:true` line in a top-level `.jsonl` carries a `usage` block), so folding these 26transcripts in adds real cost that was previously invisible without double-counting anything. 27 28Some commits still have nothing logged at all — a subagent that ran in a worktree whose session 29was never re-attached, work done on another machine, or a commit written outside a Claude Code 30session. `--estimate` fills those in from the diff alone: lines changed, files touched, and a 31tokens-per-changed-line coefficient calibrated from this repo's own history (every commit on 32origin/main that already carries a measured `Cost:` trailer with a token figure, paired with its 33own diff size). The result is always clearly labelled `est.` — never mistaken for a measurement, 34and cache-read share (impossible to guess from a diff) is simply left out of the price. 35 36Prices: LiteLLM's public price table (the same source ccusage uses), cached one day under .tmp/. 37Offline with no cache, the tokens are still split and the dollar column says n/a. On a 38subscription the dollars are not a bill — they are the quota proxy AGENTS.md asks for. 39""" 40 41import argparse 42import datetime as dt 43import json 44import os 45import re 46import sqlite3 47import statistics 48import subprocess 49import sys 50import tempfile 51import urllib.request 52from collections import defaultdict 53from pathlib import Path 54 55PRICES_URL = ( 56 "https://raw.githubusercontent.com/BerriAI/litellm/main/model_prices_and_context_window.json" 57) 58 59# --------------------------------------------------------------------------------------------- 60# Diff-size estimation constants — see calibrate() for where the coefficient itself comes from. 61# --------------------------------------------------------------------------------------------- 62# Fallback only: used when origin/main carries no measured `Cost:` trailer with a token figure to 63# calibrate against (a fresh checkout of this repo, or a --self-test fixture). ~6,200 tokens per 64# changed line is this repo's own median as of 2026-09-05 (30 calibration commits) — see 65# `scripts/cost-split.py --estimate --verbose`. 66DEFAULT_TOKENS_PER_LINE = 6200 67# The model priced when a commit has no logged usage at all, so there's no model name to read off 68# a real message — the model this repo's sessions run by default today. 69DEFAULT_PRICE_MODEL = "claude-sonnet-5" 70# A touched file costs roughly this many "changed lines" worth of tokens just to open, read for 71# context and re-check, even when the diff itself is one line — so a wide, shallow diff is not 72# free the way a deep, narrow one is. 73FILE_OVERHEAD_LINES = 15 74 75TOKEN_TRAILER_RE = re.compile(r"([\d][\d,]*\.?\d*)\s*(M|k)?\s*tokens") 76USD_TRAILER_RE = re.compile(r"\$([\d.]+)") 77 78 79def sh(*args, cwd=None): 80 env = None 81 if cwd is not None: 82 # A cwd means "operate on this repo, not whatever GIT_DIR/GIT_WORK_TREE/etc. the 83 # environment already points at" — a git hook (pre-push, pre-commit) sets those for its 84 # own repo, and without stripping them a subprocess git call here follows the env instead 85 # of cwd: with GIT_DIR set but not GIT_WORK_TREE, git treats cwd as the work tree and 86 # happily commits its content into the *real* repo's history, while the real working tree 87 # never receives the files — surfacing afterward as spurious "deleted: <file>" entries. 88 # See the self-test below, which reproduces this exact scenario. 89 env = {k: v for k, v in os.environ.items() if not k.startswith("GIT_")} 90 return subprocess.run(args, check=True, capture_output=True, text=True, cwd=cwd, env=env).stdout 91 92 93def repo_root(): 94 # The main checkout even from a worktree: Claude Code keys its logs on the directory the 95 # session was started in, which for this repo's `.tmp/wt-*` worktrees is the main checkout. 96 common = Path(sh("git", "rev-parse", "--path-format=absolute", "--git-common-dir").strip()) 97 return common.parent 98 99 100def project_dir(root): 101 # Claude Code's per-project log dir: the absolute path with every "/" (and ".") turned into "-". 102 slug = "-" + str(root).strip("/").replace("/", "-").replace(".", "-") 103 return Path.home() / ".claude" / "projects" / slug 104 105 106def load_prices(root): 107 cache = root / ".tmp" / "litellm-prices.json" 108 try: 109 if cache.exists() and (dt.datetime.now().timestamp() - cache.stat().st_mtime) < 86400: 110 return json.loads(cache.read_text()) 111 with urllib.request.urlopen(PRICES_URL, timeout=15) as r: 112 data = r.read() 113 cache.parent.mkdir(exist_ok=True) 114 cache.write_bytes(data) 115 return json.loads(data) 116 except Exception: 117 return json.loads(cache.read_text()) if cache.exists() else {} 118 119 120def price(prices, model, u): 121 p = prices.get(model) 122 if not p: 123 return None 124 cc = u.get("cache_creation") or {} 125 write_5m = cc.get("ephemeral_5m_input_tokens", u.get("cache_creation_input_tokens", 0)) 126 write_1h = cc.get("ephemeral_1h_input_tokens", 0) 127 if not cc: 128 write_5m = u.get("cache_creation_input_tokens", 0) 129 return ( 130 u.get("input_tokens", 0) * p.get("input_cost_per_token", 0) 131 + u.get("output_tokens", 0) * p.get("output_cost_per_token", 0) 132 + write_5m * p.get("cache_creation_input_token_cost", 0) 133 + write_1h 134 * p.get( 135 "cache_creation_input_token_cost_above_1hr", p.get("cache_creation_input_token_cost", 0) 136 ) 137 + u.get("cache_read_input_tokens", 0) * p.get("cache_read_input_token_cost", 0) 138 ) 139 140 141def _session_message_files(pdir, session_prefix): 142 """(session_id, path) for every jsonl to read usage from: each top-level session log, plus 143 every subagent transcript Claude Code nests under <session>/subagents/*.jsonl for that 144 session. `session_prefix` filters by the *parent* session id, same as before this existed — 145 a subagent spawned inside session d802fa69 is included by `--session d802fa69`.""" 146 for f in sorted(pdir.glob("*.jsonl")): 147 sid = f.stem 148 if session_prefix and not sid.startswith(session_prefix): 149 continue 150 yield sid, f 151 for sf in sorted((pdir / sid / "subagents").glob("*.jsonl")): 152 yield sid, sf 153 154 155def worktree_branches(root): 156 """{absolute worktree path: branch name}, from `git worktree list --porcelain` run at `root` 157 (always the main checkout, per `repo_root()`) — covers the main checkout itself plus every 158 live `.tmp/wt-*`. Used to resolve a message's `cwd` to the branch it actually ran on.""" 159 out = {} 160 path = None 161 for line in sh("git", "worktree", "list", "--porcelain", cwd=root).splitlines(): 162 if line.startswith("worktree "): 163 path = line.removeprefix("worktree ") 164 elif line.startswith("branch refs/heads/") and path: 165 out[path] = line.removeprefix("branch refs/heads/") 166 path = None 167 return out 168 169 170def _worktree_branch_for_cwd(cwd, wt_branches): 171 """The branch of the worktree that contains `cwd`, or None. Longest-match: a linked worktree 172 normally lives under the main checkout's own path (`.tmp/wt-*`), so the main checkout's path 173 is itself a prefix of it and the naive first-match would pick the wrong one.""" 174 if not cwd: 175 return None 176 best = None 177 for path, branch in wt_branches.items(): 178 if cwd == path or cwd.startswith(path + "/"): 179 if best is None or len(path) > len(best[0]): 180 best = (path, branch) 181 return best[1] if best else None 182 183 184def _under(cwd, root, linked): 185 """True if `cwd` is inside the main checkout's own tree, excluding any worktree's tree — 186 currently linked (`linked`) or since removed from git's registry. The second half matters: 187 `.tmp/wt-*` is routinely removed once a stacked task's worktree is done with (confirmed 188 against this project's real logs — a since-removed worktree's `cwd` is still `.tmp/wt-*` 189 shaped, so excluding only currently-`linked` paths would misclassify it as the main 190 checkout and trust its parent's inherited branch again).""" 191 if not cwd or not (cwd == str(root) or cwd.startswith(str(root) + "/")): 192 return False 193 if any(cwd == p or cwd.startswith(p + "/") for p in linked): 194 return False 195 return not cwd.startswith(str(root) + "/.tmp/wt-") 196 197 198def messages(pdir, session_prefix, root): 199 """(timestamp, session, model, usage, branch) per assistant message, deduplicated by message 200 id.""" 201 seen = {} 202 wt_branches = worktree_branches(root) 203 linked = {path: branch for path, branch in wt_branches.items() if path != str(root)} 204 for sid, f in _session_message_files(pdir, session_prefix): 205 with open(f, encoding="utf-8") as fh: 206 for line in fh: 207 try: 208 d = json.loads(line) 209 except json.JSONDecodeError: 210 continue 211 m = d.get("message") 212 if not isinstance(m, dict) or not m.get("usage") or not d.get("timestamp"): 213 continue 214 # Claude Code writes `<synthetic>` assistant messages (interrupts, tool-result 215 # stand-ins) with a usage block of zeros and no real model — not a priced request. 216 if m.get("model") == "<synthetic>": 217 continue 218 key = m.get("id") or d.get("requestId") or d.get("uuid") 219 ts = dt.datetime.fromisoformat(d["timestamp"].replace("Z", "+00:00")) 220 # `gitBranch` is the session directory's branch at message time — right for the 221 # long-lived, branch-switching main checkout, wrong for a *linked* worktree (every 222 # line of a subagent transcript there inherits its parent's value; issue #433). 223 # Only a linked worktree's own branch overrides it; a sidechain matching neither 224 # falls to "" rather than trust that inherited branch. 225 cwd = d.get("cwd") 226 wt_branch = _worktree_branch_for_cwd(cwd, linked) 227 if wt_branch is not None: 228 branch = wt_branch 229 elif d.get("isSidechain") and not _under(cwd, root, linked): 230 branch = "" 231 else: 232 branch = d.get("gitBranch") or "" 233 seen[key] = (ts, sid, m.get("model", "?"), m["usage"], branch) 234 return sorted(seen.values(), key=lambda x: x[0]) 235 236 237def opencode_db_path(): 238 """The real store, confirmed live 2026-09-07 against opencode 1.18.25 (MIP-0013 task 2, §11 239 OQ2): `$XDG_DATA_HOME/opencode/opencode-stable.db` — not the `storage/message/*/msg_*.json` 240 files an earlier draft of that MIP assumed; those don't exist in this version. Falls back to 241 `~/.local/share/opencode` per the XDG default when the env var is unset, and tries the 242 non-`-stable` filename too in case a future/edge/nightly channel uses it.""" 243 base = Path(os.environ.get("XDG_DATA_HOME", str(Path.home() / ".local" / "share"))) 244 d = base / "opencode" 245 for name in ("opencode-stable.db", "opencode.db"): 246 p = d / name 247 if p.is_file(): 248 return p 249 return None 250 251 252def messages_opencode(root, session_prefix): 253 """(timestamp, session, model, usage, gitBranch) per assistant message — the OpenCode-harness 254 twin of `messages()` above, same tuple shape so `main()`'s attribution loop needs no branching 255 beyond which reader it calls. `gitBranch` is always `""`: OpenCode's `message`/`session` rows 256 carry no branch field (confirmed live, same session), so a message here always falls back to 257 `main()`'s time-only rule — documented, not a bug to fix, since OpenCode sessions in this repo 258 are (so far) one-project, foreground, interactive runs, not marola's multi-worktree pattern. 259 260 `usage` is shaped like Claude Code's dict (`input_tokens`/`output_tokens`/ 261 `cache_creation_input_tokens`/`cache_read_input_tokens`) so `price()` needs no change; `model` 262 is `providerID/modelID` (e.g. `ollama/llama3.2`) — `price()` looks it up in the same LiteLLM 263 table and returns `None` for a local Ollama model, same graceful "not priced" path Claude 264 Code's `<synthetic>` skip already exercises, not a special case here. 265 """ 266 db = opencode_db_path() 267 if db is None: 268 return [] 269 seen = {} 270 con = sqlite3.connect(f"file:{db}?mode=ro", uri=True) 271 try: 272 con.row_factory = sqlite3.Row 273 cur = con.cursor() 274 cur.execute( 275 "select m.id, m.session_id, m.data from message m " 276 "join session s on s.id = m.session_id " 277 "where s.directory = ?", 278 (str(root),), 279 ) 280 for row in cur.fetchall(): 281 if session_prefix and not row["session_id"].startswith(session_prefix): 282 continue 283 try: 284 d = json.loads(row["data"]) 285 except json.JSONDecodeError: 286 continue 287 if d.get("role") != "assistant" or not isinstance(d.get("tokens"), dict): 288 continue 289 t = d["tokens"] 290 cache = t.get("cache") or {} 291 created_ms = (d.get("time") or {}).get("created") 292 if created_ms is None: 293 continue 294 ts = dt.datetime.fromtimestamp(created_ms / 1000, tz=dt.UTC) 295 usage = { 296 "input_tokens": t.get("input", 0), 297 "output_tokens": t.get("output", 0), 298 "cache_creation_input_tokens": cache.get("write", 0), 299 "cache_read_input_tokens": cache.get("read", 0), 300 } 301 model = f"{d.get('providerID', '?')}/{d.get('modelID', '?')}" 302 seen[row["id"]] = (ts, row["session_id"], model, usage, "") 303 finally: 304 con.close() 305 return sorted(seen.values(), key=lambda x: x[0]) 306 307 308_equivalents = None 309 310 311def branches_with_equivalent(sha): 312 """Branches holding a commit with the same author date and subject as `sha` — the same 313 commit after a rebase, cherry-pick or `just cost-fill` rewrite gave it a new hash. Without 314 this, the branch a subagent authored on stops "containing" its own commits the moment 315 they are rewritten, and its usage would no longer match them.""" 316 global _equivalents 317 if _equivalents is None: 318 _equivalents = defaultdict(set) 319 refs = sh( 320 "git", "for-each-ref", "--format=%(refname:short)", "refs/heads", "refs/remotes/origin" 321 ) 322 for ref in refs.splitlines(): 323 ref = ref.strip() 324 if not ref or ref in ("main", "origin/main", "origin/HEAD"): 325 continue 326 name = ref[7:] if ref.startswith("origin/") else ref 327 try: 328 log = sh("git", "log", "--format=%aI%x00%s", f"origin/main..{ref}") 329 except subprocess.CalledProcessError: 330 continue 331 for line in log.splitlines(): 332 if "\x00" in line: 333 _equivalents[tuple(line.split("\x00", 1))].add(name) 334 when, subject = sh("git", "log", "-1", "--format=%aI%x00%s", sha).rstrip("\n").split("\x00", 1) 335 return set(_equivalents.get((when, subject), set())) 336 337 338def branches_containing(sha): 339 """Local and origin branch names that contain `sha` (origin/ prefix stripped).""" 340 out = set() 341 for line in sh( 342 "git", "branch", "-a", "--contains", sha, "--format=%(refname:short)" 343 ).splitlines(): 344 b = line.strip() 345 if not b or b == "origin/HEAD" or " -> " in b: 346 continue 347 out.add(b[7:] if b.startswith("origin/") else b) 348 return out 349 350 351def commits(root, stack): 352 """[(time, branch, sha, subject)] — each commit once, on the first branch that carries it.""" 353 out, known = [], set() 354 branches = [] 355 if stack: # a MIP (mip-0005 → every mip-0005/k-* branch) or one plain branch name 356 branches = [ 357 b.strip() 358 for b in sh( 359 "git", "branch", "--list", f"{stack}/*", "--format=%(refname:short)" 360 ).splitlines() 361 if b.strip() 362 ] 363 364 # mip-NNNN/k-slug sorts by k; a branch without a numeric task prefix (mip-0012/llm4s-adoption, 365 # a one-PR MIP) sorts after the numbered ones, by name. 366 def task_no(b): 367 head = b.split("/")[1].split("-")[0] 368 return (0, int(head), "") if head.isdigit() else (1, 0, b) 369 370 branches.sort(key=task_no) 371 if not branches: 372 branches = [stack] 373 else: 374 branches = [sh("git", "branch", "--show-current").strip()] 375 for b in branches: 376 # Author date, not committer date: a rebase, cherry-pick or `just cost-fill` rewrite 377 # restamps the committer date of every commit to the same second, and the time split 378 # then hands the whole branch's usage to its first commit. Author dates survive all three. 379 log = sh("git", "log", "--reverse", "--format=%H%x00%aI%x00%s", f"origin/main..{b}") 380 for line in log.splitlines(): 381 sha, when, subject = line.split("\x00") 382 if sha in known: 383 continue 384 known.add(sha) 385 out.append((dt.datetime.fromisoformat(when), b, sha, subject)) 386 return sorted(out, key=lambda c: c[0]) 387 388 389# --------------------------------------------------------------------------------------------- 390# Diff-size cost estimation — for a commit with no logged usage at all. 391# --------------------------------------------------------------------------------------------- 392 393 394def parse_token_count(cost_line): 395 """Tokens named in one `Cost:` trailer line, or None if it doesn't carry a figure usable for 396 calibration. Skips bucket-wide notes ('shared bucket', 'not split further') — the same token 397 count copy-pasted onto several sibling commits, so pairing any one of them with its own diff 398 size would just add noise — and the one commit that admits it never got a number 399 ('unmeasured').""" 400 if "shared bucket" in cost_line or "unmeasured" in cost_line: 401 return None 402 m = TOKEN_TRAILER_RE.search(cost_line) 403 if not m: 404 return None 405 n = float(m.group(1).replace(",", "")) 406 unit = m.group(2) 407 return n * 1e6 if unit == "M" else n * 1e3 if unit == "k" else n 408 409 410def diff_stats(sha, cwd=None): 411 """(changed_lines, files_touched, paths) for one commit, from `git show --numstat`.""" 412 added = deleted = files = 0 413 paths = [] 414 for line in sh("git", "show", "--numstat", "--format=", sha, cwd=cwd).splitlines(): 415 if not line.strip(): 416 continue 417 parts = line.split("\t") 418 if len(parts) != 3: 419 continue 420 a, d, path = parts 421 added += 0 if a == "-" else int(a) 422 deleted += 0 if d == "-" else int(d) 423 files += 1 424 paths.append(path) 425 return added + deleted, files, paths 426 427 428def calibrate(ref="origin/main", cwd=None): 429 """Median tokens-per-changed-line across every `ref` commit with a measured `Cost:` trailer 430 carrying a token figure, paired with that commit's own diff size (`diff_stats`). The median 431 (not a least-squares fit) is deliberate: a big-context, small-diff commit can be two orders of 432 magnitude above a routine one (this repo's own spread is ~800-440,000 tokens/line), and a 433 maintainer already eyeballs a `Cost:` trailer the same way — ignore the extremes, trust the 434 middle. Returns (coefficient, stats); stats["n"] == 0 means DEFAULT_TOKENS_PER_LINE was used. 435 """ 436 body = sh("git", "log", ref, "--format=%H%x00%B%x00END", cwd=cwd) 437 ratios = [] 438 for chunk in body.split("\x00END\n"): 439 if "\x00" not in chunk: 440 continue 441 sha, msg = chunk.split("\x00", 1) 442 sha = sha.strip() 443 tokens = None 444 for line in msg.splitlines(): 445 if line.startswith("Cost:"): 446 tokens = parse_token_count(line) 447 if tokens: 448 break 449 if not tokens: 450 continue 451 changed, _files, _paths = diff_stats(sha, cwd=cwd) 452 if changed <= 0: 453 continue 454 ratios.append(tokens / changed) 455 if not ratios: 456 return DEFAULT_TOKENS_PER_LINE, {"n": 0} 457 ratios.sort() 458 stats = { 459 "n": len(ratios), 460 "min": ratios[0], 461 "max": ratios[-1], 462 "median": statistics.median(ratios), 463 } 464 if len(ratios) >= 4: 465 q1, _, q3 = statistics.quantiles(ratios, n=4) 466 stats["p25"], stats["p75"] = q1, q3 467 else: 468 stats["p25"], stats["p75"] = ratios[0], ratios[-1] 469 return stats["median"], stats 470 471 472def print_calibration(coeff, stats, stream=sys.stdout): 473 if stats["n"] == 0: 474 print( 475 f"cost-split --estimate: no calibration history found — falling back to the " 476 f"documented default of {coeff:.0f} tokens/changed-line", 477 file=stream, 478 ) 479 return 480 print( 481 f"cost-split --estimate: calibrated {coeff:.0f} tokens/changed-line from {stats['n']} " 482 f"origin/main commit(s) with a measured Cost: trailer " 483 f"(IQR {stats['p25']:.0f}-{stats['p75']:.0f}, range {stats['min']:.0f}-{stats['max']:.0f})", 484 file=stream, 485 ) 486 487 488def category_multiplier(paths): 489 """A documented heuristic nudge, not a second regression — too few calibration commits per 490 category (docs-only, test-heavy, ...) in this repo's history to fit each separately without 491 overfitting on a handful of points.""" 492 if not paths: 493 return 1.0 494 docs_only = all(p.endswith(".md") or p.startswith("docs/") for p in paths) 495 if docs_only: 496 return 0.6 # prose from an outline is cheaper per line than code written from scratch 497 scala = any(p.endswith((".scala", ".sbt")) for p in paths) 498 tests = any("test" in p.lower() or "spec" in p.lower() for p in paths) 499 if scala and tests: 500 return 1.15 # red/green/refactor round-trips cost more than the final diff alone shows 501 return 1.0 502 503 504def estimate_tokens_for_diff(changed_lines, files, paths, coeff): 505 effective = changed_lines + FILE_OVERHEAD_LINES * files 506 return round(coeff * effective * category_multiplier(paths)) 507 508 509def estimate_usd(prices, tokens, model=DEFAULT_PRICE_MODEL): 510 """Prices the estimated tokens as if every one of them were a fresh input token (LiteLLM's 511 `input_cost_per_token`) — cache-read share cannot be estimated from a diff, so it is left out 512 entirely rather than guessed; this repo's own measured $/token (see calibrate()'s sibling 513 check in the self-test) runs about half of that rate, so the estimate leans conservative 514 rather than optimistic.""" 515 p = prices.get(model) 516 if not p: 517 return None 518 rate = p.get("input_cost_per_token") 519 if not rate: 520 return None 521 return tokens * rate 522 523 524def _tokens_str(n): 525 if n >= 100_000: 526 return f"{n / 1e6:.1f}M" 527 if n >= 1_000: 528 return f"{n / 1e3:.0f}k" 529 return f"{n:.0f}" 530 531 532def format_tokens(tokens): 533 return f"~{_tokens_str(tokens)} tokens est." 534 535 536def format_estimate_trailer(tokens, usd, changed_lines, today, stats): 537 """`stats` is calibrate()'s return. The point figure alone hides that this repo's own 538 tokens-per-line spread is ~5x (see calibrate()'s docstring) — scale calibrate()'s p25/p75 onto 539 this estimate's own token figure so the band is in the same units as the number next to it, 540 rather than inventing a new one. `stats["n"] == 0` (no calibration history) states that 541 plainly instead of printing a band it doesn't have.""" 542 usd_part = f"~${usd:.2f} est." if usd is not None else "$n/a est." 543 n = stats.get("n", 0) 544 if n and stats.get("median"): 545 lo = tokens * stats["p25"] / stats["median"] 546 hi = tokens * stats["p75"] / stats["median"] 547 band = f"IQR {_tokens_str(lo)}-{_tokens_str(hi)} tokens for this diff from {n} calibrated commits" 548 else: 549 band = "no calibration history — documented default, not a fit" 550 return ( 551 f"Cost: {usd_part} · {format_tokens(tokens)} " 552 f"(diff-size model, {changed_lines} lines, no session log, {band}) " 553 f"· scripts/cost-split.py --estimate {today}" 554 ) 555 556 557def estimate_commit(sha, prices, coeff, cwd=None): 558 """(tokens, usd, changed_lines) for one commit's diff-size estimate, or None if the commit 559 has no diff at all (an empty commit, or a merge commit numstat can't attribute).""" 560 changed, files, paths = diff_stats(sha, cwd=cwd) 561 if changed <= 0: 562 return None 563 tokens = estimate_tokens_for_diff(changed, files, paths, coeff) 564 usd = estimate_usd(prices, tokens) 565 return tokens, usd, changed 566 567 568# --------------------------------------------------------------------------------------------- 569# Self-test — scripts/*.py's usual shape (assert, print "ok", return 0/1); wired into 570# `just quality-other` and ci.yml's quality-other job. 571# --------------------------------------------------------------------------------------------- 572 573 574def self_test(): 575 # --- parse_token_count: every shape actually seen in this repo's own `Cost:` trailers --- 576 cases = [ 577 ("Cost: ~$8.98 · 8.1M tokens, 97% cache reads (claude-fable-5-1) · ...", 8_100_000), 578 ( 579 "Cost: ~289k tokens (forked subagent, claude-fable-5-1; 40 tool calls, 8 min) · ...", 580 289_000, 581 ), 582 ( 583 "Cost: shared bucket $7.03 · 10,517,273 tokens (claude-fable-5-1) — not split further", 584 None, 585 ), 586 ("Cost: ~unmeasured in-agent · fill from just claude-cost before PR", None), 587 ( 588 "Cost: ~$0 · 0 tokens booked (claude-fable-5-1) · ...", 589 0.0, 590 ), # parses, but falsy: calibrate() skips it too 591 ("Cost: ~$0.1", None), # no token figure at all 592 ("Cost: shared bucket $3.18 · 4,008,909 tokens (claude-fable-5-1) — ...", None), 593 ] 594 for line, expected in cases: 595 got = parse_token_count(line) 596 assert got == expected, f"{line!r} -> {got}, expected {expected}" 597 598 # --- estimate_tokens_for_diff / category_multiplier --- 599 coeff = 6000 600 plain = estimate_tokens_for_diff(100, 2, ["core/src/main/Foo.scala"], coeff) 601 assert plain == round(coeff * (100 + FILE_OVERHEAD_LINES * 2)), plain 602 docs = estimate_tokens_for_diff(100, 1, ["docs/FOO.md"], coeff) 603 assert docs < plain, (docs, plain) # docs-only is cheaper per line 604 tested = estimate_tokens_for_diff( 605 100, 2, ["core/src/main/Foo.scala", "core/src/test/FooSpec.scala"], coeff 606 ) 607 assert tested > plain, (tested, plain) # scala + tests costs more than scala alone 608 assert estimate_tokens_for_diff(0, 0, [], coeff) == 0 609 610 # --- estimate_usd: LiteLLM-shaped fixture, no network --- 611 prices = {"claude-sonnet-5": {"input_cost_per_token": 2e-06, "output_cost_per_token": 1e-05}} 612 usd = estimate_usd(prices, 1_000_000, model="claude-sonnet-5") 613 assert usd == 2.0, usd 614 assert estimate_usd(prices, 1_000_000, model="no-such-model") is None 615 assert estimate_usd({}, 1_000_000) is None 616 617 # --- format_estimate_trailer: always carries "est." on both figures, never bare, and now 618 # states its own band (issue #433 task 2: a point figure alone hides calibrate()'s ~5x IQR). --- 619 fitted_stats = { 620 "n": 5, 621 "median": 6000.0, 622 "p25": 3000.0, 623 "p75": 12000.0, 624 "min": 800.0, 625 "max": 40000.0, 626 } 627 trailer = format_estimate_trailer(912_345, 1.20, 210, "2026-09-05", fitted_stats) 628 assert trailer.startswith("Cost: ~$1.20 est. ·"), trailer 629 assert "tokens est." in trailer, trailer 630 assert "210 lines, no session log" in trailer, trailer 631 assert trailer.endswith("--estimate 2026-09-05"), trailer 632 633 # Parse the band back out of the rendered string — not calibrate()'s own p25/p75 formula — so 634 # this fails if the band stops bracketing the point figure or stops matching stats["n"], not 635 # only if the arithmetic happens to change. 636 def _parse_band(s): 637 m = re.search( 638 r"IQR ([\d.]+)(k|M)?-([\d.]+)(k|M)?\s*tokens for this diff from (\d+) calibrated", s 639 ) 640 assert m, s 641 scale = {None: 1, "k": 1e3, "M": 1e6} 642 return ( 643 float(m.group(1)) * scale[m.group(2)], 644 float(m.group(3)) * scale[m.group(4)], 645 int(m.group(5)), 646 ) 647 648 lo, hi, shown_n = _parse_band(trailer) 649 assert lo < hi, (lo, hi) 650 assert lo <= 912_345 <= hi, (lo, 912_345, hi) # the point figure sits inside its own band 651 assert shown_n == fitted_stats["n"], (shown_n, fitted_stats["n"]) 652 653 no_price = format_estimate_trailer(500, None, 3, "2026-09-05", fitted_stats) 654 assert no_price.startswith("Cost: $n/a est. ·"), no_price 655 656 # n == 0 (fresh checkout, no calibration history): well-formed, but reads as a documented 657 # default, never a fit — and differently from the fitted case above. 658 fresh = format_estimate_trailer(500, None, 3, "2026-09-05", {"n": 0}) 659 assert fresh.startswith("Cost: $n/a est. ·"), fresh 660 assert "tokens est." in fresh, fresh 661 assert "3 lines, no session log" in fresh, fresh 662 assert fresh.endswith("--estimate 2026-09-05"), fresh 663 assert "documented default" in fresh, fresh 664 assert "documented default" not in trailer, trailer 665 assert "IQR" not in fresh, fresh 666 667 # --- calibrate() + diff_stats() + estimate_commit(): a synthetic repo, so the fit is a known 668 # number instead of depending on this checkout's ever-growing real history --- 669 with tempfile.TemporaryDirectory() as tmp: 670 repo = Path(tmp) 671 sh("git", "init", "-q", "-b", "main", cwd=repo) 672 sh("git", "config", "user.email", "test@example.com", cwd=repo) 673 sh("git", "config", "user.name", "Test", cwd=repo) 674 675 def commit(fname, content, message): 676 (repo / fname).write_text(content) 677 sh("git", "add", fname, cwd=repo) 678 sh("git", "commit", "-q", "-m", message, cwd=repo) 679 return sh("git", "rev-parse", "HEAD", cwd=repo).strip() 680 681 # 10 changed lines (all additions), 1,000,000 tokens measured -> ratio 100,000/line. 682 commit( 683 "a.txt", 684 "\n".join(f"line {i}" for i in range(10)) + "\n", 685 "first\n\nCost: ~1.0M tokens\n", 686 ) 687 # 100 changed lines, 5,000,000 tokens measured -> ratio 50,000/line. 688 commit( 689 "b.txt", 690 "\n".join(f"line {i}" for i in range(100)) + "\n", 691 "second\n\nCost: ~5.0M tokens\n", 692 ) 693 # A "shared bucket" trailer must not enter the calibration set even though it parses. 694 commit( 695 "c.txt", "x\n", "third\n\nCost: shared bucket $1 · 9.0M tokens — not split further\n" 696 ) 697 # An un-costed commit — the one we'll estimate against the calibration from the other two. 698 target = commit( 699 "d.txt", "\n".join(f"line {i}" for i in range(20)) + "\n", "fourth: no trailer" 700 ) 701 702 coeff, stats = calibrate(ref="main", cwd=repo) 703 assert stats["n"] == 2, stats # the shared-bucket commit must be excluded 704 assert coeff == statistics.median([100_000, 50_000]), coeff 705 706 prices2 = {DEFAULT_PRICE_MODEL: {"input_cost_per_token": 2e-06}} 707 tokens, usd, changed = estimate_commit(target, prices2, coeff, cwd=repo) 708 assert changed == 20, changed 709 expected_tokens = estimate_tokens_for_diff(20, 1, ["d.txt"], coeff) 710 assert tokens == expected_tokens, (tokens, expected_tokens) 711 assert ( 712 usd == round(expected_tokens * 2e-06, 10) or abs(usd - expected_tokens * 2e-06) < 1e-9 713 ) 714 715 # A repo with zero calibratable commits falls back to the documented default. 716 with tempfile.TemporaryDirectory() as tmp2: 717 empty_repo = Path(tmp2) 718 sh("git", "init", "-q", "-b", "main", cwd=empty_repo) 719 sh("git", "config", "user.email", "test@example.com", cwd=empty_repo) 720 sh("git", "config", "user.name", "Test", cwd=empty_repo) 721 (empty_repo / "x.txt").write_text("x\n") 722 sh("git", "add", "x.txt", cwd=empty_repo) 723 sh("git", "commit", "-q", "-m", "no cost trailer here", cwd=empty_repo) 724 fallback_coeff, fallback_stats = calibrate(ref="main", cwd=empty_repo) 725 assert fallback_stats["n"] == 0 726 assert fallback_coeff == DEFAULT_TOKENS_PER_LINE 727 728 # --- regression: sh(cwd=...) must not leak into a GIT_DIR the environment already points at. 729 # A git hook (pre-push, pre-commit) sets GIT_DIR (and friends) for its own repo; before this 730 # was fixed, the calibration commits above landed for real on whatever repo triggered the 731 # hook instead of staying inside their own tempdir. Reproduced here with a decoy "outer" repo 732 # standing in for the real one, GIT_DIR pointed at it, and an inner tempdir repo built the same 733 # way self_test() builds its calibration fixtures. --- 734 with tempfile.TemporaryDirectory() as outer_tmp: 735 outer = Path(outer_tmp) 736 sh("git", "init", "-q", "-b", "main", cwd=outer) 737 sh("git", "config", "user.email", "outer@example.com", cwd=outer) 738 sh("git", "config", "user.name", "Outer", cwd=outer) 739 (outer / "seed.txt").write_text("seed\n") 740 sh("git", "add", "seed.txt", cwd=outer) 741 sh("git", "commit", "-q", "-m", "seed", cwd=outer) 742 outer_head_before = sh("git", "rev-parse", "HEAD", cwd=outer).strip() 743 744 saved_git_dir = os.environ.get("GIT_DIR") 745 try: 746 os.environ["GIT_DIR"] = str(outer / ".git") 747 with tempfile.TemporaryDirectory() as inner_tmp: 748 inner = Path(inner_tmp) 749 sh("git", "init", "-q", "-b", "main", cwd=inner) 750 sh("git", "config", "user.email", "test@example.com", cwd=inner) 751 sh("git", "config", "user.name", "Test", cwd=inner) 752 (inner / "x.txt").write_text("x\n") 753 sh("git", "add", "x.txt", cwd=inner) 754 sh("git", "commit", "-q", "-m", "should stay inside tempdir", cwd=inner) 755 finally: 756 if saved_git_dir is None: 757 os.environ.pop("GIT_DIR", None) 758 else: 759 os.environ["GIT_DIR"] = saved_git_dir 760 761 outer_head_after = sh("git", "rev-parse", "HEAD", cwd=outer).strip() 762 assert outer_head_after == outer_head_before, ( 763 "sh(cwd=...) leaked into the ambient GIT_DIR instead of staying in cwd" 764 ) 765 assert not (outer / "x.txt").exists(), "inner repo's file leaked into the outer repo" 766 767 # --- messages(): a subagent transcript's own `gitBranch` is the *parent's*, not its own — 768 # `cwd` says which worktree it actually ran in (issue #433, confirmed live 2026-09-28). --- 769 with tempfile.TemporaryDirectory() as wt_tmp: 770 mrepo = Path(wt_tmp) / "main-checkout" 771 mrepo.mkdir() 772 sh("git", "init", "-q", "-b", "main", cwd=mrepo) 773 sh("git", "config", "user.email", "test@example.com", cwd=mrepo) 774 sh("git", "config", "user.name", "Test", cwd=mrepo) 775 (mrepo / "a.txt").write_text("a\n") 776 sh("git", "add", "a.txt", cwd=mrepo) 777 sh("git", "commit", "-q", "-m", "initial", cwd=mrepo) 778 779 worktree_dir = mrepo / ".tmp" / "wt-feature" 780 sh("git", "worktree", "add", "-q", "-b", "feature-x", str(worktree_dir), cwd=mrepo) 781 (worktree_dir / "b.txt").write_text("b\n") 782 sh("git", "add", "b.txt", cwd=worktree_dir) 783 sh("git", "commit", "-q", "-m", "the commit under test", cwd=worktree_dir) 784 target_sha = sh("git", "rev-parse", "HEAD", cwd=worktree_dir).strip() 785 target_when = dt.datetime.fromisoformat( 786 sh("git", "log", "-1", "--format=%aI", cwd=worktree_dir).strip() 787 ) 788 789 pdir = Path(wt_tmp) / "project-logs" 790 sid = "ses-parent" 791 (pdir / sid / "subagents").mkdir(parents=True) 792 793 def jsonl_line(**fields): 794 return json.dumps(fields) + "\n" 795 796 usage_small = {"input_tokens": 100, "output_tokens": 20} 797 usage_big = {"input_tokens": 300_000, "output_tokens": 5_000} 798 usage_tiny = {"input_tokens": 7, "output_tokens": 3} 799 800 (pdir / f"{sid}.jsonl").write_text( 801 jsonl_line( 802 message={"id": "msg-parent", "usage": usage_small, "model": "claude-sonnet-5"}, 803 timestamp="2000-01-01T10:00:00Z", 804 gitBranch="main", 805 cwd="/no/such/worktree", 806 ) 807 + jsonl_line( 808 # cwd IS the main checkout; gitBranch names a branch other than its current one 809 # ("main" here). gitBranch is correct per-message history and must win. 810 message={ 811 "id": "msg-main-checkout", 812 "usage": usage_tiny, 813 "model": "claude-sonnet-5", 814 }, 815 timestamp="2000-01-01T10:02:00Z", 816 gitBranch="feature-x", 817 cwd=str(mrepo), 818 ) 819 ) 820 sub_a = jsonl_line( 821 message={"id": "msg-sub-a", "usage": usage_big, "model": "claude-sonnet-5"}, 822 timestamp="2000-01-01T10:05:00Z", 823 gitBranch="main", 824 cwd=str(worktree_dir), 825 isSidechain=True, 826 ) 827 sub_b = jsonl_line( 828 message={"id": "msg-sub-b", "usage": usage_small, "model": "claude-sonnet-5"}, 829 timestamp="2000-01-01T10:06:00Z", 830 gitBranch="main", 831 cwd="/no/such/worktree/either", 832 isSidechain=True, 833 ) 834 sub_c = jsonl_line( 835 # Same as msg-main-checkout, but sidechain: must also keep gitBranch, not fall to "" 836 # just because it's a sidechain. 837 message={"id": "msg-sub-c", "usage": usage_tiny, "model": "claude-sonnet-5"}, 838 timestamp="2000-01-01T10:07:00Z", 839 gitBranch="feature-x", 840 cwd=str(mrepo), 841 isSidechain=True, 842 ) 843 sub_d = jsonl_line( 844 # cwd is a subdirectory of the main checkout, not its root — real logs have these 845 # (e.g. docs/img/logo). Still the main checkout's own tree, so still gitBranch. 846 message={"id": "msg-sub-d", "usage": usage_tiny, "model": "claude-sonnet-5"}, 847 timestamp="2000-01-01T10:08:00Z", 848 gitBranch="feature-x", 849 cwd=str(mrepo / "docs" / "img" / "logo"), 850 isSidechain=True, 851 ) 852 (pdir / sid / "subagents" / "agent-x.jsonl").write_text(sub_a + sub_b + sub_c + sub_d) 853 854 got = messages(pdir, None, mrepo) 855 assert len(got) == 6, got 856 parent_branch = got[0][4] 857 main_checkout_branch = got[1][4] 858 sub_a_branch = got[2][4] 859 sub_b_branch = got[3][4] 860 sub_c_branch = got[4][4] 861 sub_d_branch = got[5][4] 862 863 # Non-sidechain, cwd matches no known worktree: existing behaviour must not regress. 864 assert parent_branch == "main", parent_branch 865 866 # Non-sidechain, cwd IS the main checkout: keep gitBranch ("feature-x"), not 867 # worktree_branches()'s "current branch" entry for it ("main") — the main checkout's 868 # current branch says nothing about what branch it was on when a given message was made. 869 assert main_checkout_branch == "feature-x", main_checkout_branch 870 871 # Sidechain, cwd resolves to the feature-x worktree: use the worktree's branch, not the 872 # inherited (misleading) parent gitBranch. 873 assert sub_a_branch == "feature-x", sub_a_branch 874 875 # Sidechain, cwd matches no known worktree: "" (time-only rule), not the parent's branch — 876 # a wrong branch would be discarded outright, "" is only less precise. 877 assert sub_b_branch == "", sub_b_branch 878 879 # Sidechain, cwd IS the main checkout: keep gitBranch, same as the non-sidechain case — 880 # a parent that was itself running in the main checkout inherits correct history. 881 assert sub_c_branch == "feature-x", sub_c_branch 882 883 # Sidechain, cwd is a subdirectory of the main checkout (not its root): still gitBranch, 884 # not "" — real logs have these and a naive exact match would send them to the wrong arm. 885 assert sub_d_branch == "feature-x", sub_d_branch 886 887 # End to end (issue #433): run the fixture through `attribute()` itself — the same 888 # function main() calls — not a re-implementation of its `ts <= when` pairing or 889 # summation. `branches_containing`/`branches_with_equivalent` read the ambient cwd (no 890 # `cwd=` param, same as `commits()`), so chdir into the synthetic repo for the call. 891 cs = [(target_when, "feature-x", target_sha, "the commit under test")] 892 saved_cwd = os.getcwd() 893 try: 894 os.chdir(worktree_dir) 895 buckets = attribute(got, cs, {}) 896 finally: 897 os.chdir(saved_cwd) 898 899 # sub_a lands on target_sha by branch (feature-x, resolved from cwd); msg-main-checkout, 900 # sub_c and sub_d land there too, by branch (their own gitBranch, "feature-x" — cwd'd in 901 # or under the main checkout); sub_b lands there by the time-only rule (its cwd matched no 902 # worktree, so its branch is ""), since cs holds only this one commit. parent (branch 903 # "main", not in {"feature-x"}) is discarded entirely — it must not appear anywhere, not 904 # even "uncommitted". 905 bucket = buckets[target_sha] 906 actual_total = bucket["in"] + bucket["out"] + bucket["cache_w"] + bucket["cache_r"] 907 tiny_total = usage_tiny["input_tokens"] + usage_tiny["output_tokens"] 908 expected_total = ( 909 usage_big["input_tokens"] 910 + usage_big["output_tokens"] 911 + usage_small["input_tokens"] 912 + usage_small["output_tokens"] 913 + 3 * tiny_total # msg-main-checkout + sub_c + sub_d, all usage_tiny 914 ) 915 # Exact match, no tolerance: every field summed here is an int token count, so there is 916 # no rounding for a tolerance to absorb. 917 assert actual_total == expected_total, ( 918 f"attribute() must land the subagent-shaped fixture's measured total on " 919 f"{target_sha[:7]} exactly (int summation, 0 tolerance): got {actual_total}, " 920 f"expected {expected_total} — a regression in the ts<=when pairing or the " 921 "summation would silently change this" 922 ) 923 uncommitted_total = sum( 924 buckets["uncommitted"][k] for k in ("in", "out", "cache_w", "cache_r") 925 ) 926 assert uncommitted_total == 0, ( 927 f"the parent message's branch ({parent_branch!r}) doesn't carry the commit under " 928 f"test and must be discarded outright, not misrouted to uncommitted: {uncommitted_total}" 929 ) 930 # attribute() just populated this module-global cache from the synthetic (about-to-vanish) 931 # repo above; reset it so a test block added later doesn't inherit it. 932 global _equivalents 933 _equivalents = None 934 935 # --- messages_opencode(): a synthetic DB built against the real schema (MIP-0013 task 2, 936 # confirmed live 2026-09-07 against opencode 1.18.25) — table/column names, tokens.{...} JSON 937 # shape, session.directory filtering. No real ~/.local/share/opencode touched. --- 938 with tempfile.TemporaryDirectory() as oc_tmp: 939 oc_root = Path(oc_tmp) 940 data_dir = oc_root / "data" / "opencode" 941 data_dir.mkdir(parents=True) 942 saved_xdg = os.environ.get("XDG_DATA_HOME") 943 os.environ["XDG_DATA_HOME"] = str(oc_root / "data") 944 try: 945 db_path = data_dir / "opencode-stable.db" 946 con = sqlite3.connect(db_path) 947 con.execute("create table session (id text, directory text)") 948 con.execute("create table message (id text, session_id text, data text)") 949 here = str(oc_root / "repo") 950 elsewhere = str(oc_root / "other-repo") 951 con.execute("insert into session values (?, ?)", ("ses_here", here)) 952 con.execute("insert into session values (?, ?)", ("ses_elsewhere", elsewhere)) 953 user_msg = json.dumps({"role": "user"}) 954 asst_msg = json.dumps( 955 { 956 "role": "assistant", 957 "modelID": "llama3.2", 958 "providerID": "ollama", 959 "tokens": {"input": 4096, "output": 8, "cache": {"read": 100, "write": 0}}, 960 "time": {"created": 1788756841410}, 961 } 962 ) 963 elsewhere_msg = json.dumps( 964 { 965 "role": "assistant", 966 "modelID": "llama3.2", 967 "providerID": "ollama", 968 "tokens": {"input": 1, "output": 1, "cache": {"read": 0, "write": 0}}, 969 "time": {"created": 1788756841410}, 970 } 971 ) 972 con.execute("insert into message values (?, ?, ?)", ("msg_user", "ses_here", user_msg)) 973 con.execute("insert into message values (?, ?, ?)", ("msg_asst", "ses_here", asst_msg)) 974 con.execute( 975 "insert into message values (?, ?, ?)", 976 ("msg_other_repo", "ses_elsewhere", elsewhere_msg), 977 ) 978 con.commit() 979 con.close() 980 981 found = messages_opencode(Path(here), None) 982 assert len(found) == 1, ( 983 found 984 ) # the user-role row and the other repo's row are both dropped 985 ts, sid, model, usage, branch = found[0] 986 assert sid == "ses_here", sid 987 assert model == "ollama/llama3.2", model 988 assert usage == { 989 "input_tokens": 4096, 990 "output_tokens": 8, 991 "cache_creation_input_tokens": 0, 992 "cache_read_input_tokens": 100, 993 }, usage 994 assert branch == "", branch # OpenCode carries no gitBranch — main()'s time-only rule 995 996 assert messages_opencode(Path(str(oc_root / "no-such-repo")), None) == [] 997 finally: 998 if saved_xdg is None: 999 os.environ.pop("XDG_DATA_HOME", None) 1000 else: 1001 os.environ["XDG_DATA_HOME"] = saved_xdg 1002 1003 assert opencode_db_path() is None or opencode_db_path().is_file() 1004 1005 print( 1006 f"cost-split self-test: ok (parser {len(cases)} cases, synthetic calibration " 1007 f"median {coeff:.0f} tokens/line from {stats['n']} commits)" 1008 ) 1009 return 0 1010 1011 1012# --------------------------------------------------------------------------------------------- 1013 1014 1015def attribute(msgs, cs, prices): 1016 """[(ts, session, model, usage, branch)], [(when, branch, sha, subject)] -> sha/"uncommitted" 1017 -> totals. A message counts toward a branch only if it was made *on* that branch (its 1018 `mbranch`), then falls into the first commit whose author date is after it. Without the 1019 branch test, every session on the machine that ran before a branch's first commit — other 1020 agents, other features — landed on that commit (seen: 335M tokens on a 500-line commit). A 1021 message with no branch (older logs, or a sidechain whose `cwd` matched no worktree) keeps the 1022 time-only rule. 1023 1024 "On that branch" means: the message's branch contains the commit — a commit authored on 1025 feat/a and now priced from feat/b (stacked on a) is still paid for by the messages made on 1026 feat/a. Computed once per commit from `git branch -a --contains`. 1027 1028 Not pure despite the signature: `branches_containing`/`branches_with_equivalent` take no 1029 `cwd` and read the process's actual working directory, so the caller must already be in (or 1030 have chdir'd into) the repo `cs`'s commits belong to. 1031 """ 1032 buckets = defaultdict( 1033 lambda: { 1034 "in": 0, 1035 "out": 0, 1036 "cache_w": 0, 1037 "cache_r": 0, 1038 "usd": 0.0, 1039 "models": set(), 1040 "sessions": set(), 1041 "priced": True, 1042 } 1043 ) 1044 contains = { 1045 sha: branches_containing(sha) | branches_with_equivalent(sha) for _w, _b, sha, _s in cs 1046 } 1047 ours = {c[1] for c in cs}.union(*contains.values()) if cs else set() 1048 for ts, session, model, u, mbranch in msgs: 1049 if mbranch and mbranch not in ours: 1050 continue 1051 target = "uncommitted" 1052 for when, _cbranch, sha, _s in cs: 1053 if mbranch and mbranch not in contains[sha]: 1054 continue 1055 if ts <= when: 1056 target = sha 1057 break 1058 b = buckets[target] 1059 b["in"] += u.get("input_tokens", 0) 1060 b["out"] += u.get("output_tokens", 0) 1061 b["cache_w"] += u.get("cache_creation_input_tokens", 0) 1062 b["cache_r"] += u.get("cache_read_input_tokens", 0) 1063 b["models"].add(model) 1064 b["sessions"].add(session[:8]) 1065 usd = price(prices, model, u) 1066 if usd is None: 1067 b["priced"] = False 1068 else: 1069 b["usd"] += usd 1070 return buckets 1071 1072 1073def build_rows(cs, buckets, order, prices, do_estimate, coeff): 1074 rows = [] 1075 for key in order: 1076 b = buckets.get(key) 1077 meta = next( 1078 ((br, subj) for _w, br, sha, subj in cs if sha == key), ("", "uncommitted (so far)") 1079 ) 1080 if not b: 1081 if key == "uncommitted" or not do_estimate: 1082 continue 1083 # No cwd override: `key` is a full sha (unambiguous from any worktree of this repo), 1084 # but if it were ever a symbolic ref like HEAD, resolving it against repo_root()'s 1085 # main-checkout path — right for session-log lookup, wrong here — would silently 1086 # answer for the wrong branch when run from a worktree. 1087 est = estimate_commit(key, prices, coeff) 1088 if est is None: 1089 continue 1090 tokens, usd, changed = est 1091 rows.append( 1092 { 1093 "commit": key[:7], 1094 "branch": meta[0], 1095 "subject": meta[1], 1096 "input": 0, 1097 "output": 0, 1098 "cache_write": 0, 1099 "cache_read": 0, 1100 "tokens": tokens, 1101 "usd": round(usd, 2) if usd is not None else None, 1102 "models": ["diff-size estimate"], 1103 "sessions": [], 1104 "estimated": True, 1105 "changed_lines": changed, 1106 } 1107 ) 1108 continue 1109 rows.append( 1110 { 1111 "commit": key[:7] if key != "uncommitted" else "-", 1112 "branch": meta[0], 1113 "subject": meta[1], 1114 "input": b["in"], 1115 "output": b["out"], 1116 "cache_write": b["cache_w"], 1117 "cache_read": b["cache_r"], 1118 "tokens": b["in"] + b["out"] + b["cache_w"] + b["cache_r"], 1119 "usd": round(b["usd"], 2) if b["priced"] else None, 1120 "models": sorted(b["models"]), 1121 "sessions": sorted(b["sessions"]), 1122 "estimated": False, 1123 } 1124 ) 1125 return rows 1126 1127 1128def main(): 1129 ap = argparse.ArgumentParser( 1130 description=__doc__.split("\n\n")[0], formatter_class=argparse.RawDescriptionHelpFormatter 1131 ) 1132 ap.add_argument( 1133 "stack", nargs="?", help="mip-NNNN: every branch of the stack (default: current branch)" 1134 ) 1135 ap.add_argument("--session", help="session id prefix (default: every session of this project)") 1136 ap.add_argument("--json", action="store_true") 1137 ap.add_argument( 1138 "--estimate", 1139 action="store_true", 1140 help="fill in commits with no logged usage from a diff-size estimate, clearly labelled", 1141 ) 1142 ap.add_argument( 1143 "--estimate-commit", 1144 metavar="SHA", 1145 help="print just the estimated Cost: trailer for one commit — no session logs needed", 1146 ) 1147 ap.add_argument( 1148 "--verbose", action="store_true", help="with --estimate, print the calibration fit" 1149 ) 1150 ap.add_argument( 1151 "--harness", 1152 choices=["claude", "opencode", "all"], 1153 default="all", 1154 help="which agent harness's session logs to read (MIP-0013 task 2; default: both)", 1155 ) 1156 ap.add_argument("--self-test", action="store_true") 1157 a = ap.parse_args() 1158 1159 if a.self_test: 1160 sys.exit(self_test()) 1161 1162 root = repo_root() 1163 1164 if a.estimate_commit: 1165 # No cwd override here either: calibrate()'s `origin/main` and estimate_commit()'s sha 1166 # argument both mean the same thing from any worktree, but running them pinned to 1167 # repo_root() would resolve a symbolic ref like `HEAD` against the *main* checkout's 1168 # branch, not the one this command was actually run from. 1169 coeff, stats = calibrate() 1170 if a.verbose: 1171 print_calibration(coeff, stats) 1172 prices = load_prices(root) 1173 est = estimate_commit(a.estimate_commit, prices, coeff) 1174 if est is None: 1175 sys.exit( 1176 f"cost-split --estimate-commit: {a.estimate_commit} has no diff to estimate from" 1177 ) 1178 tokens, usd, changed = est 1179 print(format_estimate_trailer(tokens, usd, changed, dt.date.today().isoformat(), stats)) 1180 return 1181 1182 msgs = [] 1183 if a.harness in ("claude", "all"): 1184 pdir = project_dir(root) 1185 if pdir.is_dir(): 1186 msgs += messages(pdir, a.session, root) 1187 elif a.harness == "claude": 1188 sys.exit(f"no session logs at {pdir}") 1189 if a.harness in ("opencode", "all"): 1190 msgs += messages_opencode(root, a.session) 1191 msgs.sort(key=lambda x: x[0]) 1192 cs = commits(root, a.stack.lower() if a.stack else None) 1193 if not cs: 1194 sys.exit("no commits ahead of origin/main on the selected branch(es)") 1195 prices = load_prices(root) 1196 1197 coeff, calib_stats = (None, None) 1198 if a.estimate: 1199 coeff, calib_stats = calibrate() 1200 1201 order = [c[2] for c in cs] + ["uncommitted"] 1202 buckets = attribute(msgs, cs, prices) 1203 1204 rows = build_rows(cs, buckets, order, prices, a.estimate, coeff) 1205 1206 if a.json: 1207 print(json.dumps(rows, indent=1)) 1208 return 1209 1210 if a.estimate and a.verbose: 1211 print_calibration(coeff, calib_stats) 1212 print() 1213 1214 today = dt.date.today().isoformat() 1215 print( 1216 f"{'commit':8} {'usd':>9} {'tokens':>11} {'in':>7} {'out':>7} {'cache_w':>8} {'cache_r':>9} branch / subject" 1217 ) 1218 any_estimated = False 1219 for r in rows: 1220 mark = "*" if r.get("estimated") else "" 1221 any_estimated = any_estimated or bool(mark) 1222 usd = f"${r['usd']:.2f}{mark}" if r["usd"] is not None else f"n/a{mark}" 1223 tokens = f"{r['tokens']:,}{mark}" 1224 print( 1225 f"{r['commit']:8} {usd:>9} {tokens:>11} {r['input']:>7,} {r['output']:>7,} {r['cache_write']:>8,} {r['cache_read']:>9,} {r['branch']} {r['subject'][:60]}" 1226 ) 1227 if any_estimated: 1228 print("* diff-size estimate, no session log — scripts/cost-split.py --estimate") 1229 1230 per_branch = defaultdict(lambda: [0.0, 0, 0, True, set()]) 1231 for r in rows: 1232 if r["commit"] == "-": 1233 continue 1234 pb = per_branch[r["branch"]] 1235 pb[0] += r["usd"] or 0 1236 pb[1] += r["tokens"] 1237 pb[2] += r["cache_read"] 1238 pb[3] = pb[3] and r["usd"] is not None 1239 pb[4].update(r["models"]) 1240 print("\nCost: trailers per branch (paste into the commit / PR):") 1241 for br, (usd, tokens, cached, priced, models) in per_branch.items(): 1242 dollars = f"~${usd:.2f}" if priced else "$n/a" 1243 share = f"{100 * cached / tokens:.0f}% cache reads" if tokens else "no tokens" 1244 print( 1245 f" {br}: Cost: {dollars} · {tokens / 1e6:.1f}M tokens, {share} ({', '.join(sorted(models))}) · split by branch and author date, scripts/cost-split.py {today}" 1246 ) 1247 1248 estimated_rows = [r for r in rows if r.get("estimated")] 1249 if estimated_rows: 1250 print("\nCost: trailers for commits with no session log (estimated — paste per commit):") 1251 for r in estimated_rows: 1252 trailer = format_estimate_trailer( 1253 r["tokens"], r["usd"], r["changed_lines"], today, calib_stats 1254 ) 1255 print(f" {r['commit']}: {trailer}") 1256 1257 total = sum(r["usd"] or 0 for r in rows) 1258 print(f"\nsession total ${total:.2f} for {len(rows)} buckets (list prices, LiteLLM table)") 1259 1260 1261if __name__ == "__main__": 1262 main()
80def sh(*args, cwd=None): 81 env = None 82 if cwd is not None: 83 # A cwd means "operate on this repo, not whatever GIT_DIR/GIT_WORK_TREE/etc. the 84 # environment already points at" — a git hook (pre-push, pre-commit) sets those for its 85 # own repo, and without stripping them a subprocess git call here follows the env instead 86 # of cwd: with GIT_DIR set but not GIT_WORK_TREE, git treats cwd as the work tree and 87 # happily commits its content into the *real* repo's history, while the real working tree 88 # never receives the files — surfacing afterward as spurious "deleted: <file>" entries. 89 # See the self-test below, which reproduces this exact scenario. 90 env = {k: v for k, v in os.environ.items() if not k.startswith("GIT_")} 91 return subprocess.run(args, check=True, capture_output=True, text=True, cwd=cwd, env=env).stdout
94def repo_root(): 95 # The main checkout even from a worktree: Claude Code keys its logs on the directory the 96 # session was started in, which for this repo's `.tmp/wt-*` worktrees is the main checkout. 97 common = Path(sh("git", "rev-parse", "--path-format=absolute", "--git-common-dir").strip()) 98 return common.parent
107def load_prices(root): 108 cache = root / ".tmp" / "litellm-prices.json" 109 try: 110 if cache.exists() and (dt.datetime.now().timestamp() - cache.stat().st_mtime) < 86400: 111 return json.loads(cache.read_text()) 112 with urllib.request.urlopen(PRICES_URL, timeout=15) as r: 113 data = r.read() 114 cache.parent.mkdir(exist_ok=True) 115 cache.write_bytes(data) 116 return json.loads(data) 117 except Exception: 118 return json.loads(cache.read_text()) if cache.exists() else {}
121def price(prices, model, u): 122 p = prices.get(model) 123 if not p: 124 return None 125 cc = u.get("cache_creation") or {} 126 write_5m = cc.get("ephemeral_5m_input_tokens", u.get("cache_creation_input_tokens", 0)) 127 write_1h = cc.get("ephemeral_1h_input_tokens", 0) 128 if not cc: 129 write_5m = u.get("cache_creation_input_tokens", 0) 130 return ( 131 u.get("input_tokens", 0) * p.get("input_cost_per_token", 0) 132 + u.get("output_tokens", 0) * p.get("output_cost_per_token", 0) 133 + write_5m * p.get("cache_creation_input_token_cost", 0) 134 + write_1h 135 * p.get( 136 "cache_creation_input_token_cost_above_1hr", p.get("cache_creation_input_token_cost", 0) 137 ) 138 + u.get("cache_read_input_tokens", 0) * p.get("cache_read_input_token_cost", 0) 139 )
156def worktree_branches(root): 157 """{absolute worktree path: branch name}, from `git worktree list --porcelain` run at `root` 158 (always the main checkout, per `repo_root()`) — covers the main checkout itself plus every 159 live `.tmp/wt-*`. Used to resolve a message's `cwd` to the branch it actually ran on.""" 160 out = {} 161 path = None 162 for line in sh("git", "worktree", "list", "--porcelain", cwd=root).splitlines(): 163 if line.startswith("worktree "): 164 path = line.removeprefix("worktree ") 165 elif line.startswith("branch refs/heads/") and path: 166 out[path] = line.removeprefix("branch refs/heads/") 167 path = None 168 return out
{absolute worktree path: branch name}, from git worktree list --porcelain run at root
(always the main checkout, per repo_root()) — covers the main checkout itself plus every
live .tmp/wt-*. Used to resolve a message's cwd to the branch it actually ran on.
199def messages(pdir, session_prefix, root): 200 """(timestamp, session, model, usage, branch) per assistant message, deduplicated by message 201 id.""" 202 seen = {} 203 wt_branches = worktree_branches(root) 204 linked = {path: branch for path, branch in wt_branches.items() if path != str(root)} 205 for sid, f in _session_message_files(pdir, session_prefix): 206 with open(f, encoding="utf-8") as fh: 207 for line in fh: 208 try: 209 d = json.loads(line) 210 except json.JSONDecodeError: 211 continue 212 m = d.get("message") 213 if not isinstance(m, dict) or not m.get("usage") or not d.get("timestamp"): 214 continue 215 # Claude Code writes `<synthetic>` assistant messages (interrupts, tool-result 216 # stand-ins) with a usage block of zeros and no real model — not a priced request. 217 if m.get("model") == "<synthetic>": 218 continue 219 key = m.get("id") or d.get("requestId") or d.get("uuid") 220 ts = dt.datetime.fromisoformat(d["timestamp"].replace("Z", "+00:00")) 221 # `gitBranch` is the session directory's branch at message time — right for the 222 # long-lived, branch-switching main checkout, wrong for a *linked* worktree (every 223 # line of a subagent transcript there inherits its parent's value; issue #433). 224 # Only a linked worktree's own branch overrides it; a sidechain matching neither 225 # falls to "" rather than trust that inherited branch. 226 cwd = d.get("cwd") 227 wt_branch = _worktree_branch_for_cwd(cwd, linked) 228 if wt_branch is not None: 229 branch = wt_branch 230 elif d.get("isSidechain") and not _under(cwd, root, linked): 231 branch = "" 232 else: 233 branch = d.get("gitBranch") or "" 234 seen[key] = (ts, sid, m.get("model", "?"), m["usage"], branch) 235 return sorted(seen.values(), key=lambda x: x[0])
(timestamp, session, model, usage, branch) per assistant message, deduplicated by message id.
238def opencode_db_path(): 239 """The real store, confirmed live 2026-09-07 against opencode 1.18.25 (MIP-0013 task 2, §11 240 OQ2): `$XDG_DATA_HOME/opencode/opencode-stable.db` — not the `storage/message/*/msg_*.json` 241 files an earlier draft of that MIP assumed; those don't exist in this version. Falls back to 242 `~/.local/share/opencode` per the XDG default when the env var is unset, and tries the 243 non-`-stable` filename too in case a future/edge/nightly channel uses it.""" 244 base = Path(os.environ.get("XDG_DATA_HOME", str(Path.home() / ".local" / "share"))) 245 d = base / "opencode" 246 for name in ("opencode-stable.db", "opencode.db"): 247 p = d / name 248 if p.is_file(): 249 return p 250 return None
The real store, confirmed live 2026-09-07 against opencode 1.18.25 (MIP-0013 task 2, §11
OQ2): $XDG_DATA_HOME/opencode/opencode-stable.db — not the storage/message/*/msg_*.json
files an earlier draft of that MIP assumed; those don't exist in this version. Falls back to
~/.local/share/opencode per the XDG default when the env var is unset, and tries the
non--stable filename too in case a future/edge/nightly channel uses it.
253def messages_opencode(root, session_prefix): 254 """(timestamp, session, model, usage, gitBranch) per assistant message — the OpenCode-harness 255 twin of `messages()` above, same tuple shape so `main()`'s attribution loop needs no branching 256 beyond which reader it calls. `gitBranch` is always `""`: OpenCode's `message`/`session` rows 257 carry no branch field (confirmed live, same session), so a message here always falls back to 258 `main()`'s time-only rule — documented, not a bug to fix, since OpenCode sessions in this repo 259 are (so far) one-project, foreground, interactive runs, not marola's multi-worktree pattern. 260 261 `usage` is shaped like Claude Code's dict (`input_tokens`/`output_tokens`/ 262 `cache_creation_input_tokens`/`cache_read_input_tokens`) so `price()` needs no change; `model` 263 is `providerID/modelID` (e.g. `ollama/llama3.2`) — `price()` looks it up in the same LiteLLM 264 table and returns `None` for a local Ollama model, same graceful "not priced" path Claude 265 Code's `<synthetic>` skip already exercises, not a special case here. 266 """ 267 db = opencode_db_path() 268 if db is None: 269 return [] 270 seen = {} 271 con = sqlite3.connect(f"file:{db}?mode=ro", uri=True) 272 try: 273 con.row_factory = sqlite3.Row 274 cur = con.cursor() 275 cur.execute( 276 "select m.id, m.session_id, m.data from message m " 277 "join session s on s.id = m.session_id " 278 "where s.directory = ?", 279 (str(root),), 280 ) 281 for row in cur.fetchall(): 282 if session_prefix and not row["session_id"].startswith(session_prefix): 283 continue 284 try: 285 d = json.loads(row["data"]) 286 except json.JSONDecodeError: 287 continue 288 if d.get("role") != "assistant" or not isinstance(d.get("tokens"), dict): 289 continue 290 t = d["tokens"] 291 cache = t.get("cache") or {} 292 created_ms = (d.get("time") or {}).get("created") 293 if created_ms is None: 294 continue 295 ts = dt.datetime.fromtimestamp(created_ms / 1000, tz=dt.UTC) 296 usage = { 297 "input_tokens": t.get("input", 0), 298 "output_tokens": t.get("output", 0), 299 "cache_creation_input_tokens": cache.get("write", 0), 300 "cache_read_input_tokens": cache.get("read", 0), 301 } 302 model = f"{d.get('providerID', '?')}/{d.get('modelID', '?')}" 303 seen[row["id"]] = (ts, row["session_id"], model, usage, "") 304 finally: 305 con.close() 306 return sorted(seen.values(), key=lambda x: x[0])
(timestamp, session, model, usage, gitBranch) per assistant message — the OpenCode-harness
twin of messages() above, same tuple shape so main()'s attribution loop needs no branching
beyond which reader it calls. gitBranch is always "": OpenCode's message/session rows
carry no branch field (confirmed live, same session), so a message here always falls back to
main()'s time-only rule — documented, not a bug to fix, since OpenCode sessions in this repo
are (so far) one-project, foreground, interactive runs, not marola's multi-worktree pattern.
usage is shaped like Claude Code's dict (input_tokens/output_tokens/
cache_creation_input_tokens/cache_read_input_tokens) so price() needs no change; model
is providerID/modelID (e.g. ollama/llama3.2) — price() looks it up in the same LiteLLM
table and returns None for a local Ollama model, same graceful "not priced" path Claude
Code's <synthetic> skip already exercises, not a special case here.
312def branches_with_equivalent(sha): 313 """Branches holding a commit with the same author date and subject as `sha` — the same 314 commit after a rebase, cherry-pick or `just cost-fill` rewrite gave it a new hash. Without 315 this, the branch a subagent authored on stops "containing" its own commits the moment 316 they are rewritten, and its usage would no longer match them.""" 317 global _equivalents 318 if _equivalents is None: 319 _equivalents = defaultdict(set) 320 refs = sh( 321 "git", "for-each-ref", "--format=%(refname:short)", "refs/heads", "refs/remotes/origin" 322 ) 323 for ref in refs.splitlines(): 324 ref = ref.strip() 325 if not ref or ref in ("main", "origin/main", "origin/HEAD"): 326 continue 327 name = ref[7:] if ref.startswith("origin/") else ref 328 try: 329 log = sh("git", "log", "--format=%aI%x00%s", f"origin/main..{ref}") 330 except subprocess.CalledProcessError: 331 continue 332 for line in log.splitlines(): 333 if "\x00" in line: 334 _equivalents[tuple(line.split("\x00", 1))].add(name) 335 when, subject = sh("git", "log", "-1", "--format=%aI%x00%s", sha).rstrip("\n").split("\x00", 1) 336 return set(_equivalents.get((when, subject), set()))
Branches holding a commit with the same author date and subject as sha — the same
commit after a rebase, cherry-pick or just cost-fill rewrite gave it a new hash. Without
this, the branch a subagent authored on stops "containing" its own commits the moment
they are rewritten, and its usage would no longer match them.
339def branches_containing(sha): 340 """Local and origin branch names that contain `sha` (origin/ prefix stripped).""" 341 out = set() 342 for line in sh( 343 "git", "branch", "-a", "--contains", sha, "--format=%(refname:short)" 344 ).splitlines(): 345 b = line.strip() 346 if not b or b == "origin/HEAD" or " -> " in b: 347 continue 348 out.add(b[7:] if b.startswith("origin/") else b) 349 return out
Local and origin branch names that contain sha (origin/ prefix stripped).
352def commits(root, stack): 353 """[(time, branch, sha, subject)] — each commit once, on the first branch that carries it.""" 354 out, known = [], set() 355 branches = [] 356 if stack: # a MIP (mip-0005 → every mip-0005/k-* branch) or one plain branch name 357 branches = [ 358 b.strip() 359 for b in sh( 360 "git", "branch", "--list", f"{stack}/*", "--format=%(refname:short)" 361 ).splitlines() 362 if b.strip() 363 ] 364 365 # mip-NNNN/k-slug sorts by k; a branch without a numeric task prefix (mip-0012/llm4s-adoption, 366 # a one-PR MIP) sorts after the numbered ones, by name. 367 def task_no(b): 368 head = b.split("/")[1].split("-")[0] 369 return (0, int(head), "") if head.isdigit() else (1, 0, b) 370 371 branches.sort(key=task_no) 372 if not branches: 373 branches = [stack] 374 else: 375 branches = [sh("git", "branch", "--show-current").strip()] 376 for b in branches: 377 # Author date, not committer date: a rebase, cherry-pick or `just cost-fill` rewrite 378 # restamps the committer date of every commit to the same second, and the time split 379 # then hands the whole branch's usage to its first commit. Author dates survive all three. 380 log = sh("git", "log", "--reverse", "--format=%H%x00%aI%x00%s", f"origin/main..{b}") 381 for line in log.splitlines(): 382 sha, when, subject = line.split("\x00") 383 if sha in known: 384 continue 385 known.add(sha) 386 out.append((dt.datetime.fromisoformat(when), b, sha, subject)) 387 return sorted(out, key=lambda c: c[0])
[(time, branch, sha, subject)] — each commit once, on the first branch that carries it.
395def parse_token_count(cost_line): 396 """Tokens named in one `Cost:` trailer line, or None if it doesn't carry a figure usable for 397 calibration. Skips bucket-wide notes ('shared bucket', 'not split further') — the same token 398 count copy-pasted onto several sibling commits, so pairing any one of them with its own diff 399 size would just add noise — and the one commit that admits it never got a number 400 ('unmeasured').""" 401 if "shared bucket" in cost_line or "unmeasured" in cost_line: 402 return None 403 m = TOKEN_TRAILER_RE.search(cost_line) 404 if not m: 405 return None 406 n = float(m.group(1).replace(",", "")) 407 unit = m.group(2) 408 return n * 1e6 if unit == "M" else n * 1e3 if unit == "k" else n
Tokens named in one Cost: trailer line, or None if it doesn't carry a figure usable for
calibration. Skips bucket-wide notes ('shared bucket', 'not split further') — the same token
count copy-pasted onto several sibling commits, so pairing any one of them with its own diff
size would just add noise — and the one commit that admits it never got a number
('unmeasured').
411def diff_stats(sha, cwd=None): 412 """(changed_lines, files_touched, paths) for one commit, from `git show --numstat`.""" 413 added = deleted = files = 0 414 paths = [] 415 for line in sh("git", "show", "--numstat", "--format=", sha, cwd=cwd).splitlines(): 416 if not line.strip(): 417 continue 418 parts = line.split("\t") 419 if len(parts) != 3: 420 continue 421 a, d, path = parts 422 added += 0 if a == "-" else int(a) 423 deleted += 0 if d == "-" else int(d) 424 files += 1 425 paths.append(path) 426 return added + deleted, files, paths
(changed_lines, files_touched, paths) for one commit, from git show --numstat.
429def calibrate(ref="origin/main", cwd=None): 430 """Median tokens-per-changed-line across every `ref` commit with a measured `Cost:` trailer 431 carrying a token figure, paired with that commit's own diff size (`diff_stats`). The median 432 (not a least-squares fit) is deliberate: a big-context, small-diff commit can be two orders of 433 magnitude above a routine one (this repo's own spread is ~800-440,000 tokens/line), and a 434 maintainer already eyeballs a `Cost:` trailer the same way — ignore the extremes, trust the 435 middle. Returns (coefficient, stats); stats["n"] == 0 means DEFAULT_TOKENS_PER_LINE was used. 436 """ 437 body = sh("git", "log", ref, "--format=%H%x00%B%x00END", cwd=cwd) 438 ratios = [] 439 for chunk in body.split("\x00END\n"): 440 if "\x00" not in chunk: 441 continue 442 sha, msg = chunk.split("\x00", 1) 443 sha = sha.strip() 444 tokens = None 445 for line in msg.splitlines(): 446 if line.startswith("Cost:"): 447 tokens = parse_token_count(line) 448 if tokens: 449 break 450 if not tokens: 451 continue 452 changed, _files, _paths = diff_stats(sha, cwd=cwd) 453 if changed <= 0: 454 continue 455 ratios.append(tokens / changed) 456 if not ratios: 457 return DEFAULT_TOKENS_PER_LINE, {"n": 0} 458 ratios.sort() 459 stats = { 460 "n": len(ratios), 461 "min": ratios[0], 462 "max": ratios[-1], 463 "median": statistics.median(ratios), 464 } 465 if len(ratios) >= 4: 466 q1, _, q3 = statistics.quantiles(ratios, n=4) 467 stats["p25"], stats["p75"] = q1, q3 468 else: 469 stats["p25"], stats["p75"] = ratios[0], ratios[-1] 470 return stats["median"], stats
Median tokens-per-changed-line across every ref commit with a measured Cost: trailer
carrying a token figure, paired with that commit's own diff size (diff_stats). The median
(not a least-squares fit) is deliberate: a big-context, small-diff commit can be two orders of
magnitude above a routine one (this repo's own spread is ~800-440,000 tokens/line), and a
maintainer already eyeballs a Cost: trailer the same way — ignore the extremes, trust the
middle. Returns (coefficient, stats); stats["n"] == 0 means DEFAULT_TOKENS_PER_LINE was used.
473def print_calibration(coeff, stats, stream=sys.stdout): 474 if stats["n"] == 0: 475 print( 476 f"cost-split --estimate: no calibration history found — falling back to the " 477 f"documented default of {coeff:.0f} tokens/changed-line", 478 file=stream, 479 ) 480 return 481 print( 482 f"cost-split --estimate: calibrated {coeff:.0f} tokens/changed-line from {stats['n']} " 483 f"origin/main commit(s) with a measured Cost: trailer " 484 f"(IQR {stats['p25']:.0f}-{stats['p75']:.0f}, range {stats['min']:.0f}-{stats['max']:.0f})", 485 file=stream, 486 )
489def category_multiplier(paths): 490 """A documented heuristic nudge, not a second regression — too few calibration commits per 491 category (docs-only, test-heavy, ...) in this repo's history to fit each separately without 492 overfitting on a handful of points.""" 493 if not paths: 494 return 1.0 495 docs_only = all(p.endswith(".md") or p.startswith("docs/") for p in paths) 496 if docs_only: 497 return 0.6 # prose from an outline is cheaper per line than code written from scratch 498 scala = any(p.endswith((".scala", ".sbt")) for p in paths) 499 tests = any("test" in p.lower() or "spec" in p.lower() for p in paths) 500 if scala and tests: 501 return 1.15 # red/green/refactor round-trips cost more than the final diff alone shows 502 return 1.0
A documented heuristic nudge, not a second regression — too few calibration commits per category (docs-only, test-heavy, ...) in this repo's history to fit each separately without overfitting on a handful of points.
510def estimate_usd(prices, tokens, model=DEFAULT_PRICE_MODEL): 511 """Prices the estimated tokens as if every one of them were a fresh input token (LiteLLM's 512 `input_cost_per_token`) — cache-read share cannot be estimated from a diff, so it is left out 513 entirely rather than guessed; this repo's own measured $/token (see calibrate()'s sibling 514 check in the self-test) runs about half of that rate, so the estimate leans conservative 515 rather than optimistic.""" 516 p = prices.get(model) 517 if not p: 518 return None 519 rate = p.get("input_cost_per_token") 520 if not rate: 521 return None 522 return tokens * rate
Prices the estimated tokens as if every one of them were a fresh input token (LiteLLM's
input_cost_per_token) — cache-read share cannot be estimated from a diff, so it is left out
entirely rather than guessed; this repo's own measured $/token (see calibrate()'s sibling
check in the self-test) runs about half of that rate, so the estimate leans conservative
rather than optimistic.
537def format_estimate_trailer(tokens, usd, changed_lines, today, stats): 538 """`stats` is calibrate()'s return. The point figure alone hides that this repo's own 539 tokens-per-line spread is ~5x (see calibrate()'s docstring) — scale calibrate()'s p25/p75 onto 540 this estimate's own token figure so the band is in the same units as the number next to it, 541 rather than inventing a new one. `stats["n"] == 0` (no calibration history) states that 542 plainly instead of printing a band it doesn't have.""" 543 usd_part = f"~${usd:.2f} est." if usd is not None else "$n/a est." 544 n = stats.get("n", 0) 545 if n and stats.get("median"): 546 lo = tokens * stats["p25"] / stats["median"] 547 hi = tokens * stats["p75"] / stats["median"] 548 band = f"IQR {_tokens_str(lo)}-{_tokens_str(hi)} tokens for this diff from {n} calibrated commits" 549 else: 550 band = "no calibration history — documented default, not a fit" 551 return ( 552 f"Cost: {usd_part} · {format_tokens(tokens)} " 553 f"(diff-size model, {changed_lines} lines, no session log, {band}) " 554 f"· scripts/cost-split.py --estimate {today}" 555 )
stats is calibrate()'s return. The point figure alone hides that this repo's own
tokens-per-line spread is ~5x (see calibrate()'s docstring) — scale calibrate()'s p25/p75 onto
this estimate's own token figure so the band is in the same units as the number next to it,
rather than inventing a new one. stats["n"] == 0 (no calibration history) states that
plainly instead of printing a band it doesn't have.
558def estimate_commit(sha, prices, coeff, cwd=None): 559 """(tokens, usd, changed_lines) for one commit's diff-size estimate, or None if the commit 560 has no diff at all (an empty commit, or a merge commit numstat can't attribute).""" 561 changed, files, paths = diff_stats(sha, cwd=cwd) 562 if changed <= 0: 563 return None 564 tokens = estimate_tokens_for_diff(changed, files, paths, coeff) 565 usd = estimate_usd(prices, tokens) 566 return tokens, usd, changed
(tokens, usd, changed_lines) for one commit's diff-size estimate, or None if the commit has no diff at all (an empty commit, or a merge commit numstat can't attribute).
575def self_test(): 576 # --- parse_token_count: every shape actually seen in this repo's own `Cost:` trailers --- 577 cases = [ 578 ("Cost: ~$8.98 · 8.1M tokens, 97% cache reads (claude-fable-5-1) · ...", 8_100_000), 579 ( 580 "Cost: ~289k tokens (forked subagent, claude-fable-5-1; 40 tool calls, 8 min) · ...", 581 289_000, 582 ), 583 ( 584 "Cost: shared bucket $7.03 · 10,517,273 tokens (claude-fable-5-1) — not split further", 585 None, 586 ), 587 ("Cost: ~unmeasured in-agent · fill from just claude-cost before PR", None), 588 ( 589 "Cost: ~$0 · 0 tokens booked (claude-fable-5-1) · ...", 590 0.0, 591 ), # parses, but falsy: calibrate() skips it too 592 ("Cost: ~$0.1", None), # no token figure at all 593 ("Cost: shared bucket $3.18 · 4,008,909 tokens (claude-fable-5-1) — ...", None), 594 ] 595 for line, expected in cases: 596 got = parse_token_count(line) 597 assert got == expected, f"{line!r} -> {got}, expected {expected}" 598 599 # --- estimate_tokens_for_diff / category_multiplier --- 600 coeff = 6000 601 plain = estimate_tokens_for_diff(100, 2, ["core/src/main/Foo.scala"], coeff) 602 assert plain == round(coeff * (100 + FILE_OVERHEAD_LINES * 2)), plain 603 docs = estimate_tokens_for_diff(100, 1, ["docs/FOO.md"], coeff) 604 assert docs < plain, (docs, plain) # docs-only is cheaper per line 605 tested = estimate_tokens_for_diff( 606 100, 2, ["core/src/main/Foo.scala", "core/src/test/FooSpec.scala"], coeff 607 ) 608 assert tested > plain, (tested, plain) # scala + tests costs more than scala alone 609 assert estimate_tokens_for_diff(0, 0, [], coeff) == 0 610 611 # --- estimate_usd: LiteLLM-shaped fixture, no network --- 612 prices = {"claude-sonnet-5": {"input_cost_per_token": 2e-06, "output_cost_per_token": 1e-05}} 613 usd = estimate_usd(prices, 1_000_000, model="claude-sonnet-5") 614 assert usd == 2.0, usd 615 assert estimate_usd(prices, 1_000_000, model="no-such-model") is None 616 assert estimate_usd({}, 1_000_000) is None 617 618 # --- format_estimate_trailer: always carries "est." on both figures, never bare, and now 619 # states its own band (issue #433 task 2: a point figure alone hides calibrate()'s ~5x IQR). --- 620 fitted_stats = { 621 "n": 5, 622 "median": 6000.0, 623 "p25": 3000.0, 624 "p75": 12000.0, 625 "min": 800.0, 626 "max": 40000.0, 627 } 628 trailer = format_estimate_trailer(912_345, 1.20, 210, "2026-09-05", fitted_stats) 629 assert trailer.startswith("Cost: ~$1.20 est. ·"), trailer 630 assert "tokens est." in trailer, trailer 631 assert "210 lines, no session log" in trailer, trailer 632 assert trailer.endswith("--estimate 2026-09-05"), trailer 633 634 # Parse the band back out of the rendered string — not calibrate()'s own p25/p75 formula — so 635 # this fails if the band stops bracketing the point figure or stops matching stats["n"], not 636 # only if the arithmetic happens to change. 637 def _parse_band(s): 638 m = re.search( 639 r"IQR ([\d.]+)(k|M)?-([\d.]+)(k|M)?\s*tokens for this diff from (\d+) calibrated", s 640 ) 641 assert m, s 642 scale = {None: 1, "k": 1e3, "M": 1e6} 643 return ( 644 float(m.group(1)) * scale[m.group(2)], 645 float(m.group(3)) * scale[m.group(4)], 646 int(m.group(5)), 647 ) 648 649 lo, hi, shown_n = _parse_band(trailer) 650 assert lo < hi, (lo, hi) 651 assert lo <= 912_345 <= hi, (lo, 912_345, hi) # the point figure sits inside its own band 652 assert shown_n == fitted_stats["n"], (shown_n, fitted_stats["n"]) 653 654 no_price = format_estimate_trailer(500, None, 3, "2026-09-05", fitted_stats) 655 assert no_price.startswith("Cost: $n/a est. ·"), no_price 656 657 # n == 0 (fresh checkout, no calibration history): well-formed, but reads as a documented 658 # default, never a fit — and differently from the fitted case above. 659 fresh = format_estimate_trailer(500, None, 3, "2026-09-05", {"n": 0}) 660 assert fresh.startswith("Cost: $n/a est. ·"), fresh 661 assert "tokens est." in fresh, fresh 662 assert "3 lines, no session log" in fresh, fresh 663 assert fresh.endswith("--estimate 2026-09-05"), fresh 664 assert "documented default" in fresh, fresh 665 assert "documented default" not in trailer, trailer 666 assert "IQR" not in fresh, fresh 667 668 # --- calibrate() + diff_stats() + estimate_commit(): a synthetic repo, so the fit is a known 669 # number instead of depending on this checkout's ever-growing real history --- 670 with tempfile.TemporaryDirectory() as tmp: 671 repo = Path(tmp) 672 sh("git", "init", "-q", "-b", "main", cwd=repo) 673 sh("git", "config", "user.email", "test@example.com", cwd=repo) 674 sh("git", "config", "user.name", "Test", cwd=repo) 675 676 def commit(fname, content, message): 677 (repo / fname).write_text(content) 678 sh("git", "add", fname, cwd=repo) 679 sh("git", "commit", "-q", "-m", message, cwd=repo) 680 return sh("git", "rev-parse", "HEAD", cwd=repo).strip() 681 682 # 10 changed lines (all additions), 1,000,000 tokens measured -> ratio 100,000/line. 683 commit( 684 "a.txt", 685 "\n".join(f"line {i}" for i in range(10)) + "\n", 686 "first\n\nCost: ~1.0M tokens\n", 687 ) 688 # 100 changed lines, 5,000,000 tokens measured -> ratio 50,000/line. 689 commit( 690 "b.txt", 691 "\n".join(f"line {i}" for i in range(100)) + "\n", 692 "second\n\nCost: ~5.0M tokens\n", 693 ) 694 # A "shared bucket" trailer must not enter the calibration set even though it parses. 695 commit( 696 "c.txt", "x\n", "third\n\nCost: shared bucket $1 · 9.0M tokens — not split further\n" 697 ) 698 # An un-costed commit — the one we'll estimate against the calibration from the other two. 699 target = commit( 700 "d.txt", "\n".join(f"line {i}" for i in range(20)) + "\n", "fourth: no trailer" 701 ) 702 703 coeff, stats = calibrate(ref="main", cwd=repo) 704 assert stats["n"] == 2, stats # the shared-bucket commit must be excluded 705 assert coeff == statistics.median([100_000, 50_000]), coeff 706 707 prices2 = {DEFAULT_PRICE_MODEL: {"input_cost_per_token": 2e-06}} 708 tokens, usd, changed = estimate_commit(target, prices2, coeff, cwd=repo) 709 assert changed == 20, changed 710 expected_tokens = estimate_tokens_for_diff(20, 1, ["d.txt"], coeff) 711 assert tokens == expected_tokens, (tokens, expected_tokens) 712 assert ( 713 usd == round(expected_tokens * 2e-06, 10) or abs(usd - expected_tokens * 2e-06) < 1e-9 714 ) 715 716 # A repo with zero calibratable commits falls back to the documented default. 717 with tempfile.TemporaryDirectory() as tmp2: 718 empty_repo = Path(tmp2) 719 sh("git", "init", "-q", "-b", "main", cwd=empty_repo) 720 sh("git", "config", "user.email", "test@example.com", cwd=empty_repo) 721 sh("git", "config", "user.name", "Test", cwd=empty_repo) 722 (empty_repo / "x.txt").write_text("x\n") 723 sh("git", "add", "x.txt", cwd=empty_repo) 724 sh("git", "commit", "-q", "-m", "no cost trailer here", cwd=empty_repo) 725 fallback_coeff, fallback_stats = calibrate(ref="main", cwd=empty_repo) 726 assert fallback_stats["n"] == 0 727 assert fallback_coeff == DEFAULT_TOKENS_PER_LINE 728 729 # --- regression: sh(cwd=...) must not leak into a GIT_DIR the environment already points at. 730 # A git hook (pre-push, pre-commit) sets GIT_DIR (and friends) for its own repo; before this 731 # was fixed, the calibration commits above landed for real on whatever repo triggered the 732 # hook instead of staying inside their own tempdir. Reproduced here with a decoy "outer" repo 733 # standing in for the real one, GIT_DIR pointed at it, and an inner tempdir repo built the same 734 # way self_test() builds its calibration fixtures. --- 735 with tempfile.TemporaryDirectory() as outer_tmp: 736 outer = Path(outer_tmp) 737 sh("git", "init", "-q", "-b", "main", cwd=outer) 738 sh("git", "config", "user.email", "outer@example.com", cwd=outer) 739 sh("git", "config", "user.name", "Outer", cwd=outer) 740 (outer / "seed.txt").write_text("seed\n") 741 sh("git", "add", "seed.txt", cwd=outer) 742 sh("git", "commit", "-q", "-m", "seed", cwd=outer) 743 outer_head_before = sh("git", "rev-parse", "HEAD", cwd=outer).strip() 744 745 saved_git_dir = os.environ.get("GIT_DIR") 746 try: 747 os.environ["GIT_DIR"] = str(outer / ".git") 748 with tempfile.TemporaryDirectory() as inner_tmp: 749 inner = Path(inner_tmp) 750 sh("git", "init", "-q", "-b", "main", cwd=inner) 751 sh("git", "config", "user.email", "test@example.com", cwd=inner) 752 sh("git", "config", "user.name", "Test", cwd=inner) 753 (inner / "x.txt").write_text("x\n") 754 sh("git", "add", "x.txt", cwd=inner) 755 sh("git", "commit", "-q", "-m", "should stay inside tempdir", cwd=inner) 756 finally: 757 if saved_git_dir is None: 758 os.environ.pop("GIT_DIR", None) 759 else: 760 os.environ["GIT_DIR"] = saved_git_dir 761 762 outer_head_after = sh("git", "rev-parse", "HEAD", cwd=outer).strip() 763 assert outer_head_after == outer_head_before, ( 764 "sh(cwd=...) leaked into the ambient GIT_DIR instead of staying in cwd" 765 ) 766 assert not (outer / "x.txt").exists(), "inner repo's file leaked into the outer repo" 767 768 # --- messages(): a subagent transcript's own `gitBranch` is the *parent's*, not its own — 769 # `cwd` says which worktree it actually ran in (issue #433, confirmed live 2026-09-28). --- 770 with tempfile.TemporaryDirectory() as wt_tmp: 771 mrepo = Path(wt_tmp) / "main-checkout" 772 mrepo.mkdir() 773 sh("git", "init", "-q", "-b", "main", cwd=mrepo) 774 sh("git", "config", "user.email", "test@example.com", cwd=mrepo) 775 sh("git", "config", "user.name", "Test", cwd=mrepo) 776 (mrepo / "a.txt").write_text("a\n") 777 sh("git", "add", "a.txt", cwd=mrepo) 778 sh("git", "commit", "-q", "-m", "initial", cwd=mrepo) 779 780 worktree_dir = mrepo / ".tmp" / "wt-feature" 781 sh("git", "worktree", "add", "-q", "-b", "feature-x", str(worktree_dir), cwd=mrepo) 782 (worktree_dir / "b.txt").write_text("b\n") 783 sh("git", "add", "b.txt", cwd=worktree_dir) 784 sh("git", "commit", "-q", "-m", "the commit under test", cwd=worktree_dir) 785 target_sha = sh("git", "rev-parse", "HEAD", cwd=worktree_dir).strip() 786 target_when = dt.datetime.fromisoformat( 787 sh("git", "log", "-1", "--format=%aI", cwd=worktree_dir).strip() 788 ) 789 790 pdir = Path(wt_tmp) / "project-logs" 791 sid = "ses-parent" 792 (pdir / sid / "subagents").mkdir(parents=True) 793 794 def jsonl_line(**fields): 795 return json.dumps(fields) + "\n" 796 797 usage_small = {"input_tokens": 100, "output_tokens": 20} 798 usage_big = {"input_tokens": 300_000, "output_tokens": 5_000} 799 usage_tiny = {"input_tokens": 7, "output_tokens": 3} 800 801 (pdir / f"{sid}.jsonl").write_text( 802 jsonl_line( 803 message={"id": "msg-parent", "usage": usage_small, "model": "claude-sonnet-5"}, 804 timestamp="2000-01-01T10:00:00Z", 805 gitBranch="main", 806 cwd="/no/such/worktree", 807 ) 808 + jsonl_line( 809 # cwd IS the main checkout; gitBranch names a branch other than its current one 810 # ("main" here). gitBranch is correct per-message history and must win. 811 message={ 812 "id": "msg-main-checkout", 813 "usage": usage_tiny, 814 "model": "claude-sonnet-5", 815 }, 816 timestamp="2000-01-01T10:02:00Z", 817 gitBranch="feature-x", 818 cwd=str(mrepo), 819 ) 820 ) 821 sub_a = jsonl_line( 822 message={"id": "msg-sub-a", "usage": usage_big, "model": "claude-sonnet-5"}, 823 timestamp="2000-01-01T10:05:00Z", 824 gitBranch="main", 825 cwd=str(worktree_dir), 826 isSidechain=True, 827 ) 828 sub_b = jsonl_line( 829 message={"id": "msg-sub-b", "usage": usage_small, "model": "claude-sonnet-5"}, 830 timestamp="2000-01-01T10:06:00Z", 831 gitBranch="main", 832 cwd="/no/such/worktree/either", 833 isSidechain=True, 834 ) 835 sub_c = jsonl_line( 836 # Same as msg-main-checkout, but sidechain: must also keep gitBranch, not fall to "" 837 # just because it's a sidechain. 838 message={"id": "msg-sub-c", "usage": usage_tiny, "model": "claude-sonnet-5"}, 839 timestamp="2000-01-01T10:07:00Z", 840 gitBranch="feature-x", 841 cwd=str(mrepo), 842 isSidechain=True, 843 ) 844 sub_d = jsonl_line( 845 # cwd is a subdirectory of the main checkout, not its root — real logs have these 846 # (e.g. docs/img/logo). Still the main checkout's own tree, so still gitBranch. 847 message={"id": "msg-sub-d", "usage": usage_tiny, "model": "claude-sonnet-5"}, 848 timestamp="2000-01-01T10:08:00Z", 849 gitBranch="feature-x", 850 cwd=str(mrepo / "docs" / "img" / "logo"), 851 isSidechain=True, 852 ) 853 (pdir / sid / "subagents" / "agent-x.jsonl").write_text(sub_a + sub_b + sub_c + sub_d) 854 855 got = messages(pdir, None, mrepo) 856 assert len(got) == 6, got 857 parent_branch = got[0][4] 858 main_checkout_branch = got[1][4] 859 sub_a_branch = got[2][4] 860 sub_b_branch = got[3][4] 861 sub_c_branch = got[4][4] 862 sub_d_branch = got[5][4] 863 864 # Non-sidechain, cwd matches no known worktree: existing behaviour must not regress. 865 assert parent_branch == "main", parent_branch 866 867 # Non-sidechain, cwd IS the main checkout: keep gitBranch ("feature-x"), not 868 # worktree_branches()'s "current branch" entry for it ("main") — the main checkout's 869 # current branch says nothing about what branch it was on when a given message was made. 870 assert main_checkout_branch == "feature-x", main_checkout_branch 871 872 # Sidechain, cwd resolves to the feature-x worktree: use the worktree's branch, not the 873 # inherited (misleading) parent gitBranch. 874 assert sub_a_branch == "feature-x", sub_a_branch 875 876 # Sidechain, cwd matches no known worktree: "" (time-only rule), not the parent's branch — 877 # a wrong branch would be discarded outright, "" is only less precise. 878 assert sub_b_branch == "", sub_b_branch 879 880 # Sidechain, cwd IS the main checkout: keep gitBranch, same as the non-sidechain case — 881 # a parent that was itself running in the main checkout inherits correct history. 882 assert sub_c_branch == "feature-x", sub_c_branch 883 884 # Sidechain, cwd is a subdirectory of the main checkout (not its root): still gitBranch, 885 # not "" — real logs have these and a naive exact match would send them to the wrong arm. 886 assert sub_d_branch == "feature-x", sub_d_branch 887 888 # End to end (issue #433): run the fixture through `attribute()` itself — the same 889 # function main() calls — not a re-implementation of its `ts <= when` pairing or 890 # summation. `branches_containing`/`branches_with_equivalent` read the ambient cwd (no 891 # `cwd=` param, same as `commits()`), so chdir into the synthetic repo for the call. 892 cs = [(target_when, "feature-x", target_sha, "the commit under test")] 893 saved_cwd = os.getcwd() 894 try: 895 os.chdir(worktree_dir) 896 buckets = attribute(got, cs, {}) 897 finally: 898 os.chdir(saved_cwd) 899 900 # sub_a lands on target_sha by branch (feature-x, resolved from cwd); msg-main-checkout, 901 # sub_c and sub_d land there too, by branch (their own gitBranch, "feature-x" — cwd'd in 902 # or under the main checkout); sub_b lands there by the time-only rule (its cwd matched no 903 # worktree, so its branch is ""), since cs holds only this one commit. parent (branch 904 # "main", not in {"feature-x"}) is discarded entirely — it must not appear anywhere, not 905 # even "uncommitted". 906 bucket = buckets[target_sha] 907 actual_total = bucket["in"] + bucket["out"] + bucket["cache_w"] + bucket["cache_r"] 908 tiny_total = usage_tiny["input_tokens"] + usage_tiny["output_tokens"] 909 expected_total = ( 910 usage_big["input_tokens"] 911 + usage_big["output_tokens"] 912 + usage_small["input_tokens"] 913 + usage_small["output_tokens"] 914 + 3 * tiny_total # msg-main-checkout + sub_c + sub_d, all usage_tiny 915 ) 916 # Exact match, no tolerance: every field summed here is an int token count, so there is 917 # no rounding for a tolerance to absorb. 918 assert actual_total == expected_total, ( 919 f"attribute() must land the subagent-shaped fixture's measured total on " 920 f"{target_sha[:7]} exactly (int summation, 0 tolerance): got {actual_total}, " 921 f"expected {expected_total} — a regression in the ts<=when pairing or the " 922 "summation would silently change this" 923 ) 924 uncommitted_total = sum( 925 buckets["uncommitted"][k] for k in ("in", "out", "cache_w", "cache_r") 926 ) 927 assert uncommitted_total == 0, ( 928 f"the parent message's branch ({parent_branch!r}) doesn't carry the commit under " 929 f"test and must be discarded outright, not misrouted to uncommitted: {uncommitted_total}" 930 ) 931 # attribute() just populated this module-global cache from the synthetic (about-to-vanish) 932 # repo above; reset it so a test block added later doesn't inherit it. 933 global _equivalents 934 _equivalents = None 935 936 # --- messages_opencode(): a synthetic DB built against the real schema (MIP-0013 task 2, 937 # confirmed live 2026-09-07 against opencode 1.18.25) — table/column names, tokens.{...} JSON 938 # shape, session.directory filtering. No real ~/.local/share/opencode touched. --- 939 with tempfile.TemporaryDirectory() as oc_tmp: 940 oc_root = Path(oc_tmp) 941 data_dir = oc_root / "data" / "opencode" 942 data_dir.mkdir(parents=True) 943 saved_xdg = os.environ.get("XDG_DATA_HOME") 944 os.environ["XDG_DATA_HOME"] = str(oc_root / "data") 945 try: 946 db_path = data_dir / "opencode-stable.db" 947 con = sqlite3.connect(db_path) 948 con.execute("create table session (id text, directory text)") 949 con.execute("create table message (id text, session_id text, data text)") 950 here = str(oc_root / "repo") 951 elsewhere = str(oc_root / "other-repo") 952 con.execute("insert into session values (?, ?)", ("ses_here", here)) 953 con.execute("insert into session values (?, ?)", ("ses_elsewhere", elsewhere)) 954 user_msg = json.dumps({"role": "user"}) 955 asst_msg = json.dumps( 956 { 957 "role": "assistant", 958 "modelID": "llama3.2", 959 "providerID": "ollama", 960 "tokens": {"input": 4096, "output": 8, "cache": {"read": 100, "write": 0}}, 961 "time": {"created": 1788756841410}, 962 } 963 ) 964 elsewhere_msg = json.dumps( 965 { 966 "role": "assistant", 967 "modelID": "llama3.2", 968 "providerID": "ollama", 969 "tokens": {"input": 1, "output": 1, "cache": {"read": 0, "write": 0}}, 970 "time": {"created": 1788756841410}, 971 } 972 ) 973 con.execute("insert into message values (?, ?, ?)", ("msg_user", "ses_here", user_msg)) 974 con.execute("insert into message values (?, ?, ?)", ("msg_asst", "ses_here", asst_msg)) 975 con.execute( 976 "insert into message values (?, ?, ?)", 977 ("msg_other_repo", "ses_elsewhere", elsewhere_msg), 978 ) 979 con.commit() 980 con.close() 981 982 found = messages_opencode(Path(here), None) 983 assert len(found) == 1, ( 984 found 985 ) # the user-role row and the other repo's row are both dropped 986 ts, sid, model, usage, branch = found[0] 987 assert sid == "ses_here", sid 988 assert model == "ollama/llama3.2", model 989 assert usage == { 990 "input_tokens": 4096, 991 "output_tokens": 8, 992 "cache_creation_input_tokens": 0, 993 "cache_read_input_tokens": 100, 994 }, usage 995 assert branch == "", branch # OpenCode carries no gitBranch — main()'s time-only rule 996 997 assert messages_opencode(Path(str(oc_root / "no-such-repo")), None) == [] 998 finally: 999 if saved_xdg is None: 1000 os.environ.pop("XDG_DATA_HOME", None) 1001 else: 1002 os.environ["XDG_DATA_HOME"] = saved_xdg 1003 1004 assert opencode_db_path() is None or opencode_db_path().is_file() 1005 1006 print( 1007 f"cost-split self-test: ok (parser {len(cases)} cases, synthetic calibration " 1008 f"median {coeff:.0f} tokens/line from {stats['n']} commits)" 1009 ) 1010 return 0
1016def attribute(msgs, cs, prices): 1017 """[(ts, session, model, usage, branch)], [(when, branch, sha, subject)] -> sha/"uncommitted" 1018 -> totals. A message counts toward a branch only if it was made *on* that branch (its 1019 `mbranch`), then falls into the first commit whose author date is after it. Without the 1020 branch test, every session on the machine that ran before a branch's first commit — other 1021 agents, other features — landed on that commit (seen: 335M tokens on a 500-line commit). A 1022 message with no branch (older logs, or a sidechain whose `cwd` matched no worktree) keeps the 1023 time-only rule. 1024 1025 "On that branch" means: the message's branch contains the commit — a commit authored on 1026 feat/a and now priced from feat/b (stacked on a) is still paid for by the messages made on 1027 feat/a. Computed once per commit from `git branch -a --contains`. 1028 1029 Not pure despite the signature: `branches_containing`/`branches_with_equivalent` take no 1030 `cwd` and read the process's actual working directory, so the caller must already be in (or 1031 have chdir'd into) the repo `cs`'s commits belong to. 1032 """ 1033 buckets = defaultdict( 1034 lambda: { 1035 "in": 0, 1036 "out": 0, 1037 "cache_w": 0, 1038 "cache_r": 0, 1039 "usd": 0.0, 1040 "models": set(), 1041 "sessions": set(), 1042 "priced": True, 1043 } 1044 ) 1045 contains = { 1046 sha: branches_containing(sha) | branches_with_equivalent(sha) for _w, _b, sha, _s in cs 1047 } 1048 ours = {c[1] for c in cs}.union(*contains.values()) if cs else set() 1049 for ts, session, model, u, mbranch in msgs: 1050 if mbranch and mbranch not in ours: 1051 continue 1052 target = "uncommitted" 1053 for when, _cbranch, sha, _s in cs: 1054 if mbranch and mbranch not in contains[sha]: 1055 continue 1056 if ts <= when: 1057 target = sha 1058 break 1059 b = buckets[target] 1060 b["in"] += u.get("input_tokens", 0) 1061 b["out"] += u.get("output_tokens", 0) 1062 b["cache_w"] += u.get("cache_creation_input_tokens", 0) 1063 b["cache_r"] += u.get("cache_read_input_tokens", 0) 1064 b["models"].add(model) 1065 b["sessions"].add(session[:8]) 1066 usd = price(prices, model, u) 1067 if usd is None: 1068 b["priced"] = False 1069 else: 1070 b["usd"] += usd 1071 return buckets
[(ts, session, model, usage, branch)], [(when, branch, sha, subject)] -> sha/"uncommitted"
-> totals. A message counts toward a branch only if it was made on that branch (its
mbranch), then falls into the first commit whose author date is after it. Without the
branch test, every session on the machine that ran before a branch's first commit — other
agents, other features — landed on that commit (seen: 335M tokens on a 500-line commit). A
message with no branch (older logs, or a sidechain whose cwd matched no worktree) keeps the
time-only rule.
"On that branch" means: the message's branch contains the commit — a commit authored on
feat/a and now priced from feat/b (stacked on a) is still paid for by the messages made on
feat/a. Computed once per commit from git branch -a --contains.
Not pure despite the signature: branches_containing/branches_with_equivalent take no
cwd and read the process's actual working directory, so the caller must already be in (or
have chdir'd into) the repo cs's commits belong to.
1074def build_rows(cs, buckets, order, prices, do_estimate, coeff): 1075 rows = [] 1076 for key in order: 1077 b = buckets.get(key) 1078 meta = next( 1079 ((br, subj) for _w, br, sha, subj in cs if sha == key), ("", "uncommitted (so far)") 1080 ) 1081 if not b: 1082 if key == "uncommitted" or not do_estimate: 1083 continue 1084 # No cwd override: `key` is a full sha (unambiguous from any worktree of this repo), 1085 # but if it were ever a symbolic ref like HEAD, resolving it against repo_root()'s 1086 # main-checkout path — right for session-log lookup, wrong here — would silently 1087 # answer for the wrong branch when run from a worktree. 1088 est = estimate_commit(key, prices, coeff) 1089 if est is None: 1090 continue 1091 tokens, usd, changed = est 1092 rows.append( 1093 { 1094 "commit": key[:7], 1095 "branch": meta[0], 1096 "subject": meta[1], 1097 "input": 0, 1098 "output": 0, 1099 "cache_write": 0, 1100 "cache_read": 0, 1101 "tokens": tokens, 1102 "usd": round(usd, 2) if usd is not None else None, 1103 "models": ["diff-size estimate"], 1104 "sessions": [], 1105 "estimated": True, 1106 "changed_lines": changed, 1107 } 1108 ) 1109 continue 1110 rows.append( 1111 { 1112 "commit": key[:7] if key != "uncommitted" else "-", 1113 "branch": meta[0], 1114 "subject": meta[1], 1115 "input": b["in"], 1116 "output": b["out"], 1117 "cache_write": b["cache_w"], 1118 "cache_read": b["cache_r"], 1119 "tokens": b["in"] + b["out"] + b["cache_w"] + b["cache_r"], 1120 "usd": round(b["usd"], 2) if b["priced"] else None, 1121 "models": sorted(b["models"]), 1122 "sessions": sorted(b["sessions"]), 1123 "estimated": False, 1124 } 1125 ) 1126 return rows
1129def main(): 1130 ap = argparse.ArgumentParser( 1131 description=__doc__.split("\n\n")[0], formatter_class=argparse.RawDescriptionHelpFormatter 1132 ) 1133 ap.add_argument( 1134 "stack", nargs="?", help="mip-NNNN: every branch of the stack (default: current branch)" 1135 ) 1136 ap.add_argument("--session", help="session id prefix (default: every session of this project)") 1137 ap.add_argument("--json", action="store_true") 1138 ap.add_argument( 1139 "--estimate", 1140 action="store_true", 1141 help="fill in commits with no logged usage from a diff-size estimate, clearly labelled", 1142 ) 1143 ap.add_argument( 1144 "--estimate-commit", 1145 metavar="SHA", 1146 help="print just the estimated Cost: trailer for one commit — no session logs needed", 1147 ) 1148 ap.add_argument( 1149 "--verbose", action="store_true", help="with --estimate, print the calibration fit" 1150 ) 1151 ap.add_argument( 1152 "--harness", 1153 choices=["claude", "opencode", "all"], 1154 default="all", 1155 help="which agent harness's session logs to read (MIP-0013 task 2; default: both)", 1156 ) 1157 ap.add_argument("--self-test", action="store_true") 1158 a = ap.parse_args() 1159 1160 if a.self_test: 1161 sys.exit(self_test()) 1162 1163 root = repo_root() 1164 1165 if a.estimate_commit: 1166 # No cwd override here either: calibrate()'s `origin/main` and estimate_commit()'s sha 1167 # argument both mean the same thing from any worktree, but running them pinned to 1168 # repo_root() would resolve a symbolic ref like `HEAD` against the *main* checkout's 1169 # branch, not the one this command was actually run from. 1170 coeff, stats = calibrate() 1171 if a.verbose: 1172 print_calibration(coeff, stats) 1173 prices = load_prices(root) 1174 est = estimate_commit(a.estimate_commit, prices, coeff) 1175 if est is None: 1176 sys.exit( 1177 f"cost-split --estimate-commit: {a.estimate_commit} has no diff to estimate from" 1178 ) 1179 tokens, usd, changed = est 1180 print(format_estimate_trailer(tokens, usd, changed, dt.date.today().isoformat(), stats)) 1181 return 1182 1183 msgs = [] 1184 if a.harness in ("claude", "all"): 1185 pdir = project_dir(root) 1186 if pdir.is_dir(): 1187 msgs += messages(pdir, a.session, root) 1188 elif a.harness == "claude": 1189 sys.exit(f"no session logs at {pdir}") 1190 if a.harness in ("opencode", "all"): 1191 msgs += messages_opencode(root, a.session) 1192 msgs.sort(key=lambda x: x[0]) 1193 cs = commits(root, a.stack.lower() if a.stack else None) 1194 if not cs: 1195 sys.exit("no commits ahead of origin/main on the selected branch(es)") 1196 prices = load_prices(root) 1197 1198 coeff, calib_stats = (None, None) 1199 if a.estimate: 1200 coeff, calib_stats = calibrate() 1201 1202 order = [c[2] for c in cs] + ["uncommitted"] 1203 buckets = attribute(msgs, cs, prices) 1204 1205 rows = build_rows(cs, buckets, order, prices, a.estimate, coeff) 1206 1207 if a.json: 1208 print(json.dumps(rows, indent=1)) 1209 return 1210 1211 if a.estimate and a.verbose: 1212 print_calibration(coeff, calib_stats) 1213 print() 1214 1215 today = dt.date.today().isoformat() 1216 print( 1217 f"{'commit':8} {'usd':>9} {'tokens':>11} {'in':>7} {'out':>7} {'cache_w':>8} {'cache_r':>9} branch / subject" 1218 ) 1219 any_estimated = False 1220 for r in rows: 1221 mark = "*" if r.get("estimated") else "" 1222 any_estimated = any_estimated or bool(mark) 1223 usd = f"${r['usd']:.2f}{mark}" if r["usd"] is not None else f"n/a{mark}" 1224 tokens = f"{r['tokens']:,}{mark}" 1225 print( 1226 f"{r['commit']:8} {usd:>9} {tokens:>11} {r['input']:>7,} {r['output']:>7,} {r['cache_write']:>8,} {r['cache_read']:>9,} {r['branch']} {r['subject'][:60]}" 1227 ) 1228 if any_estimated: 1229 print("* diff-size estimate, no session log — scripts/cost-split.py --estimate") 1230 1231 per_branch = defaultdict(lambda: [0.0, 0, 0, True, set()]) 1232 for r in rows: 1233 if r["commit"] == "-": 1234 continue 1235 pb = per_branch[r["branch"]] 1236 pb[0] += r["usd"] or 0 1237 pb[1] += r["tokens"] 1238 pb[2] += r["cache_read"] 1239 pb[3] = pb[3] and r["usd"] is not None 1240 pb[4].update(r["models"]) 1241 print("\nCost: trailers per branch (paste into the commit / PR):") 1242 for br, (usd, tokens, cached, priced, models) in per_branch.items(): 1243 dollars = f"~${usd:.2f}" if priced else "$n/a" 1244 share = f"{100 * cached / tokens:.0f}% cache reads" if tokens else "no tokens" 1245 print( 1246 f" {br}: Cost: {dollars} · {tokens / 1e6:.1f}M tokens, {share} ({', '.join(sorted(models))}) · split by branch and author date, scripts/cost-split.py {today}" 1247 ) 1248 1249 estimated_rows = [r for r in rows if r.get("estimated")] 1250 if estimated_rows: 1251 print("\nCost: trailers for commits with no session log (estimated — paste per commit):") 1252 for r in estimated_rows: 1253 trailer = format_estimate_trailer( 1254 r["tokens"], r["usd"], r["changed_lines"], today, calib_stats 1255 ) 1256 print(f" {r['commit']}: {trailer}") 1257 1258 total = sum(r["usd"] or 0 for r in rows) 1259 print(f"\nsession total ${total:.2f} for {len(rows)} buckets (list prices, LiteLLM table)")