repo_stats

repo_stats — the README's CI-health, LOC and Python-coverage badges, as shields.io endpoint JSON.

scripts/repo_stats.py write --out-dir stats         --repo marola-dev/marola --run-id 123 --exclude-job repo-stats
scripts/repo_stats.py write --out-dir stats     # LOC + Python coverage (no --run-id: no CI badge)
scripts/repo_stats.py write --out-dir stats --no-python-coverage   # skip coverage.py
scripts/repo_stats.py python-coverage           # just print the measured % (no files written)
scripts/repo_stats.py --self-test               # shaping + counting rules (just quality-other)

Four files, the same shields.io "endpoint" shape ci.yml already writes for the Scala coverage ({"schemaVersion": 1, "label": ..., "message": ..., "color": ...}), published to the orphan site-data branch and copied into the site by site.yml, where the README reads them live:

ci.json               "ci steps"        19/20 green
scala-loc.json        "scala"           6,865 LOC
python-loc.json       "python"          2,877 LOC
python-coverage.json  "py-cov" 73%

CI health is step-level, not job-level: ci.yml has three jobs but ~20 named steps, and "build-test passed" hides which of them actually ran. Only steps that ran count — a step whose conclusion is skipped (most of them are gated on needs.changes.outputs.*) or still null (this job's own later steps) is left out of both numerator and denominator, so a docs-only push reports 8/8 rather than a misleading 8/20. --exclude-job drops the reporting job itself, whose steps are by definition still running while it asks.

LOC is cloc (from nix develop .#lint, locally and in ci.yml), counted over the four Scala modules and the Python trees, code lines only — blanks and comments excluded by cloc, and target/, __pycache__/, virtualenvs and node_modules/ excluded by path.

Python coverage is measured, never estimated — but read the label narrowly. marola has no pytest suite; every scripts/**/*.py is tested by its own --self-test flag, the list just quality-other runs. So this badge is statement coverage of scripts/ while those self-tests run, and nothing more: a branch a self-test never bothers to call is uncovered by construction, which is why the figure sits in the 70s rather than the 90s. dspy/ and finetune/ are out of scope — no --self-test entry point, and importing them needs torch/DSPy — so they are neither numerator nor denominator, while a scripts/*.py that grows without a self-test does count (at 0%), which is the point. Mechanically: one coverage run --parallel-mode per self-test into a temp data file, then coverage combine + coverage json.

Standard library only; cloc and coverage are the external tools, and only write needs them.

  1#!/usr/bin/env python3
  2"""repo_stats — the README's CI-health, LOC and Python-coverage badges, as shields.io endpoint JSON.
  3
  4    scripts/repo_stats.py write --out-dir stats \
  5        --repo marola-dev/marola --run-id 123 --exclude-job repo-stats
  6    scripts/repo_stats.py write --out-dir stats     # LOC + Python coverage (no --run-id: no CI badge)
  7    scripts/repo_stats.py write --out-dir stats --no-python-coverage   # skip coverage.py
  8    scripts/repo_stats.py python-coverage           # just print the measured % (no files written)
  9    scripts/repo_stats.py --self-test               # shaping + counting rules (just quality-other)
 10
 11Four files, the same shields.io "endpoint" shape ci.yml already writes for the Scala coverage
 12(`{"schemaVersion": 1, "label": ..., "message": ..., "color": ...}`), published to the orphan
 13`site-data` branch and copied into the site by site.yml, where the README reads them live:
 14
 15    ci.json               "ci steps"        19/20 green
 16    scala-loc.json        "scala"           6,865 LOC
 17    python-loc.json       "python"          2,877 LOC
 18    python-coverage.json  "py-cov" 73%
 19
 20CI health is *step*-level, not job-level: ci.yml has three jobs but ~20 named steps, and
 21"build-test passed" hides which of them actually ran. Only steps that ran count — a step whose
 22`conclusion` is `skipped` (most of them are gated on `needs.changes.outputs.*`) or still `null`
 23(this job's own later steps) is left out of both numerator and denominator, so a docs-only push
 24reports 8/8 rather than a misleading 8/20. `--exclude-job` drops the reporting job itself, whose
 25steps are by definition still running while it asks.
 26
 27LOC is `cloc` (from `nix develop .#lint`, locally and in ci.yml), counted over the four
 28Scala modules and the Python trees, code lines only — blanks and comments excluded by cloc, and
 29`target/`, `__pycache__/`, virtualenvs and `node_modules/` excluded by path.
 30
 31Python coverage is measured, never estimated — but read the label narrowly. marola has no pytest
 32suite; every `scripts/**/*.py` is tested by its own `--self-test` flag, the list `just
 33quality-other` runs. So this badge is *statement coverage of `scripts/` while those self-tests
 34run*, and nothing more: a branch a self-test never bothers to call is uncovered by construction,
 35which is why the figure sits in the 70s rather than the 90s. `dspy/` and `finetune/` are out of
 36scope — no `--self-test` entry point, and importing them needs torch/DSPy — so they are neither
 37numerator nor denominator, while a `scripts/*.py` that grows without a self-test does count
 38(at 0%), which is the point. Mechanically: one `coverage run --parallel-mode` per self-test into
 39a temp data file, then `coverage combine` + `coverage json`.
 40
 41Standard library only; `cloc` and `coverage` are the external tools, and only `write` needs them.
 42"""
 43
 44import argparse
 45import json
 46import shutil
 47import subprocess
 48import sys
 49import tempfile
 50from pathlib import Path
 51
 52SCHEMA = 1
 53
 54SCALA_PATHS = ("core", "local", "cli")
 55PYTHON_PATHS = ("dspy", "finetune", "scripts")
 56EXCLUDE_DIRS = ("target", "__pycache__", ".venv", "venv", "node_modules")
 57
 58SCALA_COLOR = "DC322F"  # = the README's hand-written Scala badge
 59PYTHON_COLOR = "3776AB"  # = python.org's brand blue, as used by shields' own python logo
 60
 61# Every `python3 <script> --self-test` line of justfile's `quality-other`, in its order. Keeping
 62# the two lists equal is what makes the badge honest: the number below is exactly what that gate
 63# already runs, not a second, friendlier suite. (The `.sh` self-tests in the same recipe are not
 64# Python and cannot contribute statements.)
 65SELF_TEST_SCRIPTS = (
 66    "scripts/smoke_record.py",
 67    "scripts/benchmark_gate.py",
 68    "scripts/cost-split.py",
 69    "scripts/repo_stats.py",
 70    "scripts/arxiv_digest.py",
 71    "scripts/awesome_agentic_digest.py",
 72    "scripts/pr_label_nlp.py",
 73    "scripts/lib/req_merge.py",
 74    "scripts/lib/uses_merge.py",
 75    "scripts/lib/mip_index_merge.py",
 76    "scripts/ocr-post.py",
 77    "scripts/mip_graph.py",
 78    "scripts/lib/tasks_issues.py",
 79    "scripts/strip_external_scripts.py",
 80    "scripts/workflow_runners.py",
 81    "scripts/analyze_training.py",
 82    "scripts/site_live_check.py",
 83)
 84# Measured tree. `dspy/`/`finetune/` are excluded on purpose — see the module docstring.
 85COVERAGE_SOURCE = "scripts"
 86
 87# A step that reached one of these actually executed; anything else (`skipped`, `neutral`, or a
 88# null conclusion for a step still queued/running) is not evidence either way and is not counted.
 89RAN = ("success", "failure", "cancelled", "timed_out")
 90GREEN = ("success",)
 91
 92
 93# ---------------------------------------------------------------------------------------------
 94# Pure shaping — everything below self-tests without a network call or a `cloc` on PATH.
 95
 96
 97def badge(label: str, message: str, color: str) -> dict:
 98    """One shields.io endpoint document (https://shields.io/badges/endpoint-badge)."""
 99    return {"schemaVersion": SCHEMA, "label": label, "message": message, "color": color}
100
101
102def count_steps(jobs: list[dict], exclude_job: str | None = None) -> tuple[int, int]:
103    """(green, ran) over every step of every job of one run — skipped/unfinished steps ignored."""
104    green = ran = 0
105    for job in jobs:
106        if exclude_job and job.get("name") == exclude_job:
107            continue
108        for step in job.get("steps") or []:
109            conclusion = step.get("conclusion")
110            if conclusion not in RAN:
111                continue
112            ran += 1
113            if conclusion in GREEN:
114                green += 1
115    return green, ran
116
117
118def ci_badge(green: int, ran: int) -> dict:
119    """Green only when every step that ran passed — a single red step is not a rounding error."""
120    if ran == 0:
121        return badge("ci steps", "no data", "lightgrey")
122    color = "brightgreen" if green == ran else "yellow" if green >= 0.9 * ran else "red"
123    return badge("ci steps", f"{green}/{ran} green", color)
124
125
126def loc_badge(label: str, code: int, color: str) -> dict:
127    return badge(label, f"{code:,} LOC", color)
128
129
130def parse_cloc(payload: str, language: str) -> int:
131    """Code lines for one language out of `cloc --json`; 0 when it found none of that language."""
132    return int(json.loads(payload).get(language, {}).get("code", 0))
133
134
135def coverage_badge(percent: float | None) -> dict:
136    """Statement coverage of `scripts/` under its own self-tests. Same thresholds as ci.yml's
137    Scala badge (red < 50 ≤ yellow < 80 ≤ green) so the two read on one scale. The label is
138    `py-cov`, paired with ci.yml's `sc-cov` — short enough that the two badges sit side by side
139    without wrapping, and still distinguishable at a glance, which is the only thing the label
140    has to do."""
141    if percent is None:
142        return badge("py-cov", "no data", "lightgrey")
143    color = "red" if percent < 50 else "yellow" if percent < 80 else "green"
144    return badge("py-cov", f"{percent:.0f}%", color)
145
146
147def parse_coverage_json(payload: str) -> float:
148    """The overall statement percentage out of `coverage json` (`totals.percent_covered`)."""
149    return float(json.loads(payload)["totals"]["percent_covered"])
150
151
152def coverage_run_argv(exe: list[str], script: str, data_file: Path) -> list[str]:
153    """One instrumented self-test run. `--parallel-mode` keeps the ten runs from overwriting each
154    other's data file; `--source` fixes the measured tree so a `scripts/*.py` that no self-test
155    imports still lands in the denominator at 0% instead of vanishing from the report."""
156    return [
157        *exe,
158        "run",
159        "--parallel-mode",
160        f"--data-file={data_file}",
161        f"--source={COVERAGE_SOURCE}",
162        script,
163        "--self-test",
164    ]
165
166
167# ---------------------------------------------------------------------------------------------
168# The two effectful sources: this run's jobs (GitHub API, via `gh`) and `cloc`.
169
170
171def fetch_jobs(repo: str, run_id: str) -> list[dict]:
172    """Every job of one workflow run, steps included. Needs `actions: read` on the token."""
173    jobs: list[dict] = []
174    page = 1
175    while True:
176        url = f"repos/{repo}/actions/runs/{run_id}/jobs?per_page=100&page={page}"
177        out = subprocess.run(
178            ["gh", "api", "-H", "Accept: application/vnd.github+json", url],
179            capture_output=True,
180            text=True,
181            check=True,
182        ).stdout
183        batch = json.loads(out).get("jobs") or []
184        jobs.extend(batch)
185        if len(batch) < 100:
186            return jobs
187        page += 1
188
189
190def cloc_code(paths: tuple[str, ...], language: str, root: Path) -> int:
191    """Code lines of `language` under `paths`; a path that does not exist is simply skipped."""
192    if not shutil.which("cloc"):
193        raise SystemExit(
194            "repo_stats: `cloc` is not on PATH (nix develop has it, and ci.yml takes it from nix develop .#lint)"
195        )
196    present = [p for p in paths if (root / p).exists()]
197    if not present:
198        return 0
199    out = subprocess.run(
200        [
201            "cloc",
202            "--json",
203            "--quiet",
204            f"--exclude-dir={','.join(EXCLUDE_DIRS)}",
205            f"--include-lang={language}",
206            *present,
207        ],
208        capture_output=True,
209        text=True,
210        check=True,
211        cwd=root,
212    ).stdout
213    # cloc prints nothing at all when no file of that language survived the filters.
214    return parse_cloc(out, language) if out.strip() else 0
215
216
217def coverage_exe(which=shutil.which, has_module=None) -> list[str]:
218    """How to invoke coverage.py here, as an argv prefix.
219
220    Two shapes: nix's `python3Packages.coverage` puts a wrapped `coverage` on PATH but *not* on
221    this interpreter's import path, while a pip or distro install does the opposite. Prefer the
222    executable, fall back to `-m`, fail loudly if neither.
223    """
224    if has_module is None:
225
226        def has_module() -> bool:
227            import importlib.util
228
229            return importlib.util.find_spec("coverage") is not None
230
231    if which("coverage"):
232        return ["coverage"]
233    if has_module():
234        return [sys.executable, "-m", "coverage"]
235    raise SystemExit(
236        "repo_stats: coverage.py is not installed (nix develop has it, and ci.yml takes it "
237        "from nix develop .#lint) — or pass --no-python-coverage"
238    )
239
240
241def python_coverage(root: Path) -> float:
242    """Run every self-test under coverage.py and return the combined statement percentage.
243
244    The data files live in a temp directory, so a run leaves no `.coverage*` behind in the repo.
245    A self-test that *fails* aborts the measurement rather than quietly reporting a smaller
246    number — `just quality-other` is the gate for that, and a green badge over a red self-test
247    would be worse than no badge.
248    """
249    exe = coverage_exe()
250    with tempfile.TemporaryDirectory() as tmp:
251        data_file = Path(tmp) / ".coverage"
252        for script in SELF_TEST_SCRIPTS:
253            subprocess.run(
254                coverage_run_argv(exe, script, data_file),
255                cwd=root,
256                check=True,
257                capture_output=True,
258                text=True,
259            )
260        subprocess.run(
261            [*exe, "combine", f"--data-file={data_file}", tmp],
262            cwd=root,
263            check=True,
264            capture_output=True,
265            text=True,
266        )
267        report = Path(tmp) / "coverage.json"
268        subprocess.run(
269            [*exe, "json", f"--data-file={data_file}", "-o", str(report)],
270            cwd=root,
271            check=True,
272            capture_output=True,
273            text=True,
274        )
275        return parse_coverage_json(report.read_text())
276
277
278def write(out_dir: Path, badges: dict[str, dict]) -> list[Path]:
279    out_dir.mkdir(parents=True, exist_ok=True)
280    written = []
281    for name, doc in badges.items():
282        path = out_dir / name
283        path.write_text(json.dumps(doc) + "\n")
284        written.append(path)
285    return written
286
287
288def collect(args) -> dict[str, dict]:
289    root = Path(args.root)
290    badges = {
291        "scala-loc.json": loc_badge("scala", cloc_code(SCALA_PATHS, "Scala", root), SCALA_COLOR),
292        "python-loc.json": loc_badge(
293            "python", cloc_code(PYTHON_PATHS, "Python", root), PYTHON_COLOR
294        ),
295    }
296    if args.run_id:
297        green, ran = count_steps(fetch_jobs(args.repo, args.run_id), args.exclude_job)
298        badges["ci.json"] = ci_badge(green, ran)
299    if not args.no_python_coverage:
300        badges["python-coverage.json"] = coverage_badge(python_coverage(root))
301    return badges
302
303
304# ---------------------------------------------------------------------------------------------
305
306
307def self_test() -> int:
308    jobs = [
309        {
310            "name": "build-test",
311            "steps": [
312                {"name": "checkout", "conclusion": "success"},
313                {"name": "compile+test", "conclusion": "success"},
314                {"name": "coverage", "conclusion": "skipped"},  # main-only, on a PR
315            ],
316        },
317        {
318            "name": "quality-other",
319            "steps": [
320                {"name": "ruff", "conclusion": "success"},
321                {"name": "actionlint", "conclusion": "failure"},
322                {"name": "hadolint", "conclusion": "skipped"},
323                {"name": "docker compose config", "conclusion": None},  # never reached
324            ],
325        },
326        {
327            "name": "repo-stats",  # the reporting job itself: excluded, steps still running
328            "steps": [
329                {"name": "checkout", "conclusion": "success"},
330                {"name": "badges", "conclusion": None},
331            ],
332        },
333    ]
334    assert count_steps(jobs, "repo-stats") == (3, 4), count_steps(jobs, "repo-stats")
335    # Without the exclusion the reporting job's own finished steps leak in.
336    assert count_steps(jobs) == (4, 5), count_steps(jobs)
337    # A job with no steps at all (queued, or `steps` absent) contributes nothing, never crashes.
338    assert count_steps([{"name": "changes"}, {"name": "x", "steps": None}]) == (0, 0)
339
340    # Skipped steps stay out of the denominator: a docs-only push is 2/2, not 2/9.
341    docs_only = [
342        {
343            "name": "quality-other",
344            "steps": [{"conclusion": "success"}] * 2 + [{"conclusion": "skipped"}] * 7,
345        }
346    ]
347    assert count_steps(docs_only) == (2, 2)
348    assert ci_badge(*count_steps(docs_only))["message"] == "2/2 green"
349
350    assert ci_badge(20, 20) == {
351        "schemaVersion": 1,
352        "label": "ci steps",
353        "message": "20/20 green",
354        "color": "brightgreen",
355    }
356    assert ci_badge(19, 20)["color"] == "yellow", "one red step out of twenty: not green, not red"
357    assert ci_badge(17, 20)["color"] == "red"
358    assert ci_badge(0, 0) == {
359        "schemaVersion": 1,
360        "label": "ci steps",
361        "message": "no data",
362        "color": "lightgrey",
363    }
364
365    assert loc_badge("scala", 6865, SCALA_COLOR)["message"] == "6,865 LOC"
366    assert loc_badge("python", 0, PYTHON_COLOR)["message"] == "0 LOC"
367
368    cloc_json = json.dumps(
369        {
370            "header": {"cloc_version": "2.10"},
371            "Scala": {"nFiles": 84, "blank": 1050, "comment": 1673, "code": 6865},
372            "SUM": {"blank": 1050, "comment": 1673, "code": 6865, "nFiles": 84},
373        }
374    )
375    assert parse_cloc(cloc_json, "Scala") == 6865
376    assert parse_cloc(cloc_json, "Python") == 0, "a language cloc did not find is 0, not an error"
377
378    # --- Python coverage: shaping, parsing, the argv builder and how coverage.py is located.
379    assert coverage_badge(73.0) == {
380        "schemaVersion": 1,
381        "label": "py-cov",
382        "message": "73%",
383        "color": "yellow",
384    }
385    assert coverage_badge(80.0)["color"] == "green", "the ci.yml Scala badge's own boundary"
386    assert coverage_badge(49.9)["color"] == "red"
387    assert coverage_badge(100.0)["message"] == "100%", "no decimals on the badge"
388    assert coverage_badge(None) == {
389        "schemaVersion": 1,
390        "label": "py-cov",
391        "message": "no data",
392        "color": "lightgrey",
393    }
394    # Both coverage badges must be distinguishable at a glance — this is the whole reason the
395    # label is not just "coverage" like ci.yml's Scala one used to be. `py-cov` here pairs with
396    # `sc-cov` in ci.yml; if one is renamed the other has to follow.
397    assert coverage_badge(73.0)["label"] != ci_badge(1, 1)["label"]
398
399    cov_json = json.dumps(
400        {
401            "meta": {"version": "7.15.4"},
402            "files": {"scripts/repo_stats.py": {"summary": {"percent_covered": 75.0}}},
403            "totals": {
404                "covered_lines": 1328,
405                "num_statements": 1820,
406                "percent_covered": 72.96703296703296,
407            },
408        }
409    )
410    assert round(parse_coverage_json(cov_json), 2) == 72.97
411    assert coverage_badge(parse_coverage_json(cov_json))["message"] == "73%"
412
413    argv = coverage_run_argv(["coverage"], "scripts/mip_graph.py", Path("/tmp/x/.coverage"))
414    assert argv[:2] == ["coverage", "run"]
415    assert "--parallel-mode" in argv, "ten runs into one data file need parallel mode"
416    assert "--data-file=/tmp/x/.coverage" in argv, "data files stay out of the repo"
417    assert f"--source={COVERAGE_SOURCE}" in argv
418    assert argv[-2:] == ["scripts/mip_graph.py", "--self-test"]
419    assert coverage_run_argv([sys.executable, "-m", "coverage"], "s.py", Path("d"))[1] == "-m"
420
421    assert coverage_exe(which=lambda _: "/usr/bin/coverage") == ["coverage"], "prefer the exe"
422    assert coverage_exe(which=lambda _: None, has_module=lambda: True) == [
423        sys.executable,
424        "-m",
425        "coverage",
426    ], "nix's coverage is on PATH; Ubuntu's python3-coverage is only importable"
427    try:
428        coverage_exe(which=lambda _: None, has_module=lambda: False)
429        raise AssertionError("a missing coverage.py must fail loudly, not report 0%")
430    except SystemExit as exc:
431        assert "--no-python-coverage" in str(exc), str(exc)
432
433    # The badge is only honest while this list is exactly justfile's; drift is the failure mode.
434    root = Path(__file__).resolve().parent.parent
435    for script in SELF_TEST_SCRIPTS:
436        assert (root / script).exists(), f"{script} is in SELF_TEST_SCRIPTS but not on disk"
437    justfile = (root / "justfile").read_text()
438    for script in SELF_TEST_SCRIPTS:
439        assert f"python3 {script} --self-test" in justfile, f"{script} left quality-other"
440    in_recipe = {
441        line.split()[1]
442        for line in justfile.splitlines()
443        if line.strip().startswith("python3 scripts/") and line.strip().endswith("--self-test")
444    }
445    assert in_recipe == set(SELF_TEST_SCRIPTS), sorted(in_recipe ^ set(SELF_TEST_SCRIPTS))
446
447    with tempfile.TemporaryDirectory() as tmp:
448        out = Path(tmp) / "stats"
449        paths = write(out, {"ci.json": ci_badge(20, 20), "scala-loc.json": loc_badge("s", 1, "x")})
450        assert [p.name for p in paths] == ["ci.json", "scala-loc.json"]
451        assert json.loads((out / "ci.json").read_text())["message"] == "20/20 green"
452        # Rewriting replaces rather than appends — every run publishes a whole document.
453        write(out, {"ci.json": ci_badge(1, 2)})
454        assert json.loads((out / "ci.json").read_text())["message"] == "1/2 green"
455
456    write_args = build_parser().parse_args(["write", "--out-dir", "x"])
457    assert write_args.repo == "marola-dev/marola", write_args.repo
458
459    print(
460        "repo_stats self-test: ok (step counting, badge shaping, cloc/coverage parsing, "
461        "the coverage argv + exe resolution, SELF_TEST_SCRIPTS vs. justfile, write, "
462        "the --repo default)"
463    )
464    return 0
465
466
467def build_parser() -> argparse.ArgumentParser:
468    ap = argparse.ArgumentParser(
469        description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter
470    )
471    ap.add_argument("--self-test", action="store_true")
472    sub = ap.add_subparsers(dest="cmd")
473    w = sub.add_parser("write", help="write the badge JSONs into --out-dir")
474    w.add_argument("--out-dir", required=True, type=Path)
475    w.add_argument("--root", default=".", help="repo root the LOC paths are relative to")
476    w.add_argument(
477        "--repo", default="marola-dev/marola", help="owner/name, for the CI-health badge"
478    )
479    w.add_argument("--run-id", help="workflow run to report on; omitted = LOC badges only")
480    w.add_argument("--exclude-job", help="job name to leave out (the reporting job itself)")
481    w.add_argument(
482        "--no-python-coverage",
483        action="store_true",
484        help="skip the coverage.py run (no coverage.py installed, or LOC/CI badges only)",
485    )
486    c = sub.add_parser(
487        "python-coverage", help="print the measured statement %% of scripts/ and exit"
488    )
489    c.add_argument("--root", default=".", help="repo root the self-test paths are relative to")
490    return ap
491
492
493def main(argv: list[str]) -> int:
494    ap = build_parser()
495    args = ap.parse_args(argv)
496    if args.self_test:
497        return self_test()
498    if args.cmd == "python-coverage":
499        percent = python_coverage(Path(args.root))
500        print(f"{percent:.1f}% ({coverage_badge(percent)['message']} on the badge)")
501        return 0
502    if args.cmd != "write":
503        ap.print_help()
504        return 2
505    badges = collect(args)
506    for path in write(args.out_dir, badges):
507        print(f"{path}: {path.read_text().strip()}")
508    return 0
509
510
511if __name__ == "__main__":
512    sys.exit(main(sys.argv[1:]))
SCHEMA = 1
SCALA_PATHS = ('core', 'local', 'cli')
PYTHON_PATHS = ('dspy', 'finetune', 'scripts')
EXCLUDE_DIRS = ('target', '__pycache__', '.venv', 'venv', 'node_modules')
SCALA_COLOR = 'DC322F'
PYTHON_COLOR = '3776AB'
SELF_TEST_SCRIPTS = ('scripts/smoke_record.py', 'scripts/benchmark_gate.py', 'scripts/cost-split.py', 'scripts/repo_stats.py', 'scripts/arxiv_digest.py', 'scripts/awesome_agentic_digest.py', 'scripts/pr_label_nlp.py', 'scripts/lib/req_merge.py', 'scripts/lib/uses_merge.py', 'scripts/lib/mip_index_merge.py', 'scripts/ocr-post.py', 'scripts/mip_graph.py', 'scripts/lib/tasks_issues.py', 'scripts/strip_external_scripts.py', 'scripts/workflow_runners.py', 'scripts/analyze_training.py', 'scripts/site_live_check.py')
COVERAGE_SOURCE = 'scripts'
RAN = ('success', 'failure', 'cancelled', 'timed_out')
GREEN = ('success',)
def badge(label: str, message: str, color: str) -> dict:
 98def badge(label: str, message: str, color: str) -> dict:
 99    """One shields.io endpoint document (https://shields.io/badges/endpoint-badge)."""
100    return {"schemaVersion": SCHEMA, "label": label, "message": message, "color": color}

One shields.io endpoint document (https://shields.io/badges/endpoint-badge).

def count_steps(jobs: list[dict], exclude_job: str | None = None) -> tuple[int, int]:
103def count_steps(jobs: list[dict], exclude_job: str | None = None) -> tuple[int, int]:
104    """(green, ran) over every step of every job of one run — skipped/unfinished steps ignored."""
105    green = ran = 0
106    for job in jobs:
107        if exclude_job and job.get("name") == exclude_job:
108            continue
109        for step in job.get("steps") or []:
110            conclusion = step.get("conclusion")
111            if conclusion not in RAN:
112                continue
113            ran += 1
114            if conclusion in GREEN:
115                green += 1
116    return green, ran

(green, ran) over every step of every job of one run — skipped/unfinished steps ignored.

def ci_badge(green: int, ran: int) -> dict:
119def ci_badge(green: int, ran: int) -> dict:
120    """Green only when every step that ran passed — a single red step is not a rounding error."""
121    if ran == 0:
122        return badge("ci steps", "no data", "lightgrey")
123    color = "brightgreen" if green == ran else "yellow" if green >= 0.9 * ran else "red"
124    return badge("ci steps", f"{green}/{ran} green", color)

Green only when every step that ran passed — a single red step is not a rounding error.

def loc_badge(label: str, code: int, color: str) -> dict:
127def loc_badge(label: str, code: int, color: str) -> dict:
128    return badge(label, f"{code:,} LOC", color)
def parse_cloc(payload: str, language: str) -> int:
131def parse_cloc(payload: str, language: str) -> int:
132    """Code lines for one language out of `cloc --json`; 0 when it found none of that language."""
133    return int(json.loads(payload).get(language, {}).get("code", 0))

Code lines for one language out of cloc --json; 0 when it found none of that language.

def coverage_badge(percent: float | None) -> dict:
136def coverage_badge(percent: float | None) -> dict:
137    """Statement coverage of `scripts/` under its own self-tests. Same thresholds as ci.yml's
138    Scala badge (red < 50 ≤ yellow < 80 ≤ green) so the two read on one scale. The label is
139    `py-cov`, paired with ci.yml's `sc-cov` — short enough that the two badges sit side by side
140    without wrapping, and still distinguishable at a glance, which is the only thing the label
141    has to do."""
142    if percent is None:
143        return badge("py-cov", "no data", "lightgrey")
144    color = "red" if percent < 50 else "yellow" if percent < 80 else "green"
145    return badge("py-cov", f"{percent:.0f}%", color)

Statement coverage of scripts/ under its own self-tests. Same thresholds as ci.yml's Scala badge (red < 50 ≤ yellow < 80 ≤ green) so the two read on one scale. The label is py-cov, paired with ci.yml's sc-cov — short enough that the two badges sit side by side without wrapping, and still distinguishable at a glance, which is the only thing the label has to do.

def parse_coverage_json(payload: str) -> float:
148def parse_coverage_json(payload: str) -> float:
149    """The overall statement percentage out of `coverage json` (`totals.percent_covered`)."""
150    return float(json.loads(payload)["totals"]["percent_covered"])

The overall statement percentage out of coverage json (totals.percent_covered).

def coverage_run_argv(exe: list[str], script: str, data_file: pathlib.Path) -> list[str]:
153def coverage_run_argv(exe: list[str], script: str, data_file: Path) -> list[str]:
154    """One instrumented self-test run. `--parallel-mode` keeps the ten runs from overwriting each
155    other's data file; `--source` fixes the measured tree so a `scripts/*.py` that no self-test
156    imports still lands in the denominator at 0% instead of vanishing from the report."""
157    return [
158        *exe,
159        "run",
160        "--parallel-mode",
161        f"--data-file={data_file}",
162        f"--source={COVERAGE_SOURCE}",
163        script,
164        "--self-test",
165    ]

One instrumented self-test run. --parallel-mode keeps the ten runs from overwriting each other's data file; --source fixes the measured tree so a scripts/*.py that no self-test imports still lands in the denominator at 0% instead of vanishing from the report.

def fetch_jobs(repo: str, run_id: str) -> list[dict]:
172def fetch_jobs(repo: str, run_id: str) -> list[dict]:
173    """Every job of one workflow run, steps included. Needs `actions: read` on the token."""
174    jobs: list[dict] = []
175    page = 1
176    while True:
177        url = f"repos/{repo}/actions/runs/{run_id}/jobs?per_page=100&page={page}"
178        out = subprocess.run(
179            ["gh", "api", "-H", "Accept: application/vnd.github+json", url],
180            capture_output=True,
181            text=True,
182            check=True,
183        ).stdout
184        batch = json.loads(out).get("jobs") or []
185        jobs.extend(batch)
186        if len(batch) < 100:
187            return jobs
188        page += 1

Every job of one workflow run, steps included. Needs actions: read on the token.

def cloc_code(paths: tuple[str, ...], language: str, root: pathlib.Path) -> int:
191def cloc_code(paths: tuple[str, ...], language: str, root: Path) -> int:
192    """Code lines of `language` under `paths`; a path that does not exist is simply skipped."""
193    if not shutil.which("cloc"):
194        raise SystemExit(
195            "repo_stats: `cloc` is not on PATH (nix develop has it, and ci.yml takes it from nix develop .#lint)"
196        )
197    present = [p for p in paths if (root / p).exists()]
198    if not present:
199        return 0
200    out = subprocess.run(
201        [
202            "cloc",
203            "--json",
204            "--quiet",
205            f"--exclude-dir={','.join(EXCLUDE_DIRS)}",
206            f"--include-lang={language}",
207            *present,
208        ],
209        capture_output=True,
210        text=True,
211        check=True,
212        cwd=root,
213    ).stdout
214    # cloc prints nothing at all when no file of that language survived the filters.
215    return parse_cloc(out, language) if out.strip() else 0

Code lines of language under paths; a path that does not exist is simply skipped.

def coverage_exe(which=<function which>, has_module=None) -> list[str]:
218def coverage_exe(which=shutil.which, has_module=None) -> list[str]:
219    """How to invoke coverage.py here, as an argv prefix.
220
221    Two shapes: nix's `python3Packages.coverage` puts a wrapped `coverage` on PATH but *not* on
222    this interpreter's import path, while a pip or distro install does the opposite. Prefer the
223    executable, fall back to `-m`, fail loudly if neither.
224    """
225    if has_module is None:
226
227        def has_module() -> bool:
228            import importlib.util
229
230            return importlib.util.find_spec("coverage") is not None
231
232    if which("coverage"):
233        return ["coverage"]
234    if has_module():
235        return [sys.executable, "-m", "coverage"]
236    raise SystemExit(
237        "repo_stats: coverage.py is not installed (nix develop has it, and ci.yml takes it "
238        "from nix develop .#lint) — or pass --no-python-coverage"
239    )

How to invoke coverage.py here, as an argv prefix.

Two shapes: nix's python3Packages.coverage puts a wrapped coverage on PATH but not on this interpreter's import path, while a pip or distro install does the opposite. Prefer the executable, fall back to -m, fail loudly if neither.

def python_coverage(root: pathlib.Path) -> float:
242def python_coverage(root: Path) -> float:
243    """Run every self-test under coverage.py and return the combined statement percentage.
244
245    The data files live in a temp directory, so a run leaves no `.coverage*` behind in the repo.
246    A self-test that *fails* aborts the measurement rather than quietly reporting a smaller
247    number — `just quality-other` is the gate for that, and a green badge over a red self-test
248    would be worse than no badge.
249    """
250    exe = coverage_exe()
251    with tempfile.TemporaryDirectory() as tmp:
252        data_file = Path(tmp) / ".coverage"
253        for script in SELF_TEST_SCRIPTS:
254            subprocess.run(
255                coverage_run_argv(exe, script, data_file),
256                cwd=root,
257                check=True,
258                capture_output=True,
259                text=True,
260            )
261        subprocess.run(
262            [*exe, "combine", f"--data-file={data_file}", tmp],
263            cwd=root,
264            check=True,
265            capture_output=True,
266            text=True,
267        )
268        report = Path(tmp) / "coverage.json"
269        subprocess.run(
270            [*exe, "json", f"--data-file={data_file}", "-o", str(report)],
271            cwd=root,
272            check=True,
273            capture_output=True,
274            text=True,
275        )
276        return parse_coverage_json(report.read_text())

Run every self-test under coverage.py and return the combined statement percentage.

The data files live in a temp directory, so a run leaves no .coverage* behind in the repo. A self-test that fails aborts the measurement rather than quietly reporting a smaller number — just quality-other is the gate for that, and a green badge over a red self-test would be worse than no badge.

def write(out_dir: pathlib.Path, badges: dict[str, dict]) -> list[pathlib.Path]:
279def write(out_dir: Path, badges: dict[str, dict]) -> list[Path]:
280    out_dir.mkdir(parents=True, exist_ok=True)
281    written = []
282    for name, doc in badges.items():
283        path = out_dir / name
284        path.write_text(json.dumps(doc) + "\n")
285        written.append(path)
286    return written
def collect(args) -> dict[str, dict]:
289def collect(args) -> dict[str, dict]:
290    root = Path(args.root)
291    badges = {
292        "scala-loc.json": loc_badge("scala", cloc_code(SCALA_PATHS, "Scala", root), SCALA_COLOR),
293        "python-loc.json": loc_badge(
294            "python", cloc_code(PYTHON_PATHS, "Python", root), PYTHON_COLOR
295        ),
296    }
297    if args.run_id:
298        green, ran = count_steps(fetch_jobs(args.repo, args.run_id), args.exclude_job)
299        badges["ci.json"] = ci_badge(green, ran)
300    if not args.no_python_coverage:
301        badges["python-coverage.json"] = coverage_badge(python_coverage(root))
302    return badges
def self_test() -> int:
308def self_test() -> int:
309    jobs = [
310        {
311            "name": "build-test",
312            "steps": [
313                {"name": "checkout", "conclusion": "success"},
314                {"name": "compile+test", "conclusion": "success"},
315                {"name": "coverage", "conclusion": "skipped"},  # main-only, on a PR
316            ],
317        },
318        {
319            "name": "quality-other",
320            "steps": [
321                {"name": "ruff", "conclusion": "success"},
322                {"name": "actionlint", "conclusion": "failure"},
323                {"name": "hadolint", "conclusion": "skipped"},
324                {"name": "docker compose config", "conclusion": None},  # never reached
325            ],
326        },
327        {
328            "name": "repo-stats",  # the reporting job itself: excluded, steps still running
329            "steps": [
330                {"name": "checkout", "conclusion": "success"},
331                {"name": "badges", "conclusion": None},
332            ],
333        },
334    ]
335    assert count_steps(jobs, "repo-stats") == (3, 4), count_steps(jobs, "repo-stats")
336    # Without the exclusion the reporting job's own finished steps leak in.
337    assert count_steps(jobs) == (4, 5), count_steps(jobs)
338    # A job with no steps at all (queued, or `steps` absent) contributes nothing, never crashes.
339    assert count_steps([{"name": "changes"}, {"name": "x", "steps": None}]) == (0, 0)
340
341    # Skipped steps stay out of the denominator: a docs-only push is 2/2, not 2/9.
342    docs_only = [
343        {
344            "name": "quality-other",
345            "steps": [{"conclusion": "success"}] * 2 + [{"conclusion": "skipped"}] * 7,
346        }
347    ]
348    assert count_steps(docs_only) == (2, 2)
349    assert ci_badge(*count_steps(docs_only))["message"] == "2/2 green"
350
351    assert ci_badge(20, 20) == {
352        "schemaVersion": 1,
353        "label": "ci steps",
354        "message": "20/20 green",
355        "color": "brightgreen",
356    }
357    assert ci_badge(19, 20)["color"] == "yellow", "one red step out of twenty: not green, not red"
358    assert ci_badge(17, 20)["color"] == "red"
359    assert ci_badge(0, 0) == {
360        "schemaVersion": 1,
361        "label": "ci steps",
362        "message": "no data",
363        "color": "lightgrey",
364    }
365
366    assert loc_badge("scala", 6865, SCALA_COLOR)["message"] == "6,865 LOC"
367    assert loc_badge("python", 0, PYTHON_COLOR)["message"] == "0 LOC"
368
369    cloc_json = json.dumps(
370        {
371            "header": {"cloc_version": "2.10"},
372            "Scala": {"nFiles": 84, "blank": 1050, "comment": 1673, "code": 6865},
373            "SUM": {"blank": 1050, "comment": 1673, "code": 6865, "nFiles": 84},
374        }
375    )
376    assert parse_cloc(cloc_json, "Scala") == 6865
377    assert parse_cloc(cloc_json, "Python") == 0, "a language cloc did not find is 0, not an error"
378
379    # --- Python coverage: shaping, parsing, the argv builder and how coverage.py is located.
380    assert coverage_badge(73.0) == {
381        "schemaVersion": 1,
382        "label": "py-cov",
383        "message": "73%",
384        "color": "yellow",
385    }
386    assert coverage_badge(80.0)["color"] == "green", "the ci.yml Scala badge's own boundary"
387    assert coverage_badge(49.9)["color"] == "red"
388    assert coverage_badge(100.0)["message"] == "100%", "no decimals on the badge"
389    assert coverage_badge(None) == {
390        "schemaVersion": 1,
391        "label": "py-cov",
392        "message": "no data",
393        "color": "lightgrey",
394    }
395    # Both coverage badges must be distinguishable at a glance — this is the whole reason the
396    # label is not just "coverage" like ci.yml's Scala one used to be. `py-cov` here pairs with
397    # `sc-cov` in ci.yml; if one is renamed the other has to follow.
398    assert coverage_badge(73.0)["label"] != ci_badge(1, 1)["label"]
399
400    cov_json = json.dumps(
401        {
402            "meta": {"version": "7.15.4"},
403            "files": {"scripts/repo_stats.py": {"summary": {"percent_covered": 75.0}}},
404            "totals": {
405                "covered_lines": 1328,
406                "num_statements": 1820,
407                "percent_covered": 72.96703296703296,
408            },
409        }
410    )
411    assert round(parse_coverage_json(cov_json), 2) == 72.97
412    assert coverage_badge(parse_coverage_json(cov_json))["message"] == "73%"
413
414    argv = coverage_run_argv(["coverage"], "scripts/mip_graph.py", Path("/tmp/x/.coverage"))
415    assert argv[:2] == ["coverage", "run"]
416    assert "--parallel-mode" in argv, "ten runs into one data file need parallel mode"
417    assert "--data-file=/tmp/x/.coverage" in argv, "data files stay out of the repo"
418    assert f"--source={COVERAGE_SOURCE}" in argv
419    assert argv[-2:] == ["scripts/mip_graph.py", "--self-test"]
420    assert coverage_run_argv([sys.executable, "-m", "coverage"], "s.py", Path("d"))[1] == "-m"
421
422    assert coverage_exe(which=lambda _: "/usr/bin/coverage") == ["coverage"], "prefer the exe"
423    assert coverage_exe(which=lambda _: None, has_module=lambda: True) == [
424        sys.executable,
425        "-m",
426        "coverage",
427    ], "nix's coverage is on PATH; Ubuntu's python3-coverage is only importable"
428    try:
429        coverage_exe(which=lambda _: None, has_module=lambda: False)
430        raise AssertionError("a missing coverage.py must fail loudly, not report 0%")
431    except SystemExit as exc:
432        assert "--no-python-coverage" in str(exc), str(exc)
433
434    # The badge is only honest while this list is exactly justfile's; drift is the failure mode.
435    root = Path(__file__).resolve().parent.parent
436    for script in SELF_TEST_SCRIPTS:
437        assert (root / script).exists(), f"{script} is in SELF_TEST_SCRIPTS but not on disk"
438    justfile = (root / "justfile").read_text()
439    for script in SELF_TEST_SCRIPTS:
440        assert f"python3 {script} --self-test" in justfile, f"{script} left quality-other"
441    in_recipe = {
442        line.split()[1]
443        for line in justfile.splitlines()
444        if line.strip().startswith("python3 scripts/") and line.strip().endswith("--self-test")
445    }
446    assert in_recipe == set(SELF_TEST_SCRIPTS), sorted(in_recipe ^ set(SELF_TEST_SCRIPTS))
447
448    with tempfile.TemporaryDirectory() as tmp:
449        out = Path(tmp) / "stats"
450        paths = write(out, {"ci.json": ci_badge(20, 20), "scala-loc.json": loc_badge("s", 1, "x")})
451        assert [p.name for p in paths] == ["ci.json", "scala-loc.json"]
452        assert json.loads((out / "ci.json").read_text())["message"] == "20/20 green"
453        # Rewriting replaces rather than appends — every run publishes a whole document.
454        write(out, {"ci.json": ci_badge(1, 2)})
455        assert json.loads((out / "ci.json").read_text())["message"] == "1/2 green"
456
457    write_args = build_parser().parse_args(["write", "--out-dir", "x"])
458    assert write_args.repo == "marola-dev/marola", write_args.repo
459
460    print(
461        "repo_stats self-test: ok (step counting, badge shaping, cloc/coverage parsing, "
462        "the coverage argv + exe resolution, SELF_TEST_SCRIPTS vs. justfile, write, "
463        "the --repo default)"
464    )
465    return 0
def build_parser() -> argparse.ArgumentParser:
468def build_parser() -> argparse.ArgumentParser:
469    ap = argparse.ArgumentParser(
470        description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter
471    )
472    ap.add_argument("--self-test", action="store_true")
473    sub = ap.add_subparsers(dest="cmd")
474    w = sub.add_parser("write", help="write the badge JSONs into --out-dir")
475    w.add_argument("--out-dir", required=True, type=Path)
476    w.add_argument("--root", default=".", help="repo root the LOC paths are relative to")
477    w.add_argument(
478        "--repo", default="marola-dev/marola", help="owner/name, for the CI-health badge"
479    )
480    w.add_argument("--run-id", help="workflow run to report on; omitted = LOC badges only")
481    w.add_argument("--exclude-job", help="job name to leave out (the reporting job itself)")
482    w.add_argument(
483        "--no-python-coverage",
484        action="store_true",
485        help="skip the coverage.py run (no coverage.py installed, or LOC/CI badges only)",
486    )
487    c = sub.add_parser(
488        "python-coverage", help="print the measured statement %% of scripts/ and exit"
489    )
490    c.add_argument("--root", default=".", help="repo root the self-test paths are relative to")
491    return ap
def main(argv: list[str]) -> int:
494def main(argv: list[str]) -> int:
495    ap = build_parser()
496    args = ap.parse_args(argv)
497    if args.self_test:
498        return self_test()
499    if args.cmd == "python-coverage":
500        percent = python_coverage(Path(args.root))
501        print(f"{percent:.1f}% ({coverage_badge(percent)['message']} on the badge)")
502        return 0
503    if args.cmd != "write":
504        ap.print_help()
505        return 2
506    badges = collect(args)
507    for path in write(args.out_dir, badges):
508        print(f"{path}: {path.read_text().strip()}")
509    return 0