Mmcp.market

Agent Cost Report skill

by thedotmack·thedotmack/claude-mem·95k stars·Apache-2.0

Believable agent cost report for any period, default the last 7 full days PT, not counting today. Measured tokens from Claude Code transcripts priced at OpenRouter list prices (ESTIMATED), measured provider spend when a sanctioned source exists, note-taker cost separate, Timing-style HTML/PDF plus report.json, line-items.csv, evidence.json.

A97/100content scan

Is the Agent Cost Report skill safe?

Clean: nothing in its files matched our rules. We read 45 files in the folder on 2026-09-28.

  • lowSKILL.md:1

    The name should be 1 to 64 lowercase letters, digits or hyphens.

    Agent Cost Report

Install the Agent Cost Report skill

A skill is a folder. Copy it into your agent's skills folder and the agent loads it when the task matches its description.

git clone --depth 1 https://github.com/thedotmack/claude-mem.git /tmp/claude-mem
mkdir -p ~/.claude/skills
cp -r /tmp/claude-mem/claude-mem-cursor/skills/agent-cost-report ~/.claude/skills/agent-cost-report
available in every project

In the Claude apps, zip the folder and upload it from the Skills settings. The folder on GitHub

The instructions your agent would load

SKILL.md as published, without the frontmatter. Read it on GitHub

Agent Cost Report

Claude-Mem / Claude Code skill. Runtime is the scripts/ pipeline (transcripts → tokens → dollars → Timing-style report) plus a progressive Mem Search review pass that confirms the drafted labels. The Notion draft is SPEC history only — never the product, never the runtime, never the ship vehicle.

Resolve the absolute directory containing this SKILL.md; all helper paths are relative to that directory. ${CLAUDESKILLDIR} is the shortcut: python3 "${CLAUDESKILLDIR}/scripts/acr.py" …. Python 3.9+ standard library only; no pip installs. The look lives in scripts/acr/render.py, never here.

Purpose

Turn Claude-Mem activity into a manager-readable cost and failure report.

Product idea: a reusable skill that searches Claude-Mem via Mem Search, reconstructs real units of work, assigns cost and failure categories, and renders a printable report.

Insight north star

The headline is dollars, to two decimals, labeled. The dollars come from Claude Code transcripts (exact per-reply token usage) priced at OpenRouter public list prices, so they are ESTIMATED. Measured provider spend appears only when a sanctioned source gives it. Directly under the dollars: what the mistakes cost, what shipped, and both on one time axis.

Questions the report must answer

  1. What work was completed?
  2. What did each outcome cost?
  3. What was wasted through looping, hedging, wrong turns, rework, or poor routing?
  4. Were any unauthorized actions attempted?
  5. What should the manager change next?

Primary unit = cost per completed outcome (not cost per observation).

When to use

  • "Agent cost report" / "cost per outcome" / "failure economics" / "was this session worth it" / "what did the agents cost this week"
  • After a real Mem session dig when leadership needs outcome economics
  • Sample / ship packs that need self-contained HTML + JSON + CSV + evidence

Memory dig mechanics: the claude-mem mem-search skill (progressive recall).

Default scope (ALWAYS)

  • Unless the user names a specific session / range / project, the window is the last 7 full days in PT, not counting today: end = the PT midnight that started today (exclusive), start = end − 7 days. The default window never contains a partial day (G3, Alex 2026-09-25).
  • Explicit windows: --start YYYY-MM-DD --end YYYY-MM-DD (PT calendar days, end exclusive). An explicit --end later than today marks the last day "partial, generated HH:MM PT".
  • One session: --session . One project plus a period: --project --start … --end … (worktrees of the project are included).
  • The report Scope strip shows the PT range, and Details list every session id in scope.

Progressive Mem Search (ALWAYS)

Follow the claude-mem mem-search three layers. Keep spend light.

  1. Search — get an index of IDs (titles, types, token hints).
  2. Timeline — only around anchors you care about.
  3. Observations — get_observations for the filtered IDs you will cite as evidence.

Recipe:

  1. Resolve scope (default: the last 7 full PT days).
  2. Search → collect IDs.
  3. Timeline for thin context only.
  4. Observations for intended / actual / outcome / waste / rework / blocked / unauthorized / status.
  5. Group into named work items + failure events.
  6. Calculate line-item costs.
  7. Render HTML + optional PDF (+ json/csv/evidence).
  8. Keep evidence IDs in the appendix — do not dump entire timelines into the main report.

ADHD process bullets:

  • Search first → pick IDs → timeline only if context is thin → fetch only needed obs.
  • Work-item titles are invented for managers ("Restore search after Chroma crash-loop"); observation titles stay evidence-only.
  • Evidence appendix lists obs IDs + short titles; main sections stay outcome-first.

In this skill the search pass is the review step (see Recipe step 4): the pipeline drafts categories and failure types from keywords (labelsource: keyword); the orchestrator confirms or changes each line item's category and failuretype from its cited evidence IDs and applies the result with review --apply. Items left unreviewed keep the "draft label" mark and the footer counts them.

Work categories (ALWAYS)

Feature · Bug fix · Incident · Maintenance · Investigation · Experiment

Failure / waste types (ALWAYS)

Looping · Hedging · Wrong turn · Rework · Regression · Premature completion · Unauthorized action · Suboptimal path · Duplicate work · Blocked work · Missed requirement · Unnecessary escalation · Context re-read · Model thrash · Fan-out waste · Recovery after miss

Rework lock (ALWAYS): Rework lives only under failuretype — never as a work category. Keep category as the intended job type; set failuretype: Rework when rework occurred.

A line item can have a work category and a failure_type (e.g. Maintenance + Looping).

Cost model

Measured tokens come from Claude Code transcripts (~/.claude/projects//.jsonl, assistant replies deduped on (message.id, requestId)); Codex transcripts are read the same way. Prices are OpenRouter public list prices per million tokens, fetched at run time and saved with the report.

agent_cost_i (per reply, micro-dollars) =
      input × price.input + output × price.output
    + cache_write_5m × price.cache_write + cache_write_1h × price.cache_write_1h
    + cache_read × price.cache_read                    # cache_write_1h = listed rate, else 2 × input

agent_estimated_usd        = Σ agent_cost_i over every reply in the window (matched or not)   # ESTIMATED, the headline
extrapolated_unmeasured    = observer tokens of sessions with no transcript here × (measured $ per observer token)
                           # EXTRAPOLATED (low confidence); "all sessions measured" when nothing remains
observer_note_taker_est    = note-taker (observer) tokens, deduped per reply, × its input list price
                           # priced separately, never agent cost, never in the headline
mistakes_estimated_usd     = Σ agent_cost_i over the same-session union of wasted turns, each turn once   # low figure
cost_per_completed_outcome = Σ attributed $ for status ∈ {shipped, completed} / count(those work items)
waste_rate                 = Σ wasted_cost / Σ attributed $     recovery_share = Σ recovery_cost / Σ attributed $

Line items are sessions: attributedusd = estimated (transcript on this box) or extrapolated (no transcript). wastedcost and recovery_cost come from the behavior pass (below), one union set for the ribbon, the line items and the mistakes line. The upper bound (redo windows plus project-wide fallback) stays in Details.

Unauthorized blocked: directcost $0, riskexposure high, actionstatus blocked. Always keep riskexposure non-dollar unless real cash/remediation is at stake — never invent risk dollars.

confidence — high when tokens + model + outcome are clear; medium when allocation across obs is judgmental; low when evidence is thin.

Unpriced models (not in the price list, or a negative "variable" price) are listed by name with their tokens and add nothing; they are never priced at zero silently.

Behavior metrics (heuristic until reviewed)

A second pass over the same transcripts tags every user turn human / bot / unknown (relayed agent prompts are never Alex's words), finds frustration episodes, and runs the pattern detectors from the Frustration Arc study: invented human gates, broke working things, wrong or expensive model, over-engineering, did something not asked, fake output, false "done", wrong tool or contact, bad outbound (incidents × recipients, never dollars), memory or rule loss, jargon, unclear cause, plus tool errors and hedging (Alex's definition: a caveat given when the answer was already available). Four summary tiles; everything else in Details. Every count is labeled heuristic until reviewed or classified. The optional classifier (--classify) is off by default, capped at $2.00 per run, uses only a regular inference OPENROUTERAPIKEY, and its spend is shown separately.

Money labeling (ALWAYS)

costbasis on each line item: estimatedusage, measured_provider, or extrapolated. Dollars are shown to two decimals everywhere ($109.25); every figure carries its label and basis.

Line-item schema

Each row in line-items.csv / report.json.line_items:

Deliverables (ALWAYS)

acr.py writes one directory containing:

  1. report.html — self-contained (inline CSS, no script, no external resources), Timing-style
  2. report.pdf — from report.print.html with headless google-chrome; when Chrome is missing the run says "PDF skipped, HTML is canonical"
  3. report.json — window, scope, spend, totals, byday, bymodel, bydevice, lineitems, wins, behavior, timeline
  4. line-items.csv — one row per work item
  5. evidence.json — observation IDs cited, short titles, observer tokens, model, session ids
  6. labels.review.json — the drafted labels for the review pass

Manager-facing HTML sections (required order)

  1. Hero — dollars with its tag (ESTIMATE / MEASURED), the basis sentence, "N things finished, about $X each", the extrapolated line, "Measured provider spend: unavailable" when it is, and the story sentence (computed, never hand-written)
  2. Wins vs mistakes — the cost-of-mistakes line (low figure), wins shipped (merged PRs, published versions, praise; cost "unmeasured" until sessions are linked to PRs), and two timelines on one PT-day axis
  3. Cost ribbon — every piece of work as a block sized by cost, waste hatched, in-progress striped
  4. Where the money went (donut by kind of work) · Day by day (stacked bars, empty days say "no agent work") · How much was useful (the ring, the behavior strip with at most four tiles)
  5. What got done — one folded row per work item, most expensive first, draft labels marked
  6. Worth your attention — at most three, computed
  7. Details & evidence (folded; open in the PDF) — wins, mistakes by day, behavior pattern table with the upper bound, rule effectiveness, spend and pricing, models, failure accounting, line items, devices, unmatched transcripts, unpriced models, label counts, the honesty rules, links to the CSV and JSON

Truthfulness ALWAYS rules

Positive truth rules only:

  1. Always label estimate vs measured (costbasis = estimatedusage or measured_provider).
  2. Always say measured spend unavailable when unmatched — never "$0 spent" for unknown.
  3. Always use completed outcomes as the unit (cost per completed outcome).
  4. Always keep risk_exposure non-dollar unless real cash/remediation is at stake.
  5. Always invent human work-item titles; observation titles are evidence-only.
  6. Always progressive Mem Search; evidence IDs live in the appendix.
  7. Always put Rework under failure_type only.
  8. Always default scope to most recent sessions unless the user names another.

Added by the rebuild, same spirit: the note-taker's cost is separate and never in the headline; keyword labels are drafts until reviewed; Grok Bot and Mac figures are "unavailable" or "extrapolated (low confidence)", never $0 and never guessed; the live database is only ever opened read-only (every run works on a snapshot).

Recipe (orchestrator)

All commands from the skill directory; OUT is one run directory (for example /tmp/acr-weekly/run). Every step is idempotent.

  1. Prices — python3 scripts/acr.py prices --out $OUT (OpenRouter list prices; --prices to run offline; no key needed).
  2. Collect — python3 scripts/acr.py collect --out $OUT (default window) or … --start YYYY-MM-DD --end YYYY-MM-DD, … --session , … --project --start … --end ….
  3. Rollup — python3 scripts/acr.py rollup --out $OUT with the same period flags. Writes report.json, line-items.csv, evidence.json, labels.review.json. Options: --no-gh (skip the read-only gh pr view confirmation of merged PRs), --classify (see above), --rules-dir (rule effectiveness).
  4. Review — for each line item in labels.review.json, getobservations on its evidenceids (not whole timelines), confirm or change category, failuretype, status, and the title; write the reviewed file with reviewedby (model id or Alex); apply with python3 scripts/acr.py review --apply --out $OUT. Items left unreviewed stay marked "draft label".
  5. Render — python3 scripts/acr.py render --in $OUT/report.json --out $OUT --print.
  6. PDF — python3 scripts/acr.py pdf --out $OUT (skipped with a message when Chrome is missing).
  7. Optional, needs Alex's go — measured OpenRouter spend: with a regular inference key in the environment, OPENROUTERAPIKEY=… python3 scripts/acr.py measure-openrouter --out $OUT before step 3. Without the key the file says unavailable and nothing else changes.
  8. Optional, needs Alex's go — Mac export merge: python3 scripts/acr.py rollup … --device-usage (see Gaps and gates).
  9. Return absolute paths, the skill slug agent-cost-report, the headline dollars with their label, the mistakes line, the wins, PDF yes/no, and how the report answers the five product questions.

Gaps and gates

Each of these needs Alex's explicit go; the report shows the honest fallback until then.

  • OpenRouter key (needs Alex) — a regular inference OPENROUTERAPIKEY supplied as an environment variable through the house's secure secret flow; never a management or provisioning key, never read from a settings file, never printed or written to any output. The per-key endpoint gives UTC day/week/month snapshots, so the figure is shown with its bucket label and switches the hero to MEASURED only when the bucket fully covers the report window. A range-by-day source (/api/v1/activity) exists but needs a management key, so it is not implemented (G5).
  • Mac transcripts (needs Alex) — sessions from Alex's Mac have no transcript on the box, so their cost is EXTRAPOLATED (low confidence) until a device export is merged. Either Alex runs, on the Mac, from a plain checkout of this skill (system python3 3.9+): python3 scripts/acr.py collect --export-device mac --start YYYY-MM-DD --end YYYY-MM-DD --out ~/acr-export and shares ~/acr-export/device-usage-mac.json (ids, timestamps, token counts, model names; no prompt or observation text, no settings, no paths); or, only after Alex's explicit go for that specific run (window, machine, destination named), an orchestrating agent runs the same command on the registered Mac through the house's registered-machine tooling and copies only device-usage-mac.json to the box. A go for one run is not a go for the next (G6 covered Sep 18–26 only). Merge with --device-usage; joined sessions flip to estimated_usage with device: mac.
  • Grok Bot (ALWAYS unavailable, no seat count) — Grok Bot is Cursor's cloud agent. Cursor shows its weekly usage only on the plan screen; there is no API or export, and the house has no xAI key for it. The report says exactly "Grok Bot usage: unavailable" (G4): no dollars, no guessed per-seat price, no seat count. A figure Alex supplies by hand would be entered as measuredmanual with enteredby: Alex and the date, never inferred; that would be a new decision.
  • Wins from GitHub (read-only) — besides box transcripts and ship observations, rollup lists merged PRs with gh pr list --state merged for the repos in --wins-repo (default thedotmack/claude-mem), filtered by merge time inside the PT window and deduplicated by repo and PR number. When gh is missing, unauthenticated or rate-limited, that source reads "unavailable" in Details and the report still completes; it is never a zero. --no-gh turns every gh call off.
  • Win cost (unmeasured until R4) — sessions are not linked to PRs today, so every win reads "unmeasured (session not linked to PR)". Once commits carry a Claude-Session: trailer (G13: trailer only, bare id), PR and publish wins get ESTIMATED · session-linked costs. There is no fallback estimate (G11).
  • Merge to main, npm publish, /version-bump — each needs Alex's separate go. Pushing the work branch and opening the PR are routine.

Keeping copies in sync

The plugin directory is the source of truth. python3 scripts/acr.py sync-check compares SKILL.md, CHECKSUMS.txt and scripts/** against the house copy (~/agent-data/workflows/agent-cost-report/) and the four mirror plugins (claude-mem-cursor, claude-mem-grok-bot, cowork, openclaw, each under skills/agent-cost-report/) and exits 1 on drift; --write copies plugin → destinations and re-checks (G2). CHECKSUMS.txt is sha256sum -c compatible.

More skills from thedotmack/claude-mem

  • AAgent Cost ReportBelievable agent cost report for any period, default the last 7 full days PT, not counting today. Measured tokens from Claude Code transcripts priced at OpenRouter list prices (ESTIMATED), measured provider spend when a sanctioned source exists, note-taker cost separate, Timing-style HTML/PDF plus report.json, line-items.csv, evidence.json.
  • AAgent Cost ReportBelievable agent cost report for any period, default the last 7 full days PT, not counting today. Measured tokens from Claude Code transcripts priced at OpenRouter list prices (ESTIMATED), measured provider spend when a sanctioned source exists, note-taker cost separate, Timing-style HTML/PDF plus report.json, line-items.csv, evidence.json.
  • AAgent Cost ReportBelievable agent cost report for any period, default the last 7 full days PT, not counting today. Measured tokens from Claude Code transcripts priced at OpenRouter list prices (ESTIMATED), measured provider spend when a sanctioned source exists, note-taker cost separate, Timing-style HTML/PDF plus report.json, line-items.csv, evidence.json.
  • AAgent Cost ReportBelievable agent cost report for any period, default the last 7 full days PT, not counting today. Measured tokens from Claude Code transcripts priced at OpenRouter list prices (ESTIMATED), measured provider spend when a sanctioned source exists, note-taker cost separate, Timing-style HTML/PDF plus report.json, line-items.csv, evidence.json.
  • AbabysitWatch a pull request or review cycle until it is ready to merge. Use when asked to babysit, monitor, or keep checking PR comments, reviews, and CI until all actionable issues are resolved.
  • Accs-alignRun the CCS Align seat's hourly breathing cycle — prove the local claude-mem worker is healthy, pull needle observations through search → timeline → get_observations, land them in a seat-owned middle cache via atomic grab → append → filter exclude-marks → replace, manage exclude marks, and walk house → project → seat rules to detect conflicts (SHADOW_HOUSE, DENY_ALLOW, DRIFT, CLOCK_HEADER) with an append-only rules-report.md. Use when asked to run CCS Align, breathe the alignment seat, refresh the middle cache, exclude or restore an observation, walk rules, check rules conflicts, or check the Worker Watch board.
  • Aclaude-mem-installUse this when setting up claude-mem on Cursor: local or remote worker, local host-login observer or remote cmem.ai inference.
  • Aclaude-mem-installUse this when setting up claude-mem on Grok Bot: local worker plus CMEM Pro observer (default), optional host-login observer, or remote cmem.ai. No Cursor required.
  • Acloud-syncSet up or check claude-mem cloud sync with cmem.ai Pro. Use when the user says "set up cloud sync", "sync my memories", "cmem pro", "cloud backup", "sync status", or wants their memory database backed up or synced to their cmem.ai account.
  • Adesign-isAudit a design against Dieter Rams' ten "Good design is..." principles, then hand off a /make-plan prompt for one of three outcomes — new design, refine design, or redesign. Use when the user says "audit this design", "design review", "check this UI against Rams", "is this UI good", "critique this design", "design audit", or asks for a critique that should lead to a plan.
  • Ado
  • AdoExecute a phased implementation plan using subagents. Use when asked to execute, run, or carry out a plan — especially one created by make-plan.

All agent skills → · MCP servers