Mmcp.market

weak-agent-test skill

by kklimuk·kklimuk/docx-cli·214 stars·MIT

Run the weak-agent adversarial test harness against docx-cli. Spawns weak exercise agents (Haiku by default, Sonnet to probe, or a local agent harness's pre-produced runs) to perform real document tasks over six scenarios — five editing (MNDA form-fill + font fidelity, invoice table-edit/restructure + logo replace, résumé styling, contract redlining + commenting, contract finalize via accept/reject + comment reply/resolve) and one authoring (T. S. Eliot poetry journal: multi-column, verse, footnotes, links, figure) — renders every result with Word, has opus judge them against ground-truth rubrics, measures each exercise's tool economy, token cost, wall-clock, and correctness (from transcripts for Claude, the exercise.json ledger for the local harness), and synthesizes a prioritized ergonomics report. Use when the user says 'adversarial review', 'test docx-cli with weak agents', 'run the haiku harness', 'weak agent test', or wants to re-run yesterday's adversarial process.

A100/100content scan

Is the weak-agent-test skill safe?

Clean: nothing in its files matched our rules. We read 22 files in the folder on 2026-09-28.

No findings.

Install the weak-agent-test skill

A skill is a folder. Copy it into your agent's skills folder and the agent loads it when the task matches its description.

git clone --depth 1 https://github.com/kklimuk/docx-cli.git /tmp/docx-cli
mkdir -p ~/.claude/skills
cp -r /tmp/docx-cli/.claude/skills/weak-agent-test ~/.claude/skills/weak-agent-test
available in every project

In the Claude apps, zip the folder and upload it from the Skills settings. The folder on GitHub

The instructions your agent would load

SKILL.md as published, without the frontmatter. Read it on GitHub

Adversarial review — weak-agent harness for docx-cli

This harness answers one question: can weak agents actually use docx-cli to get real work done, and what should we fix first? It runs the weak-agent-test workflow (.claude/workflows/weak-agent-test.js), which fans out one weak exercise agent per scenario (Haiku by default — swappable to Sonnet via args.model), renders every output with Microsoft Word, grades each against ground-truth criteria with an opus judge, and has opus synthesize a prioritized improvement report. Exercise agents do NOT self-report tool counts — every tool-economy and token number is measured after the run (agents under-count their own calls ~2×, so self-reports were dropped): from the agent transcripts for the Claude arms, from each scenario's exercise.json ledger for the local arm. Both roll up into the same Run-metrics table (tokens, wall-clock, tool split, correctness) via exercise-metrics.ts.

The test corpus is bundled with this skill under scenarios/, one folder per scenario, named after its key (scenarios/mnda/, scenarios/invoice/, …). Each scenario folder is self-describing and holds everything that scenario needs:

the goal, the data, the intent — and no tool vocabulary (no docx commands, locators, or OOXML terms), because discovering which features deliver the outcome is part of what's measured,

  • task.md — the AGENT-FACING request, written as a human delegating the work:

The stage step withholds it from the agent's run workspace, and the judge reads it from the pristine source — the agent never sees the answer key,

  • criteria.md — the JUDGE-ONLY grading rubric (the precise, tool-specific checks).

their output fresh),

  • the fixture .docx to work on (edit scenarios only; authoring scenarios create

scenarios).

  • assets/ — any additional inputs (data files, images; empty for most edit

The workflow's SCENARIOS manifest holds only the per-scenario routing metadata (key, bucket label, edit/author kind, the doc filename); whether a baseline gets rendered is DERIVED from the kind (every edit scenario has a pristine source, so it gets one — see hasBaseline()), not a stored field. The actual request/criteria/fixture/assets all live in the folder. The skill is therefore self-contained and travels with its test corpus. To change what a scenario tests, edit the files in its folder. (Heavy, ephemeral run outputs — edited docx, renders, reviews, the report — are dumped to ./tmp/docx-weak-agent-test//, never into the repo.)

Staging is ONE code path for every backend: scripts/stage-scenario.ts copies a scenario folder, strips the judge-only criteria.md, and verifies the inputs landed. The workflow's Stage agent runs it per scenario; the local corpus runner imports it.

Each run produces, under the timestamped run dir, one result folder per scenario (named after its key) plus the run-level report and metrics:

<RUN_DIR>/
  REPORT.md            ← synthesized report; the Metrics phase appends the measured
                          run-metrics section (local: in-run; Claude: your post-run pass)
  exercise-metrics.md  ← measured per-exercise-agent tokens/time/tool split
  exercise-metrics.json
  <key>/               ← one per scenario; the worked-on copy lives here
    task.md  assets/   ← (criteria.md is withheld from this copy — judge-only)
    <doc>.docx         ← the edited/authored document
    renders/output/    ← the OUTPUT: Word-rendered page PNGs + read.md (markdown read view)
    renders/baseline/  ← the pristine "before": page PNGs + read.md (every EDIT scenario;
                          absent only for the authored eliot-journal — no source to diff)
    review.md          ← the judge's saved review for this task (written in-run)
    verdict.json       ← the judge's structured verdict incl. taskSuccess (written in-run
                          by the judge — the correctness source the Metrics phase reads)
    metrics.json       ← this task's measured tokens/time/tool split + correctness
                          (local: in-run Metrics phase; Claude: your post-run pass)

The render step fires the moment each task finishes (for both arms) and produces, for the OUTPUT and — whenever a pristine source exists (every edit scenario) — its BASELINE "before", BOTH deliverables in each render dir: the page PNGs AND a read.md (the markdown read view of that doc). The judge reads all four (output PNGs + read.md, baseline PNGs + read.md) to compare before/after both visually and textually. The workflow's render step is idempotent: for the local backend the corpus runner already produced the SAME artifacts at the SAME paths as it went, so the render step just reuses them (re-rendering only anything missing) — no double-render; for the Claude backend nothing is pre-rendered, so it does the full Word render. Either way the judge grades Word-rendered PNGs (the local harness runs on the mac, where the corpus's default render engine IS Word).

Steps

Run these in order from the repo root. Do NOT skip the build — the global docx on PATH is a stale binary; the harness must test the CURRENT working tree.

1. Preflight — ALWAYS rebuild (mandatory gate)

The whole harness is meaningless if it tests a stale binary, so the build is a hard gate, not an optional step. Always run bun run build:binary, even if dist/docx already exists — never reuse a prior build. Abort the whole run if any check below fails.

REPO="$(git rev-parse --show-toplevel)"
cd "$REPO"
SCENARIOS_DIR="$REPO/.claude/skills/weak-agent-test/scenarios"   # this skill's bundled corpus (one folder per scenario)

# Word must be installed (this harness renders with Word, not LibreOffice).
test -d "/Applications/Microsoft Word.app" || echo "WARNING: Microsoft Word not found — render phase will fail."

# (1) Build the CURRENT working tree into a fresh standalone binary. Abort on failure.
bun run build:binary || { echo "BUILD FAILED — abort"; exit 1; }
BINARY="$REPO/dist/docx"

# (2) Hard gate: the fresh binary must match package.json's version AND have `render`.
# `--version` prints "docx X.Y.Z"; take the 2nd space-delimited field. NOTE: use `cut`,
# NOT an awk field reference — a literal dollar-N positional token gets clobbered by
# slash-command positional-arg substitution when this skill runs with arguments, mangling
# the gate. Keep this whole block free of dollar-N tokens for the same reason.
EXPECTED="$(bun -e 'console.log(require("./package.json").version)')"
GOT="$("$BINARY" --version | cut -d' ' -f2)"
echo "built docx $GOT (package.json: $EXPECTED)"
[ "$GOT" = "$EXPECTED" ] || { echo "VERSION MISMATCH ($GOT != $EXP

If the version mismatches or render is missing, the build did not reflect the working tree — stop and fix it before running. Do not proceed on a stale binary.

First-run note: Word-for-Mac rendering triggers a one-time macOS Automation

permission prompt for the controlling terminal. If the render phase fails on a

fresh machine, grant it under System Settings → Privacy & Security → Automation and

re-run.

2. Make an isolated run workspace (under ./tmp/)

Create an empty timestamped ./tmp/ run dir per workflow run. Do NOT copy the scenarios here — the workflow's Stage phase runs scripts/stage-scenario.ts for only the active scenarios, seeding one subfolder per scenario ($RUN_DIR//), so originals stay untouched, the repo stays clean, and a single-scenario run doesn't drag the whole corpus along:

TS="$(date +%Y.%m.%d-%H%M%S)"
RUN_DIR="./tmp/docx-weak-agent-test/$TS"
mkdir -p "$RUN_DIR"   # empty; the workflow's Stage phase seeds one subfolder per active scenario from $SCENARIOS_DIR
echo "RUN_DIR=$RUN_DIR"

3. Launch the workflow (up to 3 concurrently)

Invoke the Workflow tool with scriptPath pointing at the workflow file and pass the absolute paths as args:

Workflow({
  scriptPath: "<REPO>/.claude/workflows/weak-agent-test.js",
  args: {
    runDir: "<RUN_DIR from step 2>",
    binary: "<BINARY from step 1>",
    scenariosDir: "<SCENARIOS_DIR from step 1>",
    model: "haiku",              // the exercise model: "haiku" (default) or "sonnet"
    only: <optional scenario filter — see below>
  }
})

Exercise agent type. The exercise agents run as the repo's weak-exercise

agent type (.claude/agents/weak-exercise.md): minimal tools and no Skill

tool, so the session's skills catalog stays OUT of their context (it's a

per-turn token tax and leaks docx-cli/harness names into the

"capable-but-fresh agent" premise). The agent registry loads at SESSION

start — in a session older than that file, the workflow aborts with

"agent type 'weak-exercise' not found"; pass

exerciseAgentType: "general-purpose" to override for that session (and note

the run's base context is then ~4k tokens/turn heavier, so its token numbers

aren't comparable to weak-exercise runs).

Never resume a benchmark run whose exercise phase failed. If an exercise

agent dies (API error → that scenario reports no exercise/verdict), re-run

the WHOLE run in a FRESH run dir. resumeFromRunId replays the cached stage

step without re-copying fixtures, so re-run exercise agents would edit

already-edited documents — double redlines, double fills, unusable verdicts

(this voided run r2 on 2026-07-15, twice). The failure is worse than it

looks because the resume cache is PREFIX-based, not keyed: everything

issued AFTER the first missing/changed result re-runs live, not just the

dead agent. So a dead exercise for a MANIFEST-EARLY scenario (mnda is

first) re-runs EVERY exercise against edited docs even if you restore that

one scenario's staging state — while a dead LAST scenario (eliot-journal)

happens to resume cleanly. Don't gamble on manifest position: exercise-phase

failure → fresh run dir, no exceptions. Resume is only safe for failures at

or after the render phase (dead judge/synth), where nothing mutates

documents no matter how much of the suffix re-runs.

Running 3 at a time (the fast path to averaged numbers). The benchmark methodology is 3 runs per arm/model, and runs can go concurrently: launch up to three Workflow invocations in one message, each with its OWN RUN_DIR from step 2 (suffix the timestamp, e.g. $TS-r1, $TS-r2, $TS-r3). This is safe because the only shared mutable resource is Microsoft Word, and the CLI itself serializes Word access across processes with an advisory lock (src/core/render/engines/word-mac.ts) — concurrent runs' renders queue instead of corrupting each other. Don't go beyond ~3: renders start spending more time queueing than rendering. A haiku-vs-sonnet comparison is just two batches: three runs with model: "haiku", three with model: "sonnet" (never mix models within one run dir).

only restricts the run to a subset of scenarios (omit it to run all 6). To run a single task, pass its key as a plain string — only: "mnda". It also accepts an array (only: ["mnda", "invoice"]) or a comma/space-separated string; all forms are normalized to the same list. The keys are the folder names under $SCENARIOSDIR (run ls "$SCENARIOSDIR" if you need to confirm them); unknown keys abort the run with a "No scenarios matched" error listing the valid ones.

Use scriptPath, NOT name: "weak-agent-test". Launching by name resolves to a

More skills from kklimuk/docx-cli

  • AcommitCreate well-structured git commits from the current working tree. Use when the user says 'commit', 'save my work', 'let's commit this', 'make a commit', or any variation of wanting to commit code to git.
  • Cdocx-cliRead, edit, redline, comment on, and create Microsoft Word .docx files. Use to fill out or edit a Word doc, redline a contract with tracked changes, add/resolve comments, replace text keeping its formatting, restyle headings/fonts, edit tables, or read/extract a .docx as Markdown or text. Also BUILD a new .docx — from Markdown or programmatically (code that outputs a Word report with headings, tables, images). Not for PDF, Google Docs, Excel, PowerPoint, or .doc.
  • Asecurity-reviewReview code for security vulnerabilities. Use when the user says 'security review', 'security audit', 'check for vulnerabilities', 'pentest the code', 'OWASP check', or any variation of wanting a security assessment.

All agent skills → · MCP servers