Mmcp.market

auto-review-loop skill

by wanshuiyin·wanshuiyin/Auto-claude-code-research-in-sleep·17k stars·MIT

Autonomous multi-round research review loop. In Copilot CLI it defaults to the native complementary rubber-duck subagent with host-event model evidence; elsewhere it uses Codex, while explicit external reviewer overrides remain available. Implements fixes and re-reviews until a policy-approved positive assessment or max rounds is reached.

A100/100content scan

Is the auto-review-loop skill safe?

Clean: nothing in its files matched our rules. We read 1 file in the folder on 2026-09-28.

No findings.

Install the auto-review-loop skill

A skill is a folder. Copy it into your agent's skills folder and the agent loads it when the task matches its description.

git clone --depth 1 https://github.com/wanshuiyin/Auto-claude-code-research-in-sleep.git /tmp/Auto-claude-code-research-in-sleep
mkdir -p ~/.claude/skills
cp -r /tmp/Auto-claude-code-research-in-sleep/skills/auto-review-loop ~/.claude/skills/auto-review-loop
available in every project

In the Claude apps, zip the folder and upload it from the Skills settings. The folder on GitHub

The instructions your agent would load

SKILL.md as published, without the frontmatter. Read it on GitHub

Auto Review Loop: Autonomous Research Improvement

🔒 Do not wrap this skill in /loop, /schedule, or CronCreate. It

already loops internally (review → fix → re-review) and the reviewer carries

round-to-round memory in one threadId (codex-reply). An external timer

re-enters from the top each tick — fresh threadId, reviewer memory reset —

firing the verdict on wall-clock time instead of on artifact change: zero new

signal, full token cost. If you want to schedule something, schedule the

external wait that precedes it (experiments done → then run this once). See

shared-references/external-cadence.md.

Autonomously iterate: review → implement fixes → re-review, until an independent reviewer gives a policy-approved positive assessment or MAX_ROUNDS is reached.

Context: $ARGUMENTS

Constants

  • MAX_ROUNDS = 4
  • POSITIVETHRESHOLD: score >= 6/10 AND verdict ∈ {"ready", "almost"} — both** must hold. This matches the operative Phase-E STOP CONDITION exactly; the verdict vocabulary is {"ready", "almost", "not ready"} (a high score with a "not ready" verdict does NOT stop the loop). Earlier wording here used or and a stale verdict set ("accept"/"sufficient"/"ready for submission") — that was an internal inconsistency; the AND form is authoritative.
  • REVIEWDOC: review-stage/AUTOREVIEW.md (cumulative log) (fall back to ./AUTOREVIEW.md for legacy projects)*
  • REVIEWER_MODEL = gpt-6-astra — Default model for the Codex backend. Must be an OpenAI model (e.g., gpt-6-astra, o3, gpt-4o). Manual backend uses a model the user chooses — it must be a recognized model from a different family (OpenAI, Anthropic, Google, DeepSeek, Moonshot/Kimi, Qwen).
  • REVIEWERBACKEND — With no reviewer directive, start as auto; Step -1 runs exactly one two-call native marker/challenge probe for the first review. A bound Copilot CLI root session uses copilot-native (built-in complementary rubber-duck subagent); an unbound/non-Copilot host keeps the existing codex default. Explicit — reviewer: codex, oracle-pro, agy, or manual bypasses the probe and selects that external backend. Explicit — reviewer: copilot retains the compatibility copilot --agent drive mode and its later Codex/manual finalizer. The native path gets both actual model IDs from host session events; it never needs COPILOTCLI or caller-provided --executor-model. See shared-references/reviewer-routing.md.
  • OUTPUTDIR = review-stage/** — All review-stage outputs go here. Create the directory if it doesn't exist.
  • HUMANCHECKPOINT = false** — When true, pause after each round's review (Phase B) and present the score + weaknesses to the user. Wait for user input before proceeding to Phase C. The user can: approve the suggested fixes, provide custom modification instructions, skip specific fixes, or stop the loop early. When false (default), the loop runs fully autonomously.
  • COMPACT = false — When true, (1) read EXPERIMENT_LOG.md and findings.md instead of parsing full logs on session recovery, (2) append key findings to findings.md after each round.
  • REVIEWERDIFFICULTY = medium** — Controls how adversarial the reviewer is. Three levels:
  • medium (default): Current behavior — MCP-based review, the executor controls what context the reviewer sees.
  • hard: Adds Reviewer Memory (the reviewer tracks its own suspicions across rounds) + Debate Protocol (the executor can rebut, the reviewer rules).
  • nightmare: Everything in hard + Codex exec reviewer reads the repo directly via codex exec (the executor cannot filter what the reviewer sees) + Adversarial Verification (the reviewer independently checks if code matches claims).

⚠️ Nightmare + Manual incompatibility: If REVIEWERBACKEND = manual and REVIEWERDIFFICULTY = nightmare, STOP with:

"difficulty: nightmare requires Codex CLI / codex exec and is not compatible with --reviewer: manual. Use difficulty: hard, or switch reviewer to codex."

💡 Override: /auto-review-loop "topic" — compact: true, human checkpoint: true, difficulty: hard

Reviewer Calling Convention

When calling the reviewer, branch on REVIEWER_BACKEND:

If no --reviewer: directive was supplied: Set REVIEWERBACKEND to auto. At Step -1 of the first round, resolve copilotnativeevidence.py using the canonical four-layer helper chain. Generate a fresh binding rreview<8-random-hex> and invoke marker, wait, then invoke challenge as two distinct root Bash calls. Put the literal binding and concrete resolved helper path in both calls; Copilot Bash calls do not share variables. If the challenge binds, set REVIEWERBACKEND to copilot-native and use that same challenge for the first review. Do not issue a second activation challenge in Phase A. If it exits 3 because no current Copilot root session is bound, use codex. Explicit reviewer directives bypass this probe. If the helper is missing, native acceptance is unavailable; use Codex only if that external backend is positively available, otherwise emit REVIEWUNAVAILABLE.

If REVIEWERBACKEND = copilot-native: Read the challenge nonce and host-reported executor model. Invoke the host's native task tool with agenttype: rubber-duck; do not start a subprocess and do not specify a reviewer model. The prompt contains the exact standalone ARISREVIEWNONCE= line, artifact/diff paths, the output contract, and (round 2+) review-stage/REVIEWERMEMORY.md. It contains no executor summary or fix narrative. After the task completes, invoke copilotnative_evidence.py verify to create the evidence and raw-response artifacts. The verifier must observe one successful linked rubber-duck lifecycle and known, different host-reported model families.

Pass the evidence to both reviewgate.py --native-evidence and savetrace.sh --backend copilot-native --native-evidence. A qualifying native positive may stop directly; no external finalizer is needed. A native negative continues with a fresh marker/challenge/subagent next round. Every verdict-bearing native call—including a hard-mode rebuttal ruling—gets one unique artifact set and exactly one challenge. Missing, same/unknown-family, malformed, stale, or mismatched evidence is never a verdict. If native complementary dispatch is unavailable, fall back only to a positively available opposite-family backend: Anthropic/Google executor → Codex; OpenAI executor → manual with a reported non-OpenAI model. Otherwise emit REVIEW_UNAVAILABLE. Full protocol: shared-references/reviewer-routing.md.

If REVIEWERBACKEND = copilot: Require --executor-model: if not provided → emit REVIEWUNAVAILABLE. Determine executor family from --executor-model (see reviewer-routing.md). Router picks opposite-family profile:

Verify the profile file exists at .github/agents/.agent.md. If missing → REVIEWUNAVAILABLE. Read its model: field into REVIEWERMODEL, derive reviewerfamily from that model string, and verify it differs from executorfamily. Pass the same value through subprocess --model; never trust a caller-supplied family label or profile-only pinning under an Auto session. Identity assurance: --executor-model is caller-declared routing input, not runtime attestation. Record executormodelsource: caller-declared, the derived familyrelation, and independenceverified: unverified. A pair of different model strings must never be promoted to independently verified. Capability gate: copilot --help must advertise --model, --effort, and --allow-tool; otherwise emit REVIEWUNAVAILABLE. Use the copilot --agent subprocess (documented Copilot CLI form) with the selected profile, --model "$REVIEWERMODEL", --effort xhigh, and --allow-tool=read for each review call. Multi-round: each round is a fresh copilot --agent call with the same profile; reviewer memory is carried via review-stage/REVIEWERMEMORY.md artifact. If copilot CLI is unavailable → REVIEWUNAVAILABLE for that drive round; do not silently substitute another transport. A later positive Copilot verdict still requires the separately documented Codex/manual finalizer. See shared-references/reviewer-routing.md for the full copilot contract.

  • executor_family=openai → profile="aris-reviewer-claude" (anthropic)
  • executor_family=anthropic → profile="aris-reviewer-openai" (openai)
  • executor_family=google → profile="aris-reviewer-openai" (openai, default cross)
  • executorfamily=unknown → REVIEWUNAVAILABLE (fail closed).

If REVIEWERBACKEND = codex: Use mcpcodexcodex for new review threads. Use mcpcodex__codex-reply for follow-up rounds (reuse threadId).

If REVIEWERBACKEND = manual: Use mcpmanualreviewreview for new review threads with: prompt: [exact same prompt that would go to Codex] config: {"modelreasoningeffort": "xhigh", "executormodel": "", "requirereviewermodel": true} Save the returned threadId. Use mcpmanualreviewreviewreply for follow-up rounds with: threadId: [saved manual-review threadId] prompt: [follow-up prompt] config: {"modelreasoningeffort": "xhigh", "executormodel": "", "requirereviewermodel": true} A verdict-bearing manual response MUST begin with Reviewer-Model: . Derive reviewerfamily from that model identity. Missing, unknown, or same-family identity cannot acquit; for a mandatory escalation, emit REVIEWUNAVAILABLE rather than guessing.

Prompt fidelity: the manual review task must be exactly the same text that Codex would receive; the transport may add only the required Reviewer-Model: response-format instruction. Review tracing applies to every backend. Native traces are populated from the revalidated host-event artifact rather than caller model declarations.

State Persistence (Compact Recovery)

Long-running loops may hit the context window limit, triggering automatic compaction. To survive this, persist state to review-stage/REVIEW_STATE.json after each round:

{
  "run_id": "run_20260713_a1b2c3d4",
  "round": 2,
  "threadId": null,
  "reviewer_profile": "rubber-duck",
  "reviewer_backend": "copilot-native",
  "executor_model": "claude-sonnet-4.6",
  "executor_model_source": "host-session-event",
  "executor_family": "anthropic",
  "requested_reviewer_model": null,
  "reported_reviewer_model": "gpt-5.5",
  "reviewer_model_source": "host-session-event",
  "reviewer_family": "openai",
  "family_relation": "different",
  "identity_assurance": "host_event_verified",
  "independence_verified": true,
  "native_evidence_id": "cne_0123456789abcdef0123456789abcdef",
  "native_evidence_path": "review-stage/COPILOT_NATIVE_run_20260713_a1b2c3d4_ROUND_2_REVIEW.evidence.json",
  "requires_external_acquittal": false,
  "status": "in_progress",
  "difficulty": "medium",
  "last_score": 5.0,
  "last_verdict": "not ready",
  "pending_experiments": ["screen_name_1"],
  "timestamp": "2026-03-13T21:00:00"
}
  • runid — Globally unique per invocation. Generated on fresh start as run<8-char-hex> (e.g., run20260713a1b2c3d4). Preserved across round writes. On resume, read from state file unchanged. This binds all round state, reviewer-memory appends, and acquittal receipts to one run so a stale completed state from a previous invocation cannot leak into the current run's acquittal check.

When REVIEWERBACKEND = copilot-native, save the evidence ID/path and the host-event executor/reviewer models, derived families, and sources. Each round is a fresh rubber-duck subagent and therefore gets a fresh evidence artifact; there is no persistent child handle. When REVIEWERBACKEND = compatibility copilot, retain reviewerprofile, requested model, caller-declared executor model, independenceverified: "unverified", and the external-finalizer obligation. For codex save its MCP threadId; for manual save threadId and the reported reviewer identity. On resume, use reviewerbackend to select the continuation mechanism and preserve requiresexternal_acquittal.

Write this file at the end of every Phase E (after documenting the round). Overwrite each time — only the latest round's state matters. The run_id field MUST persist unchanged across overwrites within the same run.

On completion (positive assessment or max rounds), set "status": "completed" so future invocations don't accidentally resume a finished loop.

Append-Only External-Finalizer Receipt

Whenever a Copilot path hands the verdict to an external backend—after a positive compatibility-drive review or after a pre-verdict native dispatch failure—maintain an append-only finalizer log at review-stage/ACQUITTAL_LOG.jsonl. Each line records the Codex/manual reviewer that completed that run. A successful native rubber-duck round never needs or writes this receipt; its evidence sidecar is the acceptance record. The historical filename is retained for compatibility:

{"run_id":"run_20260713_a1b2c3d4","round":3,"backend":"codex","effort":"xhigh","verdict":"ready","score":7.5,"executor_model":"claude-sonnet-4-5","executor_model_source":"caller-declared","executor_family":"anthropic","reviewer_model":"gpt-6-astra","reviewer_model_source":"requested","reviewer_family":"openai","family_relation":"different","identity_assurance":"caller_declared","independence_verified":"unverified","trace_id":"auto-review-loop/2026-07-13_run03","timestamp":"2026-07-13T14:22:00Z"}

Rules (non-negotiable):

Why this exists: REVIEW_STATE.json is overwritten each round. The log preserves evidence that a compatibility drive verdict or failed native attempt did not terminate by itself. A successful copilot-native verdict instead uses its host-event evidence sidecar.

Output Protocols

Follow these shared protocols for all output files:

- Output Versioning Protocol — write timestamped file first, then copy to fixed name

- Output Manifest Protocol — log every output to MANIFEST.md

- Output Language Protocol — respect the project's language setting

Workflow

Initialization

  1. Check for review-stage/REVIEWSTATE.json (fall back to ./REVIEWSTATE.json if not found — legacy path):
  • If neither path exists: fresh start (normal case, identical to behavior before this feature existed)
  • Generate runid: run<8-char-hex> (e.g., run20260713a1b2c3d4). Use date +%Y%m%d and 8 random hex characters. This run_id persists across all round writes and binds acquittal receipts to this invocation.
  • If it exists AND status is "completed": fresh start (previous loop finished normally — but its ACQUITTALLOG.jsonl entries are retained as an audit trail with their own runid, and are NOT valid for the current run's stop gate)
  • Generate a new runid** for this invocation.
  • If it exists AND status is "inprogress" AND timestamp is older than 24 hours: fresh start** (stale state from a killed/abandoned run — delete the file and start over)
  • Generate a new runid** for this invocation.
  • If it exists AND status is "inprogress" AND timestamp is within 24 hours: resume**
  • Read the state file to recover runid, round, threadId (or evidence/profile fields for Copilot backends), reviewerbackend, lastscore, pendingexperiments
  • Legacy backward compat: if reviewerbackend is absent from the state file, default to codex (pre-copilot-era states did not record this field). If requiresexternalacquittal is absent, default it to false; a legacy default-Codex run must not inherit the stricter Copilot-finalizer state. If runid is absent from the state file (pre-runid era), generate a new runid and log: "No runid in legacy state file; assigned run<...> for this resume."
  • Read review-stage/AUTOREVIEW.md to restore full context of prior rounds (fall back to ./AUTOREVIEW.md)
  • If pending_experiments is non-empty, check if they have completed (e.g., check screen sessions)
  • Resume from the next round (round = saved round + 1)

REVIEWER_BACKEND to auto. Step -1 of Round 1 performs activation and uses that same challenge for the review. Explicit reviewer directives initialize their selected backend and bypass activation. Do not use environment heuristics.

  1. Read project narrative documents, memory files, and any prior review documents. When COMPACT = true and compact files exist: read findings.md + EXPERIMENTLOG.md instead of full review-stage/AUTOREVIEW.md and raw logs — saves context window.
  2. Read recent experiment results (check output directories, logs)
  3. Identify current weaknesses and open TODOs from prior reviews
  4. Initialize round counter = 1 (unless recovered from state file)
  5. Create/update review-stage/AUTO_REVIEW.md with header and timestamp
  6. If this is a fresh run with no explicit reviewer directive, initialize

Loop (repeat up to MAX_ROUNDS)

Step -1 — Resolve the automatic backend and prepare one native challenge:

marker call followed by one root challenge call. Use binding rreview<8-random-hex> and output review-stage/COPILOTNATIVEROUNDREVIEW.challenge.json. A bound challenge sets REVIEWERBACKEND to copilot-native and NATIVECHALLENGE to that path. Exit 3/unbound sets REVIEWER_BACKEND

  • If REVIEWER_BACKEND is auto, resolve the native helper and run one root

More skills from wanshuiyin/Auto-claude-code-research-in-sleep

  • Aablation-plannerUse when main results pass result-to-claim (claim_supported=yes or partial) and ablation studies are needed for paper submission.
  • Aablation-plannerUse when main results pass result-to-claim (`claim_supported = yes` or `partial`) and ablation studies are needed for paper submission. A secondary Codex agent designs ablations from a reviewer's perspective; the local executor reviews feasibility and implements.
  • AalphaxivQuick single-paper lookup via AlphaXiv LLM-optimized summaries with tiered source fallback. Use when user says "explain this paper", "summarize paper", pastes an arXiv/AlphaXiv URL, or provides a bare arXiv ID for quick understanding - not for broad literature search.
  • AalphaxivQuick single-paper lookup via AlphaXiv LLM-optimized summaries with tiered source fallback. Use when user says "explain this paper", "summarize paper", pastes an arXiv/AlphaXiv URL, or provides a bare arXiv ID for quick understanding - not for broad literature search.
  • Aanalyze-resultsAnalyze ML experiment results, compute statistics, generate comparison tables and insights. Use when user says "analyze results", "compare", or needs to interpret experimental data.
  • Aanalyze-resultsAnalyze ML experiment results, compute statistics, generate comparison tables and insights. Use when user says \"analyze results\", \"compare\", or needs to interpret experimental data.
  • AarxivSearch, download, and summarize academic papers from arXiv. Use when user says "search arxiv", "download paper", "fetch arxiv", "arxiv search", "get paper pdf", or wants to find and save papers from arXiv to the local paper library.
  • AarxivSearch, download, and summarize academic papers from arXiv. Use when user says \"search arxiv\", \"download paper\", \"fetch arxiv\", \"arxiv search\", \"get paper pdf\", or wants to find and save papers from arXiv to the local paper library.
  • Aauto-paper-improvement-loopAutonomously improve a generated paper via GPT-6-Astra xhigh review → implement fixes → recompile, for 2 rounds. Use when user says \"改论文\", \"improve paper\", \"论文润色循环\", \"auto improve\", or wants to iteratively polish a generated paper.
  • Aauto-paper-improvement-loopAutonomously improve a generated paper via Claude review through claude-review MCP → implement fixes → recompile, for 2 rounds. Use when user says \"改论文\", \"improve paper\", \"论文润色循环\", \"auto improve\", or wants to iteratively polish a generated paper.
  • Aauto-paper-improvement-loopAutonomously improve a generated paper via Gemini review through gemini-review MCP → implement fixes → recompile, for 2 rounds. Use when user says \"改论文\", \"improve paper\", \"论文润色循环\", \"auto improve\", or wants to iteratively polish a generated paper.
  • Aauto-paper-improvement-loopAutonomously improve a generated paper via GPT-6-Astra xhigh review → implement fixes → recompile, for 2 rounds. Use when user says \"改论文\", \"improve paper\", \"论文润色循环\", \"auto improve\", or wants to iteratively polish a generated paper.

All agent skills → · MCP servers