Mmcp.market

auto-paper-improvement-loop skill

by wanshuiyin·wanshuiyin/Auto-claude-code-research-in-sleep·17k stars·MIT

Autonomously improve a generated paper via GPT-6-Astra xhigh review → implement fixes → recompile, for 2 rounds. Use when user says \"改论文\", \"improve paper\", \"论文润色循环\", \"auto improve\", or wants to iteratively polish a generated paper.

A100/100content scan

Is the auto-paper-improvement-loop skill safe?

Clean: nothing in its files matched our rules. We read 1 file in the folder on 2026-09-28.

No findings.

Install the auto-paper-improvement-loop skill

A skill is a folder. Copy it into your agent's skills folder and the agent loads it when the task matches its description.

git clone --depth 1 https://github.com/wanshuiyin/Auto-claude-code-research-in-sleep.git /tmp/Auto-claude-code-research-in-sleep
mkdir -p ~/.claude/skills
cp -r /tmp/Auto-claude-code-research-in-sleep/skills/auto-paper-improvement-loop ~/.claude/skills/auto-paper-improvement-loop
available in every project

In the Claude apps, zip the folder and upload it from the Skills settings. The folder on GitHub

The instructions your agent would load

SKILL.md as published, without the frontmatter. Read it on GitHub

Auto Paper Improvement Loop: Review → Fix → Recompile

🔒 Do not wrap this skill in /loop, /schedule, or CronCreate. It

already loops internally (review → fix → recompile) with its own round

structure and a deliberate fresh-reviewer bias guard each round (no

codex-reply). Re-asking it to "improve the paper" on a

wall-clock timer produces no new signal — quality changes when the review

changes, not when the clock ticks — and a timed re-run that also accepts its

own output to decide when to stop crosses into self-acquittal

(acceptance-gate.md). Schedule the external wait that precedes it, not the

improvement loop. See

shared-references/external-cadence.md.

Autonomously improve the paper at: $ARGUMENTS

Context

This skill is designed to run after Workflow 3 (/paper-plan → /paper-figure → /paper-write → /paper-compile). It takes a compiled paper and iteratively improves it through external LLM review.

Unlike /auto-review-loop (which iterates on research — running experiments, collecting data, rewriting narrative), this skill iterates on paper writing quality — fixing theoretical inconsistencies, softening overclaims, adding missing content, and improving presentation.

Constants

  • MAXROUNDS = 2** — Two rounds of review→fix→recompile. Empirically, Round 1 catches structural issues (4→6/10), Round 2 catches remaining presentation issues (6→7/10). Diminishing returns beyond 2 rounds for writing-only improvements.
  • REVIEWERMODEL = gpt-6-astra** — Model used via Codex MCP for paper review.
  • REVIEWERBIASGUARD = true — When true, every review round uses a fresh mcpcodexcodex thread with no prior review context. Never use mcpcodexcodex-reply for review rounds. Set to false only for deliberate debugging of the legacy behavior. Empirical evidence: running the same paper with codex-reply + "since last round we did X" prompts inflated scores from real 3/10 → fake 8/10 across multiple rounds; switching to fresh threads recovered the true 3/10 assessment.
  • REVIEWLOG = PAPERIMPROVEMENTLOG.md** — Cumulative log of all rounds, stored in paper directory.
  • HUMANCHECKPOINT = false** — When true, pause after each round's review and present score + weaknesses to the user. The user can approve fixes, provide custom modification instructions, skip specific fixes, or stop early. When false (default), runs fully autonomously.
  • EDITWHITELIST = null — Optional path to a YAML/JSON whitelist file constraining which paths and operations the fix-implementation step may touch. When null (default), all edits proceed unconstrained. When set via — edit-whitelist (also accepts — editwhitelist ), the loop loads the file at startup and consults it before each edit; rejected edits are logged to PAPERIMPROVEMENTLOG.md rather than silently dropped. See "Optional: Edit Whitelist" below.

💡 Override: /auto-paper-improvement-loop "paper/" — human checkpoint: true

Optional: Style reference (— style-ref: , opt-in)

Lets the user steer structural fixes only during improvement (section reordering hints, paragraph length nudges, figure density adjustments) toward a reference paper. Default OFF — when the user does not pass — style-ref, do nothing differently from before.

Only when — style-ref: appears in $ARGUMENTS, run the helper FIRST, before the loop starts:

# Resolve $STYLE_HELPER via the canonical strict-safe chain (see
# shared-references/integration-contract.md §2). Policy A — gate:
# unresolved helper means --style-ref cannot be satisfied, so abort.
cd "$(git rev-parse --show-toplevel 2>/dev/null || pwd)" || exit 1
if [ -z "${ARIS_REPO:-}" ] && [ -f .aris/installed-skills.txt ]; then
    ARIS_REPO=$(awk -F'\t' '$1=="repo_root"{print $2; exit}' .aris/installed-skills.txt 2>/dev/null) || true
fi
if [ -z "${ARIS_REPO:-}" ] && [ -f "$HOME/.aris/repo" ]; then
    ARIS_REPO=$(cat "$HOME/.aris/repo" 2>/dev/null) || true
fi
STYLE_HELPER=".aris/tools/extract_paper_style.py"
[ -f "$STYLE_HELPER" ] || STYLE_HELPER="tools/extract_paper_style.py"
[ -f "$STYLE_HELPER" ] || { [ -n "${ARIS_REPO:-}" ] && STYLE_HELPER="$ARIS_REPO/tools/extract_paper_style.py"; }
[ -f "$STYLE_HELPER" ] || {
  echo "ERROR: extract_paper_style.py not resolved at .aris/tools/, tools/, \$ARIS_REPO/tools/, or via ~/.aris/repo." >&2
  echo "       Fix: rerun bash tools/install_aris.sh or smart_update.sh (refreshes ~/.aris/repo), export ARIS_REPO, or copy the helper to tools/." >&2
  echo "       --style-ref cannot be satisfied; aborting." >&2
  exit 1
}
STYLE_STATUS=0
CAC

Sources accepted: local TeX dir / file, local PDF, arXiv id, http(s) URL. Overleaf URLs/IDs are rejected — clone via /overleaf-sync setup first and pass the local clone path.

Strict rules (full contract in tools/extractpaperstyle.py docstring):

  • Use styleprofile.md only during the fix-implementation phase, to nudge structural choices when applying reviewer feedback. Reviewer feedback always takes precedence; style ref is tie-breaker for how to apply a fix, not whether* to apply it.
  • Never copy prose, claims, examples, or terminology from anything reachable through the cache when implementing fixes.
  • Never pass — style-ref (or the cache contents) to the GPT-6-Astra reviewer sub-agent. The Reviewer Independence Protocol below requires reviewers see only the artifact and the user's prompt — leaking the style ref would contaminate the review with author-side context. This is the most critical invariant in this skill.

Optional: Edit Whitelist (— edit-whitelist , opt-in)

Lets the caller hard-constrain which files and operations the fix-implementation step (Step 3 / Step 6) is allowed to touch. Default OFF — when the user does not pass — edit-whitelist (or the alias — editwhitelist), the loop applies all reviewer-driven edits without restriction, exactly as before.**

This is the parameter that upstream pipelines (e.g. /resubmit-pipeline Phase 2) use to enforce text-only resubmit microedits: no .bib mutations, no .sty / .bst mutations, no edits to prior-submission directories, no new \cite{...}, no new theorem environments, no new numerical claims.

Schema

The whitelist file is YAML or JSON. All four sections are optional:

allowed_paths:
  - sec/*.tex
  - main.tex
  - figures/*.tex
forbidden_paths:
  - "**/*.bib"
  - "**/*.sty"
  - "**/*.bst"
  - "../OldSubmission/**"
forbidden_operations:
  - new_cite              # blocks \cite{...}, \citep{...}, \citet{...}, \citeauthor{...} additions
  - new_bibitem           # blocks \bibitem{...} additions
  - new_theorem_env       # blocks \begin{theorem|lemma|proposition|corollary} additions
  - numerical_claim       # blocks adding new numbers / percentages / metrics
forbidden_deletions:      # operations that block REMOVALS, not additions
  - delete_existing_cite  # blocks removal of \cite{...} from the body (use citation-audit --soft-only instead)
  - delete_theorem_env    # blocks removal of an existing \begin{theorem|...} block
requires_user_approval_for:  # operations that don't auto-reject but pause for explicit user OK
  - rewrite_abstract      # paraphrasing the entire abstract triggers a checkpoint
  - rewrite_intro_first_para
  - delete_section
max_edits_per_round: 30   # hard cap on number of accepted edits per round (rejections are not counted; if cap is hit, remaining proposed edits are deferred to the next round with a warning)
rationale: "Resu

Resolution rules

  • allowedpaths empty AND forbiddenpaths empty → whitelist is a no-op (advisory: the file is loaded and rationale echoed to the log, but no path filtering is applied).
  • allowedpaths empty, forbiddenpaths non-empty → all paths NOT matched by forbidden_paths are mutable.
  • allowedpaths non-empty, forbiddenpaths empty → only paths matching allowed_paths are mutable.
  • Both non-empty → an edit is allowed iff the target matches allowedpaths AND does NOT match forbiddenpaths. forbidden_paths always wins on overlap.
  • forbiddenoperations missing or empty** → no operation-level guard; only path-level filtering applies.

Glob semantics

Use bash extglob / Python fnmatch.fnmatch semantics. matches any depth (zero or more directory segments). Patterns are matched against the path relative to the paper directory (e.g. paper/sec/intro.tex matches sec/.tex when paper-directory is paper/).

Forbidden-operation detectors

For each candidate edit's diff (the new lines being added — deletions are exempt), the loop runs these regex checks and rejects if any forbidden operation matches:

Behavior at loop start (before Round 1 fix-implementation)

  1. If — edit-whitelist is present in $ARGUMENTS, set EDIT_WHITELIST = .
  2. Load the file (yaml.safe_load; if it fails, fall back to json.loads). On load failure, abort the loop with a clear error — do NOT silently proceed unconstrained.
  3. Echo rationale (if present) into PAPERIMPROVEMENTLOG.md under a new "Edit Whitelist" preamble section so the audit trail records why edits were constrained.

Behavior during fix-implementation (Steps 3 and 6)

Before applying each proposed edit:

  1. Resolve target file path relative to the paper directory.
  2. Path check: if allowedpaths is non-empty, target must match at least one pattern. Then if forbiddenpaths is non-empty, target must NOT match any pattern. If either fails → reject as path violation.
  3. Operation check: build the unified diff (or just the set of newly-added lines) for the proposed edit. For each entry in forbidden_operations, run its detector on the added lines. If any detector matches → reject as operation violation.
  4. If all checks pass, apply the edit normally.
  5. If rejected, append an entry to PAPERIMPROVEMENTLOG.md under a ## Rejected by edit_whitelist (Round N) heading with this schema:
- file: <relative path>
     reason: path | operation
     pattern: <the offending forbidden_path glob, OR the offending forbidden_operation name + the matched substring>
     reviewer_concern: <the original Round-N weakness that motivated this edit>
  1. Continue with the remaining edits in the round. Do NOT abort the whole round on a single rejection.

End-of-round surfacing

At the end of each round (after the recompile, before moving to the next round), if any edits were rejected during that round's fix step:

  • Print a one-line summary to the round's checkpoint output: Edit whitelist rejected N edits this round (M path, K operation). See PAPERIMPROVEMENTLOG.md "Rejected by edit_whitelist (Round N)".
  • If HUMAN_CHECKPOINT = true, include the rejection list in the checkpoint shown to the user before they approve next-round fixes.

Example invocations

# Resubmit-pipeline Phase 2 caller (text-only mode):
/auto-paper-improvement-loop "paper/" — edit-whitelist .resubmit/edit_whitelist.yaml

# Aliased form is accepted:
/auto-paper-improvement-loop "paper/" — edit_whitelist .resubmit/edit_whitelist.yaml

# Combined with other flags:
/auto-paper-improvement-loop "paper/" — human checkpoint: true — edit-whitelist constraints.yaml

Rationale

Without a whitelist, the loop's reviewer-driven fix step is free to add citations, introduce new theorem environments, or tweak numerical claims — all of which are reasonable for first-submission polish but forbidden in resubmit / camera-ready / rebuttal-only modes where the paper structure is frozen by external constraint. Routing those constraints through a first-class parameter (rather than relying on the LLM to "remember" not to do them) makes the constraint enforceable, auditable via PAPERIMPROVEMENTLOG.md, and visible to the user at each round's checkpoint.

Inputs

  1. Compiled paper — paper/main.pdf + LaTeX source files
  2. All section .tex files — concatenated for review prompt

State Persistence (Compact Recovery)

If the context window fills up mid-loop, Claude Code auto-compacts. To recover, this skill writes PAPERIMPROVEMENTSTATE.json after each round:

{
  "current_round": 1,
  "threadId": "019ce736-...",
  "last_score": 6,
  "status": "in_progress",
  "timestamp": "2026-03-13T21:00:00"
}

On startup: if PAPERIMPROVEMENTSTATE.json exists with "status": "inprogress" AND timestamp is within 24 hours, read it + PAPERIMPROVEMENT_LOG.md to recover context, then resume from the next round. Otherwise (file absent, "status": "completed", or older than 24 hours), start fresh.

After each round: overwrite the state file. On completion: set "status": "completed".

Reviewer Independence Protocol

The reviewer must be context-naive on every round. Prior-round summaries, fix lists, and executor explanations are not evidence; they are a source of confirmation bias. If the reviewer is told what changed, scores tend to drift upward even when the manuscript itself has not materially improved.

More skills from wanshuiyin/Auto-claude-code-research-in-sleep

  • Aablation-plannerUse when main results pass result-to-claim (claim_supported=yes or partial) and ablation studies are needed for paper submission.
  • Aablation-plannerUse when main results pass result-to-claim (`claim_supported = yes` or `partial`) and ablation studies are needed for paper submission. A secondary Codex agent designs ablations from a reviewer's perspective; the local executor reviews feasibility and implements.
  • AalphaxivQuick single-paper lookup via AlphaXiv LLM-optimized summaries with tiered source fallback. Use when user says "explain this paper", "summarize paper", pastes an arXiv/AlphaXiv URL, or provides a bare arXiv ID for quick understanding - not for broad literature search.
  • AalphaxivQuick single-paper lookup via AlphaXiv LLM-optimized summaries with tiered source fallback. Use when user says "explain this paper", "summarize paper", pastes an arXiv/AlphaXiv URL, or provides a bare arXiv ID for quick understanding - not for broad literature search.
  • Aanalyze-resultsAnalyze ML experiment results, compute statistics, generate comparison tables and insights. Use when user says "analyze results", "compare", or needs to interpret experimental data.
  • Aanalyze-resultsAnalyze ML experiment results, compute statistics, generate comparison tables and insights. Use when user says \"analyze results\", \"compare\", or needs to interpret experimental data.
  • AarxivSearch, download, and summarize academic papers from arXiv. Use when user says "search arxiv", "download paper", "fetch arxiv", "arxiv search", "get paper pdf", or wants to find and save papers from arXiv to the local paper library.
  • AarxivSearch, download, and summarize academic papers from arXiv. Use when user says \"search arxiv\", \"download paper\", \"fetch arxiv\", \"arxiv search\", \"get paper pdf\", or wants to find and save papers from arXiv to the local paper library.
  • Aauto-paper-improvement-loopAutonomously improve a generated paper via Claude review through claude-review MCP → implement fixes → recompile, for 2 rounds. Use when user says \"改论文\", \"improve paper\", \"论文润色循环\", \"auto improve\", or wants to iteratively polish a generated paper.
  • Aauto-paper-improvement-loopAutonomously improve a generated paper via Gemini review through gemini-review MCP → implement fixes → recompile, for 2 rounds. Use when user says \"改论文\", \"improve paper\", \"论文润色循环\", \"auto improve\", or wants to iteratively polish a generated paper.
  • Aauto-paper-improvement-loopAutonomously improve a generated paper via GPT-6-Astra xhigh review → implement fixes → recompile, for 2 rounds. Use when user says \"改论文\", \"improve paper\", \"论文润色循环\", \"auto improve\", or wants to iteratively polish a generated paper.
  • Aauto-review-loopAutonomous multi-round research review loop. In Copilot CLI it defaults to the native complementary rubber-duck subagent with host-event model evidence; elsewhere it uses Codex, while explicit external reviewer overrides remain available. Implements fixes and re-reviews until a policy-approved positive assessment or max rounds is reached.

All agent skills → · MCP servers