auto-paper-improvement-loop skill
Autonomously improve a generated paper via Gemini review through gemini-review MCP → implement fixes → recompile, for 2 rounds. Use when user says \"改论文\", \"improve paper\", \"论文润色循环\", \"auto improve\", or wants to iteratively polish a generated paper.
Is the auto-paper-improvement-loop skill safe?
Clean: nothing in its files matched our rules. We read 1 file in the folder on 2026-09-28.
No findings.
Install the auto-paper-improvement-loop skill
A skill is a folder. Copy it into your agent's skills folder and the agent loads it when the task matches its description.
git clone --depth 1 https://github.com/wanshuiyin/Auto-claude-code-research-in-sleep.git /tmp/Auto-claude-code-research-in-sleep mkdir -p ~/.claude/skills cp -r /tmp/Auto-claude-code-research-in-sleep/skills/skills-codex-gemini-review/auto-paper-improvement-loop ~/.claude/skills/auto-paper-improvement-loop
In the Claude apps, zip the folder and upload it from the Skills settings. The folder on GitHub
The instructions your agent would load
SKILL.md as published, without the frontmatter. Read it on GitHub
Override for Codex users who want Gemini, not a second Codex agent, to act as the reviewer. Install this package after skills/skills-codex/*.
Auto Paper Improvement Loop: Review → Fix → Recompile
Gemini overlay assurance: reviewindependence: cross-family and acceptancestatus: accepted.
Autonomously improve the paper at: $ARGUMENTS
Context
This skill is designed to run after Workflow 3 (/paper-plan → /paper-figure → /paper-write → /paper-compile). It takes a compiled paper and iteratively improves it through external LLM review.
Unlike /auto-review-loop (which iterates on research — running experiments, collecting data, rewriting narrative), this skill iterates on paper writing quality — fixing theoretical inconsistencies, softening overclaims, adding missing content, and improving presentation.
Constants
- MAXROUNDS = 2** — Two rounds of review→fix→recompile. Empirically, Round 1 catches structural issues (4→6/10), Round 2 catches remaining presentation issues (6→7/10). Diminishing returns beyond 2 rounds for writing-only improvements.
- REVIEWERMODEL = gemini-review — Gemini reviewer invoked through the local gemini-review MCP bridge. Set GEMINIREVIEW_MODEL if you need a specific Gemini model override.
- REVIEWLOG = PAPERIMPROVEMENTLOG.md** — Cumulative log of all rounds, stored in paper directory.
- HUMANCHECKPOINT = false** — When true, pause after each round's review and present score + weaknesses to the user. The user can approve fixes, provide custom modification instructions, skip specific fixes, or stop early. When false (default), runs fully autonomously.
💡 Override: /auto-paper-improvement-loop "paper/" — human checkpoint: true
Inputs
- Compiled paper — paper/main.pdf + LaTeX source files
- All section .tex files — concatenated for review prompt
State Persistence (Compact Recovery)
If the context window fills up mid-loop, Codex auto-compacts. To recover, this skill writes PAPERIMPROVEMENTSTATE.json after each round:
{
"current_round": 1,
"thread_id": "019ce736-...",
"last_score": 6,
"status": "in_progress",
"timestamp": "2026-03-13T21:00:00"
}On startup: if PAPERIMPROVEMENTSTATE.json exists with "status": "inprogress" AND timestamp is within 24 hours, read it + PAPERIMPROVEMENT_LOG.md to recover context, then resume from the next round. Otherwise (file absent, "status": "completed", or older than 24 hours), start fresh.
After each round: overwrite the state file. On completion: set "status": "completed".
Workflow
Step 0: Preserve Original
cp paper/main.pdf paper/main_round0_original.pdfStep 1: Collect Paper Text
Concatenate all section files into a single text block for the review prompt:
# Collect all sections in order
for f in paper/sections/*.tex; do
echo "% === $(basename $f) ==="
cat "$f"
done > /tmp/paper_full_text.txtStep 2: Round 1 Review
Send the full paper text to Gemini review:
mcp__gemini-review__review_start:
prompt: |
You are reviewing a [VENUE] paper. Please provide a detailed, structured review.
Judge claim calibration in BOTH directions. Recommend narrowing only when the
current scope or modality exceeds the evidence; do not ask for extra hedges
around a supported result. Flag stacked hedges, self-defence ("we do not
claim"), instruction confessions ("we do not address X"), and generic caveats
outside Limitations as writing defects to remove. Tone fixes must never alter
facts, negation, modality, scope, comparison direction, or numbers.
Also flag narrative defects: a progress-report structure ("we first tried
A, then B"), a story built on a metric the method loses, results narrated
as defeats ("underperforms", "fails to surpass") instead of explained as a
goal difference or tradeoff, experiments with no argumentative duty, an
abstract or introduction that opens on background or implementation
instead of problem -> gap -> idea -> strongest result, and a conclusion
that ends on new self-negation. The fix is reframing and cutting where
the evidence supports the reframing; a genuine weaknessAfter this start call, immediately save the returned jobId and poll mcpgemini-reviewreview_status with a bounded waitSeconds until done=true. Treat the completed status payload's response as the reviewer output, and save the completed threadId for any follow-up round.
Save the returned jobId, poll mcpgemini-reviewreview_status until done=true, then save the completed threadId for Round 2.
Step 2b: Human Checkpoint (if enabled)
Skip if HUMANCHECKPOINT = false.**
Present the review results and wait for user input:
📋 Round 1 review complete.
Score: X/10 — [verdict]
Key weaknesses (by severity):
1. [CRITICAL] ...
2. [MAJOR] ...
3. [MINOR] ...
Reply "go" to implement all fixes, give custom instructions, "skip 2" to skip specific fixes, or "stop" to end.Parse user response same as /auto-review-loop: approve / custom instructions / skip / stop.
Step 3: Implement Round 1 Fixes
Parse the review and implement fixes by severity:
Priority order:
- CRITICAL fixes (assumption mismatches, internal contradictions)
- MAJOR fixes (overclaims, missing content, notation issues)
- MINOR fixes (if time permits)
Before applying any fix: calibrate claims to evidence and state them directly; generic caveats belong in Limitations only; writing instructions are never manuscript content; tone edits never change what the paper knows.
Common fix patterns:
Step 4: Recompile Round 1
cd paper && latexmk -C && latexmk -pdf -interaction=nonstopmode -halt-on-error main.tex
cp main.pdf main_round1.pdfVerify: 0 undefined references, 0 undefined citations.
Step 5: Round 2 Review
Start a fresh mcpgemini-reviewreview_start — do not reuse the Round 1 thread, do not send fix summaries. The reviewer re-reads the recompiled paper cold; executor notes are not evidence, and "since your last review, we implemented..." is exactly the self-report path that lets fixes be graded on their description instead of their substance:
mcp__gemini-review__review_start:
prompt: |
You are reviewing a [VENUE] paper (Round 2 — fresh read of the current
manuscript; judge only what is on the page).
Judge claim calibration in BOTH directions. Recommend narrowing only when the
current scope or modality exceeds the evidence; do not ask for extra hedges
around a supported result. Flag stacked hedges, self-defence ("we do not
claim"), instruction confessions ("we do not address X"), and generic caveats
outside Limitations as writing defects to remove. Tone fixes must never alter
facts, negation, modality, scope, comparison direction, or numbers.
Also flag narrative defects: a progress-report structure ("we first tried
A, then B"), a story built on a metric the method loses, results narrated
as defeats ("underperforms", "fails to surpass") instead of explained as a
goal difference or tradeoff, experiments with no argumentative duty, an
abstract or introduction that opens on background or implementation
instead of problem -> gap -> idea -> strongest result, and a conclusion
that ends on new self-negation. The fix is reframing and cutting where
the evidence supAfter this start call, immediately save the returned jobId and poll mcpgemini-reviewreview_status with a bounded waitSeconds until done=true. Treat the completed status payload's response as the reviewer output, and save the completed threadId for any follow-up round.
Step 5b: Human Checkpoint (if enabled)
Skip if HUMANCHECKPOINT = false.** Same as Step 2b — present Round 2 review, wait for user input.
Step 6: Implement Round 2 Fixes
Same process as Step 3. Typical Round 2 fixes:
supported claims directly, consolidate scattered generic caveats into Limitations — and do not re-soften claims the evidence already supports
- Add controlled synthetic experiments validating theory
- Re-check calibration in both directions: narrow genuine overclaims, state
- Formalize informal arguments (e.g., truncation → formal proposition)
- Make Limitations more specific only when a material limit is missing
Step 7: Recompile Round 2
cd paper && latexmk -C && latexmk -pdf -interaction=nonstopmode -halt-on-error main.tex
cp main.pdf main_round2.pdfStep 8: Format Check
After the final recompilation, run a format compliance check:
# 1. Page count vs venue limit
PAGES=$(pdfinfo paper/main.pdf | grep Pages | awk '{print $2}')
echo "Pages: $PAGES (limit: 9 main body for ICLR/NeurIPS)"
# 2. Overfull hbox warnings (content exceeding margins)
OVERFULL=$(grep -c "Overfull" paper/main.log 2>/dev/null || echo 0)
echo "Overfull hbox warnings: $OVERFULL"
grep "Overfull" paper/main.log 2>/dev/null | head -10
# 3. Underfull hbox warnings (loose spacing)
UNDERFULL=$(grep -c "Underfull" paper/main.log 2>/dev/null || echo 0)
echo "Underfull hbox warnings: $UNDERFULL"
# 4. Bad boxes summary
grep -c "badness" paper/main.log 2>/dev/null || echo "0 badness warnings"Auto-fix patterns:
If any overfull hbox > 10pt is found, fix it and recompile before documenting.
More skills from wanshuiyin/Auto-claude-code-research-in-sleep
- Aablation-plannerUse when main results pass result-to-claim (claim_supported=yes or partial) and ablation studies are needed for paper submission.
- Aablation-plannerUse when main results pass result-to-claim (`claim_supported = yes` or `partial`) and ablation studies are needed for paper submission. A secondary Codex agent designs ablations from a reviewer's perspective; the local executor reviews feasibility and implements.
- AalphaxivQuick single-paper lookup via AlphaXiv LLM-optimized summaries with tiered source fallback. Use when user says "explain this paper", "summarize paper", pastes an arXiv/AlphaXiv URL, or provides a bare arXiv ID for quick understanding - not for broad literature search.
- AalphaxivQuick single-paper lookup via AlphaXiv LLM-optimized summaries with tiered source fallback. Use when user says "explain this paper", "summarize paper", pastes an arXiv/AlphaXiv URL, or provides a bare arXiv ID for quick understanding - not for broad literature search.
- Aanalyze-resultsAnalyze ML experiment results, compute statistics, generate comparison tables and insights. Use when user says "analyze results", "compare", or needs to interpret experimental data.
- Aanalyze-resultsAnalyze ML experiment results, compute statistics, generate comparison tables and insights. Use when user says \"analyze results\", \"compare\", or needs to interpret experimental data.
- AarxivSearch, download, and summarize academic papers from arXiv. Use when user says "search arxiv", "download paper", "fetch arxiv", "arxiv search", "get paper pdf", or wants to find and save papers from arXiv to the local paper library.
- AarxivSearch, download, and summarize academic papers from arXiv. Use when user says \"search arxiv\", \"download paper\", \"fetch arxiv\", \"arxiv search\", \"get paper pdf\", or wants to find and save papers from arXiv to the local paper library.
- Aauto-paper-improvement-loopAutonomously improve a generated paper via GPT-6-Astra xhigh review → implement fixes → recompile, for 2 rounds. Use when user says \"改论文\", \"improve paper\", \"论文润色循环\", \"auto improve\", or wants to iteratively polish a generated paper.
- Aauto-paper-improvement-loopAutonomously improve a generated paper via Claude review through claude-review MCP → implement fixes → recompile, for 2 rounds. Use when user says \"改论文\", \"improve paper\", \"论文润色循环\", \"auto improve\", or wants to iteratively polish a generated paper.
- Aauto-paper-improvement-loopAutonomously improve a generated paper via GPT-6-Astra xhigh review → implement fixes → recompile, for 2 rounds. Use when user says \"改论文\", \"improve paper\", \"论文润色循环\", \"auto improve\", or wants to iteratively polish a generated paper.
- Aauto-review-loopAutonomous multi-round research review loop. In Copilot CLI it defaults to the native complementary rubber-duck subagent with host-event model evidence; elsewhere it uses Codex, while explicit external reviewer overrides remain available. Implements fixes and re-reviews until a policy-approved positive assessment or max rounds is reached.