Mmcp.market

research-pipeline skill

by wanshuiyin·wanshuiyin/Auto-claude-code-research-in-sleep·17k stars·MIT

Full end-to-end research pipeline: from a broad research direction through idea discovery, experiments, and review all the way to a polished paper PDF. Use when user says \"全流程\", \"full pipeline\", \"从找idea到投稿\", \"end-to-end research\", or wants the complete autonomous research lifecycle.

A100/100content scan

Is the research-pipeline skill safe?

Clean: nothing in its files matched our rules. We read 1 file in the folder on 2026-09-28.

No findings.

Install the research-pipeline skill

A skill is a folder. Copy it into your agent's skills folder and the agent loads it when the task matches its description.

git clone --depth 1 https://github.com/wanshuiyin/Auto-claude-code-research-in-sleep.git /tmp/Auto-claude-code-research-in-sleep
mkdir -p ~/.claude/skills
cp -r /tmp/Auto-claude-code-research-in-sleep/skills/research-pipeline ~/.claude/skills/research-pipeline
available in every project

In the Claude apps, zip the folder and upload it from the Skills settings. The folder on GitHub

The instructions your agent would load

SKILL.md as published, without the frontmatter. Read it on GitHub

Full Research Pipeline: Idea → Experiments → Submission

⏱ External cadence: non-judgmental heartbeat only. An overnight /loop /

CronCreate heartbeat may wake, detect a stalled phase (no progress, dead

process, blocked on a freed resource) and nudge it forward — it may NEVER

decide the work is good (paper good enough, proof holds, claim supported).

Every such verdict stays on its own skill's internal cadence and terminates in

the cross-model jury. A heartbeat may say "keep going," never "good enough."

See

shared-references/external-cadence.md

(overnight-pipeline rule + stall detection & forced structural pivot). At heartbeat

startup, touch the run state first each tick and register this run with the watchdog

loop type (so a silent death surfaces as STALE); unregister on completion. The

watchdog only detects — it never acquits. Each tick also record the new-finding count

via the iteration_log.py helper (resolve through the canonical

.aris/tools → tools → $ARISREPO/tools → $ARISREPO/tools via ~/.aris/repo

chain, integration-contract §2; warn-and-skip if unresolved):

python3 "$ITER_LOG" note . On the returned

pivot=structural (stale ≥ 2) the nudge must change a STRUCTURAL constraint and pick an

untried direction; on pivot=human (stale ≥ 4) flag for attention. Counting only —

never a quality verdict.

End-to-end autonomous research workflow for: $ARGUMENTS

Constants

  • AUTOPROCEED = true** — When true, every selection checkpoint is informational: report the choice and continue in the same turn. When false, ask for explicit user confirmation and end the turn at the checkpoint.
  • ARXIVDOWNLOAD = false** — When true, /research-lit downloads the top relevant arXiv PDFs during literature survey. When false (default), only fetches metadata via arXiv API. Passed through to /idea-discovery → /research-lit.
  • HUMANCHECKPOINT = false** — When true, the auto-review loops (Stage 3) pause after each round's review to let you see the score and provide custom modification instructions before fixes are implemented. When false (default), loops run fully autonomously. Passed through to /auto-review-loop.
  • REVIEWERDIFFICULTY = medium** — How adversarial the reviewer is. medium (default): standard MCP review. hard: adds reviewer memory + debate protocol. nightmare: GPT reads repo directly via codex exec + memory + debate. Passed through to /auto-review-loop.
  • CODEREVIEW = true** — GPT-6-Astra xhigh reviews experiment code before deployment. Catches logic bugs before wasting GPU hours. Set false to skip. Passed through to /experiment-bridge.
  • BASEREPO = false** — GitHub repo URL to use as base codebase. When set, /experiment-bridge clones the repo first and implements experiments on top of it. When false (default), writes code from scratch or reuses existing project files. Passed through to /experiment-bridge.
  • COMPACT = false — When true, generates compact summary files for short-context models and session recovery. Passed through to /idea-discovery and /experiment-bridge.
  • AUTOWRITE = false — When true, automatically invoke Workflow 3 (/paper-writing) after Stage 4. VENUE is needed only when Stage 5 begins — a missing venue defers paper writing; it never blocks Stages 1-4. When false (default), Stage 4 generates NARRATIVEREPORT.md and stops — user invokes /paper-writing manually.
  • VENUE = (unset) — Target venue for paper writing; bound only when Stage 5 begins. Options: ICLR, NeurIPS, ICML, CVPR, ACL, AAAI, ACM, IEEECONF, IEEEJOURNAL. No default: a missing venue defers paper writing — it never blocks Stages 1-4 and is never guessed.
  • RENDERHTML = true — When true (default), auto-render NARRATIVEREPORT.md to HTML at Stage 4 completion via /render-html. Uses --no-review (this is an internal handoff doc to /paper-writing, not a reviewer-facing final artifact — the upstream Stage 3 auto-review loop already cross-model-reviewed the claims). Set false to skip, or pass — render html: false. Non-blocking: if /render-html fails or Codex MCP is unavailable, log the failure and continue — the HTML view is a nice-to-have, not a Stage 4 prerequisite.
  • RESUMABLE = true — When true (default), the pipeline records per-stage state to .aris/runs/.json so a crashed/interrupted run can resume via /research-pipeline — resume instead of restarting. Stage status splits done (executor finished writing) from accepted (the stage's cross-model gate / deterministic verifier passed); resume re-validates any done-but-unaccepted stage. See shared-references/resumable-runs.md.

💡 Override via argument, e.g., /research-pipeline "topic" — AUTOPROCEED: false, human checkpoint: true, difficulty: nightmare, code review: false, base repo: https://github.com/org/project, autowrite: true, venue: NeurIPS.

Checkpoint execution rule

Resolve AUTO_PROCEED once from $ARGUMENTS before Stage 1 and pass that resolved value to nested workflows.

not a question. State the result and the automatically selected next action, then continue executing in the same turn. Do not ask for confirmation, request user input, sleep, wait for silence, or end the turn at a checkpoint.

  • AUTOPROCEED=true is non-blocking.** A checkpoint is a progress update,

end the turn. Resume only after an explicit reply.

  • AUTOPROCEED=false is blocking.** Present the options, ask the user, and

Never implement auto-proceed as “ask, then continue if there is no response.” Once a turn ends, silence cannot resume the pipeline. The user can still interrupt a non-blocking run at any time.

This rule governs only AUTOPROCEED-controlled selection checkpoints. If the user explicitly enables a Feishu interactive gate, that external approval or reply is an intentional blocking exception; wait for that user-controlled gate rather than treating it as a silence timeout. Feishu off/push-only modes remain non-blocking under AUTOPROCEED=true.

Overview

This skill chains the entire research lifecycle into a single pipeline:

/idea-discovery → /experiment-bridge → /auto-review-loop → /paper-writing (optional)
├── Workflow 1 ──┤├── Workflow 1.5 ──┤├── Workflow 2 ───┤ ├── Workflow 3 ──┤

It orchestrates up to four major workflows in sequence. Workflow 3 (paper writing) is optional and controlled by AUTO_WRITE.

Resumable runs (— resume )

This pipeline is long and can fail mid-run; it tracks per-stage state via run_state.py so you can resume instead of restarting (see shared-references/resumable-runs.md). Skip this whole section if RESUMABLE = false.

Resolve the helper via the canonical chain (integration-contract §2): .aris/tools/runstate.py → tools/runstate.py → $ARISREPO/tools/runstate.py → $ARISREPO/tools/runstate.py via ~/.aris/repo (warn-and-skip if unresolved — never block the pipeline).

Phases, in order: idea-discovery, experiment-bridge, auto-review-loop, summary, paper-writing.

runstate.py resume — it prints the first non-accepted phase; begin the pipeline at that stage (re-run a running/failed stage; re-audit a done-but-unaccepted stage). Otherwise derive from the direction slug + date and runstate.py start --phases "idea-discovery,experiment-bridge,auto-review-loop,summary,paper-writing".

  • At start: if — resume was passed, run

done --artifact once the stage's artifact is written.

  • Per stage: set running on entry; set

own say-so (run_state.py accept requires a recorded verdict id + reviewer):

  • Mark accepted ONLY after the stage's gate passes — never on the executor's

If AUTOWRITE = false (default), paper-writing is not part of this run: after summary is accepted, set paper-writing skipped so resume reports COMPLETE instead of pointing forever at a pending stage. Record each accept verdictid as a durable handle — the codex thread/trace id, or the path/sha of the deterministic verifier's report (e.g. the verifypaperaudits.sh output JSON) — not just the reviewer label.

A stage left done (gate failed/ambiguous, or the run crashed before the gate) is re-validated on the next resume — the acceptance obligation is never skipped.

Overnight heartbeat: stall detection → forced structural pivot

Only when an unattended heartbeat is driving this run (overnight /loop / CronCreate). Skip otherwise. Doctrine + rationale: shared-references/external-cadence.md → "Stall detection & forced structural pivot". This is a Type-A signal — it counts findings and changes direction, never judges quality.

Resolve the helper via the canonical chain (integration-contract §2), warn-and-skip if unresolved (never block the run):

ITER_LOG=".aris/tools/iteration_log.py"
[ -f "$ITER_LOG" ] || ITER_LOG="tools/iteration_log.py"
[ -f "$ITER_LOG" ] || ITER_LOG="${ARIS_REPO:-}/tools/iteration_log.py"
[ -f "$ITER_LOG" ] || { [ -z "${ARIS_REPO:-}" ] && [ -f "$HOME/.aris/repo" ] && ARIS_REPO="$(cat "$HOME/.aris/repo" 2>/dev/null)"; } || true
[ -f "$ITER_LOG" ] || ITER_LOG="${ARIS_REPO:-}/tools/iteration_log.py"
[ -f "$ITER_LOG" ] || { echo "WARN: iteration_log.py not resolved; skipping stall detection" >&2; ITER_LOG=""; }

Then, each heartbeat tick, record how many concrete new findings the current stage produced and read the returned pivot:

[ -n "$ITER_LOG" ] && python3 "$ITER_LOG" note "$ROOT" "$RUN_ID" "$STAGE" "$N_NEW_FINDINGS"
# → {"stale_count": N, "pivot": "none|structural|human"}

Act on pivot:

(frame / objective / data / representation), not a tactical parameter, and pick a direction different from every one already tried. Record the chosen frame so future ticks can avoid it: python3 "$ITERLOG" note "$ROOT" "$RUNID" "$STAGE" 0 --direction "".

  • none — keep going.
  • structural (stale ≥ 2) — the next nudge must change a structural constraint

not silently abandon).

  • human (stale ≥ 4) — stop nudging blindly; flag for human attention (escalate, do

More skills from wanshuiyin/Auto-claude-code-research-in-sleep

  • Aablation-plannerUse when main results pass result-to-claim (claim_supported=yes or partial) and ablation studies are needed for paper submission.
  • Aablation-plannerUse when main results pass result-to-claim (`claim_supported = yes` or `partial`) and ablation studies are needed for paper submission. A secondary Codex agent designs ablations from a reviewer's perspective; the local executor reviews feasibility and implements.
  • AalphaxivQuick single-paper lookup via AlphaXiv LLM-optimized summaries with tiered source fallback. Use when user says "explain this paper", "summarize paper", pastes an arXiv/AlphaXiv URL, or provides a bare arXiv ID for quick understanding - not for broad literature search.
  • AalphaxivQuick single-paper lookup via AlphaXiv LLM-optimized summaries with tiered source fallback. Use when user says "explain this paper", "summarize paper", pastes an arXiv/AlphaXiv URL, or provides a bare arXiv ID for quick understanding - not for broad literature search.
  • Aanalyze-resultsAnalyze ML experiment results, compute statistics, generate comparison tables and insights. Use when user says "analyze results", "compare", or needs to interpret experimental data.
  • Aanalyze-resultsAnalyze ML experiment results, compute statistics, generate comparison tables and insights. Use when user says \"analyze results\", \"compare\", or needs to interpret experimental data.
  • AarxivSearch, download, and summarize academic papers from arXiv. Use when user says "search arxiv", "download paper", "fetch arxiv", "arxiv search", "get paper pdf", or wants to find and save papers from arXiv to the local paper library.
  • AarxivSearch, download, and summarize academic papers from arXiv. Use when user says \"search arxiv\", \"download paper\", \"fetch arxiv\", \"arxiv search\", \"get paper pdf\", or wants to find and save papers from arXiv to the local paper library.
  • Aauto-paper-improvement-loopAutonomously improve a generated paper via GPT-6-Astra xhigh review → implement fixes → recompile, for 2 rounds. Use when user says \"改论文\", \"improve paper\", \"论文润色循环\", \"auto improve\", or wants to iteratively polish a generated paper.
  • Aauto-paper-improvement-loopAutonomously improve a generated paper via Claude review through claude-review MCP → implement fixes → recompile, for 2 rounds. Use when user says \"改论文\", \"improve paper\", \"论文润色循环\", \"auto improve\", or wants to iteratively polish a generated paper.
  • Aauto-paper-improvement-loopAutonomously improve a generated paper via Gemini review through gemini-review MCP → implement fixes → recompile, for 2 rounds. Use when user says \"改论文\", \"improve paper\", \"论文润色循环\", \"auto improve\", or wants to iteratively polish a generated paper.
  • Aauto-paper-improvement-loopAutonomously improve a generated paper via GPT-6-Astra xhigh review → implement fixes → recompile, for 2 rounds. Use when user says \"改论文\", \"improve paper\", \"论文润色循环\", \"auto improve\", or wants to iteratively polish a generated paper.

All agent skills → · MCP servers