research-refine skill
Turn a vague research direction into a problem-anchored, elegant, frontier-aware, implementation-oriented method plan via iterative Gemini review. Use when the user says \"refine my approach\", \"帮我细化方案\", \"decompose this problem\", \"打磨idea\", \"refine research plan\", \"细化研究方案\", or wants a concrete research method that stays simple, focused, and top-venue ready instead of a vague or overbuilt idea.
Is the research-refine skill safe?
Clean: nothing in its files matched our rules. We read 1 file in the folder on 2026-09-28.
No findings.
Install the research-refine skill
A skill is a folder. Copy it into your agent's skills folder and the agent loads it when the task matches its description.
git clone --depth 1 https://github.com/wanshuiyin/Auto-claude-code-research-in-sleep.git /tmp/Auto-claude-code-research-in-sleep mkdir -p ~/.claude/skills cp -r /tmp/Auto-claude-code-research-in-sleep/skills/skills-codex-gemini-review/research-refine ~/.claude/skills/research-refine
In the Claude apps, zip the folder and upload it from the Skills settings. The folder on GitHub
The instructions your agent would load
SKILL.md as published, without the frontmatter. Read it on GitHub
Override for Codex users who want Gemini, not a second Codex agent, to act as the reviewer. Install this package after skills/skills-codex/*.
Research Refine: Problem-Anchored, Elegant, Frontier-Aware Plan Refinement
Gemini overlay assurance: reviewindependence: cross-family and acceptancestatus: accepted.
Refine and concretize: $ARGUMENTS
Overview
Use this skill when the research problem is already visible but the technical route is still fuzzy. The goal is not to produce a bloated proposal or a benchmark shopping list. The goal is to turn a vague direction into a problem -> focused method -> minimal validation document that is concrete enough to implement, elegant enough to feel paper-worthy, and current enough to resonate in the foundation-model era.
Four principles dominate this skill:
- Do not lose the original problem. Freeze an immutable Problem Anchor and reuse it in every round.
- The smallest adequate mechanism wins. Prefer the minimal intervention that directly fixes the bottleneck.
- One paper, one dominant contribution. Prefer one sharp thesis plus at most one supporting contribution.
- Modern leverage is a prior, not a decoration. When LLM / VLM / Diffusion / RL / distillation / inference-time scaling naturally fit the bottleneck, use them concretely. Do not bolt them on as buzzwords.
User input (PROBLEM + vague APPROACH)
-> Phase 0 (Gemini): Freeze Problem Anchor
-> Phase 1 (Gemini): Scan grounding papers -> identify technical gap -> choose the sharpest route -> write focused proposal
-> Phase 2 (Codex/Gemini): Review for fidelity, specificity, contribution quality, and frontier leverage
-> Phase 3 (Gemini): Anchor check + simplicity check -> revise method -> rewrite full proposal
-> Phase 4 (Codex, same Gemini thread): Re-evaluate revised proposal
-> Repeat Phase 3-4 until OVERALL SCORE >= 9 or MAX_ROUNDS reached
-> Phase 5: Save full history to refine-logs/
-> Optional handoff: /experiment-plan for a detailed execution-ready experiment roadmapConstants
- REVIEWERMODEL = gemini-review — Gemini reviewer invoked through the local gemini-review MCP bridge. Set GEMINIREVIEW_MODEL if you need a specific Gemini model override.
- MAXROUNDS = 5** — Maximum review-revise rounds.
- SCORETHRESHOLD = 9** — Minimum overall score to stop.
- OUTPUTDIR = refine-logs/** — Directory for round files and final report.
- MAXLOCALPAPERS = 15 — Maximum local papers/notes to scan for grounding.
- MAXCOREEXPERIMENTS = 3 — Default cap for core validation blocks inside this skill.
- MAXPRIMARYCLAIMS = 2 — Soft cap for paper-level claims. Prefer one dominant claim plus one supporting claim.
- MAXNEWTRAINABLECOMPONENTS = 2** — Soft cap for genuinely new trainable pieces. Exceed only if the paper breaks otherwise.
Override via argument if needed, e.g. /research-refine "problem | approach" -- max rounds: 3, threshold: 9.
Output Structure
refine-logs/
├── round-0-initial-proposal.md
├── round-1-review.md
├── round-1-refinement.md
├── round-2-review.md
├── round-2-refinement.md
├── ...
├── REVIEW_SUMMARY.md
├── FINAL_PROPOSAL.md
├── REFINEMENT_REPORT.md
└── score-history.mdEvery round-N-refinement.md must contain a full anchored proposal, not just incremental fixes.
Workflow
Phase 0: Freeze the Problem Anchor
Before proposing anything, extract the user's immutable bottom-line problem. This anchor must be copied verbatim into every proposal and every refinement round.
Write:
- Bottom-line problem: What technical problem must be solved?
- Must-solve bottleneck: What specific weakness in current methods is unacceptable?
- Non-goals: What is explicitly not the goal of this project?
- Constraints: Compute, data, time, tooling, venue, deployment limits.
- Success condition: What evidence would make the user say "yes, this method addresses the actual problem"?
If later reviewer feedback would change the problem being solved, mark that as drift and push back or adapt carefully.
Phase 1: Build the Initial Proposal
Step 1.1: Scan Grounding Material
Check papers/ and literature/ first. Read only the relevant parts needed to answer:
- What mechanism do current methods use?
- Where exactly do they fail for this problem?
- Which recent LLM / VLM / Diffusion / RL era techniques are actually relevant here?
- What training objectives, representations, or interfaces are reusable?
- What details distinguish a real method from a renamed high-level idea?
If local material is insufficient, search recent top-venue/arXiv work online. Focus on method sections, training setup, and failure modes, not just abstracts.
Step 1.2: Identify the Technical Gap
Do not stop at generic research questions. Make the gap operational:
- Current pipeline failure point: where does the baseline break?
- Why naive fixes are insufficient: larger context, more data, prompting, memory bank, or stacking more modules.
- Smallest adequate intervention: what is the least additional mechanism that could plausibly fix the bottleneck?
- Frontier-native alternative: is there a more current route using foundation-model-era primitives that better matches the bottleneck?
- Core technical claim: what exact mechanism claim could survive top-venue scrutiny?
- Required evidence: what minimum proof is needed to defend that claim?
Step 1.3: Choose the Sharpest Route
Before locking the method, compare two candidate routes if both are plausible:
- Route A: Elegant minimal route — the smallest mechanism that directly targets the bottleneck.
- Route B: Frontier-native route — a more modern route that uses LLM / VLM / Diffusion / RL / distillation / inference-time scaling only if it gives a cleaner or stronger story.
Then decide:
- Which route is more likely to become a strong paper under the stated constraints?
- Which route has the cleaner novelty story relative to the closest work?
- Which route avoids contribution sprawl?
If both routes are weak, rethink the framing instead of combining them into a larger system by default.
Step 1.4: Concretize the Method First
The proposal must answer "how would we actually build this?" Prefer method detail over broad experimentation and prefer reuse over invention.
Cover:
- One-sentence method thesis: the single strongest mechanism claim.
- Contribution focus: one dominant contribution and at most one supporting contribution.
- Complexity budget: what is frozen or reused, what is new, and what tempting additions are intentionally excluded.
- System graph: modules, data flow, inputs, outputs.
- Representation design: what latent, embedding, plan token, reward signal, memory state, or alignment space is used?
- Training recipe: data source, supervision, pseudo-labeling, negatives, curriculum, losses, weighting, stagewise vs joint training.
- Inference path: how the trained components are used at test time and what signals flow where.
- Why the mechanism stays small: why a larger stack is unnecessary.
- Exact role of any frontier primitive: if you use an LLM / VLM / Diffusion / RL component, specify whether it acts as planner, teacher, critic, reward model, generator prior, search controller, or distillation source.
- Failure handling: what could go wrong and what fallback or diagnostic exists?
- Novelty and elegance argument: why this is more than naming a module and why the paper still looks focused.
If the method is still only described as "add a module" or "use a planner," it is not concrete enough.
Step 1.5: Design Minimal Claim-Driven Validation
Experiments exist to validate the method, not to dominate the document.
For each core claim, define the smallest strong experiment that can validate it:
- the claim being tested
- the necessary baseline or ablation
- the decisive metric
- the expected directional outcome
Additional rules:
- Ensure one experiment block directly supports the Problem Anchor.
- If complexity risk exists, include one simplification or deletion check.
- If a frontier primitive is central, include one necessity check showing why that choice matters.
- Default to 1-3 core experiment blocks and leave the full execution roadmap to /experiment-plan.
Step 1.6: Write the Initial Proposal
Save to refine-logs/round-0-initial-proposal.md.
Use this structure:
# Research Proposal: [Title]
## Problem Anchor
- Bottom-line problem:
- Must-solve bottleneck:
- Non-goals:
- Constraints:
- Success condition:
## Technical Gap
[Why current methods fail, why naive bigger systems are not enough, and what mechanism is missing]
## Method Thesis
- One-sentence thesis:
- Why this is the smallest adequate intervention:
- Why this route is timely in the foundation-model era:
## Contribution Focus
- Dominant contribution:
- Optional supporting contribution:
- Explicit non-contributions:
## Proposed Method
### Complexity Budget
- Frozen / reused backbone:
- New trainable components:
- Tempting additions intentionally not used:
### System Overview
[Step-by-step pipeline or ASCII graph]
### Core Mechanism
- Input / output:
- Architecture or policy:
- Training signal / loss:
- Why this is the main novelty:
### Optional Supporting Component
- Only include if truly necessary:
- Input / output:
- Training signal / loss:
- Why it does not create contribution sprawl:
### Modern Primitive Usage
- Which LLM / VLM / Diffusion / RL-era primitive is used:
- Exact role in the pipeline:
- Why it is more natural than an old-school alternative:
### Integration inPhase 2: External Method Review (Round 1)
Send the full proposal to Gemini for an elegance-first, frontier-aware, method-first review. The reviewer should spend most of the critique budget on the method itself, not on expanding the experiment menu.
mcp__gemini-review__review_start:
prompt: |
You are a senior ML reviewer for a top venue (NeurIPS/ICML/ICLR).
This is an early-stage, method-first research proposal.
Your job is NOT to reward extra modules, contribution sprawl, or a giant benchmark checklist.
Your job IS to stress-test whether the proposed method:
(1) still solves the original anchored problem,
(2) is concrete enough to implement,
(3) presents a focused, elegant contribution,
(4) uses foundation-model-era techniques appropriately when they are the natural fit.
Review principles:
- Prefer the smallest adequate mechanism over a larger system.
- Penalize parallel contributions that make the paper feel unfocused.
- If a modern LLM / VLM / Diffusion / RL route would clearly produce a better paper, say so concretely.
- If the proposal is already modern enough, do NOT force trendy components.
- Do not ask for extra experiments unless they are needed to prove the core claims.
Read the Problem Anchor first. If your suggested fix would change the problem being solved,
call that out explicitly as drift instead of treating it as a normal revision request.
After this start call, immediately save the returned jobId and poll mcpgemini-reviewreview_status with a bounded waitSeconds until done=true. Treat the completed status payload's response as the reviewer output, and save the completed threadId for any follow-up round.
CRITICAL: Save the returned jobId, poll mcpgemini-reviewreview_status until done=true, then save the completed threadId from the status result for all later rounds.
CRITICAL: Save the FULL raw response verbatim.
Save review to refine-logs/round-1-review.md with the raw response in a block.
Phase 3: Parse Feedback and Revise the Method
Step 3.1: Parse the Review
Extract:
More skills from wanshuiyin/Auto-claude-code-research-in-sleep
- Aablation-plannerUse when main results pass result-to-claim (claim_supported=yes or partial) and ablation studies are needed for paper submission.
- Aablation-plannerUse when main results pass result-to-claim (`claim_supported = yes` or `partial`) and ablation studies are needed for paper submission. A secondary Codex agent designs ablations from a reviewer's perspective; the local executor reviews feasibility and implements.
- AalphaxivQuick single-paper lookup via AlphaXiv LLM-optimized summaries with tiered source fallback. Use when user says "explain this paper", "summarize paper", pastes an arXiv/AlphaXiv URL, or provides a bare arXiv ID for quick understanding - not for broad literature search.
- AalphaxivQuick single-paper lookup via AlphaXiv LLM-optimized summaries with tiered source fallback. Use when user says "explain this paper", "summarize paper", pastes an arXiv/AlphaXiv URL, or provides a bare arXiv ID for quick understanding - not for broad literature search.
- Aanalyze-resultsAnalyze ML experiment results, compute statistics, generate comparison tables and insights. Use when user says "analyze results", "compare", or needs to interpret experimental data.
- Aanalyze-resultsAnalyze ML experiment results, compute statistics, generate comparison tables and insights. Use when user says \"analyze results\", \"compare\", or needs to interpret experimental data.
- AarxivSearch, download, and summarize academic papers from arXiv. Use when user says "search arxiv", "download paper", "fetch arxiv", "arxiv search", "get paper pdf", or wants to find and save papers from arXiv to the local paper library.
- AarxivSearch, download, and summarize academic papers from arXiv. Use when user says \"search arxiv\", \"download paper\", \"fetch arxiv\", \"arxiv search\", \"get paper pdf\", or wants to find and save papers from arXiv to the local paper library.
- Aauto-paper-improvement-loopAutonomously improve a generated paper via GPT-6-Astra xhigh review → implement fixes → recompile, for 2 rounds. Use when user says \"改论文\", \"improve paper\", \"论文润色循环\", \"auto improve\", or wants to iteratively polish a generated paper.
- Aauto-paper-improvement-loopAutonomously improve a generated paper via Claude review through claude-review MCP → implement fixes → recompile, for 2 rounds. Use when user says \"改论文\", \"improve paper\", \"论文润色循环\", \"auto improve\", or wants to iteratively polish a generated paper.
- Aauto-paper-improvement-loopAutonomously improve a generated paper via Gemini review through gemini-review MCP → implement fixes → recompile, for 2 rounds. Use when user says \"改论文\", \"improve paper\", \"论文润色循环\", \"auto improve\", or wants to iteratively polish a generated paper.
- Aauto-paper-improvement-loopAutonomously improve a generated paper via GPT-6-Astra xhigh review → implement fixes → recompile, for 2 rounds. Use when user says \"改论文\", \"improve paper\", \"论文润色循环\", \"auto improve\", or wants to iteratively polish a generated paper.