research-implement-feature skill
Build a working artifact from a plain \"implement X for me\" request: a running end-to-end spine first, then one feature per rung, with every under-determined decision written to an assumption ledger BEFORE the code that depends on it and a cross-model sweep for the ones that slipped through undeclared. Use when user says \"给我实现\", \"implement X\", \"帮我做一个能跑的\", \"先搭个原型再加功能\", \"build this feature\", \"prototype then extend\", or hands over a capability description rather than an experiment plan.
Is the research-implement-feature skill safe?
Clean: nothing in its files matched our rules. We read 1 file in the folder on 2026-09-28.
No findings.
Install the research-implement-feature skill
A skill is a folder. Copy it into your agent's skills folder and the agent loads it when the task matches its description.
git clone --depth 1 https://github.com/wanshuiyin/Auto-claude-code-research-in-sleep.git /tmp/Auto-claude-code-research-in-sleep mkdir -p ~/.claude/skills cp -r /tmp/Auto-claude-code-research-in-sleep/skills/research-implement-feature ~/.claude/skills/research-implement-feature
In the Claude apps, zip the folder and upload it from the Skills settings. The folder on GitHub
The instructions your agent would load
SKILL.md as published, without the frontmatter. Read it on GitHub
Research Implement: Feature
Build: $ARGUMENTS
This skill exists for one request shape — "just implement X for me" — where the user has a capability in mind, not an experiment plan, and does not want to be interviewed about it first.
It resolves that request the only honest way: stay autonomous, stop being silent. The skill never blocks to ask permission; it declares every decision the request left open, in a ledger, at the moment it makes it, and then a different model family goes looking for the ones it forgot to declare.
Two invariants
request and changes an interface or a meaning, it gets a ledger row — before the code that depends on it exists. A ledger reconstructed at the end of the run is not a ledger, it is a changelog, and it systematically omits exactly the assumptions the author stopped noticing.
- Declare before you act. The instant a decision is under-determined by the
Under ASK=semantic, this invariant strengthens to ask before you act for the semantic class: the ledger row is the unit of ambiguity, so a row that would have been written silently is a question that gets asked first.
from real entry point to real artifact, with stubs inside. It must run before any feature is added. Features are then added one rung at a time, each with its own acceptance check, each leaving every earlier rung green.
- Spine before features. Rung F0 is a walking skeleton: the thinnest path
Scope boundary
Relationship to /research-pipeline
/research-pipeline answers "what should we research?" and decides the question for you. This skill answers "build the thing I already decided on" and decides nothing of consequence without writing it down. Different input contracts, so they are different entry points rather than a mode flag — but they compose: a pipeline run may delegate its build stage here instead of inlining implementation, and inherits the ledger as a result.
If the target decomposes into more than the rung budget below, the scope is too large for one run. Cut to the MUST rungs and record the rest under Deferred in the build note — do not quietly grow this skill into a system build.
Constants
- EFFORT = balanced — Work intensity per shared-references/effort-contract.md. Override: — effort: max.
EFFORT never lowers the reviewer tier — a hard invariant of the effort contract.
before they are acted on.
- ASK = never — Interaction mode: which ambiguities are put to the author
ASK never changes what lands in the ledger — only who decided each row. Every row records its Source, so the record is complete in both modes.
- ASSURANCE — derived from EFFORT per the effort contract (lite/balanced → draft, max/beast → submission). Governs whether Phase 4 blocks. Override: — assurance: submission.
- BASEREPO = false** — Repo URL to build on top of. When set, clone first and implement inside it, matching its conventions. When false, extend the current project or create files in it.
- Output language — follow shared-references/output-language.md. Code, paths, config keys and ledger IDs stay English regardless.
Interaction rule (HARD CONSTRAINT)
Resolve ASK once from $ARGUMENTS before Phase 0 and hold it for the run.
ASK=never — non-blocking
Runs end-to-end with zero external approval: no AskUserQuestion, no "should I…", no "please confirm", no waiting. Framework choice, file layout, whether to overwrite, whether to install a dependency, which default to pick — all decided here, and the consequential ones logged. The author reviews the ledger and the diff after the run.
Autonomy is not permission to be vague. Every decision you make instead of asking that changes an interface or a meaning is a decision you owe the author a row for.
ASK=semantic — blocking at batch points
The run stops and ends the turn at a batch point and resumes only on an explicit reply. Never implement this as "ask, then continue if no answer arrives" — once the turn ends, silence cannot resume the run.
Batch points (the only places questions are allowed): B0, end of Phase 0, before the ladder is built · B1..Bn, start of each rung, before that rung's code · Bd, a debugging fork where the fix itself is a semantic choice ("shapes don't match: pad left or right?").
Collect the batch and ask it in one call, never one question at a time. The chosen default is always option 1, labelled (default), so accepting everything as-is is one keystroke and produces exactly what ask: never would have. An answer of "you decide" (or an Other reply that declines to choose) falls back to that default, records Source: default (deferredtoauthor), and is never re-asked. A batch point with nothing in it is skipped silently — it is not a checkpoint to announce.
Do not combine ask: semantic with /loop, CronCreate, or any overnight cadence. A blocking gate on an unattended run is a run that did nothing. Detect this at Phase 0 — if there is no interactive author, say so and stop rather than silently downgrading to never.
Acceptance-gate provenance
Per shared-references/acceptance-gate.md:
The terminating condition of the build loop is Type-A only. On a green run this skill says "the spine runs and every MUST rung's check passed". It never says the implementation is correct, the method works, or the numbers mean anything — a passing smoke test is an execution fact, not a result.
The one Type-B gate it does own is Phase 4, and it is owned for a reason: "what did I assume without saying so" is precisely the question an author cannot answer about their own work, because the assumptions they absorbed are the ones they stopped seeing. That needs a reader from a different family, not a second pass by the same one.
Artifacts
All under implement-stage/ (stage-scoped per shared-references/output-manifest.md; stage = implementation):
Create implement-stage/ if absent. Do not create a MANIFEST.md — this run produces well under the 15-artifact threshold.
The assumption ledger
Schema
implement-stage/ASSUMPTIONS.md:
# Assumption Ledger — <target>
<!-- ASK mode: never | semantic -->
| ID | Under-determined by the request | Chosen | Class | Source |
|----|--------------------------------|--------|-------|--------|
| A-001 | request says "on the benchmark", does not say which split | validation | semantic | user |
| A-002 | no tokenizer named | reuse the repo's existing `BPE-32k` | interface | default |
## Notes
Prose, only where a decision is genuinely contested: the alternative that was
rejected and why, what reversing it would cost, and the one-line override.
- **A-001** — `test` is the held-out split and `train` leaks; `validation` is the
only choice that leaves the number meaning what a reader assumes. Reversing it
is one line in `configs/eval.yaml`.Which decisions get a row. Only interface and semantic ones:
Naming, log format, file layout, and anything internal to one module that is invisible at its interface: just make the call. They do not get rows. A ledger that logs variable names buries the two rows that actually decide what the work will later claim, and turns every decision into a form.
The semantic class is the whole point. An undeclared interface assumption costs a refactor. An undeclared semantic assumption is how an implementation quietly decides what the research will later claim.
Source records who decided the row:
Under ask: semantic, a plain default row in the semantic class is exactly an ambiguity the skill did not recognise as an ambiguity in time to ask about it — which is the most interesting row in the ledger, and the first thing Phase 4 looks at. A default (deferredtoauthor) row is not that: it was recognised, asked, and handed back.
A row whose decision has no single code site is legal — say so in the Chosen cell. What is not legal is a consequential decision with no row.
Stub discipline
F0 is allowed to fake things; it is not allowed to hide that it faked them. Anything standing in for real behaviour — synthetic data, a hardcoded return, a stub model, a constant where a computation belongs — is labelled at its site:
# PLACEHOLDER: returns a fixed 0.5; real scorer lands at rung F3Two rules:
result. Prefix such values PLACEHOLDER in the artifact, or write them to *_smoke.json — never to a results path.
- A stub that produces a number* never surfaces in a path that reads like a
live.** Every stub that survives the run is listed in the final report with the rung that would retire it.
- **A rung is not green while a stub that rung was supposed to replace is still
This is shared-references/capture-antipatterns.md applied one stage earlier: a stub number that escapes into a results file is how a placeholder hardens into a cited finding.
Phase 0 — Read the request, open the ledger
it; a FILE.md#section reference → read that section; free text → use it verbatim; empty → take the topmost unchecked task from the most recent PLAN.md / TODO.md / EXPERIMENTPLAN.md in cwd.
- Resolve the target. $ARGUMENTS is, in priority order: a file path → read
More skills from wanshuiyin/Auto-claude-code-research-in-sleep
- Aablation-plannerUse when main results pass result-to-claim (claim_supported=yes or partial) and ablation studies are needed for paper submission.
- Aablation-plannerUse when main results pass result-to-claim (`claim_supported = yes` or `partial`) and ablation studies are needed for paper submission. A secondary Codex agent designs ablations from a reviewer's perspective; the local executor reviews feasibility and implements.
- AalphaxivQuick single-paper lookup via AlphaXiv LLM-optimized summaries with tiered source fallback. Use when user says "explain this paper", "summarize paper", pastes an arXiv/AlphaXiv URL, or provides a bare arXiv ID for quick understanding - not for broad literature search.
- AalphaxivQuick single-paper lookup via AlphaXiv LLM-optimized summaries with tiered source fallback. Use when user says "explain this paper", "summarize paper", pastes an arXiv/AlphaXiv URL, or provides a bare arXiv ID for quick understanding - not for broad literature search.
- Aanalyze-resultsAnalyze ML experiment results, compute statistics, generate comparison tables and insights. Use when user says "analyze results", "compare", or needs to interpret experimental data.
- Aanalyze-resultsAnalyze ML experiment results, compute statistics, generate comparison tables and insights. Use when user says \"analyze results\", \"compare\", or needs to interpret experimental data.
- AarxivSearch, download, and summarize academic papers from arXiv. Use when user says "search arxiv", "download paper", "fetch arxiv", "arxiv search", "get paper pdf", or wants to find and save papers from arXiv to the local paper library.
- AarxivSearch, download, and summarize academic papers from arXiv. Use when user says \"search arxiv\", \"download paper\", \"fetch arxiv\", \"arxiv search\", \"get paper pdf\", or wants to find and save papers from arXiv to the local paper library.
- Aauto-paper-improvement-loopAutonomously improve a generated paper via GPT-6-Astra xhigh review → implement fixes → recompile, for 2 rounds. Use when user says \"改论文\", \"improve paper\", \"论文润色循环\", \"auto improve\", or wants to iteratively polish a generated paper.
- Aauto-paper-improvement-loopAutonomously improve a generated paper via Claude review through claude-review MCP → implement fixes → recompile, for 2 rounds. Use when user says \"改论文\", \"improve paper\", \"论文润色循环\", \"auto improve\", or wants to iteratively polish a generated paper.
- Aauto-paper-improvement-loopAutonomously improve a generated paper via Gemini review through gemini-review MCP → implement fixes → recompile, for 2 rounds. Use when user says \"改论文\", \"improve paper\", \"论文润色循环\", \"auto improve\", or wants to iteratively polish a generated paper.
- Aauto-paper-improvement-loopAutonomously improve a generated paper via GPT-6-Astra xhigh review → implement fixes → recompile, for 2 rounds. Use when user says \"改论文\", \"improve paper\", \"论文润色循环\", \"auto improve\", or wants to iteratively polish a generated paper.