experiment-bridge skill
Workflow 1.5: Bridge between idea discovery and auto review. Reads EXPERIMENT_PLAN.md, implements experiment code, deploys to GPU, collects initial results. Use when user says \"实现实验\", \"implement experiments\", \"bridge\", \"从计划到跑实验\", \"deploy the plan\", or has an experiment plan ready to execute.
Is the experiment-bridge skill safe?
Clean: nothing in its files matched our rules. We read 1 file in the folder on 2026-09-28.
No findings.
Install the experiment-bridge skill
A skill is a folder. Copy it into your agent's skills folder and the agent loads it when the task matches its description.
git clone --depth 1 https://github.com/wanshuiyin/Auto-claude-code-research-in-sleep.git /tmp/Auto-claude-code-research-in-sleep mkdir -p ~/.claude/skills cp -r /tmp/Auto-claude-code-research-in-sleep/skills/skills-codex/experiment-bridge ~/.claude/skills/experiment-bridge
In the Claude apps, zip the folder and upload it from the Skills settings. The folder on GitHub
The instructions your agent would load
SKILL.md as published, without the frontmatter. Read it on GitHub
Workflow 1.5: Experiment Bridge
Implement and deploy experiments from plan: $ARGUMENTS
Overview
This skill bridges Workflow 1 (idea discovery + method refinement) and Workflow 2 (auto review loop). It takes the experiment plan and turns it into running experiments with initial results.
Workflow 1 output: This skill: Workflow 2 input:
refine-logs/EXPERIMENT_PLAN.md → implement → deploy → collect → initial results ready
refine-logs/EXPERIMENT_TRACKER.md code /run-experiment for /auto-review-loop
refine-logs/FINAL_PROPOSAL.mdConstants
- AUTODEPLOY = true** — Automatically deploy experiments after implementation. Set false to review code before deploying.
- CODEREVIEW = true** — Secondary Codex reviewer with xhigh reasoning reviews experiment code before deployment. Catches logic bugs before wasting GPU hours. Set false to skip.
- SANITYFIRST = true** — Run the sanity-stage experiment first (smallest, fastest) before launching the rest. Catches setup bugs early.
- MAXPARALLELRUNS = 4 — Maximum number of experiments to deploy in parallel (limited by available GPUs).
- BASEREPO = false** — GitHub repo URL to use as a base codebase. When set, clone it first and implement experiments on top of it.
- COMPACT = false — When true, prefer idea-stage/IDEACANDIDATES.md over the full idea-stage/IDEAREPORT.md, and append completed runs to EXPERIMENT_LOG.md.
- BACKENDS = local | ssh | vast | modal — Preserve the Claude mainline backend lifecycle. Vast.ai and Modal routes are first-class when configured; do not silently fall back to local execution if the user requested either backend.
- RESCUEONFAILURE = true — If sanity or deployment fails, run a Codex-native rescue / second opinion review before abandoning the experiment plan.
Override: /experiment-bridge "EXPERIMENT_PLAN.md" — compact: true, base repo: https://github.com/org/project
Inputs
This skill expects one or more of:
- refine-logs/EXPERIMENTPLAN.md** (best) — claim-driven experiment roadmap from /experiment-plan
- refine-logs/EXPERIMENTTRACKER.md** — run-by-run execution table
- refine-logs/FINALPROPOSAL.md** — method description for implementation context
- idea-stage/IDEACANDIDATES.md — compact idea summary (preferred when COMPACT = true) (fall back to ./IDEACANDIDATES.md if not found)
- idea-stage/IDEAREPORT.md — fallback if refine-logs don't exist (fall back to ./IDEAREPORT.md if not found)
If none exist, ask the user what experiments to implement.
Workflow
Phase 1: Parse the Experiment Plan
Read EXPERIMENT_PLAN.md and extract:
- Run order and milestones — which experiments run first (sanity → baseline → main → ablation → polish)
- For each experiment block:
- Dataset / split / task
- Compared systems and variants
- Metrics to compute
- Setup details (backbone, hyperparameters, seeds)
- Success criterion
- Priority (MUST-RUN vs NICE-TO-HAVE)
- Compute budget — total estimated GPU-hours
- Method details from FINAL_PROPOSAL.md — what exactly to implement
Present a brief summary:
📋 Experiment plan loaded:
- Milestones: [N] (sanity → baseline → main → ablation)
- Must-run experiments: [N]
- Nice-to-have: [N]
- Estimated GPU-hours: [X]
Proceeding to implementation.Research-contract fallback: if idea-stage/docs/researchcontract.md does not exist yet, create it now from templates/RESEARCHCONTRACT_TEMPLATE.md using the selected idea + claims from the experiment plan — downstream /result-to-claim and /ablation-planner read it as the claims source.
Phase 2: Implement Experiment Code
If BASEREPO is set** — clone the repo first:
git clone <BASE_REPO> base_repo/For each milestone (in order), write the experiment scripts:
- Check existing code — scan the project (or cloned base_repo/) for existing experiment scripts, model code, and data loaders. Reuse as much as possible.
- Implement missing pieces:
- Training scripts with proper argparse (all hyperparameters configurable)
- Evaluation scripts computing the specified metrics
- Data loading / preprocessing if needed
- Baseline implementations if not already present
- Fixed random seeds for reproducibility
- Results saved to JSON/CSV for later analysis
- Proper logging (wandb if configured in AGENTS.md)
- Follow the plan's run order — implement sanity-stage experiments first, then baselines, then main method, then ablations.
- Self-review before deploying:
- Are all hyperparameters from EXPERIMENT_PLAN.md reflected in argparse?
- Is the random seed fixed and controllable?
- Are results saved in a parseable format (JSON/CSV)?
- Does the code match FINAL_PROPOSAL.md's method description?
- CRITICAL: does evaluation compare predictions against dataset ground truth, never another model's output?
Phase 2.5: Fresh-Agent Code Review (same-family provisional; when CODE_REVIEW = true)
Skip this step if CODE_REVIEW is false.
Before deploying, send the experiment code to a secondary Codex reviewer with xhigh reasoning:
spawn_agent:
model: gpt-6-astra
reasoning_effort: xhigh
message: |
Review the following experiment implementation for correctness.
## Experiment Plan
[paste key sections from EXPERIMENT_PLAN.md]
## Method Description
[paste from FINAL_PROPOSAL.md]
## Implementation
[paste the experiment scripts or exact file paths plus relevant snippets]
Check for:
1. Does the code correctly implement the method described in the proposal?
2. Are all hyperparameters from the plan reflected in the code?
3. Are there logic bugs: wrong loss, wrong data split, missing eval, leakage, metric mismatch?
4. Is the evaluation metric computed against ground truth, not another model's output?
5. Are seeds, result paths, logging, and failure handling sufficient for reproducible experiments?
Output:
- BLOCKING issues that must be fixed before deployment
- NON-BLOCKING issues that can wait
- Suggested patches or checks
=== SCOPE LIMITS (these bound what you PROPOSE, never what you look for) ===
Report anything that is actually wrong here — including a rare-looking case, if
this repo actually produces it. Then keep the fix iIf BLOCKING issues are found, fix them and re-run this review once before Phase 3. Save the reviewer response and any fixes in refine-logs/EXPERIMENTCODEREVIEW.md. If reviewer delegation is unavailable, run the same checklist locally and mark the review [local-only].
Phase 3: Sanity Check (if SANITY_FIRST = true)
Before deploying the full experiment suite, run the sanity-stage experiment:
/run-experiment [sanity experiment command]Wait for completion. Verify:
- Training loop runs without errors
- Metrics are computed and saved correctly
- GPU memory usage is within bounds
- Output format matches expectations
If sanity fails → READ the traceback/stderr/logs first, then fix the code and re-run — never re-run unchanged hoping for a different outcome. (The same read-the-primary-artifact discipline applies to surprising REVIEWER verdicts: see shared-references/review-tracing.md § Debugging With Traces.) After 1–2 failed patches on the same failure, discard and reimplement the failing script cleanly from the plan — a peer move to another patch, not a last resort; delete only the attempt's own code, never the plan / tracker / data / results (per external-cadence.md, "Let a broken attempt restart, not just patch"). Two clean reimplements failing the same way put the plan or the environment in question — report that explicitly. Do not proceed to full deployment with broken code.
If the same sanity failure repeats, trigger a second opinion: summarize the plan, code diff, command, logs, backend, and failure, then ask a fresh Codex reviewer agent for a rescue diagnosis. Apply only concrete fixes grounded in the logs.
Phase 4: Deploy Full Experiments
Deploy experiments following the plan's milestone order. Route by job count and dependencies:
/run-experiment [experiment commands]For large batches (≥10 jobs), multi-seed sweeps, or teacher→student phase dependencies, use the queue scheduler:
/experiment-queue [grid spec or manifest]Auto-routing rule: if any milestone in EXPERIMENTPLAN.md declares ≥10 jobs or declares phase dependencies, route that milestone to /experiment-queue; otherwise use /run-experiment. /experiment-queue adds OOM-aware retry with backoff, stale-screen cleanup, wave-transition race prevention, phase dependency enforcement, and crash-safe state persistence in queuestate.json.
For each milestone:
- Deploy experiments in parallel (up to MAXPARALLELRUNS for /run-experiment, or max_parallel from the queue manifest for /experiment-queue)
- Use /monitor-experiment to track progress; if /experiment-queue is active, monitor queue_state.json
- Collect results as experiments complete
Backend lifecycle rules:
- Vast.ai: record instance id, SSH endpoint, mounted data/checkpoints, estimated hourly cost, and cleanup policy. If auto_destroy is configured, write the exact cleanup command before launch.
- Modal: verify app/function, image/dependencies, secrets, volumes, and output persistence before launch.
- Local/SSH: verify GPU availability, environment activation, log path, and result path before launching.
- If a backend is unreachable or misconfigured, stop with a configuration issue instead of silently switching backend.
🚦 Checkpoint (if AUTODEPLOY = false):**
🔧 Code implementation complete. Ready to deploy:
Milestone 0 (sanity): [status — passed/pending]
Milestone 1 (baseline): [N experiments, ~X GPU-hours]
Milestone 2 (main method): [N experiments, ~X GPU-hours]
Milestone 3 (ablations): [N experiments, ~X GPU-hours]
Total estimated: ~X GPU-hours on [N] GPUs
Deploy now? Or review the code first?Phase 5: Collect Initial Results
As experiments complete:
- Parse output files (JSON/CSV/logs) for key metrics
- Training quality check — if W&B data is available, invoke /training-check to detect NaN, loss divergence, plateaus, or overfitting. If W&B is not configured, skip silently.
- Update refine-logs/EXPERIMENTTRACKER.md** — fill in Status and Notes columns
- Check success criteria from EXPERIMENT_PLAN.md — did each experiment meet its bar?
- Write initial results summary:
# Initial Experiment Results
**Date**: [today]
**Plan**: refine-logs/EXPERIMENT_PLAN.md
## Results by Milestone
### M0: Sanity — PASSED
- [result]
### M1: Baselines
| Run | System | Key Metric | Status |
|-----|--------|-----------|--------|
| R001 | baseline_1 | X.XX | DONE |
### M2: Main Method
| Run | System | Key Metric | Status |
|-----|--------|-----------|--------|
| R003 | our_method | X.XX | DONE |
### M3: Ablations
...
## Summary
- [X/Y] must-run experiments completed
- Main result: [positive/negative/inconclusive]
- Ready for /auto-review-loop: [YES/NO]
## Next Step
→ /auto-review-loop "[topic]"Phase 5.5: Write Compact Log (when COMPACT = true)
More skills from wanshuiyin/Auto-claude-code-research-in-sleep
- Aablation-plannerUse when main results pass result-to-claim (claim_supported=yes or partial) and ablation studies are needed for paper submission.
- Aablation-plannerUse when main results pass result-to-claim (`claim_supported = yes` or `partial`) and ablation studies are needed for paper submission. A secondary Codex agent designs ablations from a reviewer's perspective; the local executor reviews feasibility and implements.
- AalphaxivQuick single-paper lookup via AlphaXiv LLM-optimized summaries with tiered source fallback. Use when user says "explain this paper", "summarize paper", pastes an arXiv/AlphaXiv URL, or provides a bare arXiv ID for quick understanding - not for broad literature search.
- AalphaxivQuick single-paper lookup via AlphaXiv LLM-optimized summaries with tiered source fallback. Use when user says "explain this paper", "summarize paper", pastes an arXiv/AlphaXiv URL, or provides a bare arXiv ID for quick understanding - not for broad literature search.
- Aanalyze-resultsAnalyze ML experiment results, compute statistics, generate comparison tables and insights. Use when user says "analyze results", "compare", or needs to interpret experimental data.
- Aanalyze-resultsAnalyze ML experiment results, compute statistics, generate comparison tables and insights. Use when user says \"analyze results\", \"compare\", or needs to interpret experimental data.
- AarxivSearch, download, and summarize academic papers from arXiv. Use when user says "search arxiv", "download paper", "fetch arxiv", "arxiv search", "get paper pdf", or wants to find and save papers from arXiv to the local paper library.
- AarxivSearch, download, and summarize academic papers from arXiv. Use when user says \"search arxiv\", \"download paper\", \"fetch arxiv\", \"arxiv search\", \"get paper pdf\", or wants to find and save papers from arXiv to the local paper library.
- Aauto-paper-improvement-loopAutonomously improve a generated paper via GPT-6-Astra xhigh review → implement fixes → recompile, for 2 rounds. Use when user says \"改论文\", \"improve paper\", \"论文润色循环\", \"auto improve\", or wants to iteratively polish a generated paper.
- Aauto-paper-improvement-loopAutonomously improve a generated paper via Claude review through claude-review MCP → implement fixes → recompile, for 2 rounds. Use when user says \"改论文\", \"improve paper\", \"论文润色循环\", \"auto improve\", or wants to iteratively polish a generated paper.
- Aauto-paper-improvement-loopAutonomously improve a generated paper via Gemini review through gemini-review MCP → implement fixes → recompile, for 2 rounds. Use when user says \"改论文\", \"improve paper\", \"论文润色循环\", \"auto improve\", or wants to iteratively polish a generated paper.
- Aauto-paper-improvement-loopAutonomously improve a generated paper via GPT-6-Astra xhigh review → implement fixes → recompile, for 2 rounds. Use when user says \"改论文\", \"improve paper\", \"论文润色循环\", \"auto improve\", or wants to iteratively polish a generated paper.