experiment-audit skill
Audit experiment integrity before claiming results. Uses fresh-agent GPT-6-Astra review (same-family provisional in the base Codex mirror) to check for fake ground truth, score normalization fraud, phantom results, and insufficient scope. Use when user says \"审计实验\", \"check experiment integrity\", \"audit results\", \"实验诚实度\", or after experiments complete before writing claims.
Is the experiment-audit skill safe?
Clean: nothing in its files matched our rules. We read 1 file in the folder on 2026-09-28.
No findings.
Install the experiment-audit skill
A skill is a folder. Copy it into your agent's skills folder and the agent loads it when the task matches its description.
git clone --depth 1 https://github.com/wanshuiyin/Auto-claude-code-research-in-sleep.git /tmp/Auto-claude-code-research-in-sleep mkdir -p ~/.claude/skills cp -r /tmp/Auto-claude-code-research-in-sleep/skills/skills-codex/experiment-audit ~/.claude/skills/experiment-audit
In the Claude apps, zip the folder and upload it from the Skills settings. The folder on GitHub
The instructions your agent would load
SKILL.md as published, without the frontmatter. Read it on GitHub
Experiment Audit: Fresh-Agent Integrity Verification
Codex assurance: base semantic audit results record
reviewindependence: same-family and acceptancestatus: provisional.
Deterministic evidence checks may be accepted; unavailable reviewer calls emit
BLOCKED/ERROR rather than a provisional PASS.
Audit experiment integrity for: $ARGUMENTS
Why This Exists
LLM agents can produce fraudulent experimental results through:
- Fake ground truth — creating synthetic "reference" from model outputs, then reporting high agreement as performance
- Score normalization — dividing metrics by the model's own max to get 0.99+
- Phantom results — claiming numbers from files that don't exist or functions never called
- Insufficient scope — reporting 2-scene pilots as "comprehensive evaluation"
These are NOT intentional deception — they are failure modes of optimizing agents that lack integrity constraints. This skill adds that constraint.
Core Principle
The executor (Codex) collects file paths. A fresh Codex reviewer reads code and judges integrity. The executor does NOT participate in integrity judgment; this base route is same-family/provisional.
This follows shared-references/reviewer-independence.md and shared-references/experiment-integrity.md.
Constants
- REVIEWERBACKEND = codex — Default: Codex reviewer agent (spawnagent, ultra — deep-audit tier). Override with — reviewer: oracle-pro for GPT-5.5 Pro via Oracle MCP. See shared-references/reviewer-routing.md.
Workflow
Step 1: Collect Artifacts (Executor — Codex)
Locate and list these files WITHOUT reading or summarizing their content:
Scan project directory for:
1. Evaluation scripts: *eval*.py, *metric*.py, *test*.py, *benchmark*.py
2. Result files: *.json, *.csv in results/, outputs/, logs/
3. Ground truth paths: look in eval scripts for data loading (dataset paths, GT references)
4. Experiment tracker: EXPERIMENT_TRACKER.md, EXPERIMENT_LOG.md
5. Paper claims: NARRATIVE_REPORT.md, paper/sections/*.tex, PAPER_PLAN.md
6. Config files: *.yaml, *.toml, *.json configs with metric definitionsDO NOT summarize, interpret, or explain any file content. Only collect paths.
Step 2: Send to Reviewer (GPT-6-Astra via Codex MCP)
Pass ONLY file paths and the audit checklist to the reviewer. The reviewer reads everything directly.
spawn_agent:
model: gpt-6-astra
reasoning_effort: ultra
message: |
You are an experiment integrity auditor. Start from the assumption that the
evaluation is compromised somewhere — your job is to find where. Be
adversarial. Trust nothing the author tells you — verify everything
yourself. Read ALL files listed below and check for the following fraud
patterns.
Files to read:
- Evaluation scripts: [list paths]
- Result files: [list paths]
- Experiment tracker: [list paths]
- Paper claims: [list paths]
- Config files: [list paths]
## Audit Checklist
### A. Ground Truth Provenance
For each evaluation script:
1. Where does "ground truth" / "reference" / "target" come from?
2. Is it loaded from the DATASET, or generated/derived from MODEL OUTPUTS?
3. If derived: is it explicitly labeled as proxy evaluation?
4. Are official eval scripts used when available for this benchmark?
FAIL if: GT is derived from model outputs without explicit proxy labeling.
### B. Score Normalization
For each metric computation:
1. Is any metric divided by max/min/mean of the model's OWN output?
2. Are raw scores Step 3: Parse and Write Report (Executor — Codex)
Parse the reviewer's response and write EXPERIMENT_AUDIT.md:
# Experiment Audit Report
**Date**: [today]
**Auditor**: GPT-6-Astra ultra (fresh same-family agent, read-only, provisional)
**Project**: [project name]
## Overall Verdict: [PASS | WARN | FAIL]
## Integrity Status: [pass | warn | fail]
## Checks
### A. Ground Truth Provenance: [PASS|WARN|FAIL]
[details + file:line evidence]
### B. Score Normalization: [PASS|WARN|FAIL]
[details]
### C. Result File Existence: [PASS|WARN|FAIL]
[details]
### D. Dead Code Detection: [PASS|WARN|FAIL]
[details]
### E. Scope Assessment: [PASS|WARN|FAIL]
[details]
### F. Evaluation Type: [real_gt | synthetic_proxy | ...]
[classification + evidence]
## Action Items
- [specific fixes if WARN or FAIL]
## Claim Impact
- Claim 1: [supported | needs qualifier | unsupported]
- Claim 2: ...Also write EXPERIMENT_AUDIT.json for machine consumption:
{
"audit_skill": "experiment-audit",
"verdict": "WARN",
"reason_code": "scope_exceeds_evidence",
"summary": "Two-scene evaluation supports a qualified claim only.",
"audited_input_hashes": {"results/eval.json": "sha256:<hash>"},
"trace_path": ".aris/traces/experiment-audit/2026-04-10_run01/",
"agent_id": "agent_019f...",
"verdict_id": "agent_019f...",
"executor_model": "codex-gpt-6-astra",
"executor_family": "openai",
"reviewer_model": "gpt-6-astra",
"reviewer_family": "openai",
"reviewer_reasoning": "ultra",
"review_independence": "same-family",
"acceptance_status": "provisional",
"generated_at": "2026-04-10T00:00:00Z",
"date": "2026-04-10",
"auditor": "gpt-6-astra-ultra",
"overall_verdict": "warn",
"integrity_status": "warn",
"checks": {
"gt_provenance": {"status": "pass", "details": "..."},
"score_normalization": {"status": "warn", "details": "..."},
"result_existence": {"status": "pass", "details": "..."},
"dead_code": {"status": "pass", "details": "..."},
"scope": {"status": "warn", "details": "..."},
"eval_type": "real_gt"
},
"claims": [
{"id": "C1", "impact": "supported"},
{"id": "C2", "impact": "nStep 4: Print Summary
🔬 Experiment Audit Complete
GT Provenance: ✅ PASS — real dataset GT used
Score Normalization: ⚠️ WARN — boundary metric uses self-reference
Result Existence: ✅ PASS — all files exist, numbers match
Dead Code: ✅ PASS — all metric functions called
Scope: ⚠️ WARN — 2 scenes, paper says "comprehensive"
Overall: ⚠️ WARN
See EXPERIMENT_AUDIT.md for details.Integration with Other Skills
Automatic in /research-pipeline (advisory, never blocks)
When integrated into the pipeline, this skill runs automatically after /experiment-bridge and before /auto-review-loop:
/experiment-bridge → results ready
↓
/experiment-audit (automatic, advisory)
├── PASS → continue normally
├── WARN → print ⚠️ warning, continue, tag claims as [INTEGRITY: WARN]
└── FAIL → print 🔴 alert, continue, tag claims as [INTEGRITY CONCERN]
↓
/auto-review-loop → proceeds with integrity tags visible to reviewerNever blocks the pipeline. Even on FAIL, the pipeline continues — but claims carry visible integrity tags.
Read by /result-to-claim (if exists)
if EXPERIMENT_AUDIT.json exists:
read integrity_status
attach to verdict: {claim_supported: "yes", integrity_status: "warn"}
if integrity_status == "fail":
downgrade verdict display: "yes [INTEGRITY CONCERN]"
else:
verdict as normal, integrity_status = "unavailable"
mark as "provisional — no integrity audit"Read by /paper-write (if exists)
if EXPERIMENT_AUDIT.json exists AND integrity_status == "fail":
add footnote to affected claims: "Note: integrity audit flagged concerns with this evaluation"Key Rules
- Reviewer independence: executor collects paths, reviewer judges. Period.
- Never block: warn loudly, never halt the pipeline.
- File-as-switch: no EXPERIMENT_AUDIT.md = skill was never run = zero impact on existing behavior.
- Review class: base Codex is same-family provisional; only an overlay may claim cross-family accepted.
- Honest about limits: the audit catches common patterns, not all possible fraud. It is a safety net, not a guarantee.
Acknowledgements
Motivated by community-reported integrity issues (#57, #131) where executor agents created fake ground truth and self-normalized scores.
Review Tracing
After each reviewer agent call, save the trace following shared-references/review-tracing.md (Policy C — forensic; never silently skip). Use savetrace.sh (resolved per the chain in shared-references/integration-contract.md §2) or write files directly to .aris/traces//run/. Respect the --- trace: parameter (default: full).
More skills from wanshuiyin/Auto-claude-code-research-in-sleep
- Aablation-plannerUse when main results pass result-to-claim (claim_supported=yes or partial) and ablation studies are needed for paper submission.
- Aablation-plannerUse when main results pass result-to-claim (`claim_supported = yes` or `partial`) and ablation studies are needed for paper submission. A secondary Codex agent designs ablations from a reviewer's perspective; the local executor reviews feasibility and implements.
- AalphaxivQuick single-paper lookup via AlphaXiv LLM-optimized summaries with tiered source fallback. Use when user says "explain this paper", "summarize paper", pastes an arXiv/AlphaXiv URL, or provides a bare arXiv ID for quick understanding - not for broad literature search.
- AalphaxivQuick single-paper lookup via AlphaXiv LLM-optimized summaries with tiered source fallback. Use when user says "explain this paper", "summarize paper", pastes an arXiv/AlphaXiv URL, or provides a bare arXiv ID for quick understanding - not for broad literature search.
- Aanalyze-resultsAnalyze ML experiment results, compute statistics, generate comparison tables and insights. Use when user says "analyze results", "compare", or needs to interpret experimental data.
- Aanalyze-resultsAnalyze ML experiment results, compute statistics, generate comparison tables and insights. Use when user says \"analyze results\", \"compare\", or needs to interpret experimental data.
- AarxivSearch, download, and summarize academic papers from arXiv. Use when user says "search arxiv", "download paper", "fetch arxiv", "arxiv search", "get paper pdf", or wants to find and save papers from arXiv to the local paper library.
- AarxivSearch, download, and summarize academic papers from arXiv. Use when user says \"search arxiv\", \"download paper\", \"fetch arxiv\", \"arxiv search\", \"get paper pdf\", or wants to find and save papers from arXiv to the local paper library.
- Aauto-paper-improvement-loopAutonomously improve a generated paper via GPT-6-Astra xhigh review → implement fixes → recompile, for 2 rounds. Use when user says \"改论文\", \"improve paper\", \"论文润色循环\", \"auto improve\", or wants to iteratively polish a generated paper.
- Aauto-paper-improvement-loopAutonomously improve a generated paper via Claude review through claude-review MCP → implement fixes → recompile, for 2 rounds. Use when user says \"改论文\", \"improve paper\", \"论文润色循环\", \"auto improve\", or wants to iteratively polish a generated paper.
- Aauto-paper-improvement-loopAutonomously improve a generated paper via Gemini review through gemini-review MCP → implement fixes → recompile, for 2 rounds. Use when user says \"改论文\", \"improve paper\", \"论文润色循环\", \"auto improve\", or wants to iteratively polish a generated paper.
- Aauto-paper-improvement-loopAutonomously improve a generated paper via GPT-6-Astra xhigh review → implement fixes → recompile, for 2 rounds. Use when user says \"改论文\", \"improve paper\", \"论文润色循环\", \"auto improve\", or wants to iteratively polish a generated paper.