senpi-qa skill
QA the omo Senpi adapter (packages/omo-senpi, packages/senpi-task) against the REAL senpi binary in strict isolation, and write every artifact to the one canonical evidence path .omo/evidence/omo-senpi-adapter/<slug>/. The live drivers under packages/omo-senpi/scripts/qa/ create their own isolated SENPI_CODING_AGENT_DIR and ignore the caller's, so the real ~/.senpi/agent is never written. Ships scripts/resolve-evidence-dir.mjs, which is the ONLY sanctioned way to pick an evidence directory: it rejects traversal, separators, absolute paths, and stray roots such as local-ignore/qa-evidence. Use whenever someone changes anything under packages/omo-senpi or packages/senpi-task, or wants to QA, smoke-test, verify, or debug the Senpi adapter, the task/team engine, the DAG, task RPC, or skill delivery. Triggers: senpi qa, qa senpi, senpi-qa, test senpi adapter, verify senpi task, senpi task e2e, senpi team e2e, task dag qa, live senpi driver, senpi evidence path.
Is the senpi-qa skill safe?
Clean: nothing in its files matched our rules. We read 3 files in the folder on 2026-09-28.
No findings.
Install the senpi-qa skill
A skill is a folder. Copy it into your agent's skills folder and the agent loads it when the task matches its description.
git clone --depth 1 https://github.com/code-yeongyu/oh-my-openagent.git /tmp/oh-my-openagent mkdir -p ~/.claude/skills cp -r /tmp/oh-my-openagent/.agents/skills/senpi-qa ~/.claude/skills/senpi-qa
In the Claude apps, zip the folder and upload it from the Skills settings. The folder on GitHub
The instructions your agent would load
SKILL.md as published, without the frontmatter. Read it on GitHub
Senpi QA
QA the omo Senpi adapter (packages/omo-senpi/) and the task engine (packages/senpi-task/) by driving the REAL senpi binary. Unit tests never count as live QA here: bun run test:senpi is the package gate, the drivers in packages/omo-senpi/scripts/qa/ are the harness proof.
Golden rules
.omo/evidence/omo-senpi-adapter//. Pick it with scripts/resolve-evidence-dir.mjs and nothing else — a hand-typed path is how runs end up somewhere like local-ignore/qa-evidence/ or a .qa-evidence/ at the worktree root, which is outside the ignored root and gets committed by accident (#8703).
- Evidence lives at exactly one path. Every artifact goes under
tracked-evidence audit test fails the build if any evidence path is tracked. Never git add -f an artifact; the PR body carries the summary and the decisive excerpts.
- Evidence stays local. .omo/evidence/ is gitignored and the
isolated SENPICODINGAGENT_DIR and deliberately IGNORE a caller-provided one, so ~/.senpi/agent is never used as the sandbox. Report the driver's realSenpiUntouched / changed-path fields and the isolated agent-dir path; treat a whole-directory digest as supporting evidence, not proof by itself.
- The real agent dir stays untouched. The live drivers build their own
report SKIP or FAIL in their final JSON rather than degrading to the real home. A SKIP is not a pass — say so in the evidence README.
- No binary means SKIP, not silence. When senpi is absent the live drivers
happen, which means no commit and no push. The file proves the run on the machine that made it; it is not something the commit carries.
- The captured JSON is the evidence. No file on disk means the QA did not
Resolve the evidence directory first
ev="$(node .agents/skills/senpi-qa/scripts/resolve-evidence-dir.mjs \
--repo-root "$(git rev-parse --show-toplevel)" --slug <YYYYMMDD>-<short-slug>)"
mkdir -p "$ev"The resolver returns an absolute path and creates nothing, so the caller decides when the directory appears. A slug is ONE relative segment of lowercase letters, digits, and hyphens (20260820-senpi-qa-contract). Separators, ./.., traversal, absolute paths, and a non-git root are rejected with a non-zero exit and a message naming the offending slug.
Router: pick your case
Point a driver's output at the resolved directory, e.g.:
TASK_E2E_OUT_DIR="$ev/live-task-dag" SENPI_BIN="$(command -v senpi)" \
node packages/omo-senpi/scripts/qa/task-e2e.mjsPackage gate
tsgo --noEmit -p packages/omo-senpi/tsconfig.json
bun run test:senpiWrite the evidence README
Every run leaves $ev/README.md a reviewer can read without rerunning anything. The required sections are the repo-wide evidence rules in the root AGENTS.md (what was tested / observed / why it is enough / what was omitted). For Senpi, record the driver's changed-path/isolation fields and sandbox agent-dir path. Some drivers report sandbox paths without removing them; the caller must delete every task-owned sandbox and verify child PIDs are terminal before writing the cleanup receipt.
More skills from code-yeongyu/oh-my-openagent
- Aast-grepSearches and rewrites code by AST shape across 25 languages. Use when the target is a syntax pattern (every call/class/import shaped like X, a codemod, a YAML rule) rather than literal text; for plain strings, comments, or filenames, use rg.
- AbrowserDrives a real browser through the omowright library from the js eval kernel: sites the user is already signed into, forms and clicks, JS-rendered pages, screenshots, web QA, extension popups, a human handoff for login, CAPTCHA or OTP, and a browser you own for scraping, bot-scored targets, network capture and QA traces. Use for any interactive browser task; not for a plain search or an unblocked static fetch.
- Acodex-qaQA the omo Codex Light edition (lazycodex / packages/omo-codex) itself, in strict isolation so ONLY our plugin is exercised, never the user's real ~/.codex. The first-party method drives the real `codex app-server` against an isolated CODEX_HOME plus a LOCAL mock model (no real API call), and proves a plugin hook fired by asserting hook/started + hook/completed notifications. Also: isolated install verification, per-component hook probes, a tmux TUI smoke, and runtime log observation (RUST_LOG / logs SQLite / /debug-config). Ships tested helper scripts each with a --self-test. Use whenever someone changes anything under packages/omo-codex or wants to QA, smoke-test, verify, or debug the Codex plugin, its hooks/components, the installer/config.toml, the app-server flow, or the Codex TUI. Triggers: codex qa, qa codex, codex-qa, test codex plugin, verify codex hook, codex app-server, lazycodex qa, isolated CODEX_HOME, prove codex hook fired, codex tui test.
- Acoding-agent-sessionsFinds, reads, and reconstructs coding-agent sessions across Codex, Claude, OpenCode, OMO/Senpi, and other local agent logs. Use when asked to find or search past sessions, transcripts, or subagent runs, or to recover what an earlier session did.
- Acomment-checkerUse when Codex needs to understand or respond to automatic comment-checker feedback emitted after an edit-like PostToolUse hook.
- Adag-libraryStores a DAG definition once and re-runs it by name, instead of pasting the definition into every run. Use when the user wants to save a DAG, run a saved one, or schedule the same multi-agent graph repeatedly.
- Ddata-scientistProcesses and analyzes data with resident-kernel engines (DuckDB, Polars) and one-shot tools. Use for CSV/parquet/JSON analysis, group-by/join/aggregation, time series, distributions, cleaning, or plotting a dataset.
- AdebuggingRuns a hypothesis-driven debugging loop across any language or binary, escalating to orthogonal oracle angles and locking the fix with a failing test. Use for crashes, silent failures, hangs, wrong responses, memory leaks, async misbehavior, or reverse engineering.
- Adev-browserBrowser automation with persistent page state. Use when users ask to navigate websites, fill forms, take screenshots, extract web data, test web apps, or automate browser workflows. Trigger phrases include "go to [url]", "click on", "fill out the form", "take a screenshot", "scrape", "automate", "test the website", "log into", or any browser interaction request.
- CfrontendBuilds, styles, and polishes web UI and UX. Use for any frontend, page, component, styling, layout, animation, or visual-quality task, or when asked to make an interface look or feel a certain way.
- AfrontendBuilds, styles, and polishes web UI and UX. Use for any frontend, page, component, styling, layout, animation, or visual-quality task, or when asked to make an interface look or feel a certain way.
- Aget-unpublished-changesCompare HEAD with the latest published npm versions and list all unpublished changes by release layer. Triggers: unpublished changes, changelog, what changed, whats new.