codex-qa skill
QA the omo Codex Light edition (lazycodex / packages/omo-codex) itself, in strict isolation so ONLY our plugin is exercised, never the user's real ~/.codex. The first-party method drives the real `codex app-server` against an isolated CODEX_HOME plus a LOCAL mock model (no real API call), and proves a plugin hook fired by asserting hook/started + hook/completed notifications. Also: isolated install verification, per-component hook probes, a tmux TUI smoke, and runtime log observation (RUST_LOG / logs SQLite / /debug-config). Ships tested helper scripts each with a --self-test. Use whenever someone changes anything under packages/omo-codex or wants to QA, smoke-test, verify, or debug the Codex plugin, its hooks/components, the installer/config.toml, the app-server flow, or the Codex TUI. Triggers: codex qa, qa codex, codex-qa, test codex plugin, verify codex hook, codex app-server, lazycodex qa, isolated CODEX_HOME, prove codex hook fired, codex tui test.
Is the codex-qa skill safe?
Clean: nothing in its files matched our rules. We read 16 files in the folder on 2026-09-28.
No findings.
Install the codex-qa skill
A skill is a folder. Copy it into your agent's skills folder and the agent loads it when the task matches its description.
git clone --depth 1 https://github.com/code-yeongyu/oh-my-openagent.git /tmp/oh-my-openagent mkdir -p ~/.claude/skills cp -r /tmp/oh-my-openagent/.agents/skills/codex-qa ~/.claude/skills/codex-qa
In the Claude apps, zip the folder and upload it from the Skills settings. The folder on GitHub
The instructions your agent would load
SKILL.md as published, without the frontmatter. Read it on GitHub
Codex QA
QA the omo Codex Light edition (packages/omo-codex/, shipped as lazycodex). We exercise OUR plugin in a REAL Codex while touching nothing of the user's setup: an isolated CODEX_HOME + a local mock model means no real API call and the real ~/.codex is never read or written. Each helper script ships a --self-test that asserts its scenario against the live machine, so the scripts are both the QA tools and their own regression checks.
Verified against codex-cli 0.140.0 (node, jq, tmux, bun on macOS). Confirm with codex --version; check a flag with codex --help.
Golden rules (read before running anything)
CODEXHOME (created by cqamkisolatedhome) and a LOCAL mock model provider (cqastartmock). Never QA against the real ~/.codex, never hit a real model API. The bundled scripts enforce this; if you run codex by hand, export CODEXHOME="$(mktemp -d)/codex"; mkdir -p "$CODEXHOME" FIRST (a set CODEX_HOME must already exist or codex hard-errors).
- QA ONLY our plugin. Everything that spawns codex uses an isolated
~/.codex/config.toml before and after and asserts it is unchanged. If you script by hand, do the same.
- Prove the real home stayed clean. Every script shasums
Bash scripts bypass it and get the real binary; never rely on the interactive alias. See references/isolation.md.
- The interactive codex is a shell function that injects --profile quotio.
stream (hook/started / hook/completed), not log scraping. See references/app-server.md.
- The first-party way to prove a hook fired is the app-server notification
.omo/evidence/-/ (no evidence file == the QA did not happen). That directory is gitignored: the files stay local, the PR body carries the summary and decisive excerpts, and nothing under it is ever committed.
- The captured JSON / pane IS the evidence — write it under
Setup
cd <this-skill-dir> # .agents/skills/codex-qa
bash scripts/lib/common.sh --self-check # confirm deps + isolation harnessDocker is the default QA surface. Run this QA inside a disposable container that has the latest codex and a copy of your config, with the host ~/.codex untouched: script/agent/qa-docker.sh (see references/docker-qa.md). The local scripts below are the fallback for when Docker is unavailable or on Windows.
Router: pick your case
Scripts index (each is its own regression test)
To tell a dev dogfood build apart from a published one on a REAL ~/.codex (NOT the isolated QA home), the repo ships bun run install:codex-dev, which stamps the plugin version as dev — visible as the (OmO dev) hook-status prefix every turn and as a [DEV] badge in omo get-local-version. Use it to confirm which build is loaded during manual dogfooding; it writes to the real home, so it is NEVER part of the isolated QA flow above.
When TUI visual QA evidence is needed, follow docs/reference/web-terminal-visual-qa.md: render the TUI through the real xterm.js web terminal - NEVER the tmux capture-pane frame, which degrades color and CJK width:
node script/qa/web-terminal-visual-qa.mjs --title "Codex TUI QA" \
--command "codex" --input "{Enter}" \
--evidence-dir .omo/evidence/<slug>/codex-web-terminalThe helper runs a real pty, renders it in xterm.js under Chrome, and writes terminal.txt, terminal-ansi.txt, terminal.png (true color), and metadata.json (--from-file replays a saved raw stream). Use that artifact set for TUI visual QA; use app-server-drive.sh --plugin for assertion-grade hook behavior.
Match QA to your change scope
hook-unit-probe.sh for the exact stdout, THEN app-server-drive.sh --plugin to prove the live wiring. See components-hooks.md.
- Component / hook logic (packages/omo-codex/plugin/components/*):
install-verify.sh.
- Installer / config.toml (packages/omo-codex/src/install/*):
app-server-drive.sh --plugin, and tui-smoke.sh --plugin if the TUI path matters.
- Anything that affects a live session (hooks, agents, MCP wiring):
Capturing evidence
ev=".omo/evidence/$(date +%Y%m%d)-codex-qa-<slug>"; mkdir -p "$ev"
bash scripts/app-server-drive.sh --plugin > "$ev/app-server-drive.json" 2>&1
bash scripts/install-verify.sh --self-test > "$ev/install-verify.txt" 2>&1On /debugging
There is no /debugging command in Codex. To observe a run: the app-server notification stream (above), RUSTLOG=debug on the app-server's stderr, the logs SQLite under $CODEXHOME, the TUI's /debug-config, and the codex debug … subcommands. See logging-debug.md.
More skills from code-yeongyu/oh-my-openagent
- Aast-grepSearches and rewrites code by AST shape across 25 languages. Use when the target is a syntax pattern (every call/class/import shaped like X, a codemod, a YAML rule) rather than literal text; for plain strings, comments, or filenames, use rg.
- AbrowserDrives a real browser through the omowright library from the js eval kernel: sites the user is already signed into, forms and clicks, JS-rendered pages, screenshots, web QA, extension popups, a human handoff for login, CAPTCHA or OTP, and a browser you own for scraping, bot-scored targets, network capture and QA traces. Use for any interactive browser task; not for a plain search or an unblocked static fetch.
- Acoding-agent-sessionsFinds, reads, and reconstructs coding-agent sessions across Codex, Claude, OpenCode, OMO/Senpi, and other local agent logs. Use when asked to find or search past sessions, transcripts, or subagent runs, or to recover what an earlier session did.
- Acomment-checkerUse when Codex needs to understand or respond to automatic comment-checker feedback emitted after an edit-like PostToolUse hook.
- Adag-libraryStores a DAG definition once and re-runs it by name, instead of pasting the definition into every run. Use when the user wants to save a DAG, run a saved one, or schedule the same multi-agent graph repeatedly.
- Ddata-scientistProcesses and analyzes data with resident-kernel engines (DuckDB, Polars) and one-shot tools. Use for CSV/parquet/JSON analysis, group-by/join/aggregation, time series, distributions, cleaning, or plotting a dataset.
- AdebuggingRuns a hypothesis-driven debugging loop across any language or binary, escalating to orthogonal oracle angles and locking the fix with a failing test. Use for crashes, silent failures, hangs, wrong responses, memory leaks, async misbehavior, or reverse engineering.
- Adev-browserBrowser automation with persistent page state. Use when users ask to navigate websites, fill forms, take screenshots, extract web data, test web apps, or automate browser workflows. Trigger phrases include "go to [url]", "click on", "fill out the form", "take a screenshot", "scrape", "automate", "test the website", "log into", or any browser interaction request.
- CfrontendBuilds, styles, and polishes web UI and UX. Use for any frontend, page, component, styling, layout, animation, or visual-quality task, or when asked to make an interface look or feel a certain way.
- AfrontendBuilds, styles, and polishes web UI and UX. Use for any frontend, page, component, styling, layout, animation, or visual-quality task, or when asked to make an interface look or feel a certain way.
- Aget-unpublished-changesCompare HEAD with the latest published npm versions and list all unpublished changes by release layer. Triggers: unpublished changes, changelog, what changed, whats new.
- Agit-masterHandles git work: atomic commits, rebase, squash, blame, bisect, reflog, and history questions. Use whenever a task needs a commit or a git-history investigation; skip for ordinary code edits.